Kinetic energy density functional machine learning model

By using an atomic-based kinetic density functional machine learning model in molecular systems, the problems of high computational complexity and limited extrapolation performance of density functional theory in large-scale molecular systems are solved, achieving higher accuracy and lower computational complexity, and significantly accelerating molecular science computation.

CN121569345APending Publication Date: 2026-02-24MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380100460.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-08-11
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing density functional theory (DFT) methods have high computational complexity when simulating large-scale molecular systems, and machine learning methods have limited extrapolation performance for molecular systems, especially for molecules with highly concentrated and steeply varied electron densities, which are difficult to approximate accurately.

Method used

We employ an atom-based kinetic energy density functional machine learning model (KEDF) and use the Graphormer neural network architecture for nonlocal computation. We optimize the density coefficients using the training dataset to generate predicted values ​​of electron non-interaction kinetic energies. By combining local system transformation and gradient labeling techniques, we reduce computational complexity and improve extrapolation performance.

Benefits of technology

It achieves higher accuracy and lower computational complexity on molecular systems than traditional methods, can accurately recover molecular shell structures, and achieves significant acceleration on large-scale systems. Its extrapolation performance is superior to end-to-end mapping methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121569345A_ABST
    Figure CN121569345A_ABST
Patent Text Reader

Abstract

There is provided a computing system comprising processing circuitry configured to receive inference-time molecular structure data for an inference-time molecular structure, the inference-time molecular structure data comprising an inference-time atom-based density expansion coefficient p and data indicative of an atom attribute of each atom in the inference-time molecular structure, molecular structure data during reasoning is input into a trained kinetic energy density functional machine learning model TS, theta (p,), thereby generating a predicted value of the electron non-interaction kinetic energy TS (p,) for the density function rho (r), where the trained machine learning model has been trained on a training data set, and where the trained machine learning model has been trained on the training data set. The training data set includes an atom attribute of each atom in the training-time molecular structure as an input, and an atom-based density expansion coefficient p for the training-time molecular structure, and includes a corresponding truth value for the non-interaction kinetic energy TS (p,), and outputting a predicted value for the non-interaction kinetic energy TS (p,) for the inference-time molecular structure.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] 1 Introduction Density functional theory (DFT) is a quantum chemical method for solving molecular energies and properties. It is widely adopted due to its popular accuracy-efficiency tradeoff and has facilitated numerous scientific discoveries. In solving molecular energies and properties… In the common Kohn-Sham formula for electron-bearing molecular systems, by maintaining A single-electron wavefunction called an orbit Most of the kinetic energy can be explicitly covered, and after further explicitly covering most of the internal potential energy with its classical counterpart (i.e., the Hartley energy), the remaining unknown portion, known as the exchange correlation (XC) energy, constitutes only a small fraction and is therefore easier to approximate. However, optimizing the N-orbital deviates from the original idea of ​​DFT-optimized single-electron density, and this immediately increases the computational complexity by an order of N ( Figure 1A Simulating large-scale systems for real-world applications is increasingly undesirable in the current era. For this reason, there is growing interest in methods that follow the original spirit of DFT, now known as Orbit-Free DFT (OFDFT).

[0002] Given a well-developed XC functional, the core task in OFDFT is to approximate the non-interacting kinetic energy (the kinetic energy calculated from the orbit) as a density functional (KEDF). Classical approximations employ local density values ​​or semi-local density features (derivatives of density) to calculate energy. Given the limited accuracy of OFDFT and the fewer semi-local features (e.g., kinetic and exchange energy densities), explicit nonlocal calculations (convolution / double integration) have become indispensable, and modern developments primarily follow this form. The parameters in these constructs are determined by the theory of uniform electron gas (UEG). Many successes have been achieved for material systems, but accuracy remains limited for molecules, mainly due to the highly concentrated electron density within molecules and its distance from the UEG.

[0003] Attempts have been made to approximate KEDF using machine learning by leveraging data to compensate for theoretical mismatches. Ridge regression is a known method, and deep neural networks have also been explored recently. However, these methods use regular grids (or plane wave bases) to represent density, which remains unsuitable for molecular systems because steep density changes close to the atom require high resolution, but at the cost of many unnecessary grid points in low-density regions. Even with irregular grids, grids typically require tens of thousands of times more electrons. This is particularly unacceptable in nonlocal computations. Furthermore, few of these machine learning methods demonstrate their performance in extrapolating molecules much larger than the training data (distribution out-generalization). Extrapolating larger molecules for such machine learning methods is a significant technical challenge for OFDFT. Summary of the Invention

[0004] To address the problems discussed herein and others, a computational system is provided, comprising processing circuitry configured to receive inference-time molecular structure data for an inference-time molecular structure, the inference-time molecular structure data including inference-time atomic-based density expansion coefficients and data indicating the atomic properties of each atom in the inference-time molecular structure; inputting the inference-time molecular structure data into a trained kinetic energy density functional machine learning model to generate predicted values ​​of the electron non-interacting kinetic energy for the density function, wherein the trained machine learning model has been trained on a training dataset including, as input, the atomic properties of each atom in the training-time molecular structure, and the atomic-based density expansion coefficients for the training-time molecular structure, the atomic-based density expansion coefficients indicating the values ​​of the electron charge density characterized at the atomic basis, and including the corresponding true values ​​for the non-interacting kinetic energy; and outputting the predicted values ​​of the non-interacting kinetic energy for the inference-time molecular structure.

[0005] This summary is provided to present, in a simplified form, the selection of concepts further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to embodiments that address any or all the shortcomings pointed out in any part of this disclosure. Attached Figure Description

[0006] Figure 1A A schematic overview of two conventional DFT methods and their order of complexity is shown.

[0007] Figure 1B A schematic diagram of a computational system configured with a trained kinetic energy density functional (KEDF) machine learning model, according to an example of this disclosure, is shown. The model receives inference-time molecular structure data and outputs a prediction of non-interacting kinetic energies at inference time.

[0008] Figure 1C This diagram illustrates how a KEDF machine learning model can be trained to predict non-interacting kinetic energies that can be used to calculate the ground-state energy of a molecular system, using density coefficients and kinetic energies and gradients calculated from each of multiple self-consistent field (SCF) approximation loops.

[0009] Figure 2 Three graphs are shown in (a)-(c), which show a comparison between the predictions of the KEDF machine learning model and the KSDFT calculations on several test datasets.

[0010] Figure 3 Three graphs are shown in (a)-(c) that illustrate the computational efficiency and extrapolation capability of a trained KEDF machine learning model used for OFFDFT prediction in a system known as M-OFDFT, compared to regular KSDFT.

[0011] Figure 4 Four graphs are shown in (a)-(d), illustrating the performance of the trained KEDF model in M-ODFT compared to other conventional methods on different prediction tasks for the 10-residue Chignolin protein.

[0012] Figure 5 Density optimization curves of trained KEDF models applied to M-ODFT of the QM9 molecule are shown, with different initialization functions for comparison.

[0013] Figure 6 A schematic diagram of the transformer neural network architecture of the KEDF machine learning model (called the Graphormer architecture) is shown.

[0014] Figure 7 The diagram illustrates the splicing of auxiliary bases of different atom types in an example KEDF machine learning model.

[0015] Figure 8 It shows Figure 1B The diagram illustrates the computational system that generates the training dataset and trains the KEDF machine learning model during training, and outputs the trained KEDF machine learning model for use during inference.

[0016] Figures 9A to 9C An example of a method for using a kinetic energy density functional (KEDF) machine learning model during inference and training, according to this disclosure, is shown.

[0017] Figure 10This includes plots showing the changes in coefficient scale after local system transformation. For each coefficient dimension μ, the height of the depicted bar represents the relative change in scale, with negative values ​​indicating a reduction in coefficient scale due to local system transformation. These plots demonstrate that local system transformation significantly reduces the coefficient scale across multiple basis dimensions and atom types.

[0018] Figure 11 This includes a graph showing the gradient scale changes after a local frame transform. For each coefficient dimension μ, the height of the bar represents the relative change in gradient scale. In the graph, the local frame transform significantly reduces the gradient scale across a wide range of base dimensions and atom types.

[0019] Figure 12 Includes graphs showing the gradient and density coefficient scales of the QM9 dataset. The L∞ metric is used to measure the maximum scale of the data. (a)-(b) present the gradient scales before and after the dimensionality rescaling transformation. This technique can significantly reduce gradient scaling across various coefficient dimensions. (c) shows the coefficient scales after the dimensionality rescaling transformation.

[0020] Figure 13 The graph shows the gradient scale variation without being centralized by the atomic reference model, demonstrating that the atomic reference model can significantly reduce the gradient scale.

[0021] Figure 14 Includes a pair of plots showing the curves of total energy and gradient norm during density optimization of molecules randomly selected from the QM9 dataset. Figure 14 In the table, 1stLocGradNorm represents the step size of the first local minimum of the gradient norm, 1stLocEngUpd represents the step size of the first local minimum of a single-step energy update, and GlobEngUpd represents the step size of the global minimum of a single-step energy update. E, E and ∥ pE∥ represents the optimized total energy, the true total energy, and the norm of the energy gradient, respectively. 1stLoGradNorm represents the step size for the first local minimum of the gradient norm, 1stLocEngUpd represents the step size for the first local minimum of a single-step energy update, and GlobEngUpd represents the step size for the global minimum of a single-step energy update. E, E and ∥ pE∥ represents the norm of the optimized total energy, the true total energy, and the energy gradient, respectively.

[0022] Figure 15 This is a graph showing the composition ratios of different QMug containers.

[0023] Figure 16 Includes two figures, showing additional visualizations of the optimized density around different atoms in the ethanol molecule.

[0024] Figure 17 Includes a pair of graphs showing the accuracy of HF force prediction with respect to the two extrapolation settings.

[0025] Figure 18 The results of ablation studies, including a pair of figures, illustrate the nonlocality of the model architecture. Figure (a) shows the accuracy of energy prediction on the tested ethanol molecule, and Figure (b) shows the accuracy of HF force prediction on the tested ethanol molecule.

[0026] Figure 19 It shows that it can be achieved Figures 1B-18 A schematic diagram of an example computing environment for the computing system and method shown. Detailed Implementation

[0027] This paper describes a computational system and method for OFDFT, referred to as M-OFDFT (M stands for molecule), which is particularly suitable for such molecules using a KEDF machine learning model. For an effective density representation of the molecule, expansion coefficients on the atomic basis are used in the KEDF machine learning model. The atomic basis naturally adapts to the density pattern in the molecule because both are concentrated around the atoms, as... Figure 1B As shown. This consistency allows for a much smaller representation dimension, where typically only a few dozen electrons are sufficient for the number of bases. The KEDF model receives inputs of class (e.g., data indicating atom numbers and electron counts) and atom positions, i.e., the structure of the molecule, as well as density coefficients attributed to each atom. As... Figure 1B As shown, the KEDF model is implemented using a transformer-based neural network architecture (referred to as Graphormer in this paper), which performs explicit nonlocal computation and has better generalization and extrapolation potential than conventional models, utilizing more data and parameters. (See reference...) Figure 6 Graphormer is described in Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng, Guolin Ke, Di He, Yanming Shen, and Tie-Yan Liu, “Do transformers really perform badly for graph representation?”, Advances in Neural Information Processing Systems 34:28877-28888, 2021, and its publications are incorporated herein by reference.

[0028] Learning density functionals is a more challenging task than regular machine learning. The model is used to construct an objective function to optimize the density coefficients, therefore it needs to know the fine-grained energy landscape and maintain optimization of the orbits, as simply fitting point-to-point data will not achieve sufficiently accurate results. Figure 1C As shown, to address this challenge, the inventors of this application reviewed the data generation process and managed to generate multiple coefficient data points for each molecular structure, and also provided gradient labels for each data point. As explained in further detail below, they also selected appropriate density initialization to ensure that the optimization remained in a region where the model was reliable. Optimizing the training objective is also challenging due to the geometric variability of the input data and the large range of gradient data. The inventors accordingly used local systems to orient the atomic bases to reduce geometric variability while simultaneously enhancing chemical environment invariance, and introduced a series of techniques to achieve efficient optimization.

[0029] The practical utility and appeal of the computational systems and methods described in this paper (including the KEDF machine learning model and M-ODFT) are demonstrated in the following ways. First, M-OFDFT achieves accuracy on a molecular scale that classical KEDF-OFDFT has never reached, including those as large as proteins. Chemical precision is achieved on both molecular scales and larger-scale molecules, with errors still significantly smaller than those of classical KEDF. The optimized density also reveals a clear shell structure, which is considered difficult for OFDFT. This demonstrates that the functional is capable of achieving working OFDFT for molecular systems. Second, M-ODFT actually shows better performance than KEDF. Lower observation time scaling The observation time scaling is improved. Absolute time is always smaller, achieving a 5.7x speedup, particularly on Chignolin (684 electrons) and a 27.4x speedup on the Protein B system (2,750 electrons), where KSDFT is near the tolerance margin. Third, extrapolation studies rarely achieved by other machine learning functions have been shown to large-scale systems. Here, the functional achieves non-increasing per-atom energy error with increasing numbers of heavy atoms in the molecule (exponent -0.11), which is rare in direct energy prediction (exponent 0.235 using the same architecture). This demonstrates the advantage of using machine learning for molecular science with more fundamental concepts. The learning objective of better extrapolation than learning end-to-end mappings is also consistent with observations in machine learning. Given these benefits, the computational systems and methods described in this paper push the boundaries of the accuracy scalability tradeoff in quantum chemistry and can be beneficial for a wide range of molecular science and engineering problems.

[0030] Turn now Figure 1BThe detailed description illustrates a computing system 10 including a computing device 12. The computing device 12 includes processing circuitry 14 linked by a communication bus, associated volatile memory 16, and a non-volatile storage device 18. The non-volatile storage device stores instructions that, when executed by the processing circuitry 14 using portions of the memory 16, cause the processing circuitry 14 to perform the functions described herein. In this way, the processing circuitry 14 is configured to perform these functions.

[0031] As shown in the figure, during inference, the processing circuit system 14 is configured to receive data including the molecular structure during inference. Inference-time molecular data 20, including the atomic-based density expansion coefficient p. Inference-time molecular structure. This is a data structure that includes data describing the physical molecular system 22 in terms of its structure. More specifically, the inference-time molecular data 20 includes indicators of the molecular structure at inference time. Data on the atomic properties of each atom Z. For example, the atomic property can be one or more of atomic position X and atomic number Z. Any suitable encoding for the data can be used to indicate the atomic position X and atomic number Z. The atomic-based density expansion coefficient p represents the electron density functional as the expansion coefficient on the atomic basis. ρ ®, as shown in the figure. These density expansion coefficients p can be simply referred to as density coefficients, and can be provided by the density feature preprocessing module described below after preprocessing.

[0032] Processing circuit 14 is also configured to input inference-time molecular structure data 20 into a trained kinetic energy density functional (KEDF) machine learning (ML) model T. S、θ (p, )24, thereby generating the electron non-interacting kinetic energy T of density function ρ(r). S、θ The predicted value of 26 can be determined based on or by the atomic-based density expansion factor p. See below for reference. Figure 8 To describe in more detail, for example, a trained KEDF ML model (or simply a KEDF model) has been trained on a training dataset that includes atomic properties (such as molecular structure at training time). (One or more of the atomic positions X and / or atomic numbers Z of each atom in the matrix) and the molecular structure at training time The atomic-based density expansion factor p indicates the value of the electron charge density expressed on the atomic basis and includes the non-interacting kinetic energy. T S (p, The corresponding ground truth value. The training dataset may also include the kinetic energy density unfolded gradient. p TS (p, The truth value of ). In an example configuration, the trained KEDF ML model 20 includes neural networks 28, such as Figure 8 The diagram shows a transformer-based neural network. An output layer 30 is provided to handle molecular structures. Sum the energy contributions of each atom.

[0033] The processing circuitry is also configured to output the molecular structure during inference. 20 non-interacting kinetic energy T S,θ The predicted value is 26. The processing circuitry can also be configured to target the molecular structure during inference. Backpropagation to output kinetic energy density unfolded gradient p T S (p, The processing circuit can also be configured to perform orbitless density functional theory (DFT) calculations using the predicted values ​​of the kinetic energy density expansion gradient.32 An example form of OFDFT calculation for determining the ground state energy is shown in [reference to a document]. Figure 2 As shown in the figure, and further described below.

[0034] 2 Results 2.1 Orbital-free density functional theory of molecules In OFDFT, the energy functional is compared with KSDFT: The same method is used to decompose, where It is the external potential energy originating from the nucleus within the molecule. It is the classical internal potential energy (Hartley energy) that covers the main part of the energy of the electrical interaction between electrons. It is the non-interacting part of the kinetic energy, and it exchanges correlated (XC) energy. The remaining kinetic and internal potential energies are considered, and an accurate explicit approximation is available. In KSDFT, the non-interacting kinetic energy can be calculated directly from the orbit. However, as a density functional, its expression is not available. Therefore, the focus in OFDFT is to develop an expression for KEDF. An accurate approximation.

[0035] As mentioned above, M-ODFT uses atomic-based (atom-centric contracted Gaussian with spherical harmonics) representations of effective density in molecules. The effective density representation is concentrated at the atoms, with the density distribution surrounding the atoms, such as... Figure 1B As shown. They are even designed to simulate nuclear cusp conditions to accurately describe abrupt changes near the nucleus, which is difficult for grid or plane-wave bases. Therefore, the number of atomic bases required is usually much smaller than that of grid bases. 103), which is particularly advantageous for the nonlocal calculations indispensable to KEDF. Furthermore, atomic basis functions are typically located in shell structures, i.e., the basis functions are concentrated at radii different from the nucleus. This helps OFDFT methods more easily discover dense shell structures in molecules, which are considered difficult without orbitals. Below, density can be expanded as: Since atomic bases depend on the structure of the molecule. That is, the positions of all atoms (composition) in a molecule. (Conformation) and atomic number The density is entirely determined by the tuples Description. Therefore, the KEDF model follows the form of directly outputting the integral value rather than the kinetic energy density on the grid. .

[0036] Figure 1B and Figure 1C An overview of the computational system (also known as M-OFDFT) of this disclosure is shown, while Figure 1A The conventional method is shown. For example... Figure 1A As shown, conventional KSDFT solves the Schrödinger equation by maintaining N single-electron wavefunctions, while OFDFT relies on only a single single-electron density function, thus reducing the computational complexity by an order of N. Figure 1B The diagram illustrates a computational system 10 that uses the expansion coefficient p on the atomic basis to represent electron density. The KEDF model, based on a nonlocal Graphormer architecture, takes the density coefficient feature (or simply density coefficient) p and the molecular structure M as inputs to predict kinetic energy. T S,θ .exist Figure 1C In this study, the trained KEDF model is shown as the objective function used to optimize the density coefficient p.

[0037] Due to the lack of guiding theory far removed from a unified KEDF, the adoption and training of data for... The machine learning model generates data on a specific set of molecular structures to assess the relevance in deployment. However, learning the density function is more challenging than regular machine training because the trained KEDF model is used as the objective function to optimize the density coefficient p for a given molecular structure, such as... Figure 1C As shown. To capture the structure of each molecule. The functional landscape of p in coefficient space, the inventors manage the KSDFT computation process to generate each molecular structure Multiple p samples, and also calculate gradient labels This leads to the pattern The following data. This article also introduces some techniques for handling geometrically variable data and for large-scale gradient data for effective training.

[0038] for The model architecture, note that the coefficients in p can be attributed to each atom based on the center of the corresponding basis function, and the input can be seen in the graph structure. Where the atom type and coefficients define node features, and the relative positions of atom pairs define edge features, such as... Figure 1B As shown. Therefore, the inventors used a Graphormer-based model as... Graph neural networks, such as Figure 1B , Figure 6 and Figure 8 As shown, nonlocality is overridden by a self-attention layer in the transformer architecture from the KEDF model, which explicitly computes the interaction between each pair of node features. To better accommodate the graph structure, edge features are also involved in the corresponding node pair interactions, which is shown to improve graph theory performance. The interaction results are used to update node features, and after several layers of such updates, the final node features are aggregated to give... Output. Based on the ablation studies conducted, the nonlocal conception is crucial for the desired extrapolation performance of reducing per-atom error as the system size increases.

[0039] In training After modeling, a given molecular structure can be solved from the density optimization process. The ground state energy and density (see) Figure 1C ): .

[0040] Here, from Explicit calculation and And it can be computed on the grid in a conventional manner. The optimization is solved using gradient descent because self-consistent iteration is unnatural for density. Normalization constraints are used in the optimization. ,in Since the gradient is linear, it is projected onto the admissible plane before updating the coefficients. (1) Where ε is the step size. It is worth noting that, since the operation is performed directly on the density, the calculation... , and The complexity of the terms is square root compared to those in KSDFT. , (for semilocal pure XC functionals) and Complexity. The model has The complexity is therefore, the complexity of the OFDFT in each optimization iteration of this disclosure is O(n). This is more than usual. KSDFT achieves cost reduction for orders of N.

[0041] 2.2 M-OFDFT Achieves Chemical Accuracy in Molecular Systems The inventors investigated the performance of M-OFDFT on molecules with sizes similar to those in the training data. This study considered two scenarios. First, data was generated on a collection of ethanol conformations from the MD17 dataset, and second, data was generated on molecules with equilibrium conformations from the QM9 dataset. The former investigates generalization in conformational space, while the latter investigates it in chemical space. Each dataset was divided into three parts for training and validation of the KEDF model and for testing M-OFDFT using the trained model. For ease of training, APBE KEDF was used as a reference, and the Graphormer model (see details) was also allowed. Figure 6 Learning about residues.

[0042] M-OFDFT was evaluated based on the mean absolute error (MAE) in energy and the Feynman-Hellman (HF) force. Results were 0.18 kcal / mol and 1.18 kcal / mol / Å for ethanol, and 0.93 kcal / mol and 2.91 kcal / mol / Å for QM9. M-OFDFT achieved chemical accuracy (MAE of 1 kcal / mol energy) in both cases.

[0043] To demonstrate the importance of OFDFT for these molecular results, the performance was compared with that of classical KEDF, which includes a precisely defined Thomas-Fermi (TF) KEDF with a density gradient within a uniform electron gas limit. The corrections are given, where vW represents von Weizsäcker KEDF, experimental variant TF+vW, and reference method APBE. Note that different KEDFs may have different absolute energy deviations, requiring comparison of MAE in relative energies. On the ethanol dataset, relative energies are obtained between the conformation and equilibrium conformation. On QM9, since each molecule has only one conformation, the C7H of the 6,095 isomers from the QM9 dataset is used. 10 The relative energies acquired between pairs of O2 molecules can be viewed as different conformations of the same "molecule". Figure 2 The model predictions for the test dataset are shown compared to KSDFT computation. Figure 2 The comparison of energy and HF force predictions between M-OFDFT and classical KEDF (with 95% confidence interval bars) is shown in (a)-(b). Figure 2A visualization of the density optimized by various methods is shown at (c). The curve plots the integrated density on spheres with different radii, centered on the oxygen atom in the ethanol molecule. Figure 2 As shown in (a) and (b), M-OFDFT achieves chemical accuracy with respect to relative energy and is two orders of magnitude more accurate than OFDFT using classical KEDF.

[0044] As a qualitative study of M-OFDFT, the density optimized on ethanol using these methods is... Figure 2 As seen in (c), the main peaks on the left and right sides show the density of nuclear electrons in oxygen and bonded carbon, while the secondary peak in the middle reflects the density of electrons covalently bonded to hydrogen and carbon. M-OFDFT accurately recovers this shell structure, which is considered difficult for OFDFT, and the optimized density is almost identical to that of KSDFT. In contrast, the classical KEDF of APBE does not align well with the true density around the covalent bonds. These results demonstrate that the inventors have developed a working OFDFT for molecular systems.

[0045] 2.3 M-ODFT has lower experimental time complexity than KSDFT. As mentioned, OFDFT achieves this when using atomic basis numbers to represent density. Its time complexity is lower than that of KSDFT. This advantage of M-OFDFT lies in Figure 3 The results were experimentally validated in (a). Both M-OFDFT and KSDFT were tested on a subset of molecular structures from the QMugs dataset with 106 to 696 electrons. Runtime values ​​were grouped and averaged within each bin of multiple electrons with a width of 20. The absolute runtime of M-OFDFT was consistently shorter than that of KSDFT, achieving up to 6.7x speedup relative to KSDFT. The complexity of M-OFDFT is lower than... And it is indeed N orders smaller in complexity than KSDFT.

[0046] Utilizing sophisticated atomic orbitals and numerical optimizations, KSDFT can typically process molecular systems of up to approximately 500 atoms using modern PC clusters. To demonstrate the scaling advantages of M-OFDFT, models were tested on two large-scale systems rarely encountered in the context of KSDFT computation: (1) a 47-residue (709-atom) peripheral subunit binding domain BBL-H142W (PDB ID: 2WXC), and (2) a 47-residue (738-atom) K5I / K39V double mutant of the albumin binding domain of protein B (PDB ID: 1PRB). All atomic geometries for each system were extracted from Krespen Lindorff-Larsen, Stefano Piana, Ron O Dror, and David E Shaw, “How fast-folding proteins fold,” Science, 334(6055): 517-520, 2011. The results show that M-OFDFT takes 0.41 hours and 0.45 hours to compute the two test systems, respectively, achieving speedups of 25.6x and 27.4x compared to KSDFT. PySCF software was used for KSDFT computation with default settings.

[0047] 2.4 M-OFDFT extrapolates well to larger molecular scales OFDFT has the advantage of handling large systems at affordable costs, for which other quantum chemical methods, such as KSDFT, become prohibitively expensive. However, data for such systems is scarce for the same reason. Therefore, despite the technical challenges, machine learning-based OFDFT methods are generally desired, which can extrapolate to molecules at a scale much larger than what is seen in training (out-of-distribution generalization). In studies of M-OFDFT, this allows for the use of KEDF ML models (Graphormer models – see...) Figure 8 (Details) The KEDF and XC energies are learned completely together. This is to completely eliminate the extra computation on the grid (for APBE KEDF references and PBE XC functionals) in their gradient descent. Memory costs become prohibitive for macromolecules. Note that this modification does not result in a significant loss of accuracy in studies at a scale.

[0048] To assess the importance of the extrapolation behavior, the error of M-OFDFT was compared with a natural variant of a machine learning-based quantum chemistry method that directly predicts molecular structures in an end-to-end manner. The ground-state energy. This is referred to in this paper as M-PES, representing the potential energy surface of the molecule, and is fairly compared using the same Graphormer model architecture. To investigate the effect of density features on extrapolation, the inventors also considered, in addition to In addition, there is a variant with MINAO density characteristics in its input, referred to in this paper as M-PES-Den.

[0049] Figure 3 The computational efficiency and extrapolation capability of M-OFDFT are shown. Figure 3 At (a), the experimental time cost of M-OFDFT and KSDFT for molecules at various scales is shown. At (b) and (c), the extrapolation performance on QMugs is shown. Figure 3 The shaded areas represent 95% confidence intervals. At (b), all models have been trained on molecules with fewer than 15 heavy atoms and extrapolated to larger QMugs systems with increased molecular scales. At (c), all models have been trained on a series of QMugs datasets with increased molecular scales (up to 30 heavy atoms) and extrapolated to 50 QMugs conformations with 50–60 heavy atoms. The black dashed line represents the performance of M-OFDFT trained on the first QMugs dataset.

[0050] The inventors investigated extrapolation on the QMugs dataset, which contains molecules much larger than those in QM9, which have no more than nine heavy atoms. To construct the extrapolation task, a KEDF model was trained on QMugs molecules with no more than 15 heavy atoms. Since there are very few QMugs molecules with no more than nine heavy atoms, all QM9 molecules were also included in the training. To investigate extrapolation trends, the remaining QMugs molecules were grouped into bins of width 5 according to their number of heavy atoms, and the per-atom energy (MAE) of M-OFDFT was evaluated from 50 molecular structures in each bin as the scale increased.

[0051] The result is Figure 3As shown in (b), the per-atom MAE of M-OFDFT is not only orders of magnitude smaller in absolute value than its end-to-end counterparts M-PES and M-PES-Den, but also remains non-increasing or even decreases (negative exponentially) with increasing molecular scale. This is unexpected for energy prediction models based on common local structures, since errors in each local environment are not reduced by larger global environments, and almost always, predictions for each local environment cannot cover the effects from global interactions, which become stronger with increasing molecular size and thus increase per-atom errors on amortization. The non-increasing per-atom MAE is not achieved by directly using a nonlocal model in end-to-end energy prediction, as shown by the M-PES curve, and the similar increasing trend of the M-PES-Den curve indicates that even the use of density feature enhancements does not trigger a non-increasing trend.

[0052] Using this evidence, the inventors hypothesize that superior extrapolation performance stems from modeling the density functional as an objective function of the target density and energy, rather than directly modeling a mapping to the target. In this way, some of the complexity in the mapping is transferred to the optimization process, where the objective function can be simpler and easier to extrapolate. This is particularly true for the KEDF model, which is more sensitive to molecular structures. The dependence of electrons on the basis is only used to locate and characterize density, not as a result of the entire interaction process between electrons and atoms, which appears very complex and subtle if the electron descriptor is hidden. It is worth noting that such conclusions have also been observed in machine learning, indicating potential general conclusions for better extrapolation from the objective function of the target.

[0053] To further demonstrate the importance of the extrapolation capability of M-OFDFT, the amount of additional large-scale molecular data required for end-to-end counterparts to achieve extrapolation performance comparable to M-OFDFT was investigated. To this end, 50 QMugs molecules with 50–60 heavy atoms were selected as the extrapolation baseline, and the model was trained on a series of molecular datasets with increasing molecular sizes. Specifically, QMugs molecules with no more than 30 heavy atoms were grouped into bins of width 5, and larger-scale molecules were gradually incorporated into the training group. Similar molecule counts were maintained across all training datasets to avoid the influence of data size. Figure 3 As shown in (c), both the M-PES and M-PES-Den models require at least twice the number of additional molecules (30 vs. 15) at larger scales to achieve energy prediction accuracy comparable to M-OFDFT (0.068 kcal / mol). These results support the superior extrapolation performance of M-OFDFT. Furthermore, the high efficiency of M-OFDFT in learning density functionals from available molecules holds great promise for extending quantum chemical methods to unaffordable large-scale systems.

[0054] In recent years, accelerating first-principles calculations of large-scale biomolecular systems has become an attractive task in quantum chemistry, particularly in the context of protein structure. To analyze the generalization ability of M-OFDFT for protein systems, a 10-residue Chignolin protein (containing 168 atoms) was chosen as the model system. This protein has been extensively studied experimentally and has a rich conformational structure derived from molecular dynamics (MD) simulations. For the extrapolation task, a KEDF model was trained on short peptide fragments, each containing no more than 5 residues, obtained from approximately 1,000 Chignolin conformations. The model was then extrapolated to the same protein conformation. Fragments were constructed by “cutting out” subsequences of the Chignolin sequence, including discontinuous subsequences with residue gaps. These were chosen because they are long enough to sample important long-range effects, yet short enough to allow for easy reference to KSDFT calculations within a reasonable timeframe. Both the Chignolin conformation and the fragments were neutralized, as the focus was on a discrete molecular system where all atoms or functional groups remain neutral and all electrons are paired. In this task, external potential energy is also introduced into the learning objective because it can increase the training stability of the KEDF model.

[0055] The result is Figure 4 Presented in (a). Figure 4 The extrapolation of the performance on Chignolin is shown. Figure 4 The energy prediction accuracy of 1,000 Chignolin conformations using three methods with increased training scale in (a). Figure 4 In (b), the relative energy prediction accuracy of the three methods for the 50 Chignolin conformation is compared with APBE. All deep learning models were trained using fragments here. Figure 4 Overall, fine-tuning performance on larger-scale molecules is shown at (c) and (d). Figure 4 The energy is shown at (c), and Figure 4 The Hellman-Feynman (HF) force prediction accuracy for three methods with / without a fine-tuned model on the 500 Chignolin conformation is shown at (d). The error reduction rate of the fine-tuned model on the “start-from-nothing” baseline is listed.

[0056] Significantly, as the scale of the training segments increases, M-OFDFT consistently outperforms both M-PES and M-PES-Den in terms of energy prediction accuracy. Furthermore, the performance of M-OFDFT trained using all segments is compared to that of the classic KEDF APBE on 50 random Chignolin conformations. Figure 4As shown in (b), the method disclosed in this paper achieves a significantly lower relative energy per atom (MAE) compared to APBE (0.098 kcal / mol vs. 0.684 kcal / mol), highlighting the method's strong extrapolation capability for biomolecular systems.

[0057] Extrapolating to molecules outside the scale remains a significant challenge for machine learning-driven quantum chemistry methods, particularly for large-scale systems with extremely scarce training labels. One potential strategy to improve extrapolation performance is to pre-train models on a large number of computationally readily available small molecules and then fine-tune them on a smaller sample set of the system of interest. There are technical challenges in designing efficient pre-trained frameworks that can effectively utilize small molecules to extract transferable knowledge. To evaluate this transfer learning strategy, the extrapolation performance of M-OFDFT and end-to-end correspondents on 50 random Chignolin conformations was compared after fine-tuning all models on 500 Chignolin conformations. Models trained de novo on the same 500 conformations were used as a “de novo” baseline. Figure 4 As depicted in sections (c) and (d), the disclosed method achieves greater performance gains than the "blind-to-the-nose" model in both energy and force accuracy. Notably, M-OFDFT achieves significant error reductions of 35.4% and 40.8% in force and energy predictions, respectively. This result implies that the disclosed concept can utilize transferable and physical knowledge from readily available protein fragments more effectively than end-to-end learning paradigms.

[0058] 3. Conclusions and Discussion This paper describes a computational system and method, called M-OFDFT, a trackless density functional theory that successfully acts on molecules. Kinetic energy density functionals (KEDFs) approximating OFFDFT are considered challenging (more difficult than exchange-correlation functionals), especially for the electron density space of molecular systems. By utilizing methods such as... Figure 6 The machine learning models based on the Graphormer architecture shown can achieve this approximation more accurately, as demonstrated on a wide range of molecular systems. With the achievement of accuracy on molecular systems, M-OFDFT provides the scalability of OFFDFT for solving molecular systems. The experimental cost of M-OFDFT is orders of magnitude lower than that of the popular KSDFT, and it is also significantly more cost-effective on experimental scales. The advantages become more pronounced when the order is reduced by more than N compared to KSDFT. Therefore, M-OFDFT is considered to have pushed the precision-cost trade-off boundary in quantum chemistry.

[0059] This disclosure also demonstrates the benefits of using appropriate quantum chemical conceptions when leveraging machine learning. M-OFDFT has qualitatively shown better extrapolation performance on a larger scale compared to the direct ground-state energy prediction formulas M-PES and M-PES-Den, even though they also consume [resources / capacities]. This reduces complexity and achieves lower in-scale verification errors. M-OFDFT can also provide ground state densities capable of computing various properties. KEDF models can also be used to construct quantum embeddings for more accurate multi-scale computations.

[0060] Compared to existing work on learning KEDF, this approach represents density on an atomic basis, which significantly reduces the feature dimensionality, thus enabling an efficient nonlocal KEDF model that captures long-range effects to enhance accuracy. Some other methods for learning XC functionals also employ atomic basis, but the lack of molecular structure inputs does not allow for the interaction of density components hosted on different atoms. To inform the model of the optimization landscape, gradient labels and labels for multiple densities on the same molecular structure are used to enrich the training data. Comparisons with other methods for supervising the optimization behavior of functional models show that these other methods are less efficient. Furthermore, some other methods require applying a projection to the manifold within the distribution at each step of density optimization to achieve stability, while M-OFDFT only requires initializing the distribution (see [link to M-OFDFT]). Figure 5 This indicates better optimization behavior.

[0061] While statistical guarantees have been established for intradistributed generalization of machine learning models, reliable extrapolation remains an open challenge, which has long hindered the application of machine learning in the scientific field. This disclosure achieves superior extrapolation performance to ground-state energy prediction by learning a density functional with simpler and more general underlying rules, making it easier to extrapolate. The density optimization process also takes over a portion of the complexity in the dependence of ground-state properties on molecular structures.

[0062] Incorporating the exact properties of KEDFs into machine learning models can further improve extrapolation. Geometric invariance is built into the model, which significantly improves accuracy. However, the benefits from these properties are not always direct, as some may introduce more training challenges or require specific architectures with unexpected capacity constraints. For example, one could try using the von Weizsäcker KEDF as a reference KEDF, which incorporates positive properties into residue KEDF models. However, for The resulting gradient labels are too large to effectively train ML models. Attempts were also made to learn KEDF directly (without reference) and a universal functional with convex surrogates, but the results again fell short of expectations due to the more challenging training problem. KEDF also has scaling properties, but it cannot be transformed into an exact equation on an atomic basis.

[0063] From a machine learning perspective, extrapolation can be further improved with more data and a larger model. It's worth noting that the transformer architecture already enables the underlying model of the language, and the Graphormer model itself (see...) Figure 6 It has been shown to have significant extrapolation performance in predicting molecular properties. Therefore, better extrapolation can be expected by utilizing larger datasets and model sizes, and it has been used as the base model for KEDF.

[0064] 4 methods The following outlines the methods for training and using KEDF models. As mentioned above, learning density functionals presents more challenges than conventional machine learning. This comes from two sides: (1) The model is not used as an end-to-end predictor, but rather as an objective to optimize density. This requires the model to know the energy landscape in the density coefficient space, as well as appropriate density optimization methods. This paper addresses these challenges. (2) The properties of the objective functional itself create difficulties in model training (e.g., geometric variability and huge gradient range). Such concerns were considered when designing the model and developing techniques to address the training difficulties.

[0065] 4.1 Training the KEDF model As a density functional, the KEDF model is used for a given molecular structure ( Figure 1C and Figure 8 ) Optimize the density coefficient p, and therefore need to know how to train for a fixed density coefficient p. It varies with p. To achieve this, the training data first needs to be given for each... Multiple p-values ​​are provided Tags, i.e., following the form This can be achieved through each molecular structure The KSDFT calculation utilizes solutions from intermediate steps of the self-consistent field (SCF) iteration. The basic principle is that the task in each SCF step k is to solve for the non-interacting fermionic system with effective individual potential energy, whose ground state can be solved exactly using orbitals. Therefore, the non-interacting kinetic energy can be accurately calculated from these orbitals. Furthermore, the corresponding density coefficients can be calculated from these orbitals through density fitting. Specifically, this paper employs well-homogeneous auxiliary bases whose β parameter is considered as 2.5 as a density basis. Recognized pure XC functionals (i.e., those dependent only on density, not orbitals), such as PBE, are used to ensure that the data pairs represent density functionals. Spin-restricted calculations are performed on the 6–31g(2df,p) basis set, which is sufficient to render the molecules under consideration uncharged in near-stable conformations and involves only light atoms (at most fluorine). The training loss function of the model is: (2) However, although the energy predicted by the trained KEDF model is accurate at the ground-state density coefficients, subsequent density optimization using the model still resulted in a decrease in energy. This indicates that only from multiple SCF steps... The label does not yet adequately represent the energy landscape in the coefficient space used for density optimization. The inventors recognized that the optimization curve should be static at the ground-state density, and thus considered providing a gradient... Further supervision provides more information about the energy landscape and directly impacts the optimization process. Fortunately, gradient labels can also be obtained from each SCF step, also due to the molecular structure. The solution for each SCF step k is the ground state of the non-interacting fermions in the effective monomer potential. Then, the fitted density coefficients from the solution... This represents the ground state density of the system, which allows it to undergo normalization constraints. Total energy Minimize, where It is the effective monomer potential energy in vector form on the atomic basis of density.

[0066] This means It provides gradient supervision over the subspace of the normalized density coefficients. An efficient implementation is needed to compute the unprojected gradient labels from the solution of each SCF step. The corresponding gradient loss function is: (3) An ablation study was conducted on the impact of multi-step data and gradient supervision.

[0067] 4.2 Geometric Invariance The second major challenge beyond conventional machine learning tasks is the large variability in data displayed. In the model... In the input features, the atomic coordinates X require a specific coordinate system, the choice of which introduces arbitrary variability in the input features, also known as SE(3) isovariance. Furthermore, in addition to the radial component (contracted Gaussian), the atomic basis also has an angular component that depends on the coordinate system. The coefficient feature p then also has this variability. It is then expected that the features or model change with respect to the coordinate system in order to learn the fundamental dependence on density, regardless of the geometric representation.

[0068] For atomic coordinates X, the Graphormer model is naturally SE(3)-invariant with respect to its shifts and rotations because the model only utilizes the relative distances of the atomic pairs for later processing, which is naturally invariant. For density coefficients p, possible approaches include choosing a standard global coordinate system (e.g., by performing PCA on the atomic positions) or constructing a model that is naturally invariant with respect to shifts and rotations of the global coordinate system. However, these approaches do not produce invariant features for the same density distribution pattern. For example, consider the coefficients corresponding to the atomic bases on the three hydrogen atoms in the methyl group. Since the three HC bonds are symmetrical, the density on the bonds is the same (ignoring the influence from nearby unbonded atoms), and therefore they have the same contribution to the kinetic energy from themselves. However, since the three bonds have different orientations relative to the common coordinate system, the coefficients are different. To further eliminate this unnecessary variability, this paper adopts a local system to give each atom its own coordinate system for its atomic bases and coefficients, which depends on the atom's local chemical environment, essentially the bonding pattern. A similar idea has been considered in many previous attempts, where the x-axis is a unit vector. From the central atom to the nearest neighbor atom, in It is the second closest atom to that atom The absence The coordinates on, and := × Following this, hydrogen atoms near the central atom are excluded to obtain a more stable local structure. Since the local system is chosen based on the relative positions of the atoms, it is equivalent to global shifts and rotations, making the density coefficient SE(3) invariant. Furthermore, The directional coefficients reflect the density of the most effective (shortest length) bonds on the atom. This results in identical density coefficients on the three hydrogen atoms in the (fully symmetrical) methyl group, even if the HC bonds have different orientations, and variations in the coefficients reflect variations in the HC bond density. The scale (standard deviation) of the atomic density coefficients and gradient labels is significantly reduced with the use of local systems, particularly for H atoms, where most coefficient dimensions exhibit a scale-down ratio >60%. Furthermore, this approach immediately reduces training loss and stabilizes the optimization process, leading to significant improvements in the density optimization phase.

[0069] 4.3 Learning with a large gradient range Even after removing geometric variability, the data remains enormous. Worse still, conventional data normalization techniques are unsuitable for this task because gradient labels also exist, and minimizing their scale conflicts with minimizing the scale of the coefficients (dividing the scale into coefficients means multiplying the scale by the gradient). Having both large coefficients and a large gradient implies a function with large Lipshitz coefficients. To address this challenge, a series of techniques have been introduced to achieve efficient training, including density coefficients, reparameterization of the atomic reference model, and dimensionality rescaling.

[0070] Rescaling by dimension First, consider dimensional rescaling to allow for greater flexibility in data normalization. Specifically, for each coefficient dimension, attempt to normalize its gradient data points to the target scale. If the resulting coefficient standard deviation is greater than the chosen maximum scale, rescale that dimension to normalize the coefficients to the maximum scale. Obtain the rescaling factor for dimension μ. The process can be summarized as follows: (4) in, It is the target scale of the gradient. It is the largest scale of the coefficients. It is the infinite norm of the gradient (i.e., (measurement), and It is the standard deviation of the coefficients. Each coefficient dimension will be independently determined by... Then, the corresponding gradient scale that the KEDF model needs to fit will be reduced. (In most cases, ≥1).

[0071] Natural reparameterization However, even using this compromise approach of coefficient scaling and gradient scaling, simultaneously reducing both scales remains challenging. Two techniques are introduced before applying dimensionality rescaling. First, because different bases have varying importance in representing density, different coefficient dimensions have different sensitivities to the density function, resulting in some dimensions having particularly large Lipshitz coefficients with respect to energy. Therefore, a natural reparameterization of the density coefficients is performed. The principle is to make the Euclidean metric in the reparameterized coefficient space reflect the L2 metric in the density function space, thus balancing the sensitivity across coefficient dimensions and allowing for better dimensionality rescaling results to facilitate learning. On atomic bases, this principle induces density coefficients... The metric is given by , where W is the overlap matrix of the density basis. Therefore, the density can be expressed by . Reparameterization, where M satisfies This reparameterization also leads to natural gradient descent in density optimization, which has been shown to converge faster than standard gradient descent.

[0072] Atomic reference model To further improve the coefficient gradient scaling tradeoff, the absolute scale of the coefficients can be reduced by subtracting the statistical average (data centered), but this does not reduce the absolute scale of the gradient. Therefore, an atomic reference model is introduced for the linear kinetic energy of the density coefficients. For molecular structures... The simple dependency corresponds to the density coefficients of the same atom type / species / chemical element sharing the same linear weights. The linear model takes the linear weights for each type... It is constructed as a statistical average of the gradient data for each atomic type, and fitted to each class of deviation on the training dataset. As shown below: .

[0073] This effectively centers the gradient labels for the main Graphormer model to learn.

[0074] Density optimization During the deployment phase, common density initialization is out-of-distribution in terms of training density, since the latter is a solution to an effective non-interacting system, while the former is a superposition of isolated atom densities. Two approaches circumvent this problem: using initialization as a solution to a non-interacting system, or training a model that projects the initialized density distribution in-distribution (specifically towards the ground-state density). Appropriate stopping criteria have been designed for both intra- and large-scale systems.

[0075] After training, the KEDF model Used to solve the structure by minimizing the total electron energy relative to the density coefficient p. Given the ground-state density and energy of the molecule, gradient descent is used to optimize p because it is unnatural to characterize the process as an SCF iteration, and gradient descent also works more stably in KSDFT than SCF iteration. Note that the density normalization constraint is handled by projecting the standard gradient g onto the admissible plane before updating the coefficients.

[0076] Figure 5Typical density optimization curves for QM9 molecules with different initializations are shown. Initialization with the ground-state density (“Ground-state density” in the figure) results in a fixed line. All other curves excluding MINAO consistently remain above the KSDFT energy. Notably, in the final step, projecting MINAO yields an energy only slightly higher than (0.60 kcal / mol) of the KSDFT. The energy convergence obtained from Hückel initialization is 5.34 kcal / mol, while the energy convergence achieved using MINAO+1SCF initialization is 1.90 kcal / mol.

[0077] For initialization, note that common methods like Minimal Atomic Orbitals (MINAO) yield densities in the distribution from the training data, as they do not perfectly correspond to solutions to the eigenvalue problem. This mismatch leads to unreliable results even in previous simple tasks and methods with kernels. Therefore, consider initializations that are actually eigenvalue solutions (e.g., Hückel initialization). Alternatively, one could consider using a machine learning model to project the MINAO density onto the space of eigenvalue solutions (projected MINAO).

[0078] Special care needs to be taken when initializing the density. For example... Figure 5 As shown, using all the training techniques disclosed in this paper, the optimization curve from the ground-state density remains static. However, the result from the MINAO density is undesirable, exhibiting typical KSDFT initialization: the curve stabilizes for a period but has gaps below the ground-state energy and even diverges later. Hints about the problem can be gleaned from the curve of the density after the first SCF iteration from MINAO (MINAO+1 SCF), which, although starting with the same degree of deviation from the ground-state density in terms of energy as MINAO, converges stably and approaches the true value well. This observation suggests that the problem with MINAO may be due to its out-of-distribution nature: the model is trained on the density of eigensols from the effective single-electron Hamiltonian matrix, while the MINAO density comes from the superposition of orbitals on isolated atoms. This inspires us to adopt an in-distribution initial density.

[0079] The first choice is the Hückel initialization, which comes from the solution of the eigenvalue problem of a simple Hamiltonian matrix, and is therefore an in-distribution density. According to Figure 5 Although the initialization seems much worse than MINAO in terms of total energy, it surprisingly converges very well to the ground state density.

[0080] The second option, known as projected MINAO, uses a model to correct the MINAO density towards the in-distribution density. In practice, an additional output branch is introduced into the model to predict the required correction to the input coefficients oriented towards the ground-state coefficients. Therefore, a density loss is introduced to also train this branch. Figure 5As shown, after correction, the density optimization curve does indeed converge well and gives accurate energy results. The results appear better than using Hückel initialization because the density corrector provides a starting point close to the ground state, where the model more confidently approximates the functional (since the near-convergent SCF solution is distributed around the ground state). Note that density optimization remains effective even after correction, suggesting the potential to give better results than the end-to-end "structure → ground state density → ground state energy" pipeline.

[0081] Another problem to be solved in density optimization is an appropriate stopping criterion. In principle, the optimization process converges to the rest point (zero projected gradient) of the total energy functional. However, in practice, due to discretization errors, the exact zero gradient point may be missed. There, alternatively, the result at the step with the smallest single-step energy update after a selected number of gradient descent steps is chosen. However, this criterion is not suitable for large-scale systems because the model exhibits more or less some extrapolation error, and even small errors around the ground state can lead to optimization toward unreliable regions of the model. Therefore, in this case, the result after long optimization is unreliable. Therefore, the method disclosed in this paper stops when the norm of the projected gradient first stops decreasing. If no such step exists within the selected number of steps, the step where the energy update first stops decreasing can be chosen. If still no such step exists, the step with the smallest energy update within the selected number of steps can be chosen. The priority of the projected gradient norm is higher than that of the energy update because the former is more sensitive to extrapolation and thus provides a safer stopping point.

[0082] Figure 8 The computing device 12 of the computing system 12 during training is shown. The processing circuitry 14 of the computing device 12 is configured to train a kinetic density functional machine learning model on a training dataset 56 generated at least in part by the following operations. T S,θ (p, During training, based on the orbital values ​​(e.g., wavefunction coefficients C) calculated in intermediate iterations of the self-consistent field (SCF) algorithm evaluation for the training-time molecular structure data 50, the atomic-based density expansion coefficients p and the non-interacting kinetic energy are calculated for the intermediate iterations. T S (p) Gradient of non-interacting kinetic energy p T S(p). Processing circuit 14 is also configured to store density unfolding coefficients and training-time molecular structure as training-time inputs, and to store non-interacting kinetic energies and kinetic energy density unfolding gradients as ground truth outputs in training dataset 54. These values ​​can be processed by the density feature preprocessing module in the manner described herein, either before being stored in the training dataset or before being input into the KEDF model 20.

[0083] like Figure 6 As shown, the kinetic energy density functional machine learning model T S,θ (p, It can have a graph neural network transformer architecture that can be configured to implement a non-local attention mechanism. It should be understood that the kinetic density functional machine learning model... T S,θ (p, This includes transformer neural networks with a multi-head attention mechanism, which is trained using backpropagation of the input and ground truth output during training. For example... Figure 8 As shown, an example of a multi-head attention mechanism is configured to compute the prediction of non-interacting kinetic energy given a density unfolding coefficient p as input. T S (p) represents the attention and loss for the prediction task, and is configured to compute the loss relative to the input density unfolding coefficients via backpropagation for the prediction task of predicting the kinetic energy density unfolding gradient. Each loss is typically measured by comparing the predicted value with the true value of the input. The backpropagation algorithm can be used to train a KEDF model based on the loss function.20

[0084] like Figure 8 As shown, the training-time molecular structure data 50 includes data indicating atomic properties, such as atomic coordinates X and atomic numbers Z. Furthermore, the molecular structure data 50 may indicate atomic properties, such as the number of electrons and orbital basis sets in the training-time molecular structure. The SCF algorithm 52 is performed at least in part by iterating through multiple SCF loops. In this paper, the SCF loop is also referred to as a step and is denoted by an index k. In each SCF loop, the processing circuit is configured to: calculate multiple orbital coefficients, expand the orbital basis set using these multiple orbital coefficients to determine the orbital value for each electron, construct a matrix approximating the single-electron energy operator of a given quantum system, such as the Fock matrix F, calculate the updated orbital coefficients C using the constructed matrix, calculate the non-interacting kinetic energy based on the updated orbital coefficients, and calculate the gradient value of the non-interacting kinetic energy. p T S (p, Based on the orbital coefficients C, the density expansion coefficients p are calculated on the atomic basis. The values ​​of the non-interacting kinetic energy density expansion coefficients, non-interacting kinetic energy and kinetic energy density expansion gradients are stored. The total energy and / or density matrix are determined to have converged to the convergence threshold. If converged, the density matrix of the molecular structure during training is updated based on the orbital coefficients and the iteration is stopped.

[0085] During training, molecular structure data can be stored in a graph data structure, where atomic positions and atomic numbers are represented as node features in the graph data structure, and atomic positions are used to calculate edge values ​​in the graph data structure. For example, in... Figure 8 As shown in the density feature preprocessing module, density coefficients can be expressed in a local atomic reference frame, which is equivalent to global shifts and rotations, and keeps the density expansion coefficients SE(3)-invariant. Furthermore, density coefficients can be naturally reparameterized so that the Euclidean metric in the reparameterized coefficient space tends to the L2 metric in the density function space. Additionally, density coefficients can be rescaled dimensionally. An atomic reference model can be used to linearly kinetic the density coefficients, such that density coefficients corresponding to the same atom type, species, and / or chemical element share the same linear weights.

[0086] Now go to Figures 9A-9C This illustrates a computerized method 100 according to an example implementation. The method can be implemented using the hardware and software elements described herein or other suitable hardware and software. Method 100 is implemented at two distinct times: inference time and training time. Training time typically occurs before inference time. Steps 102-112 are implemented during inference time, and steps 114-156 are implemented during training time. (Continuing...) Figure 9A Method 100 includes receiving molecular structure at 102 during reasoning. The received inference-time molecular structure data includes the inference-time atomic-based density expansion coefficient p and the inference-time molecular structure. Data on the atomic properties (such as atomic position X and atomic number) of each atom Z in the matrix. At 104, the method involves feeding the molecular structure data into a trained kinetic energy density functional machine learning model. T S,θ (p, This generates the density function. ρ (r) electron non-interaction kinetic energy T S (p, The predicted value of ) can be based on or determined by the atomic-based density expansion coefficient p. As shown at 106, a trained kinetic energy density machine learning model has been trained on a training dataset prior to inference, which includes atomic properties (such as molecular structure at training time) as input. The atomic position X and / or atomic number Z of each atom in the molecular structure during training. The atomic-based density expansion factor p. As shown at 108, the atomic-based density expansion factor p indicates the value of the electron charge density expressed on the atomic basis and includes the non-interacting kinetic energy. T S (p, The corresponding ground truth value of ), and as shown at 109, the training dataset can also include the kinetic energy density unfolded gradient. p T S (p, The method includes outputting the predicted values ​​of the non-interacting kinetic energies of the molecular structure during inference. At 111, the method includes outputting the predicted values ​​of the kinetic density expansion gradients from a trained kinetic density functional machine learning model via backpropagation during inference. At 112, the method includes performing orbital-free density functional theory (DFT) calculations using the predicted values ​​of the kinetic density expansion gradients. Figure 1C An example equation for this calculation is shown in the figure.

[0087] Figure 9BThe steps performed during training are illustrated. As shown, method 100 also includes generating a training dataset at 114. The training dataset can be generated at least in part by the steps shown at 116-128. At 116, the method includes calculating the atomic-based density expansion coefficients, non-interacting kinetic energy, and gradients of non-interacting kinetic energy in intermediate iteration steps based on orbital values ​​calculated in the self-consistent field (SCF) algorithm evaluation of training-time molecular structure data. At 118, the method may include performing density feature preprocessing, as shown in steps 120-126. At 120, the method may include expressing the density coefficients in a local atomic reference frame that is equivariant to global shifts and rotations, and keeping the density expansion coefficients SE(3)-invariant, as described herein. At 122, the method may include reparameterizing the density coefficients naturally so that the Euclidean metric in the reparameterized coefficient space tends to the L2 metric in the density function space, as described herein. At 124, the method may include rescaling the density coefficients dimensionally, as described herein. At position 126, the method may include employing an atomic reference model with linear kinetic energy over density coefficients, such that density coefficients corresponding to the same atom type, species, and / or chemical element share the same linear weights, as described herein. It should be understood that during inference, molecular structure data can be stored in a graph data structure, where atomic positions and atomic numbers are represented as node features, and atomic positions are used to compute edge values ​​in the graph data structure.

[0088] At point 128, the method includes storing the calculated density expansion coefficients and the training-time molecular structure as training-time inputs, and storing the non-interacting kinetic energy and the kinetic energy density expansion gradient as ground truth outputs from the training dataset. At point 130, the method further includes training a trained kinetic energy density functional machine learning model on the training dataset. T S,θ (p, ).

[0089] At position 132, the method may include incorporating a kinetic density functional machine learning model. T S,θ (p, The method is configured to have a graph neural network transformer architecture configured to implement a nonlocal attention mechanism. At 134, the method may include configuring the nonlocal attention mechanism as a multi-head attention mechanism, which is trained using backpropagation of the input and ground truth output during training. As shown at 136, the multi-head attention mechanism is configured to compute attention and loss for a prediction task of predicting non-interacting kinetic energy given density unfolding coefficients as input, and as shown at 138, is configured to compute the loss relative to the input density unfolding coefficients via backpropagation for a prediction task of predicting the kinetic energy density unfolding gradient loss. At 140, the method may include outputting a trained kinetic energy functional machine learning model. T S,θ (p, ).

[0090] Turn now Figure 9C Step 116 of method 100 can be implemented by performing step 116A. It should be understood that the molecular structure data at training time may include data indicating atomic coordinates and atom numbers. The molecular structure data may also include several electrons in the molecular structure at training time, as well as the orbital basis set. At 116A, the method includes performing the SCF algorithm at least in part by iterating through multiple SCF loops, and in each SCF loop performing the following operations: at 142, calculating multiple orbital coefficients, the orbital basis set being expanded from the multiple orbital coefficients to determine the orbital value for each electron; at 144, constructing a matrix approximating the single-electron energy operator for a given quantum system (e.g., the Fock matrix F) based on the orbital coefficients; at 146, calculating the updated orbital coefficients C using the Fock matrix; and at 148, calculating the non-interacting kinetic energy based on the updated orbital coefficients. T S (p, At 150, calculate the non-interacting kinetic energy. T S (p, gradient value p T S (p, At position 152, the density expansion coefficients p are calculated based on the updated orbital coefficients C. At position 154, the non-interacting kinetic energy density expansion coefficients, non-interacting kinetic energies, and kinetic energy density expansion gradients are stored. At position 156, it is determined whether the total energy and / or density matrix has converged to within the convergence threshold. If convergence is achieved, the density matrix of the molecular structure at training time is updated based on the orbital coefficients, and the iteration stops. In this way, training data can be generated during the execution of the SCF algorithm when KSDFT is performed.

[0091] Furthermore, computational methods may include using trained kinetic density functional machine learning models. T S,θ (p, Iterative optimization of molecular structure during inference The density expansion coefficient p. This can be achieved by: in each iteration of multiple iterations, feeding the molecular structure data and the density expansion coefficient p into a trained kinetic energy density functional machine learning model. T S,θ (p, This generates the density function. ρ (r) electron non-interaction kinetic energy T S,θ (p, The predicted value is then obtained by passing the kinetic energy density functional machine learning model through the expansion coefficients p relative to the input atomic-based density. T S,θ (p, Backpropagation is used to generate the predicted kinetic energy density unfolded gradient. p T S (p) calculates the sum of the Hart case energy, external potential energy, and exchange-correlated energy, relative to the input atomic-based density expansion coefficient p, and adds it to the calculated gradient. p T S (p); Project the gradient onto a linear subspace of density expansion coefficients that determine the normalized density; Update the density expansion coefficients p using the projected gradient; and Determine whether the total energy change from the previous loop or the normalized value of the projected gradient is greater than the value in the previous loop, and if it is greater, stop the iteration.

[0092] The computerized method may further include: initializing the density expansion coefficients by adding the MINAO initial density to the increment predicted using a trained initialization machine learning model, wherein the trained initialization machine learning model takes the inference-time molecular structure and the standard MINAO initial density expansion coefficients as input, wherein the trained initialization machine learning model has been trained on a training dataset including atomic positions X and indicators of the training-time molecular structure. The data of the atomic number Z of each atom in the data, and the molecular structure during training. The atomic-based density expansion coefficient p indicates the value of the electron charge density expressed on the atomic basis, and includes the corresponding true value relative to the increment of the density expansion coefficient at the SCF convergence step.

[0093] appendix Mechanism of density functional theory The following are details about DFT, including the formula for OFDFT on an atomic basis, and the mechanism and details of using KSDFT to generate values ​​and gradient data for learning KEDF. Atomic units are used throughout.

[0094] Table 1: Symbols

[0095] A.1 Basic Conception of Density Functional Theory For simplicity, the following conception applies to spinless fermions, thus considering only the spatial state. For the constrained Cohen-Shen calculations used for data generation, a pair of electrons with opposite spins sharing a common spatial orbital are employed, which is equivalent to replicating the orbital in the following conception. The mechanism of the DFT can be more intuitively introduced under Levy's constrained search formula. The ground-state N-electron Schrödinger equation is equivalent to the N-electron wavefunction under the following variational principle. Optimization issues: (5) in It is a kinetic energy operator. It is the electron-electron Coulomb interaction (internal potential), and From a given molecular structure The external potential of a single atom generated by the electrostatic field of the specified atomic nucleus ,in and : (6) Although the optimization problem is precisely defined, directly optimizing the N-electron wavefunction is computationally very challenging. Specifying the wavefunction... ψ And the assessment of energy already requires exponential cost in principle, because ψ It is its dimension A function that increases with N. To generate a more easily optimized problem, it is then desirable to optimize a functional with decreasing single-electron density. (7) It has an intuitive physical interpretation of charge density even in the classical view, and more importantly, the cost of specifying the density is in principle constant (relative to N), because density is constant in its dimension. The function on the surface. This is the starting point of density functional theory (DFT), first formally verified by Hornberger and Cohen.

[0096] In terms of density, the external potential energy (i.e., the last term in equation (5)) is already an explicit density functional, since the external potential is a single entity: (8) For other energy terms, using density as an intermediate term, the optimization problem in equation (5) can be equivalently performed at two levels: (9) (10) Here, the density functional is defined by the result of performing a constrained search on the first-order optimization problem in equation (9). (11) It is called a universal functional because it is independent of the system specification (i.e., M or ...). It consists of the kinetic energy of electrons and internal potential energy. The optimization objective is then formally transformed into a density functional. As shown in equation (10).

[0097] To perform practical calculations, variants of kinetic and internal potential energy that allow for explicit calculation or have known properties are introduced to cover... U [ ρ The main part of the corresponding energy in [ ]. The internal potential energy is covered by its classical version, which assumes no correlation, and is called the Hartley energy: (12) The kinetic energy is covered by the kinetic energy density functional (KEDF), which is constrained in a manner similar to that of the universal functional: (13) The remainder of the universal functional is called the commutative correlation (XC) functional: E XC [ ρ ] := U [ ρ ] T S [ ρ ] E H [ ρ ], It is also a density functional by definition. Under this decomposition, the density optimization problem in equation (10) becomes: (14) here E eff [ ρThe term ] is defined for future convenience and named according to the interpretation of its variation of effective potential, which will be detailed later. It is approximated using carefully designed explicit expressions or machine learning models. T S [ ρ ]and E XC [ ρ The function [] can be used for practical calculations. This is the concept of orbital-free density functional theory (OFDFT). In fact, the object to be optimized is electron density, which is It is a function on a constant-dimensional space, thus greatly reducing the computational complexity of the original variational problem equation (5). With proper design... T S [ ρ ]and E XC [ ρ Under approximation, the complexity on an atomic basis is advantageously O(N). 2 ).

[0098] Considering KEDF T S [ρ] is proportional to the functional E of XC XC [ρ] is more challenging; Cohen and Shen used an equivalent formula from KEDF to allow for accurate calculation, at the cost of increased complexity. Alternative formulas optimize the deterministic wavefunction. The deterministic wavefunction for N electrons is derived from the wavefunction of N single electrons. Functions, also known as tracks, follow this form:

[0099] Given orbit .

[0100] The equivalent optimization problem that runs parallel to equation (13) is: (15) in Given a standard orthogonal orbit (16) In equation (15), the second equation applies to calculations involving spinless fermions or confined spin electrons because the density is normalized to N, and a set of (non-collinear) functions can always be orthogonalized using, for example, a Grace-Schmidt process. This equivalence can be determined from T as the non-interacting part of the kinetic energy. S To understand this, we can use the interpretation of [ρ]. In reality, for non-interacting systems, only kinetic energy and external potential energy exist (i.e., taking...). Therefore, the two-level optimization parallel to equation (9) becomes: (17) Therefore, T S[ρ] is the ground-state kinetic energy of a non-interacting system with ground-state density ρ. On the other hand, the ground-state wavefunction of a non-interacting system is generally deterministic (at least when the ground state is non-degenerate [Theorem 4.6]). Therefore, the optimization can be rewritten as: (18) Its indicator equation (15).

[0101] Returning to the main problem, in the variational problem equation (14), we utilize information about T. S This knowledge of [ρ] leads to:

[0102] It can be converted into direct trajectory optimization using a single-level optimization: (19) This is the formula for Cohen-Shen density functional theory (KSDFT). With decades of development of the XC functional approximation, KSDFT has achieved remarkable success and become the most popular method in quantum chemistry. In its formula, the object to be optimized is a set of orbitals. , it is The N function on. This is still better than optimization. The N-electron wavefunction is cheaper, but requires optimization that is N times greater. The function, KSDFT, has a complexity at least O(N) times greater than OFDFT. On an atomic basis, KSDFT has at least O(N) complexity. 3 (Using density fitting) without the complexity of further approximation. The machine learning model described in this paper provides a chance to approximate the KEDF accurately enough to match the successful XC functional approximation. This enables accurate and practical OFDFT calculations with lower complexity, improving the accuracy-scalability trade-off in quantum chemistry.

[0103] A.2 KSDFT calculation generates KEDF labels The reason why the KSDFT calculation process can provide KEDF values ​​and gradient labels will now be explained. Details of calculations on the atomic basis are deferred to supplementary section A.4. First, a typical algorithm for solving optimization problems in KSDFT is described. To determine the orbitals... The optimal solution requires equation (19) relative to each orbit. φ i Energy functional E Changes in [Φ]:

[0104] Item V eff[ρ[Φ] As relative to density ρ The changes occur because orbits are determined solely by the density they define. ρ [Φ] Determine except T S E other than energy eff (Defined only in equation (14)). Through equation (16), Then, equation (20) is given. Combining equation (20) with the variation of the orthogonal constraint yields the optimality equation of problem equation (19): (twenty two) In the derivation, only the normalization constraint is used. i | i =1 Lagrange multiplier ε i It was imposed because of the equation (22) obtained. It is an Hermitian operator The eigenstates are called the Fock operator, and are therefore naturally orthogonal in the non-degenerate general case. These equations are analogous to the Schrödinger equations for N non-interacting fermions, where V eff[ρ[Φ] ](r) as The function on it acts as an effective monomeric external potential, hence the name.

[0105] Notice, V eff[ρ[Φ] Previously unknown, because it depends on the solution of the orbit. Therefore, fixed-point iteration is used: starting from a set of initial orbits. We begin by constructing the Fock operator using the results from previous iterations. (twenty three) After this derivation, Take as ,Right now, The eigenvalue problem of the orbit in the current iteration: (twenty four) The iteration stops when "self-consistency" is achieved, that is, the characteristic state solution in the current step is obtained. With definition The track in the previous steps Consistent (up to acceptable error). This is the Self-Consistent Field (SCF) method.

[0106] An important fact about SCF is that in each iteration k, the solution... The ground state of a non-interacting system with N fermions is precisely defined, where the N fermions are represented by the external electric potential V. extIn effective single potential China Mobile. In fact, the variational problem equation (18) describing a non-interactive system can be reformulated as a single-level optimization, such as: (25) Its changes and therefore The solution yields the same equation (24). This reveals the relationship between the SCF solution and the KEDF solution: orbital This solution achieves the minimum non-interacting kinetic energy in all orthogonal orbits. The minimum non-interacting kinetic energy Leading to the same density ρ [Φ ( k)] Therefore, (25) can be further minimized; it can also be seen from the equivalence of equation (18); by using the KEDF equation (15) as an alternative for the non-interacting ground state kinetic energy, therefore: (26) This instructs each SCF iteration to generate a target for T S The label. Furthermore, since the non-interacting variation problem equation (25) is equivalent to its two-level optimization form equation (18), which in turn is equivalent to the density optimization form using the KEDF equation (17) (which explains the alternative KEDF form equation (15)), the solution from each SCF iteration... density ρ [Φ ( k)] Minimize equation (17). Therefore, it satisfies the condition of having a Lagrange multiplier (chemical potential). µ (k) The normalized constraint equation (17) (taking V) ext for The equation for the change of ) (Euler's equation): (27) When density ρ When expanded on the basis (see supplementary section A.4.3), KEDF The change is related to the gradient relative to the density coefficient. Therefore, each SCF iteration also generates a value for... T S The gradient label is applied until the projection.

[0107] It is worth noting that when SCF iterates k effective potential no At that time, these independent variables still hold true, because the derivation from equation (25) to equation (27) only requires It is a single potential. This allows for a more flexible data generation process because in common DFT computation settings, such as when using the "Direct Inversion in Iterative Subspace" (DIIS) method, Actually deviated This results in more stable and faster convergence. It also shows that even when the XC functional used in data generation is inaccurate, the generated values ​​and the convergence with respect to T... S The gradient labels of [ρ] remain accurate because, as long as it is pure (i.e., depends only on the density feature), the XC functional still gives an effective single-unit potential to define the non-interacting system equations (24) or (25). In this sense, KEDF data generation is easier than XC functional data generation.

[0108] A.3 Equations on the Atomic Base For practical calculations, KSDFT typically uses atomic bases. To expand the orbitals, we can perform SCF iterations in equation (24) for the molecular system. The expansion gives: (28) It transforms the solution for eigenfunctions into a common problem of solving for the eigenvectors of a matrix. On the other hand, as emphasized in Introduction 1 and Results 2.1, the goal is to [achieve this goal on an atomic basis]. The above represents the density to achieve an efficient OFDFT implementation. (29) The left-hand side of equations (28) and (29) can also be expressed as i,C or Φ C and ρ p This highlights the dependence on the coefficients. Note the orbital basis. and density base All depend on the molecular structure M, because the position and type of each basis function are determined by the coordinates x of the corresponding atom. (a) and the number of atoms Z (a) Confirmed. However, the developments in this subsection are for a given molecular system M, therefore its appearance is omitted for density or orbital representations. Generally, the base... B and M The number of electrons increases with the number of electrons. N Linear increase, i.e. O ( B ) = O ( M ) = O ( N ).

[0109] A.3.1 KSDFT on atomic basis For the SCF iteration in equation (24), using the orbital expansion in equation (28), it becomes: , Pair each functional equation with its basis functions. η α (r) The integral is given as follows: It can then be formulated as a generalized eigenvalue problem in matrix form:

[0110] Here, F (k) It is called the Fock matrix, and S is the overlap matrix of the orbital basis.

[0111] To show the expression for the Fock matrix, we can first express the expression for the density defined by the orbital coefficients from equations (16) and (28): (32) The density matrix is ​​defined as corresponding to the orbital coefficients: (33) Note that equation (16) requires orthogonal orbits, therefore equation (32) requires C to satisfy the corresponding orthogonality constraint shown in equation (43) below. The orbit coefficient solution C in the SCF iteration. (k) This constraint is satisfied, as explained later.

[0112] In the derivation of the SCF iteration, Taken as ,Right now This allows for the application of equation (21) and Will Explicitly computed as The following series of definitions are introduced:

[0113] In definition For the sake of simplicity, the notation for the Coulomb integral will be used: drdr. (37) In fact, S, T, V ext Coulomb integral The integral in can be analytically computed as a basis in Gaussian type orbitals (GTOs). V XCC The integration is performed digitally on an orthogonal grid, as is commonly used in DFT calculations. After convergence, the total energy can be calculated from the orbital coefficients C using the following equation:

[0114] in, , , They are flattened Γ, T, and V respectively. ext The vector. The term E is calculated again by numerically integrating C over an orthogonal grid. XC [ρ C ].

[0115] computational complexity Note that V comes from equation (35). HC The construction of E from equation (40) H (C) evaluation requires O(B) 4 )=O(N 4 The complexity is O(N). Even when using [method / method], the complexity is reduced to O(N). 2 When performing density fitting, the complexity of each SCF iteration of KSDFT is also O(N). 3 Because the complexity of density fitting itself is O(N), 3 (See Supplementary Section A.4.1).

[0116] Orbital standard orthogonality At the atomic level, orbital orthogonality i | j = δ ij Become Or in matrix form, C SC=I. (43) As mentioned after equation (22), only the normalization constraint needs attention, since the orbitals are eigenstates of the Hermitian operator and are therefore orthogonal if non-degenerate. This property is transferred to problems in matrix form: (44) To fully fill the constraint, it is only necessary to fill each eigenvector of the problem in equation (30). Normalization into form ; Explicitly .

[0117] Relationship derived from direct gradient Therefore, the gradient of the energy functional with coefficients can also be obtained: The matrix form of the optimization equation of the SCF iterative problem equation (30) under the basis is directly derived from equation (19). Its gradient is related to the change in the orbital functional by integrating over the basis: (45) This change is given by equation (20), which is , It transforms the gradient into matrix form: C E(C) = 2F C C. (46) For orbital orthogonality constraints, as mentioned above, only the normalization constraints need to be explicitly handled. This is achieved by introducing a Lagrange multiplier ε for each constraint in equation (44). i And give the gradient for the corresponding Lagrange term. This leads to the optimality equation in matrix form: F C C = SCε. (47) By constructing the corresponding fixed-point iteration, equation (30) is derived.

[0118] Accelerating and stabilizing SCF iterations As mentioned above, for more stable and faster convergence, it can be combined with F C ( k 1) and F is used in different ways (k) and V eff (k) Direct inversion in the Iterative Subspace (DIIS) method is one option for this. In DIIS, the Fock matrix F in the eigenvalue problem equation (24) for each SCF iteration k is used in the previous step. (k) Take the standard Fock matrix F C ( k ′ ) , k ′ <k As a weighted mixture:

[0119] in, Positive and normalized weights Due to normalization, the kinetic energy part T of the matrix remains unchanged, thus it is consistent with the form in equation (31) (or the operator form of equation (23)), where: (48) OFDFT on A.3.2 atomic base To solve the OFDFT optimization problem in equation (14), it is unnatural to construct a fixed-point SCF iterative process from its variation in equation (27). Therefore, density optimization based on direct gradients is performed. To this end, the energy functional of the density function in equation (14) needs to be transformed into a function of the density coefficients using the basis expansion of the density function in equation (29):

[0120] The concept of Coulomb integral (ω) µ |ω ν The density basis { is defined in equation (37). Recall that in this subsection, the density basis { ω µ} µ The dependency has been omitted from the following functions, for example, on the molecular structure M. T S (p). Using software libraries (e.g., libcint in PySCF), it is possible to efficiently compute Gaussian type orbitals (GTOs). and v ext The integral, as the basis { ω µ} µ Density defined by numerical integration on orthogonal grids ρ p To calculate the item E XC [ ρ p This is similar to what is commonly used in DFT calculations. In M-OFDFT, a machine learning model is used. T S,θ (p , M) Calculated directly from coefficient p T S (p).

[0121] In order to use the learned KEDF model T S,θ (p) Performing direct optimization requires the gradient of the energy functional in equation (49), given by the following equation: (54) Then, following equation (1), this gradient is used to update the density coefficients after projection onto the linear subspace of the normalized density.

[0122] The derivation is related to the integral of the basis change. gradient p E(p) can also be derived from the relationship between gradient and change already seen in equation (45): (55) The changes given in equation (21) are compared with the basis functions { ω µ} µ Integrating this equation will restore the Hartley energy gradient in equation (54). and external energy gradient v ext The formula also applies to kinetic energy. p T S Gradient of (p) and XC energy p E XC The gradient of (p).

[0123] Automatic differentiation implementation for gradients In the implementation of M-OFDFT, automatic differentiation is used directly to evaluate the KEDF model. p T S,θ (p , The gradient of M can be easily calculated if the model is implemented using a common machine learning programming framework (such as PyTorch). For ease of calculation... p E XC (p) reimplements the PBE XC functional in PySCF using PyTorch, and also evaluates its gradient via automatic differentiation. When using the residue version of the KEDF model... T S,res,θ The basic KEDF (taken as APBE KEDF) is also implemented in this way, as detailed in equation (67) in Supplementary Section B.4 later.

[0124] computational complexity As will be detailed in Supplementary Section B, the transformer-based KEDF model for M-OFDFT has a quadratic complexity of O(A). 2 )=O(N 2 ).against E XC The PBE functional and the APBE functional for the basic KEDF are at the GGA level (generalization gradient approximation), so the amount of energy is evaluated to compute the density feature with O(M) cost at each grid point, totaling O ( MN grid ) costs, of which N grid It is the number of grid points, and then O(N) grid The costs are orthogonal. Therefore, the complexity of these energies is O(MN). grid It is also a quadratic O(N) 2), because N grid =O(N) (despite having a large pre-factor). Using automatic differentiation to evaluate the gradient results in the same order as the evaluation function, therefore also O(N). 2 Complexity. E is evaluated using equations (51) and (53). H (p) and E ext (p) and evaluating their gradients using equation (54) requires O(M) 2 )=O(N 2 Therefore, the complexity of M-OFDFT is quadratic O(N). 2 It is indeed less complex than KSDFT (as detailed in Supplementary Section A.3.1 above).

[0125] In addition to the advantages of asymptotic complexity, the fact that M-OFDFT is implemented in PyTorch allows it to effectively utilize the GPU. These factors combined contribute to the significantly higher throughput of M-OFDFT compared to KSDFT.

[0126] A.4 Tag Calculation on Atomic Bases This section details the orbital solutions used from each SCF iteration k. orbital coefficient C (k) Data tuples for learning the KEDF model The calculations are as follows. Following the previous subsection, the appearance of density and orbital representations of M has been omitted (e.g., in T...). S (p, M) is used, and the index d for different molecular systems is omitted. The k index is kept to reflect the solution based on the SCF iteration, but this does not apply to arbitrary orbital coefficients C.

[0127] A.4.1 Density Fitting This process begins at the density base. Calculate the density coefficient p (k) , used to represent the solution C from the orbital coefficients (k) The defined density. This process is called density fitting, which is also used for acceleration in KSDFT, where the atomic basis of the density is also called the auxiliary basis. Density coefficient p (k) The density needs to be expressed by equation (29). ρ p ( k) Fit to the density defined by equation (32) ρ C ( k) Note that equation (32) essentially utilizes Γ as a flattening. (k) ) The density is expanded to the paired basis {η α (r)ηβ (r)} αβ This is a classical coordinate transformation problem from paired bases to density bases. The classical approach is to minimize the standard L2 norm of the residual density:

[0128] Among them W µν := ω µ | ω ν ,L µ,αβ := ω µ | η α η β , and D αβ,γδ := η α η β | η γ η δ These are the overlap matrices of the density basis, the overlap matrix between the density basis and the paired basis, and the overlap matrix of the paired basis, respectively.

[0129] Note that this is p (k) The quadratic form of the solution is

[0130] However, the standard L2 metric on the density function may not be the most favorable metric for density fitting. Instead, energy is a directly relevant quantity. Kinetic energy is the most desirable metric because it will fit the density p. (k) With kinetic energy label T S (k) The mismatch is minimized. However, there is no explicit expression to calculate the kinetic energy from the density coefficient. Therefore, the Hartley energy and external energy are considered as metrics. (Using the XC energy requires arbitrary choice of functional approximation and is computationally more expensive.) For the Hartley energy defined in equation (12), note that it is quadratic in density, and fit p by minimizing the (2x) Hartley energy caused by the residual density:

[0131] The symbols with wavy lines are in the integral kernel. The corresponding overlap matrix on the matrix, which uses the notation of the Coulomb integral as defined in equation (37), is W. µν := ( ω µ | ω ν ),L µ,αβ := ( ω µ | η α η β ), and D αβ,γδ := ( η α η β | η γ η δ As defined in equation (36), the solution is a quadratic form. This result can be interpreted as the Hartley energy (Coulomb integral) defining a metric on the density function space.

[0132] For the external energy defined in equation (8), since it is linear in density, p is fitted by directly minimizing the difference between the defined external energies:

[0133] To combine these two metrics, the final optimization problem is a combinatorial least squares problem:

[0134] It corresponds to the overdetermined linear equation in matrix form:

[0135] This is solved directly using a least-squares solver. Normalization constraints are not explicitly considered in this transformation. This is because it has been satisfied with high accuracy due to its close fit to the original density.

[0136] A.4.2 Value Label Calculation The label of the KEDF value can be calculated using equation (26) on the atomic base by employing expression equation (39):

[0137] Among them, from equation (33) and of It is a flattened vector. The density is fitted from C using the method described above. (k) Calculate the corresponding density coefficient p (k)Due to the fitted density p (k) Possibly due to the incompleteness of finite bases, and related to C (k) The original density is defined slightly differently, therefore T S (C (k) It might not be p. (k) The optimal kinetic energy label. In fact, as mentioned above, in density fitting, there is no known way to directly find the closest approximation to T. S (C (k) The kinetic energy density coefficient p (k) Conversely, p (k) It is fitted to match the Hartley and external energies. Therefore, it is assumed that the total energy is less affected by the density fitting error than the kinetic energy, rather than taking T. S (p (k) )≈T S (C (k) Instead, take E(p) (k) )≈E(C (k) This means T S (p (k) )+E eff (p (k) )≈T S (C (k) )+E eff (C (k) Therefore, p (k) The KEDF value label is considered as:

[0138] Following the definitions and expressions of equations (38-42) and (49-53), the labels for the two variants of the functional model detailed in Supplementary Section B.4 can be computed accordingly. For KEDF T S,res The residual version, using p (k) The value at the location is used to modify the label accordingly:

[0139] For the version of learning KEDF and XC energy summation The corresponding tag is: The last term can be omitted because the Hartley energy difference and the external energy difference are p in the density fit. (k) Minimize. Therefore, the label is considered as:

[0140] A.4.3 Gradient Label Calculation On the atomic basis, after equation (50), the kinetic energy function of density is converted into the density coefficient. T S(p) := T S [ ρ p The function of ]. For its gradient p T S (p), following the facts in equation (55), it is consistent with the functional T S The change in [ρ] is related to the basic integral: .

[0141] The change corresponding to the known density is given by equation (27) above, which comes from the orbital solution in the KSDFT SCF iteration. If the error in the density fitting is omitted and... approximate The gradient can then be obtained by integrating both sides of equation (27) with the density basis:

[0142] In practice, a chemical potential μ is not required. (k) This is because the gradient is important in its projection onto the tangent space of the normalized density in order to maintain density optimization of the normalized density (see Equation (1)). The space of the normalized density is a linear space. ,because It is consistent with its tangent space. Through application... A projection onto the tangent space is given by:

[0143] For the same reason, the gradient loss function equation (3) also only matches the projected gradient of the model with the projected gradient label.

[0144] The remaining task is to evaluate v in equation (57). eff (k) Considering V eff (k) (r) may not be considered as an explicit effective potential. The complexity of (Equations (21) and (32)) can be mitigated by leveraging the orbital basis representation already available in the SCF problem (Equation (30)) when DIIS (Equation (48)) is used for SCF iterations. Complete the evaluation v eff In equation (31), it is defined as the integral of the paired orbital basis: Using the expansion factor K of the density basis onto paired orbital basis, i.e. ω µ (r) = ∑ αβ K µ,αβ η α (r) η β (r) can be accessed via Perform the conversion.

[0145] However, solving for K is unacceptable: using least squares is equivalent to solving for DK=L (or solving for...). Due to D (or ) has shape B 2 ×B 2 Therefore, the complexity is O(MB). 4 )=O(N 5 Even if it only needs to be called once for a molecular structure, the cost remains prohibitively high, even for medium-sized molecules. Furthermore, approximation... The optimality of density optimization using the learned KEDF model on the same molecular structure is not guaranteed, as explained later.

[0146] Therefore, consider another approximation to calculate the gradient in a more direct way. The approximation is that equation (27) also applies to the fitted density. ρ p ( k) : (59) in V eff{p ( k ′ )} k ′ <k It is constructed from the fitted density coefficients in previous SCF iterations. Instead of V in equation (27) eff The orbital coefficient solution from the previous SCF iteration is used. The chemical potential μ′(k) can be different, but it is not used due to the above independent variable in equation (58). The approximation holds if, for example, the density fitting error can be omitted when using a large basis set. Following the above process, the corresponding kinetic gradient after projection is given by the following equation:

[0147] Accordingly, calculation V eff{p ( k ′ )} k ′ <k A direct method is needed. This can follow a known function. Constructed in SCF iteration The relationship is accomplished by considering the matrix definition equations (31) and (34), which are weighted averages. Following this pattern, the effective potential required in equation (59) is constructed as follows: Weight They are considered to be the same weights calculated in the SCF iteration.

[0148] Its vector form on the density basis is given by the following equation:

[0149] Each v effp ( k) The calculation can be directly displayed on the screen as V. eff[ρp] Equation (21) is then executed. In this embodiment, the automatic differentiation implementation mentioned in Supplementary Section A.3.2 is used to conveniently calculate each v. effp ( k) .

[0150] See A.3.2, because the inventor noted the following facts: veffp= p E eff(p) , (63) in E eff (p) := E eff [ ρ p Defined in equation (49), and its gradient p E eff (p) is given by equation (54). This fact is again attributed to the relationship between gradient and change revealed in equation (55) and note V eff[ρ] The change in the effective energy functional is defined as the change in equation (21). As analyzed at the end of Supplementary Section A.3.2. The evaluation gradient has the same properties as the evaluation energy E. eff (p) The same complexity, which is O(M) 2 )+O(MN grid )=O(N 2 This is better than the above O(N) 5 It is much less complex.

[0151] In short, {v effp ( k ′ )} kFirst, the automatic differential following equation (63) is used to calculate v, which is used to construct v following equation (62). eff{p ( k ′ )} k ′ <k Then gradient p T S (p (k) The gradient label is given upwards from equation (60) to the projection. Since the gradient-supervised loss function equation (3) explicitly projects the gradient error, it is not necessary to project the gradient label itself before evaluating the loss (i.e., the loss is the same regardless of the gradient label being projected; because the projection is idempotent). Gradient label p T S (k) The effective potential vector of DIIS directly as a density construct: .

[0152] For KEDF T S,res and version E TXC The residue version, which also includes the XC energy as detailed in Supplementary Section B.4, produces the tag accordingly:

[0153] It also uses automatic differential calculation. p T APBE (p (k) )and p E XC (p (k) ).

[0154] Besides the convenient and efficient computation using automatic differentiation, this choice of gradient labels can also train a KEDF model that leads to the correct optimal density, because the labeling method is compatible with the density optimization process of M-OFDFT shown in equations (1) and (54). More specifically, in the SCF iterative convergence (see labeled "..."), When the quantity is "", the effective potential of DIIS is By designing a single-step, explicit effective potential V eff[ρC ] (Equations (21) and (32)). Accordingly, the effective potential vector v of the DIIS constructed by the density eff{p ( k ′ )}k ′ < (Equation (61)) converges with v effp (Equation (62)), which is derived from equation (63) p E eff (p This is consistent. This provides gradient labels to the KEDF model via equation (60), which forces the model to satisfy:

[0155] This forces the density optimization process equation (1) to converge to p. This refers to the true ground-state density coefficient. The optimality of density optimization using the learned KEDF model can then be expected.

[0156] A.5 Scaling properties on atomic bases KEDF possesses precise scaling properties, describing the change in density after uniform scaling (stretching or extrusion). At a scaling (extrusion) rate λ, the uniformly scaled density follows the following rule: The scaling properties are described below: (64) Now consider using the KEDF model T on an atomic basis. S This property of (p, M). Because the coefficient p given by equation (29) is used as... ρ p At the atomic base The following represents the input density, therefore it is first compared with the coefficients. (p) represents the scaled density on the same basis. : (65) Solve using least squares (p) gives the same as above Among them, this is And W µν := ω µ | ω ν Same as above. Then the scaling property equation (64) is transformed as follows: (66) However, in order to preserve equation (64) precisely, the expansion of equation (65) must hold exactly. But this is not the case: finite basis It is incomplete, and the scaling basis function is... It depends non-linearly on the original basis functions. Therefore, the transformed scaling property equation (66) on the atomic basis is not exact and may therefore be inapplicable to the KEDF model. T S,θ (p , M) offers many benefits. When using isothermal atomic bases, it appears possible that the form is: ωµ =( a,τ, ξ)(r) = wa,τ, ξ(r x( a )),in wa,τ, ξ(r= ( x,y,z )) := x ξ1 y ξ2 z ξ3 exp( αa, |ξ| βτ ∥r∥2), where the temperature ratio β>1, the monomial exponential parameter ξ=(ξ1, ξ2, ξ3), and the exponential parameter shared across ξ values ​​with the same |ξ| := ξ1+ ξ2+ ξ3. α a,|ξ| The index τ ranges from 0 to T. a,|ξ| Therefore, in accordance with In the case of scaling:

[0157] x , It exists in the same isothermal base but in scaled molecular structures , As long as τ > 0. For τ = 0, it corresponds to the flattest basis function, and its coefficients p µ=(a,τ=0,ξ) Approaching zero, therefore, density representation can be omitted. Therefore, in using... In the case of scaling, the scaling density can be expressed as:

[0158] The new coefficient is defined as follows:

[0159] Then, the scaling property equation (64) can be written as: .

[0160] However, this only applies to a single value that depends on the scaling of the radix (also for integers n). (This can be done by omitting contributions from the first n coefficients). Furthermore, this constraint connects the original input to a scaled configuration. It can be located far from the equilibrium structure. Used for these scaling conformations T S,θ The model's behavior may be largely irrelevant to real-world applications. Therefore, no benefit can still be gained from scaling characteristics.

[0161] Details in the BKEDF model The KEDF model takes the atomic number Z, position X, and density coefficient p of all atoms as input, which is used to construct node features encoding information about the electron density around each atom. These features are further transformed and used to predict the kinetic functional energy, and the gradient of the kinetic functional with respect to the density coefficient is obtained through automatic differentiation. Considering the essentiality of nonlocal computation for KEDF, a neural network transformer architecture with nonlocal attention is adopted, called the Graphormer architecture, such as... Figure 6 As shown, this architecture demonstrates attractive performance in predicting various properties (e.g., ground-state energy and optical gap) using molecular structure. Furthermore, distance features (as edge features) are incorporated into the interaction calculations to accommodate those not covered by the atomic features of the pair. Based on the original architecture, a node embedding layer is proposed to generate initial atomic representations from various inputs, where intra-atomic interactions are handled by MLP modules ( Figure 6 Modeling is performed in (a) and (b). Furthermore, the self-attention module in each G3D layer considers the interatomic interactions between the density coefficients hosted on any pair of atoms. Note that the nonlocal architecture has O(N) 2 While reducing complexity, it maintains better scalability than KSDFT. Furthermore, local covariance is designed to eliminate geometric variability, and several efficient training techniques are proposed to handle large gradient ranges. The exploration of different functional variants of the Graphormer model is also discussed.

[0162] B.1 Model Specification B.1.1 Auxiliary base splicing As described in Method 4, M-OFDFT uses atoms as the effective density representation in a molecule, and can be attributed to the fact that the basis coefficients of each atom can naturally be used as the density features by node in the molecular graph. Specifically, isothermal auxiliary basis sets with a β parameter of 2.5 are used as density basis sets. The number of basis functions in molecule M is on the same order of magnitude as the number of atoms A or electrons N. Each chemical element has its own basis function set, and each set T... elemThe magnitudes vary depending on the element type. Detailed orbital compositions for each element type are summarized in Supplementary Table 2. To make the coefficient vector uniform across all atoms, these basis functions are concatenated and broadcast to all atoms, resulting in a T-dimensional density coefficient vector p on any atom. a T=P elem T elem Each base ω µ This is then specified by the position of the central atom and the basis function index. Therefore, (a, τ) is used to index the τ-th basis of atom a. More specifically, in the QM9 dataset, the 477-dimensional concatenated coefficient vector consists of 20, 109, 116, 116, and 116 basis functions corresponding to elements H, C, N, O, and F, respectively. For example, given the coefficients of a hydrogen atom, they are placed in the first 20 dimensions, and a zero-value vector is used to mask the other positions.

[0163] Figure 6 The architecture of the KEDF model based on atomic numbers is shown. (a) The KEDF model features atomic number Z, position X, and density coefficient. Input predictive kinetic energy T S,θ The node embedding module generates atomic features h by combining different input features updated through a series of G3D layers. The GBF layer and the Multilayer Perceptron (MLP) module generate distance features from location X. X is incorporated into each G3D layer to model atomic interactions. The MLP module converts the updated features into atomic energies, which are summed to obtain kinetic energy. The atomic reference module provides an analytical reference for the kinetic energy. (b) The node embedding module integrates information from three sources: atomic features obtained by the atomic embedding layer, which assigns learnable weight vectors to each element, and density features processed by the shrinking gate layer along with the MLP encoding the chemical environment and the shrinking GBF embedding. (c) The G3D layer is based on the standard transformer encoder layer and uses distance features as auxiliary attention bias. (d) The density feature preprocessing module uses three techniques to transform the original density coefficients p into SE(3)-invariant and easily controllable density features. Local frame, natural parameterization, and dimensional rescaling. Figure 7 The diagram illustrates this. Zero-value masks are also used to mask predicted gradients during density optimization, avoiding the introduction of irrelevant gradient information from the masking location.

[0164] Figure 7 The assembly of auxiliary bases with different atomic types is shown. Figure 7 In this context, τ is the index of the density coefficient dimension, and p a,τ This represents the coefficient vector of element a. Note that this is illustrative, so the coefficient dimensions may not correspond precisely to the actual dimensions.

[0165] Table 2: Orbital composition of the basis set associated with each element.

[0166]

[0167] B.1.2 Node Embedding Layer like Figure 6 As shown in (b), the node embedding layer is designed to integrate various input features into an initial hidden node representation h, which is then passed to the Graphormer model to estimate the kinetic functional. Specifically, for atom a, the atom embedding layer embeds the node embedding layer by assigning a learnable weight vector to each atom type Z. (a) This is used to generate atomic feature representations. It's important to note the density coefficient vector p. a: ∈ R T , are SO(3)-isovariant geometric tensors of different orders, which cannot be directly processed by Graphormer. To solve this problem, a system consisting of several sub-modules was designed ( Figure 6 The density feature preprocessing module (d) in the middle can not only process the original coefficient vector p a: Transformation to SO(3)-invariant feature representation Furthermore, it effectively reduces the geometric variability of the electron density at the atomic center. Technical details of this module are mentioned in Method 4.2 and further explained in Supplementary Sections B.2-B.3. It was also observed that the treated... This exhibits a large numerical range, which is unfavorable for stable optimization of common neural network architectures. Therefore, applying the shrinkage function λ... co ·tanh(λ mul · To map each density coefficient to a bounded space, where λ co and λ mul These are learnable parameters. Because... Relative to the invariance of arbitrary rotations, applying any nonlinear operation (e.g., shrinking function) to this feature will not violate SO(3) symmetry. The coefficient vector is then further projected into the hidden density feature using a multilayer perceptron (MLP). Following the graphormer, the relative distance between two atoms is embedded with a set of edge-type perceptual Gaussian functions (GBFs), where the mean and variance of each basis are learnable parameters. To capture the chemical environment around atom a, the shrinking GBF embedding is used. These features are merged into auxiliary node features. The sum of these three types of features is used to form the initial node representation h. a ∈ , where D is the hidden dimension of the model.

[0168] Figure 7This illustrates the splicing of auxiliary groups of different atom types. τ is the index of the density coefficient dimension, p a,τ This refers to the coefficient vector of element a. Note that this is an illustrative representation, so the coefficient dimensions may not precisely correspond to the actual dimensions. B.1.3 Backbone Structure The KEDF model is based on the Graphormer model, which is a transformer-based graph neural network that can efficiently perform global computations on molecular graphs while retaining the powerful expressiveness of the transformer architecture. The model mainly consists of an embedding layer and several stacked Graphormer-3D (G3D) layers, where each G3D layer comprises an attention layer and a feedforward layer. Figure 6 The attention layer takes the hidden node representation h as the input token and the relative distance feature... Learnable attention biases are used as attention mechanisms. Specifically, given node features h = {h1, ..., h2}... A} and pairwise attention bias Where a and b represent atomic indices, the G3D layer assigns an attention score A to an attention head. ab The calculation is as follows:

[0169] Among them, U (Q) and U (K) is a linear projection of the node features' "query" and "key," where d is the dimension of the query and key vectors. To account for the interaction between electron density features at the atom centers, GBF embeddings are used. Projected as pairwise attentional bias .

[0170] B.2 Local System Details As mentioned in Method 4.2, by expanding the electron density over a set of atomic bases (i.e., a contracted Gaussian with spherical harmonics), the expansion coefficients p mathematically act as geometric tensors for various angular components, consistent with rotations and translations. A local framework that rotates simultaneously with the structure is employed to decouple the density characteristics from coordinate system changes and unnecessary geometrical transformations. This is to achieve the desired effect on atomic x... a , Construct a local system, selecting a pointer that points to its nearest atom. . The axis is located at the second nearest non- The orientation of atoms In the row of the cross product, and then Shaft = × Given. Note that excluding hydrogen atoms near the central atom makes the resulting local system more dependent on heavy atoms, whose positions are more stable than hydrogen atoms and reflect a more reliable local structure. The local system associated with atom a is represented as...

[0171] The density coefficient vector and the gradient vector in the local system are calculated by the following formula:

[0172] as well as .

[0173] like Figure 10 As shown, by projecting the density at the atom center onto the most efficient chemical bonds, the local system always achieves a significant reduction in the proportion (standard deviation) across various atom types, demonstrating its ability to eliminate unwanted geometrical variations caused by the flexible orientation of chemical bonds. More importantly, this method can also significantly reduce the proportion of various atom types (such as...) Figure 11 The gradient labels on (as shown) are particularly effective for the H atom, where all coefficient dimensions exhibit a reduction in gradient scale of >50%. This is crucial for mitigating large gradient ranges and making subsequent dimension-wise rescaling of modules easier (see Supplementary Section B.3.3 for further discussion). These two advantages of locality contribute to efficient optimization of KEDF models. Experimental results in Supplementary Section F.4.3 demonstrate that KEDF models can achieve a 2x reduction in training and testing errors using this technique.

[0174] B.3 Techniques for Efficient Training As mentioned in Method 4.3, the range of KEDF data is enormous, and conventional data normalization techniques are unsuitable for the task at hand because minimizing the gradient scale conflicts with minimizing the coefficient scale. For the same reason, the logarithm of the density input features cannot be used, as the gradient will be amplified accordingly. Having both large coefficients and a large gradient implies a function with large Lipschitz coefficients, which presents a significant challenge for optimizing common neural networks. It does not work to use a separate model to directly learn the gradient, as the energy model will easily overfit, and in density optimization, the gradient model unnecessarily reduces the energy from the energy model and even appears non-conservative. Another effective strategy for reducing the gradient scale is to reduce the energy scale. However, KEDF models equipped with this approach suffer from a loss of accuracy in energy prediction because increasing the energy scale or unit leads to a loss of resolution. Based on these observations, a series of detailed techniques are introduced to achieve efficient training.

[0175] B.3.1 Natural Reparameterization Direct measurement using Euclidean metrics can introduce unbalanced sensitivity into the output. Parameterization using metrics that reflect physical meaning is preferable. The coefficients represent density, and the natural measure of density is the L2 metric. Based on the basis expansion expression, this is then used to determine the density coefficients: Introduce a metric, where It is the overlap matrix of the density basis. Using this natural metric makes the scale of change of p reflect the scale of change of the represented density ρ, and the resulting density loss and gradient loss have better physical meaning and balanced sensitivity in the dimension of p.

[0176] Note that matrix W is independent of p; an alternative way to incorporate this matrix is ​​by reparameterizing the density: Let :=M p, where M satisfies MM =W. In this way, the Euclidean metric in the new parameters ( ) ( ′ ) = (p p ′ )MM (p p ′ ) = (p p ′ )W(p p ′ This restores the aforementioned natural metric. This reparameterization makes the implementation of density / gradient loss and natural gradient descent cheaper.

[0177] Select Matrix √M still faces the degree of freedom of orthogonal transformation. We choose M = √(QΛ), where the diagonal Λ and orthogonal Q give W = QΛQ. The eigenvalue decomposition (note that W is symmetric positive definite, hence such a decomposition exists) is given by √Λ, which is equivalent to the element-wise square root. It is found to outperform the use of the Choreski decomposition. The use of natural reparameterization significantly stabilizes the optimization of the KEDF model and results in lower training and validation losses, thus highlighting the necessity of balancing density coefficients with physical metrics; see Supplementary Section F.4.3 for more details.

[0178] B.3.2 Atomic Reference Model Standard gradient labels are large in scale, which poses a challenge for the gradients the model aims to capture. The goal is to subtract a gradient reference to center the gradient labels. The problem lies in the size and structure of the input p, meaning the gradient g depends on the chemical composition of the molecule (the number of atoms per element) and varies across molecules because different chemical elements have non-overlapping auxiliary bases. Therefore, for each element, the gradient mean is taken: corresponding to the average gradient dimension of the auxiliary bases for that element. The resulting element-wise gradient reference represents the T gradient of the auxiliary bases for that element. S The average response to changes in density coefficients. In this way, gradient references for molecules in structure M can be constructed based on the constructed tiling element gradient references. .

[0179] To match the new gradient, it should also be used Subtract structure Each energy label of the molecules in ,in This bias is a deviation that centers the energy labels of molecular transformations in structure M for easier learning. Faced with the same problem of dependency, therefore similarly, an energy bias is fitted for each element, and by... To construct by summing the biases of all atoms in . Add a global bias once for each molecule.

[0180] Regarding the standard deviation of the energy labels after normalization transformation, no further improvement was found. This is likely due to the trade-off between ease of learning and model resolution: errors at the (small)-scale normalization will be amplified at the original scale. In the KEDF model, after processing by local system and natural reparameterization techniques, the density coefficients are fed into the atomic reference model to provide an analytical reference for the final energy prediction. Figure 6 (a))

[0181] B.3.3 Rescaling by Dimension As described in Method 4.3, there is a trade-off when scaling the input coefficient labels: a medium (small) coefficient scale will increase the gradient scale. Under this consideration, in Equation (4), the target gradient scale... Set to 0.05 and maximum coefficient scale Set it to 50.

[0182] To illustrate the importance of rescaling by dimension, Figure 12The label scale for the kinetic functional gradient, rescaled kinetic functional gradient, and density coefficients of the QM9 dataset is shown. The energy functional was chosen as the residue energy between the non-interacting kinetic energy and the experimental kinetic functional APBE. In the context of deep learning, the Lipschitz coefficient (i.e., the maximum absolute gradient of a neural network model) is often used to indicate the model's ability to fit gradient labels. Following this convention, L... ∞ A metric is used to measure the label scale. It should also be noted that the gradient labels here have already been centralized by subtracting the mean gradient for each dimension from the atomic reference model. Therefore, the dimension-wise rescaling operation only needs to reduce the range of variation around the zero gradient mean. The combination of these two techniques can significantly improve the robustness of the optimization. For example, in Figure 12 As shown in (a), many gradient dimensions have very large scales. This leads to significant difficulties in the practical optimization of deep learning models. After rescaling by dimension, most dimensions can be rescaled to have the desired gradient scale, while the gradient scales of other dimensions remain within acceptable limits. Figure 12 (at point (b)). For example Figure 12 As shown in (c), the rescaled density coefficients also exhibit difficulty in handling the enormous range (up to 250) of neural networks, even with data normalization techniques. Therefore, a shrinkage function is introduced to compress the coefficients into a dense space, as mentioned in Supplementary Section B.1.2. In fact, KEDF models without dimensionality rescaling are very difficult to converge, and the loss curve is particularly volatile. This is because the smoothness of common neural network architectures limits the range of the output gradient of KEDF models. Applying dimensionality rescaling techniques can effectively alleviate this predicament and achieve efficient optimization.

[0183] Furthermore, atomic reference modeling techniques significantly facilitate the dimensional rescaling process. Figure 13 The use of an atomic reference model demonstrates that the variance / scale of gradient labels can be significantly reduced, especially for the dimensions corresponding to 's' atomic orbitals. Since 's' orbitals typically have large gradient averages but small variances, subtracting the gradient averages for these dimensions greatly simplifies dimensional rescaling.

[0184] Density feature preprocessing module Finally, local corpus techniques, natural reparameterization, and dimensional rescaling are combined into a density feature preprocessing module, such as ( Figure 6 As shown in (d), it transforms the density coefficient p into density characteristics. As a direct input to the Graphormer model, this makes the model easier to learn. This module eliminates unnecessary geometric variability in density features and facilitates efficient learning of a wide range of gradient labels per component. This preprocessing module is required only once for each molecular structure M during training or running density optimization. Therefore, although natural reparameterization requires three computational costs, it does not dominate the scaling of the total cost compared to iterative gradient computation of the energy term in equation (49) which includes a transformer-based KEDF model, establishing the O(N) of M-OFDFT. 2 Complexity.

[0185] B.4 Functional Variation In principle, any energy density functional (Equation (14)) in the decomposition formula of the ground state energy, except for the unknown KEDF, can be regarded as the training target of the proposed M-OFDFT framework. Here, two versions of the implementation scheme mainly used for M-OFDFT are first described, and exploration in other functional variants is described.

[0186] B.4.1 Residual KEDF In the implementation of M-OFDFT, the first functional version introduced aims to directly reduce the difficulty of modeling KEDF by using the existing KEDF as the basic functional and learning the residuals: T S ,θ (p , M) := T S,res ,θ (p) + T base(p , M). (67) Specifically, the APBE KEDF was chosen as the reference / baseline KEDF because it best fits the training data to the QM9 dataset, one of the four functionals mentioned in Supplementary Section D.2. Although von... Weizsäcker's KEDF provides a lower bound and can constrain the residual model to be positive, but using it as a baseline introduces large gradient labels on the residuals, making training difficult.

[0187] B.4.2 TXC Functional While the KEDF residual functional provides a simple learning objective with easily controllable gradient range, it has an unavoidable computational bottleneck. The gradients of the APBE KEDF reference and the PBE XC functional need to be backpropagated through the grid relative to the density coefficients. As discussed in Supplementary Section A.3.2, the memory requirement for evaluating the gradients at grid points is O(MN). grid Considering the large pre-factor N). grid (usually) 103 The computational cost of KEDF residuals in gradient descent becomes prohibitive for large-scale systems. This cost is evident when running KEDF residual functionals on the QMugs dataset. To completely eliminate computation on the grid, a second version is introduced. E TXC,θ (p , M) together study KEDF and XC energy (referred to as TXC functional) E TXC ) and.

[0188] B.4.3 Learning Other Functions In addition to the two versions mentioned above, other direct goals of machine learning model learning were explored, including directly learning KEDF. T S (i.e., without reference to KEDF), universal functionals U [ ρ (Equation (11)) and total energy functional E [ ρ [Equation (14)] Both direct training of the KEDF model and training of the general functional are difficult to optimize due to their large gradient ranges, which are even larger than the gradient ranges used to learn the residual KEDF model and the TXC model, and cannot be handled efficiently even with the proposed technique. Interestingly, although the learned functional will be inherently system-dependent, the total energy functional is learned. E θ [ ρ That is, to transfer external energy. E ext Adding the generalized functional as the objective significantly reduced the gradient range and achieved stable optimization, particularly on the protein systems under consideration. Specifically, for 1,000 Chignolin conformations, the learning... E θ [ ρ The model yields a higher learning efficiency through density optimization (MAE 0.071 kcal / mol). E TXC,θ The model yields better energy results. E TXC,θ The model still achieves reasonable results (MAE 0.102 kcal / mol), which is still much better than the classic KEDF.

[0189] C density optimization C.1 Gradient-based density optimization As described in Method 4.4, a trained KEDF model is used during density optimization. T S,θ (p ,M) minimizes the total electron energy relative to the density coefficient p. This process utilizes a gradient descent process instead of a fixed-point SCF iterative process because the optimality equations (27) or (56) for the density coefficient are difficult to convert to fixed-point iterations. The details of calculating the energy gradient are discussed in Supplementary Section A.3.2. Figure 14 Typical energy optimization curves for randomly selected molecules from the QM9 dataset are shown, along with curves for gradient scaling (measured by the L2 metric). The two curves for the two initial densities converge stably and closely approximate the true total energy.

[0190] C.2 Stop Criteria In principle, the solution of M-OFDFT can be obtained at the rest point (zero projected gradient). However, due to discretization errors, the exact zero gradient point may be missed. Therefore, a set of stopping criteria is proposed to determine the actual rest point, such as... Figure 14 As shown. For intrascale molecules such as the QM9 and ethanol datasets, GlobEngUpd It always indicates a satisfactory resting point and exhibits optimal performance on a set of evaluated molecules. Therefore, it is used as the standard criterion for molecules at the intrascale. However, for large-scale systems, this criterion is unsuitable because unavoidable extrapolation errors can cause the model to fall into unreliable density regions. Therefore, a more conservative criterion is considered. 1stLoGradNorm or 1stLocEngUpd Then, if the first two criteria do not exist, then GlobEngUpd The guidelines will be adopted.

[0191] C.3 Density Initialization To perform projection MINAO initialization for the density coefficient, an additional output branch is added. p θ (p,M) is constructed to predict the difference between a given density coefficient p and the ground-state density coefficient p for a given cellular structure M. The differences between them. Figure 6 (a) shows a branch built on top of the last G3D layer of the KEDF model. It processes the hidden representation of each atom through an MLP module to predict the corresponding coefficient difference. The branch is trained to iteratively project the density coefficient data along the SCF of the KSDFT onto the corresponding ground state coefficients: (68) The fitted MINAO or Hückel density coefficients can be obtained by applying density fitting to the original MINAO or Hückel initialized orbitals (Supplementary Section A.4.1). This process is performed only once for density initialization, so the tertiary computational cost of density fitting (and generating the Hückel initialized orbitals) does not dominate the secondary scaling of density optimization on conventional workloads. To maintain secondary scaling for very large molecules, an alternative initialization has also been proposed: generating density coefficients directly by superimposing the densities of isolated atoms using the same concept as MINAO with superpositions of isolated atomic orbitals. This is equivalent to fitting the MINAO density coefficients connected to each atom as if they were isolated, with only linear computational cost. In practice, this technique significantly reduces the time required to construct initial coefficients, further reducing the total time consumption of M-OFDFT, and only results in a slight increase in error (approximately 0.02 kcal / mol in energy per atom on the protein system of interest).

[0192] D Overall Experimental Setup D.1 Data Preparation KSDFT calculation All KSDFT calculations were performed using the PySCF package, which is open-source and has a convenient Python interface that allows for flexible customization, such as for implementing the data generation process detailed in Supplementary Section A.4. The M-OFDFT calculations described in Results 2.1 were implemented using an automatic differentiation wrapper in Python. All KSDFT calculations were performed at the PBE / 6-31G(2df,p) level. To accelerate computation, density fitting with the def2-universal-jfit basis set was implemented in calculations involving molecules with more than 30 atoms. The grid level was set to 2. The convergence tolerance was set to 1 meV. Minor modifications were made to the PySCF code to record information required for intermediate SCF steps (e.g., molecular orbitals). Direct inversion in the iterative subspace (DIIS) was enabled after the default settings. For benchmarking purposes, per-molecule statistics such as run time and number of atoms were collected along with the calculations.

[0193] initialization MINAO initialization is used for KSDFT computation, adhering to the default settings. Consider using Hückel initialization to run the KSDFT for data generation, as it is based on the eigenvalue problem solution of the simple Hamiltonian matrix, making it more similar to the in-distribution density of the ground state density. However, compared to MINAO initialization, the SCF results obtained from Hückel initialization are more diverse and difficult to optimize, leading to a significant performance degradation in density optimization.

[0194] Hardware details All KSDFT computations and preprocessing (e.g., density fitting) are performed on a cluster of 700 capable CPU servers on Azure. Each server in the cluster has 256 GiB of memory and 32 Intel Xeon Platinum 8272 CL cores, with hyper-threading disabled.

[0195] Dataset preprocessing Following the SCF run, an additional process is performed to transform the SCF results into a dataset for training and evaluation. Density coefficients are obtained using the density fitting procedure detailed in Supplementary Section A.4.1. The auxiliary basis sets are generated during the run and have… β The isothermal basis set is 2.5. Subsequently, the density coefficients obtained after Supplementary Sections A.4.3 and D.4 are used to calculate the gradient and force labels. Energy labels are calculated after Supplementary Section A.4.2. For large orbital integrals in the process, appropriate symmetries in the atomic orbital indexes are utilized to save computation and reduce memory usage. Local system transformations and natural reparameterization (detailed in Supplementary Sections B.2 and B.3.1) are applied to the coefficients to obtain the model input.

[0196] D.2 Classical Baseline To evaluate the advantages of the M-OFDFT, several classical kinetic energy density functionals were selected as baselines. The Thomas-Fermi (TF) KEDF is accurate in uniform electron gas confinement. Its form is... Feng Weizsäcker KEDF takes the following form , This is exact for both bosons and individual fermions. When the KEDF is expanded up to the first derivative of the density, the result is... Also consider the correction T TF [ρ]+T vW Variants of [ρ]. APBE reference KEDF is also considered. For fair comparison, these functionals are implemented within the same OFDFT framework as M-OFDFT. Gradient descent with the same stopping criterion is used for density optimization. Kinetic energy values ​​are evaluated on the grid. For initialization, the MINAO method, which is superior to the Hückel method, is used.

[0197] We also attempted to use other state-of-the-art OFDFT software as baselines, such as PROCESS, GPAW, ATLAS, and DFTpy. However, most of these are tailored for periodic material systems and heavily rely on complex pseudopotentials. Adapting these codes to the molecular systems of interest is challenging because they only provide local pseudopotentials for a few main group elements with no covered elements (HCNOF) and encounter difficulties in regenerating reliable pseudopotentials for these elements. Although GPAW can handle molecular systems, it was found that its OFDFT calculations for molecular systems are difficult to converge, even for small molecules. Therefore, only classical KEDF was implemented as a baseline in this framework.

[0198] D.3 M-PES baseline Results 2.4 demonstrate that M-OFDFT possesses the ability to extrapolate to macromolecular systems, a valuable advantage for quantum chemical methods. The significance of this advantage is measured by comparing M-OFDFT with two machine learning-based alternatives, M-PES and M-PES-Den, which employ the same Graphormer backbone architecture as M-OFDFT but lack gradient prediction and density correction prediction branches. Furthermore, M-PES does not utilize a node embedding layer incorporating density features. Figure 6 (b)). The model parameters were also adjusted to be comparable to those of the M-OFDFT model.

[0199] D.4 Hermann-Feynman Force Calculation To more comprehensively evaluate M-OFDFT, the Hermann-Feynman (HF) force is also used as a metric for evaluating the results of M-OFDFT. The forces experienced by atoms in a molecule are themselves important dimensions because they are directly required for geometry optimization and molecular dynamics simulations. The HF force is a simple way to estimate forces and is accurate under the constraint of complete basis functions. Due to its simplicity, the HF force can also be considered as an evaluation of the optimization density.

[0200] Herman Feynman Formally, force is the negative gradient of the total energy of the molecule, represented by atomic coordinates X=(x (1) ,··,x (A) The function E) tot (X), representing molecular conformation. For simplicity, the dependence on molecular composition Z is omitted. Total energy Including the energy of electrons in the ground state (Including interactions with the nucleus), and the energy resulting from internuclear interactions. It provides the internuclear part of the force. (69) Following the variational optimization process used to solve for the electronic ground state of the molecule in conformation X, the electron energy... It is the minimum value. In its most basic form, The Hamiltonian operator is determined by the variational problem on the N-electron wavefunction shown in equation (5). ,pass V ext,X It depends on the conformation X given in equation (6). Ground state wavefunction and energy Therefore, it also depends on X. Then it can be... The gradient renormalization is:

[0201] The equation (*) is due to It is a Hermitian operator Having real eigenvalues The eigenstates of X, and equation (#) is due to the wave function being normalized for all X. Equation (70) is the Herman-Feynman (HF) theorem. To continue the calculation, the gradient of the Hamiltonian operator in the expression can be derived as follows: And through attention It is multiplicative and consists of single units, such as equation (6). As shown, and subsequently, the electronic force can be calculated as:

[0202] In equation (7), the following is defined: This is the Hermann-Feynman (HF) force f X This expression is consistent with the electrostatic force in the classical view, indicating that "there is no mysterious quantum-mechanical force at play in the molecule." According to the expression, assessing the HF force only requires considering the ground-state electron density. A good approximation, which is available in various quantum chemical methods, including the KSDFT given by equation (32). And the M-OFDFT given by equation (29) The total force on the nucleus is the sum of the internuclear forces shown in equation (69).

[0203] Analytical ability However, when using atomic bases, approximation errors occur in the functional representation because the bases are incomplete. Therefore, the conditions of the HF theorem do not hold precisely. This makes the HF force in equation (71) only approximate the real electronic force. Furthermore, other approximations are possible. For example, it is also possible to directly employ the electron energy E, expressed with optimal coefficients on the atomic basis. X To estimate the gradient analysis Therefore, this method of estimating the electronic force is called analytical force f. ana For KSDFT, E X (C) is given by equations (38-42). (Note the basis functions) η α , η β Matrix T,D ˜ and V ext Both depend on X), and C X These are the optimal orbital coefficients. The corresponding analytical force on atom a can be expanded as:

[0204] in( x ( a) E X Take a gradient with a fixed C, and It is Jacobi , where ξ∈{1,2,3} indexes the three spatial components, and matrix operations (including transpose, matrix multiplication, and tracing) operate on indices α and i. When When accurately optimized, it satisfies the self-consistency in equation (47) and the orthogonality in equation (43). Also note that in equation (46), the second term in equation (72) disappears:

[0205]

[0206]

[0207] If there is actually no precise optimization Then we should also consider the source . contributions.

[0208] Note that only item( It is the optimal density matrix for flattening. (vector), as A portion of this corresponds to the HF force (see equation (42)). The analytical force in equation (72) The other terms in the equation are collectively referred to as driving forces. They appear as gradients of the basis functions relative to atomic coordinates, and are therefore non-zero when using atomic bases, thus distinguishing analytical forces from HF forces.

[0209] Implementation Although analytical forces are considered more accurate estimates when using atomic bases, the HF force is still used to evaluate the results because the way analytical forces are calculated differs for different methods: KSDFT and M-OFDFT require different types of Coulomb integrals under different basis sets, while M-PES and M-PES-Den only require backpropagation of gradients through a machine learning model. Furthermore, the HF force in equation (71) depends only on the density from the ground-state solution, so it can also be regarded as a relevant metric for evaluating the solution density.

[0210] For each molecular system, the true values ​​of the HF forces are considered to be those given by the KSDFT solution, where the density is obtained after density fitting (see Supplementary Section A.4.1). The HF forces via M-OFDFT are calculated directly from the optimized density. For M-PES and M-PES-Den, since they do not provide the density, only the corresponding analytical forces are available. and It is available. To convert them into HF forces, extract them from them using the Prey forces calculated from KSDFT: And for Similarly, where calculations are made from KSDFT and The error in predicting the force is measured by the mean absolute error (MAE) over each of the three spatial components of the force on each atom in each molecule of the test group.

[0211] D.5 Curve Fitting To fit Figure 3 The curves in the figure, the formula for all curves in each figure (aN) heavy +b) c +d is restricted to having the same scale, i.e., the same a. Otherwise, the flexibility of a would reduce the exponent c in the provided N. heavy The curvature of the curve reflects the range (i.e., a smaller a allows for a larger exponent c). Sharing the same a across different curves hardly hinders the proximity of the fitted data points. Based on this idea, a two-stage fitting strategy is proposed: (1) jointly fit all curves with the shared a to obtain a pre-optimized a. ′ (2) Take a ′ As an initial guess, all curves are independently refitted using the Trust Region Reflection Optimization Algorithm to obtain the post-optimized a. The search space for variable a is restricted to [a... ′ (1 ξ ) , a ′ (1 + ξ )], where ξ=0.5. The two-stage fitting strategy approximately preserves all post-optimized a This achieves the same scale and better fitting accuracy. The fitting process is performed using the SciPy package in Python.

[0212] E Implementation Details E.1 dataset To demonstrate the efficacy of the proposed method, two distinct molecular datasets were evaluated: ethanol and QM9. These datasets were specifically chosen to assess the generalization performance of M-OFDFT in both conformational and chemical spaces. Furthermore, to evaluate the extrapolation capability of M-OFDFT with larger molecular systems, two additional datasets were prepared: QMugs and Chignolin. Comprehensive supplementary information and generation details for each dataset are provided below.

[0213] E.1,1-ethanol To investigate the generalization ability of M-OFDFT in the image space, a set of 100,000 non-equilibrium ethanol geometries was randomly selected from the MD17 dataset. These geometries were then randomly divided into training, validation, and test sets using an 8:1:1 ratio.

[0214] E.1.2 QM9 The QM9 dataset is used as a popular benchmark for predicting quantum chemical properties using deep learning methods. It comprises equilibrium geometries of approximately 134k small organic molecules composed of H, C, O, N, and F atoms. These molecules represent subsets of all species with up to nine heavy atoms from the GDB-17 general chemical dataset. Because the dataset contains only equilibrium geometries, it is used as a benchmark for evaluating the generalization ability of M-OFDFT in chemical space. Furthermore, to compare M-OFDFT with the classical KEDF, C7H from the QM9 dataset is included. 10 The 6905 isomers of O2 were used as a benchmark for evaluating M-OFDFT in conformational space. Molecules were then randomly divided into training, validation, and test sets using an 8:1:1 ratio.

[0215] E.1.3QMugs The QMugs dataset, proposed by Isert et al., includes over 665k biological and pharmacologically relevant molecules extracted from the ChEMBL database, totaling approximately 2M conformational isomers. Notably, QMugs provides a much larger sample size of molecules than the QM9 and MD17 datasets, with each compound having an average of 30.6 and up to 100 heavy atoms. This allows us to investigate the extrapolation capabilities of M-OFDFT.

[0216] We categorize QMugs molecules into bins of width 5 based on the number of heavy atoms, except for the first bin, which contains molecules with 10–15 heavy atoms. To construct the extrapolation task, the union of QM9 and the first bin is split into training and validation sets at a 9:1 ratio, and M-OFDFT is tested on 50 molecular structures from each of the other bins using an increased scale. To reduce the cost associated with expensive DFT computations, only 50 molecular geometries are sampled for each test bin.

[0217] Furthermore, to investigate the amount of additional large-scale molecular data required for end-to-end counterparts to achieve extrapolation performance comparable to M-OFDFT, a series of training sets with progressively increasing molecular scales were constructed. Specifically, a subset was initially sampled from each of the first four bins of the QMugs molecule, and these subsets, along with the entire QM9 dataset, were assembled into four training sets. These training sets used a finely synthesized ratio designed to maintain the total size of all training sets while reflecting the ratios of each bin in the original QMugs dataset. The composition ratios of different bins were... Figure 15 As shown in the image.

[0218] E.1.4Chignolin Data Selection To further evaluate the extrapolation capability of M-OFDFT for large-scale biomolecules (e.g., proteins), 1,000 Chignolin conformations were sampled from extensive molecular dynamics (MD) simulations. Chignolin has been widely used in molecular dynamics studies due to its rapid folding properties and short length (TYR-TYR-ASP-PROGLU-THR-GLY-THR-TRP-TYR). First, the backbone conformation was characterized by the distance between all paired α-carbon atoms, excluding nearest-neighbor residue pairs. Then, the distance was calculated using τ... lag Time-Lapse Independent Component Analysis (TICA) was performed with a lag time of 20 ns, projecting the conformation space into a low-dimensional subspace. The first eight TICA components were clustered into 1,000 groups using the k-means algorithm. The centroid structure of each cluster was extracted using MDTraj and used as a representative structure.

[0219] Conformational neutralization This study focuses on isolated molecular systems where all atoms or functional groups remain neutral and all electrons are paired. However, protein molecules extracted from MD trajectories often contain numerous ionizable groups, such as amino and carboxyl groups in N-terminal amino acids. Several strategies have been proposed to neutralize ionizable groups in protein conformations, including adding neutral terminal caps (e.g., n-methyl and methyl acetate) and binding counterions. However, introducing additional terminal groups or ions into the extracted protein conformation can lead to significant geometric changes in the original conformation and may result in instability. To avoid this, protein conformations are neutralized by editing (adding / removing) missing / extra hydrogen atoms in ionizable groups using OpenMM. The hydrogen definition template is modified to handle hydrogen atoms not defined in standard amino acids, such as hydrogen atoms missing from the carboxyl group in C-terminal amino acids. The neutralized conformation is then locally energy-optimized using Amber with an Amber-ff14SB force field and up to 100 optimization cycles. To prevent significant geometric changes, optimization is only allowed for hydrogen atoms associated with heavy atoms involved in the hydrogen editing process. For neutralized amino acids for which there is no standard force field template, the force field parameters were constructed using the antechamber and parmchk2 tools in Amber.

[0220] Protein fragmentation To investigate the extrapolation utility of M-OFDFT from “local” fragments to “global” proteins, a dataset was generated by dicing 1,000 Chignolin conformations into protein fragments (i.e., peptides) of varying lengths. To construct these fragments, all possible peptides with sequence lengths up to 5 in the Chignolin sequence were enumerated. Specifically, residue gaps were allowed in tetrapeptides and pentapeptides, i.e., dipeptide-[gap]-dipeptide and tripeptide-[gap]-dipeptide. The extracted fragment molecules were neutralized and then used as the training and validation sets with a 9:1 ratio. It should be noted that when benchmarking the performance of M-OFDFT models trained at different molecular scales, the largest peptide length L was used. pep The training set consists of values ​​less than L. pep The sequence length of all polypeptides.

[0221] Data Filtering Preliminary analysis revealed that the range of energy and gradient labels in the fragment dataset was too large to fit, even with several efficient training techniques (see Supplementary Section B.3). Furthermore, the vast majority of data points originated from the first few steps of the SCF iterations, far from the convergence state. To mitigate this issue, data points with an SCF-enhanced energy level exceeding 500 kcal / mol of residual energy were filtered out. This strategy experimentally improves optimization robustness and prediction performance.

[0222] E.2 Model Configuration In all experimental settings, the same skeleton structure (Graphormer) was used, specifically utilizing the Graphormer-3D (G3D) encoder layer. To ensure fair comparisons, most hyperparameters (e.g., model depth and hidden dimensions) in all models were consistently maintained. A summary of all hyperparameter choices can be found in Supplementary Table 3. Notably, the learnable parameter v is located within the shrinking gate of the node embedding layer. co The initial value of is set to 10, while v mul The initial values ​​were set to 0.02 for the ethanol dataset and 0.05 for all other datasets. No extensive hyperparameter search was performed, and most hyperparameters were selected using Graphormer.

[0223] Table 3: Hyperparameters of the Graphormer model used in all methods.

[0224]

[0225] E.3 Training Hyperparameters Following Graphormer, all models were trained using the Adam optimizer and linearly decaying learning schedule. For the functional models, the peak learning rate was set to 1×10⁻⁶. -4 For the baseline models (i.e., M-PES and M-PES-Den), the peak learning rate was set to 3 × 10⁻⁶. -4 Other optimizer hyperparameters can be found in Supplementary Table 3. A warm-up phase, representing a linear increase in the learning rate, was introduced to stabilize training during the initial phase. The number of warm-up steps was set to 30k for M-OFDFT and 60k for the baseline model. The batch size was set to 256 for the ethanol dataset and 128 for all other datasets. All models were trained on an Nvidia Tesla V100 GPU.

[0226] For different molecular datasets, the number of training epochs was determined by examining the loss curves. Training was stopped once the loss failed to decrease during 20 epochs. Specifically, the functional models were trained for approximately 600 and 700 epochs on the ethanol and QM9 datasets, respectively. For the QMugs dataset, the model was trained for approximately 700 epochs, while the M-PES and M-PESDen models were trained for 2,300 epochs. Notably, Results 2.4 presents a series of QMugs datasets with increasing molecule sizes, where the number of SCF data points increases slightly with the average molecule size (larger molecules typically require more SCF iterations to converge). To ensure fair comparisons, approximately the same number of training epochs were maintained for the different datasets. In the Chignolin experiments, four Chignolin datasets with increased peptide lengths and data sizes were used. The functional models were trained for 1,400, 1,100, 800 and 750 epochs, respectively, while the M-PES and M-PES-Den models were trained for 3,000, 3,000, 1,000 and 900 epochs, respectively.

[0227] In practice, a weighted loss function is used to optimize the KEDF model: L = ξ eng L eng + ξ grad L grad + ξ den L den , Among them, L eng L grad and L den Let ξ represent the KEDF energy loss (Equation (2)), gradient loss (Equation (3)), and projected density loss (Equation (68)). eng ξ grad and ξ den These are the corresponding loss weights. The selection of loss weights was determined by independently performing a grid search on the validation set for various datasets, and the loss weights used are provided in Supplementary Table 4.

[0228] Table 4: Loss weights for various datasets and learning objectives.

[0229]

[0230] E.4 Density Optimization Hyperparameters During the deployment phase, the gradient descent step and step size ε were selected based on the performance of the evaluation dataset. For all molecular datasets and functional variants (including the traditional KEDF baseline), the gradient descent step and step size were set to 1000 and 1x10, respectively. -3 In addition to the KEDF model that is learned on the ethanol dataset, T S,res,θ and E TXC,θ In addition, the step size ε is set to 5×10 -4 The density optimization process employs a stochastic gradient descent (SGD) optimizer.

[0231] Additional Results F.1 Additional visualization of optimized density To further investigate the effectiveness of the optimized density generated by M-OFDFT, in Figure 16 The integrated density is concentrated on two other heavy atoms in the ethanol molecule. The results show that the density of M-OFDFT is closely aligned with the KSDFT density of the two atoms. In contrast, the classical KEDF of APBE struggles to accurately recover the shell structure, particularly in the secondary peaks associated with covalent bonds. In these regions, the APBE density exhibits a significant deviation from the true KSDFT density.

[0232] Table 5 shows the performance of M-OFDFT with different learning objectives and initialization configurations. The mean absolute error (MAE) of energy and Hermann-Feynman (HF) force is listed in kcal / mol and kcal / mol / A, respectively.

[0233] Table 5

[0234]

[0235] F.2 More results at different scales As explained in Method 4.4, two initialization strategies are proposed to address the out-of-distribution problem of initial density. Figure 2 The results in (a)-(b) highlight the superior performance of projective MINAO initialization. Notably, as shown in Supplementary Table 5, even without a shortcut to machine learning, M-OFDFT still achieves similar orders of magnitude chemical accuracy when using standard Hückel initialization. Because the density corrector provides a starting point close to the ground state, where the model more confidently approximates the functional, projective MINAO initialization yields better ground-state density and significantly lower HF force loss. This demonstrates the importance of incorporating the density corrector's predictive branch.

[0236] F.3 More extrapolation results As previously mentioned, the extrapolation capability of M-OFDFT was evaluated on two datasets, QMugs and Chignolin, using two end-to-end counterparts of M-PES and M-PES-Den with the same architecture as the baseline. It is important to note that while M-PES and M-PES-Den achieve better validation errors, this does not necessarily imply superior extrapolation performance for larger-scale molecules. Unlike M-OFDFT, the M-PES method relies solely on molecular structure to model the total energy, neglecting fine-grained interaction processes between electrons. Therefore, the end-to-end mapping from M to E appears highly complex and nuanced, making generalization to molecular systems with unseen structural patterns challenging. This observation is supported by a sharp increase in the generalization error of the M-PES and M-PES-Den models in energy predictions, as shown in Supplementary Table 6.

[0237] Furthermore, M-OFDFT also demonstrates excellent extrapolation performance in HF force prediction. For example... Figure 17 As shown in (a), the M-OFDFT model trained on molecules with fewer than 15 heavy atoms consistently maintains a lower HF force MAE than M-PES and M-PES-Den on larger-scale molecules. While M-PES-Den has a relatively low error increase exponent, M-OFDFT achieves a significantly lower absolute MAE. Furthermore, the disclosed method outperforms two end-to-end counterparts across various fragment datasets in Chignolin experiments.

[0238] Table 6 shows the generalization results of the M-PES and M-PES-Den baselines on the QMugs and Chignolin datasets. For QMugs, both the M-PES and M-PES-Den models were trained on molecules with fewer than 15 heavy atoms and tested on 50 QMugs conformations with 56–60 heavy atoms. For Chignolin, both models were trained on fragment conformations with all peptide lengths (2–5) and tested on 1,000 Chignolin conformations. The mean absolute error per atom (MAE) of the training set, validation set, and density-optimized energy from the large-scale system (represented as energy prediction) is listed in kcal / mol.

[0239] Table 6

[0240]

[0241] F.4 Ablation Study F.4.1 Multi-step data and gradient labels To capture the energy landscape in the density coefficient space, multiple p samples are generated for each molecular structure M, and gradient labels are calculated for additional supervision information. p T S Ablation experiments were conducted on the ethanol dataset to investigate the importance of each type of supervision by removing supervision labels during the training process. The ablation results are presented in Supplementary Table 7. Gradient label removal. p T S This leads to a significant decrease in the energy and HF force accuracy of both density initialization strategies, highlighting the importance of maintaining the optimization of M-OFDFT on the physical orbit. Incorporating multi-step p-samples is also crucial for enhancing the performance of M-OFDFT, especially for projective MINAO initialization, as the accuracy of the density corrector model depends on the magnitude of the training density.

[0242] Table 7 shows ablation studies for various data augmentation strategies. Using learned data... T S,res,θ The model evaluates all results on the test of ethanol molecules. MAE in energy and HF force is listed in kcal / mol and kcal / mol / A, respectively.

[0243] Table 7

[0244] F.4.2 Nonlocality Given the nonlocal dependence of KEDF on electron density, it is important to incorporate nonlocal computation into OFDFT, which inspires our use of Graphormer as the backbone architecture. To demonstrate the importance of nonlocality, ablation is performed by gradually decreasing the receiver field (i.e., the distance cutoff of the atom neighborhood) of the M-OFDFT model. Figure 18 As shown, the accuracy of both energy and HF force predictions tends to deteriorate as the distance cutoff decreases, thus providing experimental evidence for the importance of incorporating nonlocality into the KEDF model.

[0245] F.4.3 Density Feature Preprocessing Module The geometric variability of the input data and the large range of gradient data make optimizing the KEDF objective challenging. Therefore, a density feature preprocessing module is proposed, which implements several techniques to achieve efficient optimization. Ablation experiments are conducted to investigate the importance of local systems and natural reparameterization techniques. As shown in Supplementary Table 8, incorporating local systems leads to a considerable performance improvement compared to the baseline model without these techniques, demonstrating the ability of local systems to reduce the geometric variability of the input data.

[0246] Furthermore, since different basis functions have varying importance in representing density, some coefficient dimensions can significantly influence the density function and total energy. This is challenging for M-OFDFT because the model treats each dimension equally. During density optimization, this effect is amplified when the input density deviates from the ground state. This problem is addressed by introducing natural reparameterization to balance the influence of each coefficient dimension. The results in Supplementary Table 8 demonstrate that achieving natural reparameterization significantly improves model performance, especially for Hückel initializations far from the ground state density. Here, the energy MAE is reduced by one or two orders of magnitude, highlighting the powerful ability of natural reparameterization to balance sensitivity across coefficient dimensions.

[0247] We also attempted to examine the importance of the atomic reference model and the dimensionality rescaling technique, but found it difficult to optimize the learning objective when either technique was removed from the model. Therefore, unreasonable evaluation metrics are omitted in Supplementary Table 8. As illustrated in Supplementary Section B.3.3, the significant impact on effective training and the quantitative results for reducing geometric variability strongly support the necessity of both techniques.

[0248] Table 8 shows the results of the ablation study of the components in the density pretreatment module. (Using learned...) T S,res,θ The model evaluates all results on the test of ethanol molecules. MAE in energy and HF force is listed in kcal / mol and kcal / mol / A, respectively.

[0249] Table 8

[0250] F.4.4 Results using other training strategies To provide the model with information about the optimization landscape, the training data is enriched using gradient labels and labels corresponding to multiple densities of the same molecular structure. While alternative methods exist for supervising the optimization behavior of functional models, experiments have been conducted to indicate that they are not very effective with respect to the observed settings. In the context of learning XC functionals, Kirkpatrick et al. introduced SCF loss to regularize the target landscape, but it actually only acts as gradient supervision in the ground state (convergence orbit). In contrast, M-OFDFT utilizes multi-step data and gradient supervision to capture a more refined energy landscape, resulting in a significant performance improvement (Supplementary Table 7).

[0251] Some other studies implicitly implement energy landscape optimization by supervising the energy and density (or orbits) of the model. This makes the model optimization process quite complex. To address this challenge, some works resort to inefficient gradient-free optimization methods or backpropagation of loss through a series of eigenvalue solvers, which is both expensive and difficult to generalize. Specifically, the method of Li et al. was attempted, which involves supervising the energy and density of the model for optimization. Optimizing the model is extremely challenging because it requires backpropagating parameter gradients through a series of density optimization steps. Even when training E TXC Even after completely eliminating the grid to save costs and pre-training the model on the data to simplify optimization, the method still cannot optimize the model with more than 8 density optimization steps. Experiments on QM9 showed that the best results could provide a reasonable solution after the optimization steps used in training. However, if optimization continues, the energy does not converge, indicating that the model does not learn a physical functional (non-N-representable) but instead behaves like an end-to-end predictor by stacking K times, where K is the specified number of optimization steps in training.

[0252] Chen et al. attempted to bypass expensive backpropagation through iterative density optimization by assuming the model's optimal state is accurate. Specifically, the method uses the best density from the current model paired with the true ground-state energy label as the data pair for the next step. Since finding the minimum density after each model update is expensive, the method chooses to optimize the model and density alternately. However, this modification breaks properties, and the optimization process may not perform as expected. Furthermore, this method still focuses primarily on landscapes close to the ground state, which may not effectively guide the initial density. In the inventors' experiments on Chignolin, the alternative optimization was repeated four times, but only marginal performance improvements were observed. Specifically, the performance gains for each iteration were 0.5%, -0.06%, 1.7%, and -0.8%, respectively, meaning this strategy is ineffective in capturing energy landscapes.

[0253] F.5 Scalability Study of Large Protein Systems With the aid of sophisticated atomic orbitals and numerical optimization, KSDFT can be extended to molecular systems of up to approximately 500 atoms using modern PC clusters. To further demonstrate the scaling advantage of M-OFDFT, the published model was tested on two large-scale protein systems rarely encountered in the context of KSDFT computation (Results 2.3). Considering the considerable scale gap between the training and target molecules, the M-OFDFT model was trained on all Chignolin fragments and fine-tuned on the 800 Chignolin conformation as an alternative KEDF model, which theoretically provides the most robust extrapolation capability. The results show that M-OFDFT achieves speedups of 25.6x and 27.4x, respectively, on the two test systems compared to KSDFT computation. The per-atom energy (MAE) of M-OFDFT for the two test systems are 0.23 kcal / mol and 0.31 kcal / mol, respectively. Although higher than the prediction error for Chignolin molecules, the results still show a significant advantage over M-PES (0.36 kcal / mol and 0.63 kcal / mol, respectively).

[0254] In some implementations, the methods and processes described herein can be attached to a computing system of one or more computing devices. In particular, such methods and processes can be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.

[0255] Figure 19 A non-limiting embodiment of a computing system 1000 capable of performing one or more of the methods and processes described above is schematically illustrated. The computing system 1000 is shown in a simplified form. The computing system 1000 can embody the above description and... Figure 1B and Figure 8 The computing system 1000 is shown in the figure. The components of the computing system 1000 may be included in one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video game devices, mobile computing devices, mobile communication devices (e.g., smartphones) and / or other computing devices, as well as wearable computing devices (such as smartwatches and head-mounted augmented reality devices).

[0256] The computing system 1000 includes processing circuitry 1002, volatile memory 1004, and non-volatile storage device 1006. The computing system 1000 may optionally include a display subsystem 1008, an input subsystem 1010, a communication subsystem 1012, and / or... Figure 19 Other components not shown.

[0257] Processing circuitry typically includes one or more logic processors, which are physical devices configured to execute instructions. For example, a logic processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve technical effects, or otherwise achieve desired results.

[0258] The logical processor may include one or more physical processors configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processor of the processing circuit 1002 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Components of the processing circuit may optionally be distributed across two or more separate devices that may be remotely located and / or configured for coordinated processing. For example, aspects of the computing system disclosed herein may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In this case, it will be understood that these virtualized aspects operate on different physical logical processors on various different machines. These different physical logical processors on different machines will be understood to be collectively contained within the processing circuit 1002.

[0259] The non-volatile storage device 1006 includes one or more physical devices configured to store instructions executable by processing circuitry to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 1006 can be transformed—for example, to retain different data.

[0260] The non-volatile storage device 1006 may include removable and / or built-in physical devices. The non-volatile storage device 1006 may include optical memory, semiconductor memory, and / or magnetic memory, or other mass storage device technologies. The non-volatile storage device 1006 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that the non-volatile storage device 1006 is configured to retain instructions even when the non-volatile storage device 1006 is powered off.

[0261] Volatile memory 1004 may include a physical device including random access memory. Volatile memory 1004 is typically used by processing circuitry 1002 to temporarily store information during the processing of software instructions. It should be understood that when volatile memory 1004 is powered off, volatile memory 1004 typically does not continue storing instructions.

[0262] The processing circuitry 1002, the volatile memory 1004, and the non-volatile storage device 1006 can be integrated together into one or more hardware logic components. For example, such hardware logic components may include field-programmable gate arrays (FPGAs), application-specific integrated circuits (PASICs / ASICs), application-specific standard products (PSSPs / ASSPs), system-on-a-chip (SoCs), and complex programmable logic devices (CPLDs).

[0263] The terms "module," "program," and "engine" can be used to describe aspects of a computing system 1000 typically implemented in software by a processor to perform specific functions using portions of volatile memory, involving transformation processing specifically configured for the processor to perform these functions. Therefore, a module, program, or engine can be instantiated by executing instructions held by non-volatile storage device 1006 using portions of volatile memory 1004 via processing circuitry 1002. It should be understood that different modules, programs, and / or engines can be instantiated from the same applications, services, code blocks, objects, libraries, routines, APIs, functions, etc. Similarly, the same module, program, and / or engine can be instantiated from different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module," "program," and "engine" can encompass single or grouped executable files, data files, libraries, drivers, scripts, database records, etc.

[0264] When included, the display subsystem 1008 can be used to present a visual representation of the data stored by the non-volatile storage device 1006. The visual representation may take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data held by the non-volatile storage device, and thus transform the state of the non-volatile storage device, the state of the display subsystem 1008 can also be transformed to visually represent changes in the underlying data. The display subsystem 1008 may include one or more display devices utilizing virtually any type of technology. Such a display device may be combined with the processing circuitry 1002, the volatile memory 1004, and / or the non-volatile storage device 1006 in a shared housing, or such a display device may be a peripheral display device.

[0265] When included, the input subsystem 1010 may include or interface with one or more user input devices, such as a keyboard, mouse, touchscreen, camera, or microphone.

[0266] When included, the communication subsystem 1012 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 1012 may include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via wired or wireless local area networks or wide area networks, broadband cellular networks, etc. In some embodiments, the communication subsystem may allow the computing system 1000 to send messages to and / or receive messages from other devices via a network such as the Internet.

[0267] As used herein, “and / or” is defined as inclusive or ∨, as specified by the following truth table:

[0268] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these particular implementations or examples should not be considered limiting, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Therefore, the various actions shown and / or described may be performed in the order shown and / or described, in a different order, in parallel, or omitted. Similarly, the order of the above processes may be changed.

[0269] The subject matter of this disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems and configurations disclosed herein, as well as any and all equivalents thereof.

Claims

1. A computing system, comprising: A processing circuit and an associated non-volatile storage device, the non-volatile storage device storing instructions that, when executed by the processing circuit, cause the processing circuit to: In reasoning: Receive inference-time molecular structure data for inference-time molecular structure, the inference-time molecular structure data including inference-time atomic basis density expansion coefficients and data indicating the atomic properties of each atom in the inference-time molecular structure; The inference-time molecular structure data is input into a trained kinetic energy density functional machine learning model to generate predicted values ​​for the electron non-interaction kinetic energy for the density function. The trained machine learning model has been trained on a training dataset that includes atomic properties of each atom in the training-time molecular structure as input, and atomic-based density expansion coefficients for the training-time molecular structure. These atomic-based density expansion coefficients indicate the values ​​of the electron charge density characterized at the atomic basis and include the corresponding true values ​​for the non-interaction kinetic energy. as well as Output the predicted value of the non-interacting kinetic energy of the molecular structure for the reasoning process.

2. The computing system according to claim 1, wherein the atomic properties include one or more of atomic position and / or atomic number.

3. The computational system according to claim 1, wherein the density function is based on atomic-based density expansion coefficients.

4. The computing system according to claim 1, The training dataset also includes the ground truth values ​​of the gradient unfolded with respect to the kinetic energy density; and The trained kinetic density functional machine learning model is also configured to output a predicted value of the kinetic density unfolded gradient via backpropagation during inference.

5. The computing system of claim 4, wherein the processing circuitry is further configured to perform orbitless density functional theory (DFT) calculations using the predicted value of the kinetic energy density expansion gradient.

6. The computing system according to claim 1, wherein The processing circuitry is configured to train the kinetic energy density functional machine learning model on a training dataset, which is generated at least in part based on the following process: During training: Based on the orbital values ​​calculated in the intermediate iteration step of the self-consistent field SCF algorithm evaluation of the training-time molecular structure data for the training-time molecular structure, the atomic-based density expansion coefficient, the non-interacting kinetic energy, and the gradient of the non-interacting kinetic energy are calculated for the intermediate iteration step; and The density expansion coefficient and the molecular structure at training time are used as inputs at training time, and the non-interacting kinetic energy and the kinetic energy density expansion gradient are stored as true values ​​in the training dataset.

7. The computing system of claim 6, wherein the kinetic density functional machine learning model has a graph neural network transformer architecture configured to implement a nonlocal attention mechanism.

8. The computing system of claim 7, wherein the kinetic density functional machine learning model comprises a transformer neural network with a multi-head attention mechanism, the multi-head attention mechanism being trained using the training-time input and the ground truth output via backpropagation.

9. The computing system of claim 8, wherein the multi-head attention mechanism is configured to: calculate attention and loss for a prediction task of predicting the non-interacting kinetic energy given the density expansion coefficients as input, and is configured to: calculate loss for a prediction task of predicting the kinetic energy density expansion gradient via backpropagation with respect to the input density expansion coefficients.

10. The computing system according to claim 6, wherein The training molecular structure data includes data indicating atomic coordinates, atom numbers, the number of electrons in the training molecular structure, and orbital basis sets. The SCF algorithm is executed at least in part by iterating through multiple SCF loops, and in each SCF loop: Multiple orbital coefficients are calculated, and the orbital basis set is expanded using the multiple orbital coefficients to determine the orbital value for each electron; Based on the orbital coefficients, construct a matrix that approximates the single-electron energy operator of a given quantum system; Using the constructed matrix, calculate the updated orbit coefficients; Based on the updated orbital coefficients, the non-interacting kinetic energy is calculated; Calculate the gradient value with respect to the non-interacting kinetic energy; The density expansion coefficients are calculated on an atomic basis based on the orbital coefficients. Store the values ​​of the non-interacting kinetic energy density expansion coefficient, the non-interacting kinetic energy, and the kinetic energy density expansion gradient for the given conditions; Determine whether the total energy and / or the density matrix has converged to a convergence threshold, and if converged, update the density matrix of the molecular structure at training time based on the orbital coefficients and stop iterating.

11. The computing system of claim 10, wherein the matrix used to approximate the single-electron energy operator of a given quantum system is a Fock matrix.

12. The computing system of claim 10, wherein the molecular structure data during training is stored in a graph data structure, the atom positions and atom numbers are represented as node features in the graph data structure, and the atom positions are used to calculate edge values ​​in the graph data structure.

13. The computing system according to claim 1, wherein the density coefficient is expressed in a local atomic reference frame that corresponds to global shifts and rotations, and the density expansion coefficient SE(3) remains unchanged.

14. The computing system of claim 1, wherein the density coefficients are naturally reparameterized such that the Euclidean metric in the reparameterized coefficient space tends to the L2 metric in the density function space.

15. The computing system of claim 14, wherein the density coefficients are rescaled by dimension.

16. The computational system of claim 15, wherein the atomic reference model is employed for the kinetic energy that is linear with respect to the density coefficient, such that density coefficients corresponding to the same atom type, species and / or chemical element share the same linear weight.

17. A computerized method implemented via processing circuitry, the method comprising: In reasoning: Receive inference-time molecular structure data for inference-time molecular structure, the inference-time molecular structure data including inference-time atomic basis density expansion coefficients and data indicating the atomic properties of each atom in the inference-time molecular structure; The molecular structure data is input into a trained kinetic energy density functional machine learning model to generate predicted values ​​for the electron non-interaction kinetic energy for the density function. The trained machine learning model has been trained on a training dataset that includes atomic properties of each atom in the molecular structure at training time as input, and atomic-based density expansion coefficients for the molecular structure at training time. The atomic-based density expansion coefficients indicate the values ​​of the electron charge density characterized at the atomic basis and include the corresponding true values ​​for the non-interaction kinetic energy. as well as Output the predicted value of the non-interacting kinetic energy of the molecular structure for the reasoning process.

18. The computerized method of claim 17, wherein the atomic properties include one or more of atomic position and / or atomic number.

19. The computerized method of claim 17, wherein the density function is based on atomic-based density expansion coefficients.

20. The computerized method according to claim 17, wherein The training dataset also includes ground truth values ​​for gradient expansion with respect to kinetic energy density, and the method further includes: During inference, the predicted value of the kinetic density expansion gradient is output from the trained kinetic density functional machine learning model via backpropagation.

21. The computerized method according to claim 20, further comprising: The predicted values ​​of the kinetic energy density expansion gradient are used to perform trackless density functional theory (DFT) calculations.

22. The computerized method according to claim 21, further comprising: The training dataset is generated, at least in part, through the following methods: Based on the orbital values ​​calculated in the intermediate iteration step of the self-consistent field SCF algorithm evaluation of the training-time molecular structure data for the training-time molecular structure, the atomic-based density expansion coefficient, the non-interacting kinetic energy, and the gradient of the non-interacting kinetic energy are calculated for the intermediate iteration step; and The density expansion coefficient and the molecular structure at training time are used as inputs at training time, and the non-interacting kinetic energy and the kinetic energy density expansion gradient are stored as true values ​​in the training dataset. as well as The trained kinetic density functional machine learning model is trained on the training dataset.

23. The computerized method of claim 22, wherein the kinetic density functional machine learning model has a graph neural network transformer architecture configured to implement a nonlocal attention mechanism.

24. The computerized method of claim 23, wherein the nonlocal attention mechanism is a multi-head attention mechanism, the multi-head attention mechanism being trained using the training-time input and the truth output via backpropagation.

25. The computerized method of claim 24, wherein the multi-head attention mechanism is configured to: compute attention and loss for a prediction task of predicting the non-interacting kinetic energy given the density unfolding coefficients as input, and is configured to: compute loss for a prediction task of predicting the kinetic energy density unfolding gradient by backpropagation with respect to the input density unfolding coefficients.

26. The computerized method according to claim 22, wherein The training molecular structure data includes data indicating atomic coordinates, atom numbers, the number of electrons in the training molecular structure, and orbital basis sets. The SCF algorithm is executed at least in part by iterating through multiple SCF loops, and in each SCF loop: Multiple orbital coefficients are calculated, and the orbital basis set is expanded using the multiple orbital coefficients to determine the orbital value for each electron; Based on the orbital coefficients, construct a matrix that approximates the single-electron energy operator of a given quantum system; Using the constructed matrix, calculate the updated orbit coefficients; Based on the updated orbital coefficients, the non-interacting kinetic energy is calculated; Calculate the gradient value with respect to the non-interacting kinetic energy; Density expansion coefficients are calculated on an atomic basis based on the updated orbital coefficients; Store the values ​​of the non-interacting kinetic energy density expansion coefficient, the non-interacting kinetic energy, and the kinetic energy density expansion gradient for the given conditions; Determine whether the total energy and / or the density matrix has converged to a convergence threshold, and if converged, update the density matrix of the molecular structure at training time based on the orbital coefficients and stop iterating.

27. The computerized method according to claim 22, wherein The molecular structure data during training is stored in a graph data structure, where the atom positions and atom numbers are represented as node features in the graph data structure, and the atom positions are used to calculate edge values ​​in the graph data structure. The density coefficient is expressed in a local atomic reference frame that corresponds to global shifts and rotations, and the density expansion coefficient SE(3) remains unchanged. The density coefficients are naturally reparameterized so that the Euclidean metric in the reparameterized coefficient space tends to the L2 metric in the density function space. The density coefficients are rescaled according to dimensions; and The atomic reference model is adopted for the kinetic energy that is linear with respect to the density coefficient, such that the density coefficients corresponding to the same atomic type, species and / or chemical element share the same linear weight.

28. The computerized method according to claim 21, further comprising: The density expansion coefficients for the molecular structure at inference time are iteratively optimized using the trained kinetic energy density functional machine learning model, through each iteration in multiple iteration loops: The molecular structure data and density expansion coefficients are input into the trained kinetic energy density functional machine learning model to generate a predicted value of the electron non-interaction kinetic energy for the density function. Then, the predicted kinetic energy density expansion gradient is generated by backpropagation of the trained kinetic energy density functional machine learning model with respect to the input atomic-based density expansion coefficients. The gradient of the sum of the Hartley energy, external potential energy, and exchange-correlated energy is calculated relative to the input atomic-based density expansion coefficients and added to the calculated gradient. The gradient is projected onto a linear subspace of density expansion coefficients that determine the normalized density; The density expansion coefficients are updated using the projected gradient; as well as Determine whether the total energy change value or the projected gradient normalized value from the previous cycle is greater than the value in the previous cycle, and if it is greater than the value in the previous cycle, stop the iteration.

29. The computerized method according to claim 28, further comprising: The density expansion coefficients are initialized by adding the MINAO initial density and the standard MINAO initial density expansion coefficients. The MINAO initial density has an increment predicted using a trained initial value machine learning model that takes the inference-time molecular structure as input. The trained initial value machine learning model has been trained on a training dataset that includes atomic positions as input and data indicating the atomic number of each atom in the training-time molecular structure, as well as atomic-based density expansion coefficients for the training-time molecular structure. The atomic-based density expansion coefficients indicate the value of the electron charge density represented on the atomic basis, and include the corresponding true value of the density expansion coefficient increment relative to the SCF convergence step.

30. A computerized method implemented via processing circuitry, the method comprising: The training dataset for the kinetic energy density functional machine learning model is generated, at least in part, through the following operations: Based on the orbital values ​​calculated in the intermediate iteration step of the self-consistent field SCF algorithm evaluation using training-time molecular structure data, the atomic-based density expansion coefficients, non-interacting kinetic energies, and gradients of the non-interacting kinetic energies for the intermediate iteration step of the SCF algorithm are calculated; and The atomic-based density expansion coefficients and the molecular structure during training are used as inputs during training, and the non-interacting kinetic energy and the kinetic energy density expansion gradient are stored as true values ​​in the training dataset. Train the kinetic energy density functional machine learning model on the training dataset; as well as Output the trained kinetic energy density functional machine learning model.