Electronic structure density matrix prediction method, system, electronic device, and storage medium

CN121687220BActive Publication Date: 2026-08-11ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0008]基于上述背景,本发明提供了一种电子结构密度矩阵预测方法、系统、电子设备和存储介质,以解决现有技术中至少一个方面的问题

Benefits of technology

[0032]1)能够严格保持物理约束:通过指数参数化数学框架,从根本上保证了输出密度矩阵的所有物理约束(厄米性、迹守恒、幂等性、半正定性),确保了结果的物理正确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121687220B_ABST
    Figure CN121687220B_ABST
Patent Text Reader

Abstract

This invention provides a method, system, electronic device, and storage medium for predicting the electronic structure density matrix. The method includes the following steps: constructing a first equivalent graph neural network and a second equivalent graph neural network; modifying the output of the first equivalent graph neural network to satisfy Hermitian property and particle number conservation; in a first training phase, using atomic structure information as input, training the first equivalent graph neural network to generate an initial density matrix that satisfies idempotency; in a second training phase, freezing the parameters of the first equivalent graph neural network, training the second equivalent graph neural network to generate an antisymmetric matrix, and constructing the final density matrix using an exponential parameterization method, where is the overlap matrix; using the trained first and second equivalent graph neural networks, outputting the ground state density matrix of the target system based on the input atomic structure information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of computational physics, computational quantum chemistry, materials simulation and artificial intelligence, and specifically to a method, system, electronic device and storage medium for predicting electronic structure density matrix. Background Technology

[0002] Density functional theory (DFT) and the Hartley-Fock (HF) method are the most widely used methods for calculating electronic structure in modern computational chemistry and materials science. These methods typically rely on self-consistent field (SCF) iterative procedures to solve for the ground-state energy and electron density. A key and computationally expensive step in the SCF procedure is constructing and diagonalizing the Hamiltonian matrix (or Kohn-Sham matrix) of the system; the computational complexity of this step typically increases cubically with the system size. Even higher growth rates severely limit the computable system size and computational efficiency.

[0003] To overcome this bottleneck, academia and industry have made various attempts. One mainstream approach is to use machine learning (ML) to construct potential functions or energy models. These methods typically require a large dataset (containing numerous structures and their corresponding energy, force, and other labels) pre-computed using traditional DFT / HF methods. The machine learning model learns from this dataset to establish a mapping from atomic structure to energy. However, this "data-driven" approach has the following drawbacks:

[0004] Reliance on large-scale pre-computation datasets: The generation of datasets itself requires huge computational resources, and the accuracy and generalization ability of the model depend heavily on the quality and coverage of the dataset.

[0005] Lack of physical constraints: Traditional machine learning models, as a function fitting tool, may not strictly adhere to the fundamental physical constraints in quantum mechanics in their predictions, potentially leading to non-physical results.

[0006] Another approach is to directly use machine learning models to predict physical quantities, such as directly predicting orbital coefficients to construct the density matrix. However, this method is difficult to guarantee the equivariance of the density matrix to geometric transformations such as rotation and translation, and the trained model is difficult to generalize to systems of different sizes.

[0007] Therefore, there is an urgent need in this field for a novel electronic structure calculation method that can avoid the resource-intensive matrix diagonalization step in the traditional SCF method, does not rely on pre-calculated datasets, and can inherently guarantee all necessary physical constraints based on first principles. Summary of the Invention

[0008] Based on the above background, this invention provides an electronic structure density matrix prediction method, system, electronic device, and storage medium to solve at least one problem in the prior art. Specifically, the following technical solution is adopted:

[0009] In a first aspect, the present invention provides a method for predicting electronic structure density matrix, comprising the following steps:

[0010] Construct a first-equivariant graph neural network and a second-equivariant graph neural network;

[0011] The output of the first isomorphic graph neural network is modified to satisfy Hermitian property and particle number conservation.

[0012] In the first training phase, atomic structure information is used as input to train the first equal-variable graph neural network to generate an initial density matrix that satisfies idempotency. ;

[0013] In the second training phase, the parameters of the first-variant graph neural network are frozen, and the second-variant graph neural network is trained to generate an antisymmetric matrix. The final density matrix is ​​constructed using the exponential parameterization method. ,in It is an overlapping matrix;

[0014] Using the trained first and second equivalent graphical neural networks, the ground state density matrix of the target system is output based on the input atomic structure information.

[0015] Furthermore, the output matrices of the first and second equivariant graph neural networks are constructed in the following manner:

[0016] Irreducible tensors are constructed from the edge features output by the network, and then combined into matrix blocks using the o3.wigner_3j() function of the e3nn library. Finally, symmetric or antisymmetric operations are performed to obtain the target matrix.

[0017] Furthermore, the correction of the output of the first isomorphic graph neural network includes:

[0018] Calculate a scalar correction factor The initial output of the first-order variable graph neural network Convert to , making Precisely meet ,in It is the number of electrons in the system.

[0019] Furthermore, in the first training phase, a first loss function is constructed to train the first equivariant graph neural network to ensure idempotency. The first loss function is constructed using one of the following two schemes:

[0020] minimize ;

[0021] optimization The eigenvalues ​​of a matrix tend to be 0 or 1.

[0022] Furthermore, in the second training phase, a second loss function is constructed to train the second equivariant graph neural network. The second loss function is set based on the final density matrix. The calculated total energy of the system The training objective is to minimize the total energy of the system, and the parameters of the second equivariant graph neural network are optimized through backpropagation.

[0023] In a second aspect, the present invention provides an electronic structure density matrix prediction system, comprising:

[0024] The data preprocessing module is used to process atomic structure information and calculate the necessary integral matrix;

[0025] An isomorphic graph neural network module is used to generate the initial density matrix. and antisymmetric matrix ;

[0026] The exponential parameterization module is used to calculate the exponential parameterization based on the formula. Calculate the final density matrix;

[0027] The energy calculation module is used to calculate the total energy of the system based on the density matrix, and to optimize the network parameters in the equivariant graph neural network module using the backpropagation algorithm with the total energy as the loss function.

[0028] The results interaction module receives the input atomic structure information and outputs the ground state density matrix of the target system based on the optimized isomorphic graph neural network module.

[0029] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first aspect above.

[0030] Fourthly, the present invention provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first aspect above.

[0031] The beneficial technical effects of the present invention are as follows:

[0032] 1) It can strictly maintain physical constraints: Through the exponential parameterization mathematical framework, all physical constraints (Hermitian property, trace conservation, idempotency, positive semidefiniteness) of the output density matrix are fundamentally guaranteed, ensuring the physical correctness of the results.

[0033] 2) Correct equivariance: The density matrix is ​​directly predicted, and its equivariance is naturally guaranteed by the network architecture and exponential transformation, which can correctly describe the degenerate orbits and emergent symmetries in high symmetry systems.

[0034] 3) No need for pre-computed datasets: It adopts a "physics-driven" self-supervised learning paradigm and uses first-principles energy functionals as the loss function, completely eliminating the dependence on expensive and potentially biased pre-computed datasets.

[0035] 4) High precision and strong generalization ability: In molecules such as H2, H2O, CH4, CO, CH2O, N2, NH3, and BH3, as well as periodic systems such as graphene, diamond, boron nitride (BN), and silicon, the energy, electron density, and density matrix prediction accuracy of the method in this invention are highly consistent with the results of traditional SCF, demonstrating excellent generalization ability and extrapolation. Even on large-scale supercells not seen during training (such as 2x2x1 graphene / BN supercells and 2x2x2 silicon supercells), the prediction accuracy is far higher than the standard initial conjecture based on core Hamiltonian diagonalization.

[0036] 5) Provides high-quality initial conjectures: The model trained based on the method of this invention can serve as an efficient "initial conjecture generator", providing a nearly convergent initial density matrix for traditional SCF calculation, greatly reducing the number of SCF iterations and accelerating the overall calculation process. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the overall process and model architecture of an embodiment of the present invention.

[0038] Figure 2 This is a schematic diagram of the backpropagation process in an embodiment of the method of the present invention.

[0039] Figure 3 This is a schematic diagram illustrating the construction of a density matrix based on the features of nodes and edges in an embodiment of the present invention.

[0040] Figure 4 This is a schematic diagram illustrating the results of verifying molecular systems (water, methane) and periodic systems (graphene, diamond) in embodiments of the present invention.

[0041] Figure 5This diagram illustrates the idempotency and energy error changes during model training in this embodiment of the invention, as well as the performance during testing.

[0042] Figure 6 This is a schematic diagram showing the convergence of the model of the present invention to SCF accuracy on bulk silicon and a 2x2x2 silicon supercell doped with one carbon atom.

[0043] Figure 7 This is a schematic diagram illustrating the accuracy of the model of the present invention converging to the SCF on the primitive cell of a two-dimensional BN. Detailed Implementation

[0044] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.

[0045] See Figure 1 The first embodiment of the present invention provides a method for predicting the electronic structure density matrix, comprising the following steps:

[0046] S1. Construct the first and second equivalent graph neural networks.

[0047] To ensure the physical consistency of the model, this invention employs an isovariable graph attention network as its core architecture. See also Figure 1 (a) This network takes an atomic system as input, abstracts it into a graph structure (nodes are atoms, edges are connections between atoms), and iteratively updates the features of the nodes and edges. The isovariant graph neural network in this embodiment has the following features and mechanisms:

[0048] Equivariance: Network design ensures that its output (such as a matrix) undergoes correct covariant transformations under operations such as rotations in three-dimensional space. This can be achieved by representing features as irreducible representations of spherical harmonics and performing tensor multiplication using Clebsch-Gordan coefficients. See also... Figure 3 Because the model uses an isovariant network, its output contains irreducible tensor representations of different orders. Using the Wigner function, diagonal and off-diagonal blocks of the matrix can be constructed. By combining all the diagonal and off-diagonal blocks, the complete matrix can be obtained. The complete matrix can then be symmetric or antisymmetric to obtain the target matrix. This makes the output density matrix and antisymmetric matrix inherently possess the correct physical transformation properties, fundamentally solving the isovariability deficiency problem of traditional methods.

[0049] Message passing mechanism and attention mechanism: see Figure 1 (a) Through multi-round message passing and multi-head attention mechanisms, the network can capture remote interactions between atoms and chemical environment information, thereby accurately constructing the target matrix.

[0050] As a preferred embodiment, in this example, the output matrices of the first and second equivariant graph neural networks are constructed in the following manner:

[0051] Irreducible tensors are constructed from the edge features output by the network, and then combined into matrix blocks using the o3.wigner_3j() function of the e3nn library. Finally, symmetric or antisymmetric operations are performed to obtain the target matrix.

[0052] S2. Modify the output of the first variable graph neural network so that its output satisfies Hermitian property and particle number conservation.

[0053] As a preferred embodiment, in this example, the original output matrix of the network can be symmetrically transformed. To ensure Hermitian properties.

[0054] As a preferred embodiment, in this example, the correction of the output of the first equivalent graph neural network includes:

[0055] Calculate a scalar correction factor The initial output of the first-order variable graph neural network Convert to , making Precisely meet ,in It is the number of electrons in the system.

[0056] S3, First Training Phase, Reference Figure 1 (b) Using atomic structure information as input, train the first-order graph neural network to generate an initial density matrix that satisfies idempotency. .

[0057] As a preferred implementation, in this embodiment, a first loss function is constructed to train the first equivariant graph neural network to satisfy idempotency. The first loss function is constructed using one of the following two schemes:

[0058] minimize ;

[0059] optimization The eigenvalues ​​of a matrix tend to be 0 or 1.

[0060] The goal of this stage is to obtain a physically plausible initial conjecture that approximates the true solution and satisfies all the fundamental constraints.

[0061] S4, Second Training Phase, Reference Figure 1 (c) Freeze the parameters of the first equivalent graph neural network and train the second equivalent graph neural network. Its input is also an atomic structure, and its output is an antisymmetric matrix. The final density matrix is ​​constructed using an exponential parameterization method. This method conjectures from an initial density matrix that satisfies some constraints. Starting from, passing through an antisymmetric matrix Controlled similarity transformations are used to generate the final density matrix:

[0062]

[0063] in, It is an overlapping matrix.

[0064] This method provides mathematical guarantees during the computation process: if the initial matrix It satisfies Hermitian property, particle number conservation, idempotency, and positive semidefiniteness, and the transformation matrix... It is an antisymmetric matrix ( Then the transformed matrix All these physical constraints will be automatically maintained. This provides a solid mathematical foundation for the safe optimization of neural networks and is the key to the superiority of this method over traditional methods.

[0065] For a preferred implementation, see Figure 1 In this embodiment, a second loss function is constructed to train the second isotropic graph neural network. The second loss function is set based on the final density matrix. The calculated total energy of the system The training objective is to minimize the total energy of the system, and the parameters of the second isomorphic graph neural network are optimized through backpropagation. When training converges... This is the system's ground state density matrix. Thanks to exponential parameterization, the energy optimization process does not violate any physical constraints. This invention integrates density matrix-based DFT calculation into model training, achieving self-supervised learning without a dataset and unifying DFT calculation and model training.

[0066] S5. Using the trained first and second equivalent graph neural networks, output the ground state density matrix of the target system based on the input atomic structure information.

[0067] A second embodiment of the present invention also provides an electronic structure density matrix prediction system, comprising:

[0068] The data preprocessing module is used to process atomic structure information and calculate the necessary integral matrix;

[0069] An isomorphic graph neural network module is used to generate the initial density matrix. and antisymmetric matrix ;

[0070] The exponential parameterization module is used to calculate the exponential parameterization based on the formula. Calculate the final density matrix;

[0071] The energy calculation module is used to calculate the total energy of the system based on the density matrix, and to optimize the network parameters in the equivariant graph neural network module using the backpropagation algorithm with the total energy as the loss function.

[0072] The results interaction module receives the input atomic structure information and outputs the ground state density matrix of the target system based on the optimized isomorphic graph neural network module.

[0073] The third embodiment of the present invention provides a specific application example of the above-described electronic structure density matrix prediction method and system.

[0074] 1) System Construction and Environment Configuration

[0075] Hardware: Computing servers equipped with high-performance GPUs (such as NVIDIA V100, A100).

[0076] Software: Operating system is Linux, programming language is Python 3.8+. Main dependencies: PyTorch, PyTorchGeometric (or similar graph neural network library), e3nn (for isovariant networks), PySCF.

[0077] 2) Data preprocessing and input

[0078] Input: Structural information of the atomic system (atom types, atomic coordinates, lattice vectors).

[0079] Basis set and integral: Predefined basis sets (e.g., 6-31G for molecules, GTH-SZV for periodic systems) are used to pre-compute the overlap matrix, kernel Hamiltonian matrix, and three-center integral required for Gaussian density fitting using PySCF, which are then used as constant inputs.

[0080] 3) Network Model Construction

[0081] Network Structure: Two structurally similar Equivariant Graph Attention Networks (E-GAT) are constructed. The networks typically contain two equivariant convolutional layers, using the SiLU activation function.

[0082] Matrix construction: based on Figure 3 The process shown involves constructing irreducible tensors from the edge features of the network, and then combining them into the target matrix.

[0083] 4) Training process

[0084] Phase 1 training: Optimizer: Adam; Learning rate: 1e-3; Batch size: 8; Training period: 300 epochs (can be changed depending on the specific situation).

[0085] Second-stage training: Optimizer: Adam; Learning rate: 1e-3 (can be adjusted depending on the specific system); Loss function: Training cycle: 200 epochs (can be changed depending on the specific situation).

[0086] Numerical precision: All calculations use double-precision floating-point numbers (FP64).

[0087] 5) Application Examples

[0088] Example 1: Water molecules (H2O) and methane molecules (CH4)

[0089] Using LDA-VWN functionals and 6-31G basis sets. See [link / reference]. Figure 4 As shown in (a), (b), and (c), the final energy optimized using the method of the present invention has an error of less than 0.1 mHa compared with the SCF result, and the electron density distribution is almost indistinguishable from the SCF result.

[0090] Example 2: Graphene and Diamond

[0091] See Figure 4 As shown in (d) and (e), the final energy optimized using the method of the present invention has an error of less than 0.1 mHa compared with the SCF result, and the electron density distribution and density matrix are almost identical to the SCF result.

[0092] Example 3: Graphene Supercell Extrapolation

[0093] Using a model trained on graphene primitive cells, the density matrix and energy of a 2x2x1 supercell are directly predicted. See also Figure 5 (a) shows the error changes in idempotency and energy during the training of the two models. (b) and (c) show the direct prediction of the density matrix and energy of a 2x2x1 supercell using a model trained on graphene primitives. The prediction errors of energy and density matrices are much smaller than the initial conjectures obtained by diagonalizing the core Hamiltonian, demonstrating the model's excellent extrapolation capability.

[0094] Example 4: Bulk silicon, a 2x2x2 silicon supercell doped with one carbon atom and boron nitride (BN).

[0095] See Figure 6 and Figure 7 , Figure 6 (a) and Figure 6(b) It is shown that the model of the present invention can converge rapidly to SCF accuracy in both bulk silicon and 2x2x2 silicon supercells doped with one carbon atom. Figure 7 This demonstrates that the model of this invention converges rapidly to SCF accuracy in the unit cell of two-dimensional BN. It can be seen that the model of this invention can converge rapidly to SCF accuracy in bulk silicon, 2x2x2 silicon supercells doped with one carbon atom, and the unit cell of two-dimensional BN, further verifying its effectiveness and robustness in various periodic systems.

[0096] The fourth embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the first embodiment above.

[0097] The fifth embodiment of the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in the first embodiment above.

[0098] It should be noted that the method of this embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this embodiment, and the multiple devices will interact with each other to complete the method described.

[0099] The embodiments of this invention are intended to cover all such substitutions, modifications, and variations falling within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this invention should be included within the protection scope of this invention.

Claims

1. A method for predicting electronic structure density matrix, characterized in that, Includes the following steps: Construct a first-equivariant graph neural network and a second-equivariant graph neural network; The output of the first isomorphic graph neural network is modified to satisfy Hermitian properties and particle number conservation, including: Calculate a scalar correction factor The initial output of the first-order variable graph neural network Convert to , making Precisely meet ,in It is the number of electrons in the system; In the first training phase, atomic structure information is used as input to train the first equal-variable graph neural network to generate an initial density matrix that satisfies idempotency. ,include: The first loss function is constructed to train the first isomorphic graph neural network, ensuring idempotency. The first loss function is constructed using one of the following two schemes: Minimize ; optimization The eigenvalues ​​of the matrix tend to be 0 or 1; In the second training phase, the parameters of the first-variant graph neural network are frozen, and the second-variant graph neural network is trained to generate an antisymmetric matrix. The final density matrix is ​​constructed using the exponential parameterization method. ,in It is an overlapping matrix; Using the trained first and second equivalent graphical neural networks, the ground state density matrix of the target system is output based on the input atomic structure information. The output matrices of the first and second equivalent graphical neural networks are constructed in the following way: Irreducible tensors are constructed from the edge features output by the network, and then combined into matrix blocks using the o3.wigner_3j() function of the e3nn library. Finally, symmetric or antisymmetric operations are performed to obtain the target matrix.

2. The method according to claim 1, characterized in that, In the second training phase, a second loss function E is constructed to train the second isovariant graph neural network, based on the final density matrix. The calculated total energy of the system The training objective is to minimize the total energy of the system, and the parameters of the second equivariant graph neural network are optimized through backpropagation.

3. An electronic structure density matrix prediction system, used to implement the method as described in claim 1 or 2, characterized in that, include: The data preprocessing module is used to process atomic structure information and calculate the necessary integral matrix; An isomorphic graph neural network module is used to generate the initial density matrix. and antisymmetric matrix ; The exponential parameterization module is used to calculate the exponential parameterization based on the formula. Calculate the final density matrix; The energy calculation module is used to calculate the total energy of the system based on the density matrix, and to optimize the network parameters in the equivariant graph neural network module using the backpropagation algorithm with the total energy as the loss function. The results interaction module receives the input atomic structure information and outputs the ground state density matrix of the target system based on the optimized isomorphic graph neural network module.

4. A computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the method as claimed in claim 1 or 2.

5. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as claimed in claim 1 or 2.

Citation Information

Patent Citations

  • Energy and atom stress calculation method based on machine learning density matrix

    CN119763691A

  • A computer-implemented method of training a graph neural network

    GB201821192D0