Method, device, equipment and medium for determining quantum system state representation
By generating trajectory representations and updating network parameters through a neural network based on the attention mechanism, the problem of insufficient accuracy of existing methods in representing the state of quantum systems is solved, and high-precision analysis of complex quantum systems is achieved.
Patent Information
- Application Number
- CN202510819963.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-10-03
AI Technical Summary
Existing quantum chemical calculation methods lack accuracy when constructing quantum system state representations, making it difficult to meet the analysis needs of complex quantum systems. In particular, when dealing with high-dimensional, strongly correlated systems such as high-temperature superconducting copper oxides, traditional numerical methods have poor accuracy.
A neural network based on an attention mechanism is used to generate trajectory representations and update the neural network by minimizing energy to approximate the ground state energy of the quantum system to determine the accurate state representation.
It improves the modeling accuracy and scalability of complex quantum systems, can maintain high computational accuracy at larger system sizes, and is suitable for quantum system analysis in more scenarios.
Smart Images

Figure CN120745860A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computer technology, and more particularly, to a method, apparatus, device, and medium for determining a representation of a quantum system state. Background Art
[0002] Quantum chemistry is a computational science based on quantum mechanics that studies the microscopic structure and reaction mechanisms of matter. In its practical applications, accurately constructing a state representation of a quantum system is crucial for understanding and predicting its properties. In recent years, neural networks have been used to construct state representations of quantum systems and have shown promising results. However, current methods still have limitations in terms of accuracy and other aspects, making them inadequate for quantum chemical computations. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for determining a state representation of a quantum system is provided. The method comprises: generating at least one orbital representation for the plurality of electrons using a first neural network based on an attention mechanism based on corresponding electron configurations of a plurality of lattice points in a quantum system comprising the plurality of electrons; determining an energy of the quantum system based on the at least one orbital representation; and updating the first neural network by minimizing the energy.
[0004] In a second aspect of the present disclosure, a device for determining a state representation of a quantum system is provided. The device includes: an orbital representation generation module configured to generate at least one orbital representation for a plurality of electrons based on corresponding electron configurations of a plurality of lattice points in a quantum system comprising the plurality of electrons, using a first neural network based on an attention mechanism; an energy determination module configured to determine the energy of the quantum system based on the at least one orbital representation; and an update module configured to update the first neural network by minimizing the energy.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored on the computer-readable storage medium, and the computer-executable instructions can be executed by a processor to implement the method of the first aspect.
[0007] In a fifth aspect of the present disclosure, a computer program product is provided, which includes computer-executable instructions, which, when executed by a processor, implement the method according to the first aspect of the present disclosure.
[0008] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] The above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0010] Figure 1 A schematic diagram illustrating an example environment in which embodiments of the present disclosure can be implemented;
[0011] Figure 2 A flowchart illustrating an example process for determining a representation of a quantum system state according to some embodiments of the present disclosure is shown;
[0012] Figure 3 A schematic diagram illustrating an example architecture for determining a representation of a quantum system state according to some embodiments of the present disclosure;
[0013] Figure 4 A schematic structural block diagram of an apparatus for determining a quantum system state representation according to some embodiments of the present disclosure is shown; and
[0014] Figure 5 A block diagram of an electronic device is shown in which one or more embodiments of the present disclosure may be implemented. DETAILED DESCRIPTION
[0015] The following describes embodiments of the present disclosure in more detail with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0016] It should be noted that the titles of any section / subsection provided herein are not limiting. Various embodiments are described throughout this document, and any type of embodiment may be included under any section / subsection. Furthermore, the embodiments described in any section / subsection may be combined in any manner with any other embodiments described in the same section / subsection and / or in different sections / subsections.
[0017] In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may be included below. The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may be included below.
[0018] The embodiments of the present disclosure may involve user data, data acquisition and / or use, etc. These aspects shall comply with the corresponding laws, regulations and relevant provisions. In the embodiments of the present disclosure, all data collection, acquisition, processing, processing, forwarding, use, etc. are carried out on the premise that the user is aware of and confirms them. Accordingly, when implementing the various embodiments of the present disclosure, the types, scope of use, and usage scenarios of the data or information that may be involved should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with the relevant laws and regulations. The specific notification and / or authorization method may vary according to the actual situation and application scenario, and the scope of the present disclosure is not limited in this respect.
[0019] If this specification and the solutions in the examples involve the processing of personal information, such processing will be done only with a legitimate basis (such as with the consent of the subject of personal information or as necessary for the performance of a contract) and only within the prescribed or agreed scope. A user's refusal to process personal information other than that required for basic functions will not affect the user's use of basic functions.
[0020] Figure 1 A schematic diagram of an example environment 100 is shown in which embodiments of the present disclosure can be implemented. Figure 1 , the example environment 100 may include a terminal device 110 and an electronic device 120 .
[0021] In example environment 100, user 130 can interact with electronic device 120 via terminal device 110 and / or its attached devices. As an example, user 130 can provide input to electronic device 120 via terminal device 110 for determining a representation of the state of a quantum system. For example, the quantum system may include multiple electrons, and the input may include corresponding electronic configurations of multiple lattice points in the quantum system. The electronic configuration is also called the electronic structure or electronic arrangement. The electronic configuration may indicate, for example, the distribution of electrons at corresponding lattice points (e.g., atoms, etc.).
[0022] After receiving input for determining a quantum system state representation, electronic device 120 may determine orbital representations for multiple electrons based on the input. In a quantum system, each lattice point may be associated with multiple orbitals. For example, each lattice point may be associated with a first orbital associated with an electron spin-up and a second orbital associated with an electron spin-down. The orbital representation may be, for example, a parameterized representation or a quantized representation of the electron distribution on the orbitals associated with each lattice point.
[0023] Next, electronic device 120 constructs a wave function based on the orbital representation to describe the state of the quantum system. By continuously optimizing the wave function parameters to minimize the energy of the quantum system, electronic device 120 can approximate the ground state energy of the quantum system and determine a corresponding target wave function. Based on the target wave function, electronic device 120 can then determine a target state representation that approximates the ground state of the quantum system.
[0024] In the example environment 100, the terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including a mobile phone, a desktop computer, a laptop computer, a notebook computer, a netbook computer, a tablet computer, a media computer, a multimedia tablet, a personal communication system (PCS) device, a personal navigation device, a personal digital assistant (PDA), an audio / video player, a digital camera / camcorder, a positioning device, a television receiver, a radio broadcast receiver, an e-book device, a gaming device or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 can also support any type of interface for the user (such as "wearable" circuitry, etc.).
[0025] The electronic device 120 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, and big data and artificial intelligence platforms. The electronic device 120 may include, for example, a computing system / server, such as a mainframe, an edge computing node, a computing device in a cloud environment, and the like.
[0026] A communication connection may be established between the electronic device 120 and the terminal device 110. The communication connection may be established via a wired or wireless method. The communication connection may include, but is not limited to, a Bluetooth connection, a mobile network connection, a Universal Serial Bus connection, a Wi-Fi connection, etc., and the embodiments of the present disclosure are not limited in this respect. In the embodiments of the present disclosure, the electronic device 120 and the terminal device 110 may implement signaling interaction via the communication connection between them.
[0027] It should be understood that the above description of the structure and function of each element in environment 100 is for exemplary purposes only and does not imply any limitation on the scope of the present disclosure. For example, in addition to a physical server, electronic device 120 may also be any type of mobile terminal, fixed terminal, or portable terminal. In this case, user 130 can directly interact with electronic device 120 to perform analysis of the quantum system.
[0028] As briefly described above, in the practical application of quantum chemistry, how to accurately construct the state representation of a quantum system is the key to understanding and predicting its properties. In one approach, the state representation of a quantum system can be determined based on the Hubbard Model. The Hubbard Model can indicate quantum magnetism, Mott insulating states, non-uniform charges, spin-ordered structures, and possible unconventional superconductivity. Currently, the Hubbard Model can be solved by numerical methods to determine the state representation of a quantum system. Numerical methods may include, for example, the Density Matrix Renormalization Group (DMRG), Infinite Projected Entangled Pair States (iPEPS), and Constrained Path Auxiliary Field Quantum Monte Carlo (AFQMC), etc. However, these numerical methods have poor accuracy when dealing with complex quantum systems (such as high-temperature superconducting copper oxides), while the Hubbard Model is highly sensitive to accuracy. This makes it difficult for traditional numerical methods to meet the practical needs of the Hubbard model, and therefore a more precise solution method is urgently needed to achieve accurate analysis of complex quantum systems.
[0029] In light of this, embodiments of the present disclosure provide a scheme for determining a quantum system state representation. In this embodiment, first, for a quantum system comprising multiple electrons, a first neural network based on an attention mechanism is used to generate at least one orbital representation for the multiple electrons based on the corresponding electron configurations at multiple lattice points in the quantum system. Then, based on the at least one orbital representation, the energy of the quantum system is determined. Subsequently, the first neural network is updated by minimizing the energy.
[0030] It will be more clearly understood through the following description that according to the embodiments of the present disclosure, in the process of determining the state representation of a quantum system, a neural network based on an attention mechanism (e.g., a neural network based on a Transformer architecture) is introduced. The neural network based on the attention mechanism can be used as an ansatz of the wave function of the quantum system. By continuously optimizing the parameters of the neural network to minimize the energy of the quantum system, the ground state energy of the quantum system can be approximated, and the target wave function corresponding thereto can be determined. Subsequently, based on the target wave function, a state representation that is close to the ground state of the quantum system can be determined.
[0031] Compared with traditional numerical methods, neural networks based on attention mechanisms can demonstrate higher modeling accuracy and expression capabilities when processing complex quantum systems such as high-dimensional and strongly correlated systems (such as high-temperature superconducting copper oxides). This is conducive to the accurate analysis of complex quantum systems. Thanks to these advantages, the embodiments of the present disclosure can still maintain high computational accuracy under larger system sizes (such as quantum systems containing dozens or even hundreds of grid points). In this way, the scalability of quantum system analysis can be improved, making the embodiments of the present disclosure applicable to more scenarios.
[0032] Various example implementations of this solution will be described in detail below with reference to the accompanying drawings. Figure 2 1 shows a flow chart of an example process 200 for determining a representation of a quantum system state according to some embodiments of the present disclosure. Figure 3 A schematic diagram of an example architecture 300 for determining a quantum system state representation according to some embodiments of the present disclosure is shown. Figure 1 and Figure 3 The process 200 is described. In an embodiment of the present disclosure, the process 200 may be implemented as follows: Figure 1 At the electronic device 120 shown.
[0033] At block 210, electronic device 120 generates at least one orbital representation for a quantum system 301 comprising multiple electrons based on the corresponding electronic configurations of multiple lattice points in quantum system 301 using an attention-based neural network (e.g., first neural network 303). In example architecture 300, quantum system 301 is shown as including four lattice points 302-1, 302-2, 302-3, and 302-4. However, as needed, quantum system 301 may include more lattice points. For ease of discussion, lattice points 302-1, 302-2, 302-3, and 302-4 will be referred to individually or collectively as lattice points 302. Two orbital representations are shown, namely orbital representations 304-1 and 304-2. However, as needed, the number of orbital representations may be greater. For ease of discussion, orbital representations 304-1 and 304-2 will be referred to individually or collectively as orbital representations 304.
[0034] Quantum system 301 may refer to a microscopic system described based on quantum mechanics. Quantum system 301 may be a many-body quantum system 301. Many-body quantum system 301 may be a system comprising multiple interacting particles (e.g., electrons and atoms). In many-body quantum system 301, each particle follows the basic laws of quantum mechanics. The interactions between multiple particles may be direct (e.g., Coulomb interactions) or indirect through a medium.
[0035] In some embodiments, quantum system 301 can be modeled as a Fermi system having N lattice points 302. N is a positive integer. In this case, the particles in quantum system 301 obey Fermi-Dirac statistics. For example, for each atom in quantum system 301, the electron occupancy number in the atom's orbital is either 0 (i.e., the orbital is unoccupied) or 1 (i.e., the orbital is occupied). It should be noted that the atomic orbital can be a first orbital associated with the electron's spin-up state or a second orbital associated with the electron's spin-down state, etc. For each lattice point 302 in quantum system 301, the lattice point 302 can represent an atom or molecule in quantum system 301. Different lattice points 302 represent different atoms or molecules. In addition, in quantum system 301, each lattice point 302 can be configured with a corresponding electronic configuration. The electronic configuration can indicate, for example, the distribution of electrons configured at the lattice point 302. For example, the electronic configuration can indicate, for example, the configuration of the electron occupancy number at the lattice point 302.
[0036] In some embodiments, the quantum system 301 can be represented by the Hubbard model or other models. In some embodiments described below, the Hubbard model will be used as an example. However, it should be understood that the aspects described with respect to the Hubbard model are also applicable to other models. Taking a Fermi system with N lattice points 302 as an example, the Hubbard model used to represent the system can be expressed by formula (1):
[0037]
[0038] in represents the electron generation operator with spin state σ at the i-th lattice point 302, is the particle number operator, which represents the number of electrons with spin state σ at the i-th lattice point. <ij>represents the nearest neighbor point pair, < <ij>> represents the next nearest neighbor grid point pair, t represents the nearest neighbor transition amplitude, t′ represents the next nearest neighbor transition amplitude, is the local Coulomb repulsion term, which represents the Coulomb repulsion between two electrons with different spin directions at the same lattice point, and U represents the local Coulomb interaction strength.
[0039] In some embodiments, the nearest neighbor transition amplitude t can be set to 1 or other values as a unit of energy. The local Coulomb interaction strength U can be set to 8 or other values. When U = 8, the strong electron correlation effect in the copper oxide material can be effectively described. For simulations including next-nearest neighbor transitions, the next-nearest neighbor transition amplitude t′ can be set to -0.2 or other values. In some embodiments, the hole doping concentration δ can be set to 1 / 8 or other values to focus on the lightly doped region.
[0040] In some embodiments, the electronic device 120 may solve the Hubbard model based on a variational Monte Carlo (VMC) algorithm or other appropriate quantum Monte Carlo (QMC) algorithm to determine a state representation of the quantum system 301. In an embodiment of the present disclosure, the electronic device 120 may use a first neural network 303 as a pseudo-property for the wave function 305 of the quantum system 301. The wave function 305 may describe the state of the quantum system 301, such as a quantum state. The quantum state described by the wave function 305 determined by the neural network may also be referred to as a neural network quantum state (NQS). On this basis, by optimizing the parameters of the first neural network 303 to minimize the energy of the quantum system 301, the electronic device 120 may approximate the ground state energy of the quantum system 301 and determine a corresponding target wave function. Subsequently, the electronic device 120 may determine a state representation that approximates the true quantum state of the quantum system 301 based on the target wave function.
[0041] The first neural network 303 can be a generative neural network based on the Transformer architecture or a generative neural network based on another architecture. The electronic device 120 can use a given configuration |n> of the quantum system 301 (e.g., the corresponding electron configurations at multiple lattice points 302) as input to the first neural network 303. Based on this input, the first neural network 303 can generate at least one orbital representation 304 for multiple electrons by exploiting the influence of each lattice point 302 on each other. Subsequently, based on these orbital representations 304, the first neural network 303 can output the corresponding amplitude ψθ(n) of the wave function 305 for the given configuration |n>. θ is an adjustable parameter of the first neural network 303.
[0042] In some embodiments, in quantum system 301, each lattice point 302 in multiple lattice points 302 may be associated with at least one orbital. For example, each lattice point 302 may be associated with at least the first and second orbitals mentioned above. In some embodiments, for each orbital representation 304 in at least one orbital representation 304, the orbital representation 304 may indicate the distribution of multiple electrons in the orbitals associated with the multiple lattice points 302. Alternatively or additionally, in some embodiments, the orbital representation 304 may indicate the association relationship between each electron and other electrons in the multiple electrons. As an example, the association relationship may include interactions between multiple electrons.
[0043] In some embodiments, the trajectory representation 304 may be any suitable parameterized representation or feature vectorized representation. Figure 3 In the example architecture 300, the track representation 304 is shown as a track matrix related to the backflow effect. The quantum system 301 is shown as including four grid points 302. In the four grid points 302, each grid point 302 is associated with two tracks, namely the first track and the second track mentioned above. In this case, the quantum system 301 can be considered to involve a total of 8 tracks, 4*2, namely tracks 306-1, 306-2, ..., 306-8. In tracks 306-1, 306-2, ..., 306-8, track 306-1 and track 306-5 are the first track and the second track associated with grid point 302-1, respectively. And so on. Track 306-2 and track 306-6 are the first track and the second track associated with grid point 302-2, respectively. Track 306-3 and track 306-7 are the first track and the second track associated with grid point 302-3, respectively. Track 306 - 4 and track 306 - 8 are respectively the first track and the second track associated with grid point 302 - 4 .
[0044] Assume that the quantum system 301 contains 3 electrons, and Figure 3 As shown, three electrons are respectively located on orbital 306-2 associated with lattice point 302-2, orbital 306-5 associated with lattice point 302-1, and orbital 306-7 associated with lattice point 302-3. In this case, each row of the orbital matrix can be a quantized representation of the eigenvectors corresponding to the eight orbitals. The dimensions of each row of the quantized representation of the eigenvectors can have eight degrees of freedom corresponding to the eight orbitals, thereby indicating the correlation between electrons in different orbitals. It should be noted that the above description of the dimensions of orbital representation 304, the number of lattice points 302, and the number of electrons is merely exemplary. Depending on actual needs, the dimensions of orbital representation 304, the number of lattice points 302, and the number of electrons can also use other appropriate values. In this way, orbital representation 304 can fully reflect the interactions between electrons in quantum system 301, thereby facilitating accurate analysis of quantum system 301 with strong correlations.
[0045] In some embodiments, the first neural network 303 may include an embedding layer 3031 and at least one attention layer 3032. Based on this, the electronic device 120 may utilize the embedding layer 3031 of the first neural network 303 to generate corresponding embedded representations 307 (e.g., embedding vectors) of the plurality of grid points 302 based on the corresponding electronic configurations of the plurality of grid points 302. Subsequently, the electronic device 120 may utilize at least the at least one attention layer 3032 of the first neural network 303 to generate at least one trajectory representation 304 based on the corresponding embedded representations 307 of the plurality of grid points 302.
[0046] As an example, the embedding layer 3031 of the first neural network 303 can assign a unique embedding representation 307 to each lattice point 302 in the quantum system 301. For each lattice point 302 in the at least one lattice point 302, the lattice point 302 can be mapped into a four-dimensional physical Hilbert space consisting of electron spin states |0>, |↑>, |↓>, and |↑↓>. |0> indicates that the orbital associated with the lattice point 302 is unoccupied by an electron. |↑> indicates that the first orbital associated with the lattice point 302 (i.e., the orbital associated with the electron's spin-up state) is occupied by an electron. |↓> indicates that the second orbital associated with the lattice point 302 (i.e., the orbital associated with the electron's spin-down state) is occupied by an electron. |↑↓> indicates that both the first and second orbitals associated with the lattice point 302 are occupied.
[0047] As an example, the embedding layer 3031 of the first neural network 303 can be embedded by the matrix The four representations |0>, |↑>, |↓> and |↑↓> of each grid point 302 are encoded into a d-dimensional vector. In some embodiments, in order to take into account the spatial arrangement information (due to the permutation equivariance of the self-attention mechanism), the embedding representation 307 of each grid point 302 can be enhanced by a learnable position encoding. For example, the enhanced matrix Added to the input embedding matrix. Such an embedding matrix can be compactly represented as the matrix X 0 , X 0 =[E(n1)+P1,……,E(n N )+P N ], where E(n1) is the element in the embedding matrix E corresponding to the first grid point 302 in the N grid points 302, and P1 is the element in the enhancement matrix P corresponding to the first grid point 302 in the N grid points 302. Similarly, E(n N ) is the element in the embedding matrix E corresponding to the Nth grid point 302 among the N grid points 302, P N is the element in the enhancement matrix P corresponding to the Nth grid point 302 among the N grid points 302 .
[0048] In addition to the embedding layer 3031, the first neural network 303 may further include a plurality of attention layers 3032 connected in sequence. The first neural network 303 may embed the plurality of grid points 302 through the plurality of attention layers 3032. 0 ) to generate at least one trajectory representation 304. As an example, in the multiple attention layers 3032 of the first neural network 303, the complex and potential long-range interactions and correlations between the multiple grid points 302 can be modeled by dynamically weighting the mutual influence between the multiple grid points 302. This capability can effectively capture the physical properties of strongly correlated systems (such as the Fermi-Hubbard model).
[0049] In some embodiments, in addition to the embedding layer 3031 and at least one attention layer 3032, the first neural network 303 may further include a linear layer 3033. Based on this, the electronic device 120 may utilize the at least one attention layer 3032 to generate corresponding feature representations 308 for the plurality of grid points 302 based on the corresponding embedding representations 307 of the plurality of grid points 302. The corresponding feature representations 308 of the plurality of grid points 302 may indicate the influence of the plurality of grid points 302 on each other. Subsequently, the electronic device 120 may utilize the linear layer 3033 of the first neural network 303 to convert the corresponding feature representations 308 of the plurality of grid points 302 into at least one track representation 304.
[0050] The first neural network 303 may include L attention layers 3032. L is a positive integer. The attention layer 3032 may also be called an attention block or a Transformer block. Each attention layer 3032 may use formulas (2) and (3) to calculate the input (e.g., the matrix X 0 ) to transform:
[0051] Y (l) =X (l-1) +Attn (l) (X (l-1) ) (2)
[0052] X (l) =Y (l) +FFN (l) (Y (l) ); (3)
[0053] Where l∈[L], represents the lth attention layer3032, Attn (l) and FFN (l) Respectively represent the multi-head self-attention mechanism layer and the feedforward neural network in the l-th attention layer 3032. In some embodiments, the multi-head self-attention mechanism layer Attn (l) It can be expressed by formula (4):
[0054]
[0055] Where H represents the number of multi-head attention heads, h∈[H], Represents the learnable parameter matrix corresponding to the h-th attention head, which is used to generate the query vector, key vector, and value vector respectively. represents the projection matrix, d k =d / H, which represents the dimension of each attention head.
[0056] In some embodiments, the feed-forward neural network FFN (l) It can be expressed by formula (5):
[0057]
[0058] in Represents the feedforward neural network FFN in the lth attention layer 3032 (l) The learnable weight matrix in , σ is the activation function. In some embodiments, the activation function can adopt the Sigmoid Linear Unit (SiLU) activation function or other appropriate activation functions.
[0059] Next, the final output X of the attention layer 3032 (l) At least one track representation 304 corresponding to a given configuration |n> can be generated by mapping through a linear layer 3033. In some embodiments, the track representation 304 can be a track matrix, specifically a track matrix related to the reflow effect. In this document, such a track matrix can also be referred to as a reflow track matrix, and the reflow track matrix corresponding to a given configuration |n> can be expressed as a reflow track matrix M(n). In some embodiments, the final output X of the multiple attention layers 3032 of the first neural network 303 is (l) After the transformation by the linear layer 3033 related to the electron spin state, K return track matrices M are generated. α,k , where α∈{↑↓} represents the electron spin state (also known as the spin direction), and k∈K. K is a positive integer whose value can be determined based on actual needs. In this way, the first neural network 303 can fully exploit the mutual influence between the grid points 302, thereby generating an orbital representation 304 that can accurately reflect the interaction between the electrons.
[0060] Return to reference Figure 2 After determining the at least one orbital representation 304 , the electronic device 120 performs block 220 . At block 220 , the electronic device 120 determines the energy of the quantum system 301 based on the at least one orbital representation 304 .
[0061] As mentioned above, after determining at least one orbital representation 304, the first neural network 303 can output the amplitude ψθ(n) of the wave function 305. The electronic device 120 can determine the energy of the quantum system 301 based on the amplitude ψθ(n) and the energy operator. In some embodiments, the energy operator can be the Hamiltonian operator Or other appropriate energy operators. As an example, the energy of the quantum system 301 can be expressed as the expected value of the local energy, which can be specifically shown as formula (6):
[0062]
[0063] in is the average energy corresponding to the wave function 305 defined by the parameter θ of the first neural network 303, ψθ represents the wave function 305 parameterized by the first neural network 303, E loc (n) represents the local energy at a given configuration |n>. As an example, E loc (n) can be expressed by formula (7):
[0064]
[0065] where n' represents all possible configurations.
[0066] Since the Hilbert space grows exponentially, the calculation accuracy using formula (6) is low. Therefore, a Monte Carlo method can be used to approximate this sum to improve this problem. For example, in some embodiments, a Markov Chain Monte Carlo (MCMC) algorithm or other appropriate algorithm can be used to sample a set of configurations (e.g., occupation numbers) from the probability distribution defined by the square amplitude of the wave function 305. B represents the batch size of the sampling. Then, for the sampled configuration, the corresponding local energy E is calculated loc (n). For the Hubbard model, the number of summation terms in Equation (7) grows linearly with the system size, so the computational efficiency is high.
[0067] In some embodiments, each orbital representation 304 in the at least one orbital representation 304 may include an orbital matrix. Based on this, the first neural network 303 may determine at least one determinant representing a wave function 305 of the quantum system 301 based on the plurality of electrons and the orbital matrix of each orbital representation 304 in the at least one orbital representation 304. Next, the first neural network 303 may determine a neural network quantum state of the quantum system 301 based on the at least one determinant. Subsequently, the electronic device 120 may determine the energy of the quantum system 301 based on the neural network quantum state determined by the first neural network 303.
[0068] As mentioned above, the orbital matrix can be an orbital matrix related to the backflow effect. As an example, the determinant can be a Slater determinant or other appropriate determinant. In some embodiments, for each orbital matrix, the first neural network 303 can extract effective elements from the orbital matrix based on the number of electrons in the quantum system 301. Then, the first neural network 303 can construct a square matrix for calculating the determinant based on these effective elements. As an example, the dimension of the square matrix can be determined based on the number of electrons in the quantum system 301. Figure 3 Two matrices are shown: matrices 309 - 1 determined based on track representation 304 - 1 and 309 - 2 determined based on track representation 304 - 2 . For ease of discussion, matrices 309 - 1 and 309 - 2 will be referred to individually or collectively as matrices 309 hereinafter.
[0069] In the example architecture 300, the quantum system 301 is shown as including 3 electrons. In this case, the square matrix 309 can be a 3×3 square matrix. Similarly, if the quantum system 301 includes 4 electrons, the corresponding square matrix 309 can be a 4×4 square matrix.
[0070] As mentioned above, in the example architecture 300, each row of orbital representation 304 includes eight degrees of freedom, each corresponding to one of the eight orbitals (e.g., orbitals 306-1 through 306-8). Based on this, first neural network 303 can selectively retain Z rows of representations in orbital representation 304 based on the energies of the eight orbitals and the number of electrons in quantum system 301. The value of Z can be determined based on the number of electrons in quantum system 301. For example, in the example architecture 300, quantum system 301 is shown as including three electrons. To ensure that each electron can be filled into its corresponding orbital, first neural network 303 can choose to retain three rows of representations in orbital representation 304. In some embodiments, first neural network 303 can choose to retain the three rows of representations in orbital representation 304 corresponding to the orbitals with the lowest energy (e.g., the first three rows in orbital representation 304-1). Similarly, if quantum system 301 includes four electrons, first neural network 303 can choose to retain four rows of representations in orbital representation 304 to ensure that each electron can be filled into its corresponding orbital. In some embodiments, the electronic device 120 may choose to retain the four rows of representations corresponding to the tracks with the minimum energy in the track representation 304 .
[0071] Next, the first neural network 303 can determine the orbital occupied by each electron based on the corresponding electronic configurations of the plurality of grid points 302. Then, the first neural network 303 can further filter out the degrees of freedom corresponding to the orbital occupied by the electron from the retained Z row representation. Figure 3 . In the example architecture 300, the orbital 306-2 of the grid point 302-2 is occupied by an electron, the orbital 306-5 of the grid point 302-1 is occupied by an electron, and the orbital 306-7 of the grid point 302-3 is occupied by an electron. In this case, the first neural network 303 can further filter out the degrees of freedom corresponding to the orbital 306-2 (for example, the second column in the orbital representation 304-1), the degrees of freedom corresponding to the orbital 306-5 (for example, the fifth column in the orbital representation 304-1), and the degrees of freedom corresponding to the orbital 306-7 (for example, the seventh column in the orbital representation 304-1) from the retained 3 rows of representations. In this way, the first neural network 303 can obtain the valid content of the orbital representation 304 (or the valid elements of the orbital matrix). Then, the first neural network 303 can use the calculation module 3034 to construct a 3×3 square matrix 309 for calculating the determinant based on the valid elements of the orbital matrix. After obtaining the square matrix 309 corresponding to each trajectory representation 304 , the first neural network 303 may perform corresponding determinant calculations through the calculation module 3034 to obtain the determinant of each square matrix 309 .
[0072] In some embodiments, the first neural network 303 can determine the wave function ψ corresponding to a given configuration |n> by linearly combining the calculated determinants. As an example, the linear combination of the calculated determinants can be expressed by formula (8):
[0073]
[0074] in A square matrix representing the kth orbital matrix After determining the wave function ψ, the neural network quantum state of quantum system 301 is uniquely determined. On this basis, electronic device 120 can efficiently determine the energy of quantum system 301 in this neural network quantum state using equations such as (6) and (7).
[0075] Return to reference Figure 2 After determining the energy of the quantum system 301, the electronic device 120 executes block 230. In block 230, the electronic device 120 updates the first neural network 303 by minimizing the energy. For example, the electronic device 120 updates the parameter θ of the first neural network 303.
[0076] Following the above formulas (6) and (7), updating the parameters of the first neural network 303 by minimizing the energy of the quantum system 301 can be expressed by formula (9):
[0077]
[0078] where θ * It represents the parameters of the first neural network 303 when the energy of the quantum system 301 is minimized.
[0079] In some embodiments, the first neural network 303 can be updated (or optimized) in multiple iterations to gradually approach the ground state energy of the quantum system 301. The electronic device 120 can update the first neural network 303 by any appropriate optimization algorithm. For example, the electronic device 120 can update the first neural network 303 by a stochastic reconfiguration algorithm (SR). The stochastic reconfiguration algorithm improves the standard gradient descent method by introducing information about the geometric structure of the variational manifold, thereby achieving more stable and efficient convergence performance. The objective function of the stochastic reconfiguration algorithm can be expressed by formula (10):
[0080]
[0081] where O(n) represents the normalized logarithmic wave function gradient, also known as part of the “natural gradient”, which is used to capture the statistical information of the wave function 305 changes in the parameter update direction, ∈(n) represents the residual term related to the local energy, and E L (n) represents the local energy at a given configuration |n>. As an example, O(n) can be expressed by formula (11):
[0082]
[0083] As an example, ∈(n) can be expressed by formula (12):
[0084]
[0085] where δ represents the scaling factor.
[0086] Since O(n) may be ill-conditioned, in some embodiments, regularization can be added to improve numerical stability. It should be noted that the ill-conditioning mentioned here may refer to being extremely sensitive to small perturbations in the input data, resulting in unstable or unreliable calculation results.
[0087] Alternatively, in some embodiments, the historical changes in the parameters of the first neural network 303 over multiple iterations may be introduced into the objective function (e.g., formula (10)). Subsequently, in each of the multiple iterations, the update of the parameters of the first neural network 303 may be guided based on the historical changes in the parameters. For example, in some embodiments, in a current iteration of the multiple iterations, the electronic device 120 may determine the objective function for the current iteration (e.g., the first objective function) based on at least one reference change amount related to the change in the parameters of the first neural network 303 in at least one iteration prior to the current iteration. Subsequently, the electronic device 120 may update the parameters of the first neural network 303 based on the first objective function.
[0088] By introducing a reference variation into the first objective function, electronic device 120 can adaptively transform and scale each parameter in first neural network 303 based on the historical variation of the parameters of first neural network 303 over multiple iterations. In this way, historical information about parameter variations in first neural network 303 during the optimization process can be taken into account, allowing for dynamic adjustment of the parameter update process. This can accelerate the convergence of the parameters of first neural network 303 toward the optimal parameters.
[0089] In some embodiments, the at least one reference variation may include a first reference variation and a second reference variation. The first reference variation may indicate a direction of change in a parameter of the first neural network 303 in at least one iteration prior to the current iteration. The second reference variation may indicate a magnitude of change in a parameter of the first neural network 303 in at least one iteration prior to the current iteration.
[0090] As an example, the first reference variation may be a first-order moment estimation related to the parameter variation of the first neural network 303, and the second reference variation may be a second-order moment estimation or a higher-order moment estimation related to the parameter variation of the first neural network 303. In this case, the first objective function may be expressed by formula (13):
[0091]
[0092] Among them, φ k =μdθ k is a first-order moment estimate, representing a momentum term related to the direction of parameter change of the first neural network 303, V k =βV k-1 +(dθ k -dθ k-1 ) 2 is a second-order moment estimator, representing a momentum term related to the historical magnitude of parameter changes in the first neural network 303, and μ, β, and λ are hyperparameters. k-1 ) 1 / 4 represents the diagonal matrix constructed by raising the second-order moment estimation vector to the fourth power, which is used to adaptively scale the parameter updates element by element. Therefore, formula (13) can be further rewritten as formula (14):
[0093] dθ=diag(V) -1 / 2 O T (Odiag(V) -1 / 2 O T +λI) -1 (∈-Oφ)+φ (14)
[0094] Where diag(V k-1 ) -1 / 2 represents the inverse square root of the diagonal matrix, which is used to normalize the gradient. I represents the identity matrix, which ensures numerical stability during the matrix inversion process. ∈-Oφ represents the energy residual after considering the momentum term. In this way, it is equivalent to combining momentum and smoothing mechanisms in the direction of the natural gradient, thereby better accelerating the convergence of the parameters of the first neural network 303 to the optimal parameters. Such a parameter update algorithm can also be called a moment adaptive reconstruction solver (MARCH).
[0095] In some embodiments, before updating the parameters of the first neural network 303 in the above manner, the electronic device 120 may use random initialization to determine the initial parameters of the first neural network 303. Alternatively, in some embodiments, the electronic device 120 may pre-train the first neural network 303 based on the second neural network, thereby providing better initial parameters for the first neural network 303. Specifically, the first neural network 303 may be pre-trained in the following manner. The electronic device 120 may generate at least one orbital representation sample (e.g., a first orbital representation sample) based on the electronic configuration samples (e.g., the first electronic configuration samples) of the plurality of grid points 302 using a second neural network having a multilayer perceptron (MLP). The electronic device 120 may pre-train the first neural network 303 based on the first electronic configuration samples of the plurality of grid points 302 and the at least one first orbital representation sample.
[0096] The second neural network can be a neural network based on a neural network backflow (NNB) architecture. In the second neural network, the multilayer perceptron can have two or more hidden layers. As an example, the multilayer perceptron can process the first electron configuration sample, such as the electron occupation number n∈{0,1} 2N Mapping to the first track representation sample, such as the track matrix Orbital matrix M nnb (n) can be restricted to a 2×2 block-diagonal matrix, as shown in formula (15):
[0097]
[0098] in Represents the orbital matrix M nnb (n) in relation to spin-up orbitals, Represents the orbital matrix M nnb (n) related to spin-down orbitals, N e represents the total number of electrons in the quantum system 301. Then, based on the current occupancy number, the second neural network can obtain the orbital matrix M nnb Extract N from (n) e Row elements, forming a square matrix In some embodiments, the square matrix Φ may also be maintained in the form of 2×2 block diagonals. The square matrix Φ can be expressed by formula (16):
[0099]
[0100] Where Φ↑ represents the content of the matrix Φ related to spin-up orbitals, and Φ↓ represents the content of the matrix Φ related to spin-down orbitals. Next, based on the resulting matrix Φ, the second neural network can determine the corresponding Slater determinant. As mentioned above, the Slater determinant can be used to represent the wave function of quantum system 301.
[0101] In some embodiments, the determinant of the square matrix Φ can be used as the amplitude of the second neural network wave function 305. This process can be expressed by formula (17):
[0102] ψnnb(n))=det[Φ]; (17)
[0103] where ψ nnb (n) represents the amplitude of the second neural network wave function 305.
[0104] Next, the electronic device 120 can use the second neural network to determine a training sample set with labels. Each sample in the training sample set contains an occupation number n i and its corresponding orbital matrix M nnb (n i ). Training sample set S pre It can be expressed by formula (18)
[0105]
[0106] where N pre is the total number of samples in the training sample set.
[0107] Then, the electronic device 120 can pre-train the first neural network 303 in a supervised manner based on the training sample set. During the pre-training process, the loss function can be determined based on the mean square error of the trajectory matrix. As an example, the loss function can be expressed by formula (19):
[0108]
[0109] in represents the loss function of pre-training, M Transformer (n i ) represents the first neural network 303 according to the occupation number n i Generated trajectory matrix. After completing pre-training of the first neural network 303, the electronic device 120 can update and optimize the parameters of the first neural network 303 using the methods mentioned above. In some embodiments, the optimization of the parameters of the first neural network 303 can use an adaptive moment estimation optimizer (Adam), etc. Compared with using random initialization to determine the parameters of the first neural network 303, the pre-trained first neural network 303 is more stable and converges faster.
[0110] In some embodiments, before updating the parameters of the first neural network 303 in the above manner, the electronic device 120 may perform another type of pre-training on the first neural network 303 (which may be used alone or in combination with the pre-training mentioned above), and temporarily introduce a pinning energy term during the pre-training process. This can prevent the optimization process of the first neural network 303 from falling into a local minimum.
[0111] Specifically, first neural network 303 may be pre-trained in the following manner. Electronic device 120 may train a third neural network based on an attention mechanism using a second objective function based on a pinning energy term. The pinning energy term is related to electron spin. Subsequently, electronic device 120 may pre-train first neural network 303 based on the second objective function with the pinning energy term removed and the trained third neural network. The similarity in output between the pre-trained first neural network 303 and the trained third neural network meets a predetermined similarity requirement.
[0112] In some embodiments, the second objective function may be an energy operator (e.g., Hamiltonian operator) that introduces a pinning energy term. ). The pinning energy term can be an antiferromagnetic term. The Hamiltonian operator for the pinning energy term is introduced It can be expressed by formula (20):
[0113]
[0114] where h m = 0.2 or other values, indicating the strength of the pinning energy term, i∈{1,L x } represents the pinning energy term acting on the first and last rows (along the x direction) of the spatial grid 302, j∈{1,L y } means covering all columns in the y direction, (-1) i+j represents the coefficients that produce a checkerboard-like alternating sign, simulating an antiferromagnetic arrangement, The Pauli operator for the spin at position (i, j) in the z direction, used to describe the spin-up or spin-down state.
[0115] In some embodiments, electronic device 120 may train a third neural network under the influence of a pinning energy term. The structure of the third neural network can refer to that of the first neural network 303 and will not be further described here. This approach facilitates convergence of the wave function to a state close to the desired mode (such as an antiferromagnetic state, a stripe state, etc.). Subsequently, electronic device 120 may initialize first neural network 303 based on the third neural network. During this process, electronic device 120 may remove the influence of the pinning energy term. For example, electronic device 120 may use a Hamiltonian operator without a pinning energy term. The goal of pre-training may be to maximize the similarity (or fidelity) between the outputs of first neural network 303 and third neural network. This process can effectively "transfer" the desired physical state from the third neural network to the first neural network 303. Next, electronic device 120 may update and optimize the parameters of first neural network 303 using the aforementioned method to enable free relaxation, thereby finding the true ground state of quantum system 301.
[0116] In some embodiments, the electronic device 120 can solve the ground state of the quantum system 301 by projecting the wave function ψT determined by the first neural network 303 onto the ground state wave function ψ0 of the quantum system 301 based on the Green's function Monte Carlo algorithm. This process can be achieved by applying the imaginary time evolution operator e -τH In some embodiments, the imaginary time evolution can be iteratively implemented using a short-time approximation of the Green's function. This process can be expressed as a projection operator T = 1-τ(HE), where H represents the Hamiltonian operator, E represents the energy offset term, and τ is the time step.
[0117] However, when the Hamiltonian operator has negative off-diagonal matrix elements, negative values will appear in the projection operator T, thereby destroying the physical meaning of the wave function amplitude as a probability distribution. This is also known as the "sign problem". This makes Monte Carlo sampling difficult. In some embodiments, a fixed node approximation method can be used to solve the sign problem. It can force all weights to be positive, thereby allowing effective Monte Carlo sampling. In this case, the electronic device 120 can construct an effective Hamiltonian H eff Specifically, for any two given configurations R and R′, the electronic device 120 first determines whether there is a sign flip transition (SF), that is, whether the transition satisfies the following formula (21):
[0118] 〈R|H"R′〉ψ T (R)ψ T (R′)>0 (21)
[0119] If the product is greater than zero, it means<R|H|R'> The sign of the wave function 305 is inconsistent with the sign of the wave function 305 in the two configurations, which may cause a sign problem. Such a transition can be called an SF transition. The electronic device 120 will eff Based on this, the effective Hamiltonian H eff It can be expressed by formula (22):
[0120]
[0121] The off-diagonal elements corresponding to the sign-flip transitions are removed. To compensate, in some embodiments, the off-diagonal elements can be compensated by adding terms derived from these deleted sign-flip transitions.<R|H|R> To improve sampling efficiency, in some embodiments, an importance sampling strategy based on the wave function ψT can be adopted. Under this sampling strategy, the electronic device 120 no longer directly simulates the virtual-real evolution controlled by the projection operator T = 1-τ(HE), but considers a new distribution: f(R, t) = ψ(R, t) ψT(R). The corresponding importance sampling transfer matrix is It can be expressed as: Its explicit form can be expressed by formula (23):
[0122]
[0123] By using the fixed node approximation, we can ensure that the effective Hamiltonian H eff The energy obtained is an upper bound on the true ground state energy E0. In some embodiments, for an operator A that commutes with the Hamiltonian H (the two operators satisfy the commutative law), the mixed expectation value mixed It can be expressed by formula (24):
[0124]
[0125] In some embodiments, the electronic device 120 may also obtain a more accurate pure expected value (pure estimator) by future-walking sampling. pure It can be expressed by formula (25):
[0126]
[0127] in It can be expressed by formula (26):
[0128]
[0129] Where W(R) is the total weight of all descendants evolved from configuration R. Pure expected value pure can be estimated directly by Monte Carlo sampling. Therefore, the pure expected value pure It can also be expressed by formula (27):
[0130]
[0131] In some embodiments, the electronic device 120 can accelerate the training process of the first neural network 303 through mixed precision training. Mixed precision training accelerates the training process and reduces memory consumption by utilizing different numerical precision formats without significantly affecting model accuracy. In some embodiments, the electronic device 120 can assign different numerical precisions to different modules of the first neural network 303 to balance computational efficiency and numerical stability. For example, in order to ensure the stability of random reconstruction, double-precision floating point numbers (FP64) can be used for these calculations. For attention score calculation QK T , can be computed in single-precision floating point (FP32). In addition, all other matrix multiplications can be performed in single-precision floating point (FP32), and so on.
[0132] In some embodiments, the attention mechanism of the first neural network 303 can employ a flash attention mechanism. The flash attention mechanism can overcome the computational and memory limitations of traditional attention mechanisms, significantly improving the efficiency of the first neural network 303 in processing long sequences. Furthermore, by optimizing memory access patterns, the flash attention mechanism significantly reduces the number of read and write operations between the high-bandwidth memory of the processor and the on-chip static random access memory.
[0133] As can be clearly understood from the various embodiments described above, the embodiments of the present disclosure utilize a first neural network 303 based on an attention mechanism to achieve accurate analysis of a larger and more complex quantum system 301. Furthermore, during the optimization process of the first neural network 303, the embodiments of the present disclosure introduce a reference variation in the objective function that can indicate historical parameter changes. This approach enables faster and more stable convergence in large and complex quantum systems 301.
[0134] Through improvements in neural networks and optimization algorithms, the embodiments of the present disclosure demonstrate higher accuracy and robustness in quantum state variational optimization tasks. In particular, the embodiments of the present disclosure can effectively handle quantum system 301 models that include periodic boundary conditions and next-nearest-neighbor hopping. This type of configuration is of great significance for realistically simulating the particle-hole asymmetry observed in copper oxide high-temperature superconductors. In addition, since the neural network quantum state (NQS) form adopted in the embodiments of the present disclosure is insensitive to boundary conditions, it can be flexibly applied to modeling quantum systems 301 under a variety of boundary conditions, thereby further expanding its scope of application and generalization capabilities.
[0135] The embodiments of the present disclosure also provide corresponding devices for implementing the above methods or processes. Figure 4 FIG4 is a schematic block diagram of an apparatus 400 for determining a quantum system state representation according to some embodiments of the present disclosure. The apparatus 400 may be implemented as or included in the electronic device 120. Each module / component in the apparatus 400 may be implemented by hardware, software, firmware, or any combination thereof.
[0136] Reference Figure 4 Apparatus 400 includes an orbital representation generation module 410, an energy determination module 420, and an update module 430. Orbital representation generation module 410 is configured to generate at least one orbital representation for a plurality of electrons based on corresponding electron configurations of a plurality of lattice points in a quantum system comprising the plurality of electrons, using a first neural network based on an attention mechanism. Energy determination module 420 is configured to determine the energy of the quantum system based on the at least one orbital representation. Update module 430 is configured to update the first neural network by minimizing the energy.
[0137] In some embodiments, the orbital representation generation module 410 is further configured to: generate corresponding embedding representations of the plurality of lattice points based on the corresponding electronic configurations of the plurality of lattice points using an embedding layer of the first neural network; and generate at least one orbital representation based on the corresponding embedding representations of the plurality of lattice points using at least one attention layer of the first neural network.
[0138] In some embodiments, the track representation generation module 410 is further configured to: generate corresponding feature representations of the multiple grid points based on the corresponding embedding representations of the multiple grid points, using at least one attention layer, wherein the corresponding feature representations of the multiple grid points indicate the influence of the multiple grid points on each other; and convert the corresponding feature representations of the multiple grid points into at least one track representation using a linear layer of the first neural network.
[0139] In some embodiments, the first neural network is updated in multiple iterations, and the updating module 430 is further configured to: determine, in a current iteration of the multiple iterations, a first objective function for the current iteration based on at least one reference change amount related to a change in parameters of the first neural network in at least one iteration before the current iteration; and update the parameters of the first neural network based on the first objective function.
[0140] In some embodiments, the at least one reference variation includes a first reference variation and a second reference variation, the first reference variation indicating a direction of change of a parameter of the first neural network in at least one iteration prior to a current iteration, and the second reference variation indicating a magnitude of change of a parameter of the first neural network in at least one iteration prior to the current iteration.
[0141] In some embodiments, a grid point among the plurality of grid points is associated with at least one orbital, and for each orbital representation in the at least one orbital representation, the orbital representation indicates a distribution of the plurality of electrons in the orbitals associated with the plurality of grid points and a correlation relationship between each electron and other electrons in the plurality of electrons.
[0142] In some embodiments, the energy determination module 420 is further configured to: determine at least one determinant for representing a wave function of the quantum system based on the plurality of electrons and an orbital matrix of each orbital representation in the at least one orbital representation; determine a neural network quantum state of the quantum system based on the at least one determinant; and determine the energy of the quantum system based on the neural network quantum state.
[0143] In some embodiments, the first neural network is a pretrained neural network, and the first neural network is pretrained by at least the following means: generating at least one first orbital representation sample based on the first electronic configuration samples of the plurality of lattice points using a second neural network having a multilayer perceptron; and pretraining the first neural network based on the first electronic configuration samples of the plurality of lattice points and the at least one first orbital representation sample.
[0144] In some embodiments, the first neural network is a pretrained neural network, and the first neural network is pretrained in at least the following manner: using a second objective function based on a pinning energy term to train a third neural network based on an attention mechanism, wherein the pinning energy term is related to electron spin; and pretraining the first neural network based on the second objective function with the pinning energy term removed and the trained third neural network, wherein the similarity in output between the pretrained first neural network and the trained third neural network meets a predetermined similarity requirement.
[0145] Figure 5 1 is a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. The electronic device 500 may be used to implement, for example, Figure 1 The electronic device 120 shown or Figure 4 The device 400 shown. It should be understood that Figure 5 The illustrated electronic device 500 is merely exemplary and should not be construed as limiting the functionality and scope of the embodiments described herein.
[0146] Reference Figure 5 , electronic device 500 is in the form of a general electronic device. Components of electronic device 500 may include, but are not limited to, one or more processors 510, memory 520, storage device 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560. Processor 510 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 520. In a multi-processor system, multiple processors execute computer-executable instructions in parallel to increase the parallel processing capabilities of electronic device 500.
[0147] The electronic device 500 typically includes a plurality of computer storage media. Such media can be any available media accessible to the electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 520 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 530 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 500.
[0148] The electronic device 500 may further include additional removable / non-removable, volatile / non-volatile storage media. Figure 5 As shown in FIG, a magnetic disk drive for reading from or writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. Memory 520 may include a computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0149] The communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 500 can be implemented in a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the electronic device 500 can operate in a networked environment using a logical connection with one or more other servers, a network personal computer (PC), or another network node.
[0150] Input device 550 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 560 may be one or more output devices, such as a display, a speaker, or a printer. Electronic device 500 may also communicate with one or more external devices (not shown) via communication unit 540 as needed, such as a storage device, a display device, or the like, with one or more devices that allow a user to interact with electronic device 500, or with any device that allows electronic device 500 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0151] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0152] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0153] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine such that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0154] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0155] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0156] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is intended to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.< / ij> < / ij>
Claims
1. A method for determining a representation of a state of a quantum system, comprising: Based on corresponding electronic configurations of a plurality of lattice points in a quantum system including a plurality of electrons, generating at least one orbital representation for the plurality of electrons using a first neural network based on an attention mechanism; determining an energy of the quantum system based on the at least one orbital representation; as well as The first neural network is updated by minimizing the energy.
2. The method of claim 1 , wherein generating at least one orbital representation for the plurality of electrons comprises: generating, using an embedding layer of the first neural network, corresponding embedded representations of the plurality of lattice points based on the corresponding electronic configurations of the plurality of lattice points; as well as The at least one trajectory representation is generated using at least one attention layer of the first neural network based on the corresponding embedded representations of the plurality of lattice points.
3. The method of claim 2, wherein generating the at least one track representation using at least one attention layer of the first neural network comprises: generating, using the at least one attention layer, corresponding feature representations of the plurality of lattice points based on the corresponding embedded representations of the plurality of lattice points, wherein the corresponding feature representations of the plurality of lattice points indicate influences between the plurality of lattice points; as well as The linear layer of the first neural network is used to convert the corresponding feature representations of the plurality of grid points into the at least one track representation.
4. The method of claim 1 , wherein the first neural network is updated in a plurality of iterations, and updating the first neural network comprises: In a current iteration of the plurality of iterations, determining a first objective function for the current iteration based on at least one reference change amount related to a change in a parameter of the first neural network in at least one iteration before the current iteration; as well as Based on the first objective function, the parameters of the first neural network are updated.
5. The method according to claim 4, wherein the at least one reference variation comprises a first reference variation and a second reference variation, the first reference variation indicating a direction of change of a parameter of the first neural network in at least one iteration before a current iteration, and the second reference variation indicating a magnitude of change of a parameter of the first neural network in at least one iteration before the current iteration.
6. The method according to claim 1, wherein a grid point among the plurality of grid points is associated with at least one orbital, and for each orbital representation in the at least one orbital representation, the orbital representation indicates the distribution of the plurality of electrons in the orbitals associated with the plurality of grid points and the association relationship between each electron and other electrons among the plurality of electrons.
7. The method of claim 1 , wherein an orbital representation in the at least one orbital representation comprises an orbital matrix, and determining the energy of the quantum system comprises: determining at least one determinant for a wave function representing the quantum system based on the plurality of electrons and an orbital matrix for each of the at least one orbital representation; determining a neural network quantum state of the quantum system based on the at least one determinant; as well as Based on the neural network quantum state, an energy of the quantum system is determined.
8. The method of claim 1 , wherein the first neural network is a pre-trained neural network, and the first neural network is pre-trained by at least: generating at least one first orbital representation sample based on the first electronic configuration samples of the plurality of lattice points using a second neural network having a multilayer perceptron; and The first neural network is pre-trained based on the first electronic configuration samples of the plurality of lattice points and the at least one first orbital representation sample.
9. The method of claim 1 , wherein the first neural network is a pre-trained neural network, and the first neural network is pre-trained at least by: training a third neural network based on an attention mechanism using a second objective function based on a pinning energy term, wherein the pinning energy term is related to electron spin; and The first neural network is pre-trained based on the second objective function with the pinning energy term removed and the trained third neural network, wherein similarity in output between the pre-trained first neural network and the trained third neural network meets a similarity requirement.
10. An apparatus for determining a representation of a state of a quantum system, comprising: An orbital representation generating module is configured to generate at least one orbital representation for the plurality of electrons based on corresponding electron configurations of a plurality of lattice points in a quantum system including the plurality of electrons and using a first neural network based on an attention mechanism; an energy determination module configured to determine an energy of the quantum system based on the at least one orbital representation; as well as An updating module is configured to update the first neural network by minimizing the energy.
11. An electronic device comprising: at least one processor; as well as At least one memory is coupled to the at least one processor and stores instructions for execution by the at least one processor, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processor.
12. A computer-readable storage medium having computer-executable instructions stored thereon, wherein the computer-executable instructions can be executed by a processor to implement the method according to any one of claims 1 to 9.
13. A computer program product comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method according to any one of claims 1 to 9.
Citation Information
Cited By
Lattice Fermi processing method and electronic equipment
CN121352056A