Battery electron level large-scale generalized eigenvalue problem solving method and system
By adopting the hybrid accuracy and load balancing methods in the calculation of battery electronic energy level, the problem of low calculation efficiency of large-scale generalized eigenvalue problems is solved, and more efficient battery material evaluation is achieved, and the accuracy of battery performance prediction is improved.
Patent Information
- Application Number
- CN202510057258.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art is difficult to efficiently solve the problem of large-scale generalized eigenvalues, especially in battery electronic energy level calculations. Traditional algorithms have large amounts of calculation and high storage, and cannot fully utilize matrix characteristics, which affects the stability and electrochemical activity evaluation of battery materials.
An efficient solution method based on mixing accuracy and load balancing is adopted. By constructing the Hamiltonian matrix and overlapping matrix of battery material molecules, the integrated interval of the surrounding channel is demarcated, and parallel solution is performed in the low-precision stage. Cluster analysis is used to optimize load balancing and improve algorithm performance.
It significantly improves the calculation speed and solution performance of large-scale generalized eigenvalue problems of battery electronic energy levels, improves the efficiency of battery material calculation, and can more accurately evaluate the stability and electrochemical activity of battery material.
Smart Images

Figure CN120104937A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of batteries and relates to a battery electronic energy level calculation technology, and in particular to a method and system for solving a large-scale generalized eigenvalue problem of a battery electronic energy level. Background Art
[0002] The calculation of battery electronic energy levels usually involves a quantum mechanical description of the electronic state within the battery material. This calculation usually uses methods such as density functional theory (DFT) to predict the stability and electrochemical activity of battery materials by calculating the electronic structure. Electronic energy levels, especially the highest occupied molecular orbital energy level (HOMO) and the lowest unoccupied molecular orbital energy level (LUMO), are crucial to understanding the electrochemical reaction mechanism of the battery. The calculation of electronic energy levels in density functional theory (DFT) has a wide range of applications in materials science, especially in the fields of battery materials, semiconductors, catalysts, etc., for predicting and optimizing the physical and chemical properties of materials. By calculating the electronic structure and energy level distribution of electrode materials, DFT can evaluate its key properties such as reactivity and structural stability during the charging and discharging process, thereby screening out battery materials with high stability and long life. The core of DFT is to obtain the ground state energy and electronic structure information by calculating the electron density of the system. Its calculation method converts complex multi-electron problems into single-electron problems through the Kohn-Sham equation, so that the energy level distribution and spatial distribution of electrons can be simplified. This process requires solving a large-scale Hamiltonian matrix, whose eigenvalues represent the energy levels of different electronic states, and whose eigenvectors describe the spatial distribution of electrons. Solving the matrix eigenvalues is an indispensable step in DFT. By obtaining these eigenvalues and eigenvectors, we can gain a deep understanding of the electronic properties of materials, such as conductivity, band gap, and magnetism, thereby providing theoretical support for the structural stability evaluation and performance optimization of materials. The matrix eigenvalue problem is a classic problem in higher mathematics. At the same time, it is also a difficult problem in computational mathematics and is a very active research topic in the rapidly developing fields of computer science and numerical algebra. For small and medium-scale eigenvalue problems, after long-term and unremitting efforts by a large number of scientific researchers, a variety of effective methods such as the Jacobi algorithm, the power method, and the QR algorithm have emerged, and their solutions have basically been satisfactorily solved. However, for large-scale generalized eigenvalue problems evolved from practical problems, matrices often have the characteristics of large size, sparseness, block structure, symmetry, and positive definiteness. Generally, only a few eigenpairs need to be solved. Many traditional eigenvalue solving algorithms cannot make full use of these characteristics and are limited by large computational complexity and high storage capacity. Therefore, further research and design of efficient and stable algorithms for solving large-scale matrix generalized eigenvalue problems has high theoretical and practical significance.
[0003] FEAST algorithm (see reference: Polizzi, Eric. "Density-matrix-based algorithm forsolving eigenvalue problems." Physical Review B—Condensed Matter and Materials Physics 79.11 (2009): 115112.) is a type of algorithm that draws inspiration from the contour integral and density matrix representation in quantum mechanics and can efficiently solve generalized eigenvalue problems. Unlike the traditional Krylov subspace-based iteration method, it uses the contour integral as a projection operator to construct an invariant subspace, which is repeatedly applied in the iterative process and eventually shifts the search space to the invariant subspace containing the required eigenvectors, which greatly improves the performance of solving such problems.
[0004] In recent years, mixed precision has received more and more attention. The main advantage of using mixed precision is that compared with the default working precision (double precision), executing the algorithm at low precision can achieve significant acceleration. However, for some critical applications, the accuracy of the calculated eigenvalues and eigenvectors guaranteed by the working precision may be required. In this case, it is natural to consider using mixed precision algorithms. The FEAST algorithm uses a simple mixed precision design to achieve a maximum speedup of 1.5 times. As an algorithm for a class of internal eigenvalue problems, the FEAST algorithm needs to give the interval of the desired eigenvalue. For large-scale generalized eigenvalue problems, the search interval spans a large range and contains a large number of eigenvalues. If the large search interval can be divided into multiple smaller search intervals for parallel execution, the performance of the algorithm can be effectively improved. FEAST implements a three-level parallel system, in which Level 1 parallelism is used to divide the large search interval into multiple small search intervals for parallel execution. The resulting division of the range of the desired solution interval into different contour integral intervals has an important impact on the performance of the algorithm. If the small search interval after the division is large, the number of eigenvalues in the interval will still remain at a high level, and the parallelism of the FEAST algorithm is not fully utilized; if the small search interval after the division is small, the number of eigenvalues contained in it is small, which is not conducive to the subsequent selection of the size of the iterative subspace, affecting the convergence speed of the algorithm. In addition, the division of small search intervals cannot be determined solely based on the number of eigenvalues. This is because the eigenvalues of the matrix are clustered. If the eigenvalues that originally belong to the same eigenvalue cluster are split into two different search intervals for solution, not only will this potential space for improving the algorithm's performance be ignored, but errors may also occur in the calculation of eigenvalues in the boundary area of the two search intervals, thereby affecting the calculation results, which is intolerable.
[0005] To solve this problem, a common solution is to perform a pre-calculation phase before executing the algorithm to estimate the distribution of eigenvalues in advance. Existing methods include: using Sylvester's inertia theorem, performing two LDLT decompositions on the matrix to obtain the number of eigenvalues in the search area (see Johnson, Charles R., and Susana Furtado. "A generalization of Sylvester's law of inertia." Linear Algebra and its Applications 338.1-3 (2001): 287-290.); using unbiased matrix trace estimation and contour integral to stochastically estimate the eigenvalue density for nonlinear eigenvalue problem of analytic matrix function to find the distribution (see Maeda, Yasuyuki, Yasunori Futamura, and Tetsuya Sakurai. "Stochastic estimation method of eigenvalue density for nonlinear eigenvalue problem on the complex plane." JSIAM letters 3 (2011): 61-64.); Stochastic estimation methods for approximating the trace of the eigenprojection operator using polynomial or rational function expansions (see Di Napoli, Edoardo, Eric Polizzi, and YousefSaad. "Efficient estimation of eigenvalue counts in an interval." Numerical Linear Algebra with Applications 23.4 (2016): 674-692.). However, this series of eigenvalue estimation methods have many problems, such as high overhead and inaccurate eigenvalue estimation. Although LDLT decomposition can accurately estimate the number of eigenvalues in an interval, the cost of performing two LDLT decompositions on a large-scale matrix is high and unaffordable. Although using the trace of the matrix to estimate can get a more reasonable estimate at a lower cost, there are always the following three problems: first, the larger the search interval, the lower the accuracy. Second, the denser the eigenvalue distribution, the more difficult it is to estimate, especially for the case of clustered eigenvalue distribution. Finally, the accuracy of the estimation of eigenvalues at the edge of the interval is not high. Unreasonable estimation will lead to unbalanced workload when each sub-search interval is executed in parallel, thus affecting the overall performance of the algorithm. Summary of the invention
[0006] Technical problem to be solved by the present invention: In view of the above-mentioned problems in the prior art, a method and system for solving the large-scale generalized eigenvalue problem of battery electronic energy level are provided. The present invention aims to improve the calculation speed and solution performance of solving the large-scale generalized eigenvalue problem of battery electronic energy level, and improve the efficiency of battery electronic energy level calculation.
[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: A method for solving a large-scale generalized eigenvalue problem of battery electronic energy levels comprises the following steps: S1, construct the Hamiltonian matrix of battery material molecules and the overlap matrix , where the Hamiltonian matrix is the total energy between battery material molecules, and the overlap matrix The overlap between the molecular orbitals of the battery material is shown in the figure. According to the Hamiltonian matrix of the battery material molecule and the overlap matrix Construct the generalized eigenvalue problem shown below: , In the above formula, is an eigenvector, which is used to represent the column vector of the electron wave function. Each component in the eigenvector corresponds to the coefficient of different molecular orbital basis functions in the system, which is used to represent the state of the electron in the system. is the eigenvalue, which is used to represent the electronic energy level of the battery material; the interval where the electronic energy level is located is used as the search interval for the generalized eigenvalue problem ,in are the left and right boundaries of the search interval respectively; S2, delineate the contour integration and determine the number of integration nodes and the number of process resources; S3, the Hamiltonian matrix and the overlap matrix Copy and generate a low-precision version, and perform matrix decomposition on the matrix form of the low-precision version to adapt to the format requirements of solving the linear equation system in the subsequent stage; S4, based on search interval Construct a linear equation system in parallel on all integration nodes and solve it at a low-precision scale, and use the solution results to construct an iterative subspace Q; S5, Rayleigh–Ritz iterative update using iterative subspace Q at low precision scale; S6, Judgment Whether to set the convergence accuracy of the low-precision stage , if the convergence accuracy of the set low-precision stage is reached , then jump to step S7; otherwise, construct the right-hand side of a new linear equation system; S7, cluster analysis of the feature pairs obtained in the low-precision stage; S8, according to the cluster analysis results, the search interval Divide into multiple sub-search intervals ,in are the left and right boundaries of the ith sub-search interval, respectively. Containing one or more clustered eigenvalues, and the number of eigenvalues to be solved in each sub-search interval is equal or approximately equal to achieve load balancing, and the process resources are divided into groups according to the number of sub-search intervals, so that the number of process resources in each group is equal, and each group of process resources is used to calculate one sub-search interval; S9, each sub-search interval re-defines the contour integral and determines the number of integral nodes and process resources. and the overlap matrix Decompose the matrix in each subspace; S10, based on search interval , the linear equations of the generalized eigenvalue problem under low-precision scale are solved in parallel on all integration nodes, and the iterative subspace Q is constructed using the calculation results; S11, Rayleigh–Ritz iterative update using iterative subspace Q at high precision scale; S12, judgment Whether the convergence accuracy of the set high-precision stage is achieved , if the convergence accuracy of the set high-precision stage is reached , then jump to step S13; otherwise, construct the right-hand side of a new linear equation system and jump to step S10; S13, the main process summarizes the calculation results of each sub-search interval and outputs the final calculation results, including: battery materials in the search interval The eigenvalues within The electronic energy levels of the battery materials represented, as well as the eigenvectors The electron wave function represented by .
[0008] Optionally, the contour integral delimited in step S2 includes: As the diameter, make a semicircle on the complex plane, and select the best sample points and weights according to the Gauss-Legendre integral to maximize the accuracy of the integral. ,in It represents the value of the sample point. It represents the weight of the sample point, and the selected sample point is used as the integration node of the contour integral.
[0009] Optionally, when performing matrix decomposition on the matrix form of the low-precision version in step S3, if the matrix form of the low-precision version is a dense matrix, the corresponding linear system solver function in the LAPACK library is used to decompose the matrix form of the low-precision version; if the matrix form of the low-precision version is a sparse matrix, the corresponding linear system solver function in the MKL-PARDISO library is used to decompose the matrix form of the low-precision version; if the matrix form of the low-precision version is a banded matrix, the corresponding linear system solver function in the SPIKE library is used to decompose the matrix form of the low-precision version.
[0010] Optionally, step S4 includes: S4.1, generate a random orthogonal normalized matrix ; S4.2, based on search interval This constructs the linear system in parallel at all integration nodes: , in, For the The value of the integration node, is the overlap matrix, is the Hamiltonian matrix, For the The solution to be solved for the integration node, j is the serial number of the integration node, is a random orthogonal normalized matrix; S4.3, solve the linear system equations in parallel at each integration node to obtain the corresponding solution ; S4.4, construct the iterative subspace Q according to the following formula: , in, For the The weight of the integration node.
[0011] Optionally, the Rayleigh–Ritz iterative update using the iterative subspace Q includes: using the iterative subspace Q as a projection subspace, and transforming the generalized eigenvalue problem Projection into a simplified new generalized eigenvalue problem ,in is the new overlap matrix, is the new Hamiltonian matrix, is the new feature vector, is the new eigenvalue, and , ; Call the generalized eigenvalue solving function in the LAPACK library to solve the new generalized eigenvalue problem , the solution is completed and the new eigenvalue is obtained and the new eigenvector ; The new eigenvalue and the new eigenvector according to , = Transformed into the original generalized eigenvalue problem The eigenvalue of and the eigenvector .
[0012] Optionally, constructing the right-hand side of the new linear equation system in step S6 includes: using the solved multiple sets of eigenvectors Concatenate into a matrix , the matrix Alternative random orthogonal normalized matrix Get the right side of the new linear equation system .
[0013] Optionally, when determining the number of integration nodes and the number of process resources in step S2, the number of process resources is positively correlated with the performance of the computing device, and when redefining the contour integral and determining the number of integration nodes and the number of process resources in each sub-search interval in step S9, the number of process resources is specified within the number of process resources of the process resource group corresponding to the sub-search interval.
[0014] In addition, the present invention also provides a system for solving the large-scale generalized eigenvalue problem of the battery electronic energy level, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of the battery electronic energy level.
[0015] In addition, the present invention also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level through a processor.
[0016] In addition, the present invention also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level through a processor.
[0017] Compared with the prior art, the present invention has the following advantages: the method for solving the large-scale generalized eigenvalue problem of the battery electronic energy level of the present invention includes: and the overlap matrix Constructing a generalized eigenvalue problem based on the search interval The linear equations are constructed in parallel on all the integral nodes and solved at a low-precision scale, and the iterative subspace Q is constructed using the solution results; then the iterative subspace Q is used to perform Rayleigh–Ritz iterative updates to construct the right-hand side of the new linear equations at low-precision scales and high-precision scales in turn until the accuracy at low-precision scales and high-precision scales meets the requirements, and finally the calculation results of each sub-search interval are summarized to obtain the electronic energy level and electronic wave function of the battery material, which mainly has the following advantages: 1. The present invention further utilizes the mixed precision space existing in the original FEAST algorithm to construct a pure low-precision stage, thereby improving the calculation speed of the algorithm. 2. The present invention uses the calculation results of the low-precision stage to perform cluster analysis to find the eigenvalue clusters contained in the interval, and allocates the corresponding number of processes according to the number of eigenvalues in the cluster, while effectively avoiding the errors caused by unreasonable eigenvalue interval division, and realizes the load balancing of the FEAST parallel structure, thereby further improving the performance of the algorithm. 3. The idea of the present invention can be extended to various eigenvalue problems based on contour integration, and faster calculation speeds are achieved in various situations such as banded, sparse, and dense. In summary, the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level of the present invention can improve the calculation speed and solution performance of solving the large-scale generalized eigenvalue problem of battery electronic energy level, and improve the efficiency of battery electronic energy level calculation. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 The figure is a flow chart of the detailed process of the method according to the embodiment of the present invention.
[0019] Figure 2 This is the three-layer parallel structure of the FEAST algorithm used in the embodiment of the present invention. DETAILED DESCRIPTION
[0020] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0021] The method for solving the large-scale generalized eigenvalue problem of the electronic energy level of the battery of the present invention proposes a method for efficiently solving the large-scale generalized eigenvalue problem of the electronic energy level of the battery based on mixed precision and load balancing. The core of the method is to deeply optimize and expand the traditional FEAST algorithm, and make full use of the undeveloped mixed precision computing potential therein, thereby significantly improving the overall computing performance while maintaining the computing accuracy. Specifically, the algorithm first introduces the calculation process of the low-precision stage, while accelerating the execution of the algorithm, effectively preserving the low-precision calculation results, and laying a data foundation for subsequent high-precision calculations. In order to further improve the parallel processing capability of the algorithm, this embodiment adopts a clustering analysis method to divide the eigenvalue space into multiple subspaces with cluster distribution characteristics. The eigenvalues in each subspace can be calculated in parallel in multiple processes due to their similar characteristics, thereby giving full play to the performance potential of multi-core processors. In order to achieve the optimal allocation of computing resources, this algorithm accurately distributes the computing process in proportion according to the number and complexity of the eigenvalues in each subspace, ensuring that the computing load is balanced between the processes, and avoiding resource waste or computing bottlenecks caused by uneven load. Through the above technical means, this embodiment not only improves the calculation speed of the algorithm, but also effectively solves the load balancing problem in the multi-level parallel computing structure, and significantly improves the application effect of the algorithm on large-scale eigenvalue problems. The following will provide a more detailed explanation and analysis of the technical implementation and operation steps of the invention. Figure 1 As shown, the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level in this embodiment includes the following steps: S1, construct the Hamiltonian matrix of battery material molecules (Hamiltonian matrix) and overlap matrix (Overlap matrix), where the Hamiltonian matrix is the total energy between battery material molecules, and the overlap matrix The overlap between the molecular orbitals of the battery material is shown in the figure. According to the Hamiltonian matrix of the battery material molecule and the overlap matrix Construct the generalized eigenvalue problem shown below: , In the above formula, is an eigenvector, which is used to represent the column vector of the electron wave function. Each component in the eigenvector corresponds to the coefficient of different molecular orbital basis functions in the system, which is used to represent the state of the electron in the system. is the eigenvalue, which is used to represent the electronic energy level of the battery material; the interval where the electronic energy level is located is used as the search interval for the generalized eigenvalue problem ,in are the left and right boundaries of the search interval respectively; S2, delineate the contour integration and determine the number of integration nodes and the number of process resources; S3, the Hamiltonian matrix and the overlap matrix Copy and generate a low-precision version, and perform matrix decomposition on the matrix form of the low-precision version to adapt to the format requirements of solving the linear equation system in the subsequent stage; S4, based on search interval Construct a linear equation system in parallel on all integration nodes and solve it at a low-precision scale, and use the solution results to construct an iterative subspace Q; S5, Rayleigh–Ritz iterative update using iterative subspace Q at low precision scale; S6, Judgment Whether to set the convergence accuracy of the low-precision stage , if the convergence accuracy of the set low-precision stage is reached , then jump to step S7; otherwise, construct the right-hand side of a new linear equation system; S7, cluster analysis of the feature pairs obtained in the low-precision stage; S8, according to the cluster analysis results, the search interval Divide into multiple sub-search intervals ,in are the left and right boundaries of the i-th sub-search interval, respectively, where i ranges from 1 to k, k is the number of clusters, and each sub-search interval Containing one or more clustered eigenvalues, and the number of eigenvalues to be solved in each sub-search interval is equal or approximately equal to achieve load balancing, and the process resources are divided into groups according to the number of sub-search intervals, so that the number of process resources in each group is equal, and each group of process resources is used to calculate one sub-search interval; S9, each sub-search interval re-defines the contour integral and determines the number of integral nodes and process resources. and the overlap matrix Decompose the matrix in each subspace; S10, based on search interval , the linear equations of the generalized eigenvalue problem under low-precision scale are solved in parallel on all integration nodes, and the iterative subspace Q is constructed using the calculation results; S11, Rayleigh–Ritz iterative update using iterative subspace Q at high precision scale; S12, judgment Whether the convergence accuracy of the set high-precision stage is achieved , if the convergence accuracy of the set high-precision stage is reached , then jump to step S13; otherwise, construct the right-hand side of a new linear equation system and jump to step S10; S13, the main process summarizes the calculation results of each sub-search interval and outputs the final calculation results, including: battery materials in the search interval The eigenvalues within The electronic energy levels of the battery materials represented, as well as the eigenvectors The electron wave function represented by .
[0022] In step S1 of this embodiment, the Hamiltonian matrix Refers to the total energy between battery material molecules, the overlapping matrix Refers to the degree of overlap between the molecular orbitals of battery materials, input into the Hamiltonian matrix in practical applications and the overlap matrix Constructing the generalized eigenvalue problem The interval where the electron energy level is located is taken as the search interval [a,b] of the generalized eigenvalue problem, where is an eigenvector, which is a representation of the electron wave function. It is usually a column vector. Each component corresponds to the coefficient of different basis functions (such as molecular orbitals) in the system and is used to represent the state of the electron in the system. is the characteristic value, which represents the electronic energy level in the battery material. In addition, step S1 of this embodiment also includes setting the convergence accuracy of the low-precision stage and the convergence accuracy of the high-precision stage The convergence accuracy of the low-precision stage set in step S1 of this embodiment is For 10 -6 ; Convergence accuracy in high-precision stage The numerical precision required by the task objective is usually 10 for double precision. -10 Up to 10 -14 .
[0023] In step S2 of this embodiment, the contour integral is delineated as follows: As the diameter, a semicircle is made on the complex plane, and the optimal sample points and weights are selected according to the Gauss-Legendre quadrature to maximize the accuracy of the integral. ,in It represents the value of the sample point. It represents the weight of the sample point, and the selected sample point is used as the integral node of the contour integral. When determining the number of process resources in step S2, the selection of the number of process resources is determined by the actual machine used. The numerical experimental results show that selecting 8 to 16 integral nodes is a better choice under the requirements of accuracy and parallelism. Therefore, when determining the number of integral nodes in step S2, the value range of the number of integral nodes is [8,16].
[0024] For the Hamiltonian matrix and the overlap matrix Copy the low-precision version and select a suitable method (such as LU decomposition, QR decomposition or Cholesky decomposition, etc.) according to different matrix types (sparse, dense, banded, etc.) to perform matrix decomposition to adapt to the format requirements of solving the linear equations in the subsequent stage. In step S3 of this embodiment, when the matrix form of the low-precision version is decomposed, if the matrix form of the low-precision version is a dense matrix, the corresponding linear system solving function in the LAPACK library is used to perform matrix decomposition on the matrix form of the low-precision version; if the matrix form of the low-precision version is a sparse matrix, the corresponding linear system solving function in the MKL-PARDISO library is used to perform matrix decomposition on the matrix form of the low-precision version; if the matrix form of the low-precision version is a banded matrix, the corresponding linear system solving function in the SPIKE library is used to perform matrix decomposition on the matrix form of the low-precision version.
[0025] Step S4 is divided into three specific execution steps, which are respectively constructing a linear system in parallel on all integral nodes, solving at a low-precision scale, and constructing an iterative subspace Q in iteration. Specifically, step S4 in this embodiment includes: S4.1, generate a random orthogonal normalized matrix ; S4.2, based on search interval This constructs the linear system in parallel at all integration nodes: , in, For the The value of the integration node, is the overlap matrix, is the Hamiltonian matrix, For the The solution to be solved for the integration node, j is the serial number of the integration node, is a random orthogonal normalized matrix; S4.3, solve the linear system equations in parallel at each integration node to obtain the corresponding solution ; Select the corresponding library according to the different matrix types, such as using the corresponding linear system solving function in the LAPACK library for dense matrices, using the corresponding linear system solving function in the MKL-PARDISO library for sparse matrices, and using the corresponding linear system solving function in the SPIKE library for banded matrices; S4.4, construct the iterative subspace Q according to the following formula: , in, For the The weight of the integration node.
[0026] In this embodiment, the Rayleigh–Ritz iterative update using the iterative subspace Q includes: using the iterative subspace Q as the projection subspace, and transforming the generalized eigenvalue problem Projection into a simplified new generalized eigenvalue problem ,in is the new overlap matrix, is the new Hamiltonian matrix, is the new feature vector, is the new eigenvalue, and , ; Call the generalized eigenvalue solving function in the LAPACK library to solve the new generalized eigenvalue problem , the solution is completed and the new eigenvalue is obtained and the new eigenvector ; The new eigenvalue and the new eigenvector according to , = Transformed into the original generalized eigenvalue problem The eigenvalue of and the eigenvector .
[0027] In step S6, it is determined Whether to set the convergence accuracy of the low-precision stage The function expression is: , if If it is established, it is determined that if the convergence accuracy of the set low-precision stage is reached , jump to step S7; otherwise, construct the right-hand side of a new linear equation system. In this embodiment, constructing the right-hand side of a new linear equation system in step S6 includes: using the multiple sets of eigenvectors solved Concatenate into a matrix , the matrix Alternative random orthogonal normalized matrix Get the right side of the new linear equation system .
[0028] The purpose of clustering the feature pairs obtained in the low-precision stage in step S7 is to better balance the load and allocate resources in subsequent calculations. The eigenvalues and corresponding eigenvectors in each cluster are recorded to ensure that this information can be effectively used for calculation optimization in the mixed precision stage. When clustering the feature pairs obtained in the low-precision stage in step S7, the clustering can use common algorithms such as K-means clustering and spectral clustering to divide the calculated feature pairs (eigenvalues and eigenvectors, where the eigenvalues represent the electronic energy levels of battery materials and the eigenvectors are the representations of the corresponding electronic wave functions) into multiple clusters according to the eigenvalues. Considering the clustering distribution of eigenvalues in the generalized eigenvalue problem, this step can effectively identify the clustering distribution and reduce the calculation error.
[0029] In the calculation process of steps S3 to S6 in this embodiment, all process resources are used. Level 1 parallelism is not used at this time, and the entire search interval is solved by a contour integral. Step S8 divides the sub-search interval and the process group and corresponds them. After this process, the original entire search interval is divided into multiple smaller search sub-intervals and multiple small contour integrals. The purpose is to make full use of Level 1 parallelism in the three-layer parallel structure of the FEAST algorithm. The specific division effect needs to be different. Generally speaking, when the number of parallel resources is large, it will have a better effect to put it in Level 2 parallelism (see literature: Kestyn J, Kalantzis V, Polizzi E, et al. PFEAST: a highperformance sparse eigenvalue solver using distributed-memory linear solvers[C] / / SC'16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2016: 178-189.). Steps S9 to S12 are executed in parallel in the corresponding process groups of the multiple sub-search intervals and the calculation results in each sub-search interval are obtained respectively, until the main process summarizes and outputs the results in step S13. Figure 2 The three-layer parallel structure of the FEAST algorithm used in the embodiment of the present invention is: Figure 2The symbols are explained as follows: Level 1 parallelism: This is the highest level of parallelism in the algorithm, which means that the search interval is divided into multiple search subspaces for contour integral calculation, and the calculations on each search subspace can be run in parallel. Level 2 parallelism: This is the middle level of parallelism in the algorithm. According to the steps of the FEAST algorithm, when constructing the projection subspace Q, it is necessary to After the contour integral is delineated, each integration node calculates a , so all integration nodes can be calculated in parallel. Level 3 parallelism: This is the lowest level of parallelism, solving at each integration node If there are sufficient parallel processing resources, this step of linear system solution can use mature parallel solution algorithms, such as MKL-PARDISO library, SPIKE library, etc. 0 , P 1 , ..., P 11 : These symbols represent parallel processing units, which are processors or computing cores that distribute tasks in parallel computing. 0 , Z 1 , ..., Z 5 : These symbols represent data or tasks passed between Level 2 parallel levels. 0 , C 1 : These symbols represent data or tasks passed between Level 1 parallelism levels. The arrows indicate the direction in which data or tasks are passed from a higher level to a lower level. Figure 2 It shows how the FEAST algorithm decomposes and solves the eigenvalue problem through a hierarchical parallel structure. The parallel processing units at each level are responsible for solving the subproblems assigned to them.
[0030] In step S9, each sub-search interval re-defines the contour integral and determines the number of integral nodes, and the Hamiltonian matrix and the overlap matrix The decomposition of the matrix on each subspace adopts the same contour integral delineation method and the number of integral nodes as step S2. Note that the difference is that the number of process resources is no longer determined by the actual machine in use, but by the number of process resources in the process resource group. Specifically, when determining the number of integral nodes and the number of process resources in step S2 of this embodiment, the number of process resources is positively correlated with the performance of the computing device. When re-delineating the contour integral and determining the number of integral nodes and the number of process resources in each sub-search interval in step S9, the number of process resources is specified within the number of process resources of the process resource group corresponding to the sub-search interval.
[0031] Step S10 is based on the search interval , the linear equations of the generalized eigenvalue problem at a low-precision scale are solved in parallel on all integration nodes, and the iterative subspace Q is constructed using the calculation results; in this embodiment, the FEAST algorithm is specifically executed in parallel in multiple search subintervals to finally calculate all electronic energy levels and corresponding molecular orbital coefficients that meet the convergence conditions and search intervals.
[0032] Steps S9 to S12 are executed in parallel in the corresponding process groups of multiple sub-search intervals and the calculation results in each sub-search interval are obtained respectively. The specific calculation process is basically the same as steps S3 to S6. The only difference is that the iterative update and convergence judgment in steps S11 and S12 are performed under high precision.
[0033] The main process in step S13 summarizes the calculation results of each sub-search interval and outputs the final calculation results, including: after each process group calculates the feature pair (feature value and its corresponding feature vector) of the corresponding search sub-interval, the main process of each process group communicates the result to the main process of the global process resource, summarizes and outputs the result, including the battery material in the search interval The eigenvalues within The electronic energy levels and eigenvectors of the battery materials represented The electron wave function represented by .
[0034] This embodiment first reads the Hamiltonian matrix from a storage device or an input source and the overlap matrix , ensure that the input matrix format is correct, and store it in the memory of the computing device for subsequent processing. Next, determine the search interval [a, b] for the generalized eigenvalue problem. For example, the possible range of eigenvalues can be determined based on the specific requirements of the problem, prior knowledge, or preliminary calculation results, thereby improving the calculation efficiency and accuracy. In addition, according to the available computing resources, set the number of processes required for parallel computing. When using MPI (Message Passing Interface) for parallel computing, a reasonable selection of the number of processes helps to fully utilize the performance of multi-core processors. At the same time, according to the scale of the problem and the calculation accuracy requirements, set the number of integration nodes. The selection of the number of integration nodes requires finding a balance between calculation accuracy and calculation overhead. Then, determine the convergence accuracy ε of the algorithm, which is the accuracy requirement of the algorithm in the double precision stage. Generally speaking, the convergence accuracy requirement of the double precision stage is higher than that of the low precision stage. Finally, the convergence criterion of the pure low precision stage must be set. The maximum number of iterations in the low precision stage can be specified, or a specific convergence threshold can be set. When the low precision stage reaches the convergence threshold, the algorithm will end the iteration and enter the next stage. In the pure low-precision stage, the matrix and are converted into low-precision floating-point form and saved in memory to reduce memory usage and computational overhead. According to the type of matrix (sparse, dense or banded), select an appropriate matrix decomposition method, such as LU decomposition, QR decomposition or Cholesky decomposition, to decompose the low-precision matrix and save the decomposition results for subsequent calculations. The low-precision linear system is solved in parallel on all integration nodes, and a new iterative subspace is constructed using the decomposition results. The new iterative subspace is updated using the Rayleigh–Ritz iterative method, and after each iteration, it is checked whether it meets the convergence criteria of the low-precision stage. If the current iteration result does not meet the convergence criteria, the right-hand side of the linear system is updated and returned for the next round of calculation; if the convergence criteria are met, the low-precision eigenvalues and the corresponding low-precision eigenvectors are output as the basis for subsequent mixed-precision stage calculations. After the low-precision eigenvalue calculations are completed, these eigenvalues are clustered. The eigenvalues within each cluster should have good clustering properties, that is, the distance between the eigenvalues within the cluster is close, while the distance between the eigenvalues of different clusters is relatively far. The purpose of cluster analysis is to better balance load and allocate resources in subsequent calculations. According to the results of cluster analysis, the original search interval [a, b] is divided into multiple sub-intervals [a i ,b i], each subinterval contains one or more clustered eigenvalues. The purpose of this is to perform more detailed calculations and optimizations in each subinterval. According to the number of eigenvalues in each subinterval, computing resources are allocated in proportion to ensure load balancing. Specifically, the computing tasks of each subinterval can be allocated through MPI process communication to ensure that the computing load on each computing node is roughly equal. Construct a subcommunication domain of multiple processes, and set a sub-master process in each subcommunication domain to coordinate the allocation of computing tasks and inter-process communication in the domain. In this way, the waste of computing resources can be effectively avoided and the overall computing efficiency can be improved. In the mixed precision stage, the main process first broadcasts the matrix decomposition results saved in the low-precision stage to the sub-master processes of all subcommunication domains. The master subprocess of each subcommunication domain converts the received low-precision eigenvectors into double-precision floating-point numbers to improve the calculation accuracy. Then, the converted double-precision eigenvectors are used to generate a new iterative subspace in each subinterval. Next, the FEAST algorithm is executed in parallel in each subinterval, and the new projection subspace is used for fine eigenvalue calculation. During the iteration process, it is continuously checked whether the preset convergence accuracy ε is met. If the preset convergence accuracy is reached in a certain subinterval, the calculation task of the subinterval is completed; otherwise, the iteration continues until the convergence condition is met. In this way, the mixed precision and load balancing technology can be fully utilized to solve the large-scale generalized eigenvalue problem, significantly improving the calculation efficiency and accuracy. The method for solving the large-scale generalized eigenvalue problem of the electronic energy level of the battery in this embodiment significantly improves the efficiency and accuracy of the solution algorithm of the large-scale generalized eigenvalue problem based on the contour integral classification with FEAST as an example by introducing mixed precision calculation and load balancing strategy. This embodiment takes the calculation of the electronic energy level in the battery material as a specific application scenario to demonstrate the execution process of the algorithm. The input data includes the Hamiltonian matrix representing the electronic structure of the battery electrode material and the overlap matrix describing the overlap degree of different atomic orbitals in the material. Among them, the Hamiltonian matrix represents the energy relationship of the electrons in the electrode material, reflecting its interaction under different electric fields and chemical environments, while the overlap matrix represents the overlap between different atomic orbitals. After the algorithm calculation is completed, the output data is the electronic energy level and the molecular orbital coefficient, where the electronic energy level represents the energy state of the electrode material at different energy levels during the charge and discharge process, and the molecular orbital system describes the spatial distribution of electrons in the material. These output results can be used to evaluate the stability, conductivity and ion diffusion performance of electrode materials, and then guide the development and optimization of high-performance battery materials. It should be noted that the method of this embodiment can also be extended to all problems involving the solution of generalized eigenvalue problems in the field of high-performance computing. For example, for the vibration mode analysis problem in structural mechanics, the input data is a structural system composed of an elastic body and a mass point, and its corresponding mass matrix and stiffness matrix. The output results are the natural frequency (eigenvalue) and vibration mode (eigenvector) of the structure.
[0035] In addition, this embodiment also provides a system for solving a large-scale generalized eigenvalue problem of battery electronic energy levels, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute the method for solving a large-scale generalized eigenvalue problem of battery electronic energy levels.
[0036] In addition, this embodiment also provides a computer-readable storage medium, which stores a computer program or instruction, and the computer program or instruction is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level through a processor.
[0037] In addition, this embodiment also provides a computer program product, including a computer program or instructions, which are programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of battery electronic energy levels through a processor.
[0038] Those skilled in the art should understand that the technical solutions provided by the embodiments of the present application may be in the form of methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the process Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0039] The above is only a preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions under the concept of the present invention belong to the protection scope of the present invention. It should be pointed out that for ordinary technicians in this technical field, some improvements and modifications without departing from the principle of the present invention should also be regarded as the protection scope of the present invention.
Claims
1. A method for solving the large-scale generalized eigenvalue problem of battery electronic energy levels, characterized in that: The steps include: S1, construct the Hamiltonian matrix of battery material molecules and the overlap matrix , where the Hamiltonian matrix is the total energy between battery material molecules, and the overlap matrix The overlap between the molecular orbitals of the battery material is shown in the figure. According to the Hamiltonian matrix of the battery material molecule and the overlap matrix Construct the generalized eigenvalue problem shown below: , In the above formula, is an eigenvector, which is used to represent the column vector of the electron wave function. Each component in the eigenvector corresponds to the coefficient of different molecular orbital basis functions in the system, which is used to represent the state of the electron in the system. is the eigenvalue, which is used to represent the electronic energy level of the battery material; the interval where the electronic energy level is located is used as the search interval of the generalized eigenvalue problem ,in are the left and right boundaries of the search interval respectively; S2, delineate the contour integration and determine the number of integration nodes and the number of process resources; S3, the Hamiltonian matrix and the overlap matrix Copy and generate a low-precision version, and perform matrix decomposition on the matrix form of the low-precision version to adapt to the format requirements of solving the linear equation system in the subsequent stage; S4, based on search interval Construct a linear equation system in parallel on all integration nodes and solve it at a low-precision scale, and use the solution results to construct an iterative subspace Q; S5, Rayleigh–Ritz iterative update using iterative subspace Q at low precision scale; S6, Judgment Whether to set the convergence accuracy of the low-precision stage , if the convergence accuracy of the set low-precision stage is reached , then jump to step S7; otherwise, construct the right-hand side of a new linear equation system; S7, cluster analysis of the feature pairs obtained in the low-precision stage; S8, according to the cluster analysis results, the search interval Divide into multiple sub-search intervals ,in are the left and right boundaries of the ith sub-search interval, respectively. Containing one or more clustered eigenvalues, and the number of eigenvalues to be solved in each sub-search interval is equal or approximately equal to achieve load balancing, and the process resources are divided into groups according to the number of sub-search intervals, so that the number of process resources in each group is equal, and each group of process resources is used to calculate one sub-search interval; S9, each sub-search interval re-defines the contour integral and determines the number of integral nodes and process resources. and the overlap matrix Decompose the matrix in each subspace; S10, based on search interval , the linear equations of the generalized eigenvalue problem under low-precision scale are solved in parallel on all integration nodes, and the iterative subspace Q is constructed using the calculation results; S11, Rayleigh–Ritz iterative update using iterative subspace Q at high precision scale; S12, judgment Whether the convergence accuracy of the set high-precision stage is achieved , if the convergence accuracy of the set high-precision stage is reached , then jump to step S13; Otherwise, construct the right-hand side of a new linear equation system and jump to step S10; S13, the main process summarizes the calculation results of each sub-search interval and outputs the final calculation results, including: battery materials in the search interval The eigenvalues within The electronic energy levels of the battery materials represented, as well as the eigenvectors The electron wave function represented by .
2. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 1, characterized in that: Delineating the contour integral in step S2 includes: searching the interval As the diameter, make a semicircle on the complex plane, and select the best sample points and weights according to the Gauss-Legendre integral to maximize the accuracy of the integral. ,in It represents the value of the sample point. It represents the weight of the sample point, and the selected sample point is used as the integration node of the contour integral.
3. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 1, characterized in that: When performing matrix decomposition on the low-precision version of the matrix in step S3, if the matrix form of the low-precision version is a dense matrix, the corresponding linear system solver function in the LAPACK library is used to decompose the matrix form of the low-precision version; if the matrix form of the low-precision version is a sparse matrix, the corresponding linear system solver function in the MKL-PARDISO library is used to decompose the matrix form of the low-precision version; if the matrix form of the low-precision version is a banded matrix, the corresponding linear system solver function in the SPIKE library is used to decompose the matrix form of the low-precision version.
4. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 1, characterized in that: Step S4 includes: S4.1, generate a random orthogonal normalized matrix ; S4.2, based on search interval This constructs the linear system in parallel at all integration nodes: , in, For the The value of the integration node, is the overlap matrix, is the Hamiltonian matrix, For the The solution to be solved for the integration node, j is the serial number of the integration node, is a random orthogonal normalized matrix; S4.3, solve the linear system equations in parallel at each integration node to obtain the corresponding solution ; S4.4, construct the iterative subspace Q according to the following formula: , in, For the The weight of the integration node.
5. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 4, characterized in that: The Rayleigh–Ritz iterative update using the iterative subspace Q in step S5 includes: using the iterative subspace Q as the projection subspace, and transforming the generalized eigenvalue problem Projection into a simplified new generalized eigenvalue problem ,in is the new overlap matrix, is the new Hamiltonian matrix, is the new feature vector, is the new eigenvalue, and , ; Call the generalized eigenvalue solving function in the LAPACK library to solve the new generalized eigenvalue problem , the solution is completed and the new eigenvalue is obtained and the new eigenvector ; The new eigenvalue and the new eigenvector according to , = Transformed into the original generalized eigenvalue problem The eigenvalue of and the eigenvector .
6. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 5, characterized in that: The right-hand side of the new linear equation system constructed in step S6 includes: using the multiple sets of eigenvectors solved Concatenate into a matrix , the matrix Alternative random orthogonal normalized matrix Get the right side of the new linear equation system .
7. The method for solving the large-scale generalized eigenvalue problem of battery electronic energy level according to claim 1, characterized in that: When determining the number of integration nodes and the number of process resources in step S2, the number of process resources is positively correlated with the performance of the computing device. When redefining the contour integral in each sub-search interval and determining the number of integration nodes and the number of process resources in step S9, the number of process resources is specified within the number of process resources of the process resource group corresponding to the sub-search interval.
8. A system for solving large-scale generalized eigenvalue problems of battery electronic energy levels, comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of battery electronic energy level as described in any one of claims 1 to 7.
9. A computer-readable storage medium having a computer program or instruction stored therein, characterized in that: The computer program or instruction is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of the battery electronic energy level as claimed in any one of claims 1 to 7 through a processor.
10. A computer program product comprising a computer program or instructions, characterized in that The computer program or instruction is programmed or configured to execute the method for solving the large-scale generalized eigenvalue problem of the battery electronic energy level as claimed in any one of claims 1 to 7 through a processor.
Citation Information
Cited By
Linear scale electronic structure calculation method and system based on dyeing superposition state, and terminal
CN122369620A