GPU (Graphic Processing Unit) parallel solving method and system for linear equation group of power system

By hierarchically parallel computing the power system nodes, a sparse decomposition thread tree and a collection of sibling nodes is established, the overlapping calculation problem in parallel solving GPUs of linear equations in the power system is solved, and the calculation accuracy and efficiency are improved, which is suitable for large power grid current calculations.

CN120262428AActive Publication Date: 2025-07-04DONGFANG ELECTRONICS CO LTD

Patent Information

Application Number
CN202510747873.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-04
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

The existing GPU parallel solution method for linear equations in power systems has overlapping calculation problems, the calculation results are low in accuracy, and the parallel flow calculation execution efficiency is low.

Method used

All nodes of the power system are divided into different levels according to the distance from the initial node, and a sparse decomposition thread tree is established. Nodes at the same level are used as sets of sibling nodes, threads are allocated in the GPU, and sparse matrix is ​​calculated in parallel. The Gaussian elimination triangular decomposition algorithm is used for decomposition and previous generation back-generation calculation.

Benefits of technology

The overlapping calculation problem in sparse matrix factor decomposition and hierarchical parallel computing is solved, the calculation accuracy and parallel current computing efficiency are improved, and the online computing practicality of large power grid current computing is met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120262428A_ABST
    Figure CN120262428A_ABST
Patent Text Reader

Abstract

The invention discloses a GPU parallel solving method and system for linear equations of a power system, and relates to the field of power system automation. In order to solve the defects that an existing GPU parallel solving method for the system of linear equations of the power system has an overlapping calculation problem and is low in calculation result accuracy and low in parallel load flow calculation execution efficiency, an initial node is defined in all nodes of the power system, other nodes are divided into different hierarchies according to the distance between the other nodes and the initial node, and the calculation efficiency is improved. Establishing a sparse decomposition thread tree; child nodes with the same father node in the same hierarchy are called as brother nodes, and a brother node set is established; threads are distributed in the GPU according to the brother node set; and sparse matrix parallel solving calculation is carried out on each node of the same level. The method is mainly used for optimizing the GPU parallel solution of the system of linear equations of the power system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of power system automation, and particularly to a method and a system for parallelly solving a linear equation set of a power system by using a GPU (Graphics Processing Unit). Background Art

[0002] Currently, with the development of the scale of the power system and the promotion of the construction of the power grid regulation cloud, the scale of power grid analysis and calculation is increasing day by day. In the current cloud system of the regulation center, the modeling range ranges from 1000 kV of ultra-high voltage to the outgoing line of 10 kV feeders. The scale of calculation nodes in a province can reach tens of thousands, which greatly increases the calculation time and is difficult to meet the real-time analysis requirements of the power grid.

[0003] Actually, the core algorithms of the current power system state estimation and power flow calculation are to solve a large-scale linear equation set. How to quickly perform calculation operations on the corresponding sparse matrix is the key problem to improve the solution speed. The commonly used calculation methods include using the parallel calculation technology of a multi-core central processing unit (CPU) and the parallel calculation technology of a graphics processing unit (GPU).

[0004] However, the above methods are all restricted by the memory bandwidth and the number of physical cores. There is a natural calculation bottleneck in the parallel calculation efficiency of the traditional multi-core CPU, so it is not applicable to large-scale systems. As a special processor designed for data parallelism, the graphics processing unit (GPU) has more powerful parallel calculation capabilities due to its superior performance in floating-point operations, memory bandwidth, and a large number of cores, and has been successfully applied to many scientific calculation fields including the power system field.

[0005] Although there are also some researches on parallel calculation in the prior art, because the nodes of the sparse path graph have upward dependencies, that is, they depend on the calculation results of the upper-layer nodes, there is an overlapping calculation problem between the current parallel threads, which seriously affects the calculation result accuracy of the power flow calculation, resulting in a slower convergence speed or even non-convergence.

[0006] Therefore, there is a need for a method and a system for parallelly solving a linear equation set of a power system by using a GPU, which can solve the overlapping calculation problem in the hierarchical parallel calculation of sparse matrix factorization, improve the calculation accuracy, and improve the parallel power flow calculation efficiency of the power system. Summary of the Invention

[0007] The present invention aims to solve the defects of the existing GPU parallel solution method for linear equations in power systems, such as overlapping calculations, low accuracy of calculation results, and low execution efficiency of parallel power flow calculations. It provides a GPU parallel solution method and its solution system for linear equations in power systems, which can solve the overlapping calculation problem in hierarchical parallel calculation of sparse matrix factorization, improve the calculation accuracy, and improve the efficiency of parallel power flow calculation in power systems.

[0008] The GPU parallel solution method for linear equations in power systems according to the present invention includes the following steps: Define an initial node among all nodes in the power system, and divide the remaining nodes into different levels according to the distance from the initial node, so as to establish a sparse decomposition thread tree; Define the child nodes with the same parent node in the same level as sibling nodes, and establish a sibling node set; Allocate threads in the GPU according to the sibling node set; Perform parallel solution calculation of the sparse matrix for each node in the same level.

[0009] Further: The step of establishing the sparse decomposition thread tree includes: Describe all nodes in the power system using a weighted directed factor graph; p, j, k, and l are all nodes for power flow calculation, where p is the power generation point, and j, k, and l are the power receiving points; On the weighted directed factor graph, among the edges emitted from each node, select the edge with the smallest receiving node number as the path on the sparse decomposition path graph.

[0010] Further: In the parallel solution calculation of the sparse matrix for each node in the same level, for the sparse matrix calculations of the same-level nodes, they are calculated in parallel simultaneously. The specific process includes: On the sparse decomposition path graph of node j, the definition of the parent node number parent(j) of node j is: ; In the formula, is the element in the lower triangular matrix L of the Jacobian matrix A at the p-th row and j-th column. It is said that node p is the parent node of node j, and node j is the child node of node p; The in-degree In_Degree(i) of node i is defined as the sum of the number of all nodes connected to node i and with node numbers less than i, that is: ; Among them, is the edge weight of each edge pointing to node i from node j, where i = p, j, k, or l; The out-degree Out_Degree(i) of node i is defined as the sum of the number of all nodes connected to node i and with node numbers greater than i, that is: ; where is the edge weight of the edge from node i to node j.

[0011] Furthermore: In the parallel solution calculation of the sparse matrix for each node at the same level, only after the child nodes are calculated, the parent node starts to calculate, so as to realize the layering of the sparse decomposition path graph. The specific process of the layering includes: Initialize the levels of all nodes; ; In the formula, level(i) represents the i-th level; Traverse all nodes, and assign the nodes with in-degree 0 to level 0; ; For the nodes with in-degree greater than 0, define the level of node i as: ; where is the column number of the non-zero non-diagonal element in the i-th row of the lower triangular matrix L.

[0012] Furthermore: The solution process of the linear equation for each node includes: Decompose the linear equation based on the triangular decomposition algorithm of Gaussian elimination to obtain a triangular matrix; Use the decomposed triangular matrix to perform forward substitution and back substitution calculations to solve the unknown variables of the linear equation.

[0013] Furthermore: In the forward substitution process, the forward substitution calculations are performed in parallel according to the different layer sets to which the nodes belong; The triangular decomposition algorithm and the forward substitution calculations start from layer i = 0 to layer i = n - 1, and the back substitution calculations end from layer i = n - 1 to layer i = 0; In each layer, the number of GPU threads is allocated according to the number of sibling node sets at the current level.

[0014] Furthermore: The triangular decomposition algorithm includes the following steps: Use the triangular decomposition algorithm to decompose the Jacobian matrix formed by the linear equations into an upper triangular matrix and a lower triangular matrix; Solve the unknown variables of the nodes, and the unknown variables include the voltage amplitude and voltage phase angle of the nodes; Update and correct the diagonal elements of the upper triangular matrix; At the same time, correct the non-zero elements of the upper triangular matrix.

[0015] Furthermore: The operation processes of the forward substitution and back substitution are both: ; wherein, is the i-th element of the unknown variable; is the unbalance; is the element at the i-th row and j-th column in the upper triangular matrix; Update the unknown variables of the nodes of the weighted directed factor graph through forward substitution and back substitution calculations.

[0016] The parallel solving system for implementing the GPU parallel solving method of the linear equations of the power system disclosed by the present invention includes a sparse decomposition thread tree establishment module, a node definition module, a thread allocation module, and a GPU; The sparse decomposition thread tree establishment module is used to generate a sparse decomposition thread tree for all nodes in the power system; The node definition module is used to define the nodes of the sparse decomposition thread tree to form a sibling node set; The thread allocation module is used to match threads according to the sibling node set of the sparse decomposition thread tree; The GPU performs parallel calculations according to threads at the same level.

[0017] Furthermore: The GPU includes a triangular decomposition module and a forward-back substitution calculation module; The triangular decomposition module is used to decompose the linear equation based on the triangular decomposition algorithm of Gaussian elimination to obtain a triangular matrix; The forward-back substitution calculation module is used to perform forward substitution and back substitution calculations using the decomposed triangular matrix to solve the unknown variables of the linear equation.

[0018] The beneficial effects of the present invention are: By classifying all nodes into different levels according to the distance from the initial node and establishing a sparse decomposition thread tree, the sparse matrix calculations of nodes at the same level can be performed in parallel simultaneously, which can solve the overlapping calculation problem of matrix calculations in the process of solving the linear equations of the power flow of the power system in the CPU+GPU heterogeneous mode, improve the execution efficiency of the parallel power flow calculation of the power system, and thus improve the online calculation practicability of the power flow calculation of large power grids.

[0019] The present invention proposes a GPU parallel computing method for matrix calculation of power flow linear equations based on a sparse decomposition path thread tree of sibling nodes. All nodes are divided into different levels according to the distance from the initial node, and a sparse decomposition thread tree is established. The sparse matrix calculations of nodes at the same level can be performed in parallel simultaneously. Search for BroNode nodes among nodes at the same level on the sparse decomposition path tree. The child nodes with the same parent node in the same layer are called sibling nodes BroNode, and a BroNodeSet set is established. When programming for GPU parallel computing, threads are allocated according to the BroNodeSet set, and the factorization and forward substitution calculations of the nodes in the BroNodeSet set are sequentially executed in the same thread. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of a sparse decomposition path thread tree based on BroNode after layering; Figure 2 It is a schematic diagram of a multi-threaded concurrent computing process based on BroNode; Figure 3 It is a schematic diagram of a triangular decomposition algorithm based on Gaussian elimination; Figure 4 It is a schematic diagram of a node of an elimination operation with p as the axis; Figure 5 It is a schematic diagram of the original matrix A. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] The following are only preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention. The following embodiments are only used to explain the present invention and cannot be construed as a limitation of the present invention. The protection scope of the present invention should be subject to the protection scope of the claims. The embodiments of the present invention are described in detail below. For the convenience of describing the present invention and simplifying the description, the technical terms used in the specification of the present invention should be interpreted in a broad sense, including but not limited to conventional substitution schemes not mentioned in the present application, and including both direct implementation methods and indirect implementation methods.

[0022] Embodiment 1 Combined with Figures 1 - 5 To illustrate this embodiment, a GPU parallel solution method for linear equations of a power system disclosed in this embodiment includes the following steps: Define an initial node among all nodes of the power system, and divide the remaining nodes into different levels according to the distance from the initial node, so as to establish a sparse decomposition thread tree; The steps of establishing the sparse decomposition thread tree include: Such as Figure 4 AndFigure 5 As shown, all nodes of the power system are described by a weighted directed factor graph; p, j, k, and l are all nodes for power flow calculation, where p is the power generation node, and j, k, and l are the power consumption nodes; On the said weighted directed factor graph, among the edges emitted from each node, the edge with the smallest receiving node number is taken as the path on the sparse decomposition path graph.

[0023] The child nodes with the same parent node in the same level are called sibling nodes, and a sibling node set is established; Threads are allocated in the GPU according to the said sibling node set; Parallel solution calculations of sparse matrices are performed for each node in the same level.

[0024] In the sparse matrix solution calculation for each said node, the sparse matrix calculations for sibling nodes are performed simultaneously in parallel. The specific process includes: On the sparse decomposition path graph of node j, the definition of the parent node number parent(j) of node j is: ; In the formula, is the element in the p-th row and j-th column of the lower triangular matrix L in the Jacobian matrix A. It is said that node p is the parent node of node j, and node j is the child node of node p; The in-degree In_Degree(i) of node i is defined as the sum of the number of nodes connected to node i and with node numbers less than i, that is: ; Among them, is the edge weight of each edge from node j pointing to node i, where i = p, j, k, or l; The out-degree Out_Degree(i) of node i is defined as the sum of the number of nodes connected to node i and with node numbers greater than i, that is: ; Among them, is the edge weight of the edge from node i to node j.

[0025] In the sparse matrix solution calculation for each said node, only after the child node calculation is completed, the parent node starts to calculate, so as to realize the layering of the sparse decomposition path graph. The specific layering process includes: Initialize the levels of all nodes; ; In the formula, level(i) represents the i-th level; Traverse all nodes, and assign the nodes with in-degree 0 to level 0; ; For nodes with in-degree greater than 0, the level of node i is defined as: ; where is the column number of the non-zero non-diagonal elements in the i-th row of the lower triangular matrix L.

[0026] According to the above analysis, the basic steps of GPU-based matrix factorization and forward / backward substitution parallel calculation are as follows, with n-1 layers: (1) Layer all nodes according to the above method; (2) Starting from i = 0 to i = n-1 layer (factorization and forward calculation) or from i = n-1 layer to i = 0 layer (backward substitution calculation), allocate corresponding GPU threads according to the number of nodes in each layer at each layer and perform parallel calculation.

[0027] As Figure 5 shown, the core problems of power system state estimation and power flow calculation can be reduced to solving a large-scale linear algebraic equation system , where A is the Jacobian matrix, x is the state quantity in the power system, including voltage, amplitude and phase angle, and b is the unbalance quantity, to obtain the unknown state quantity x. The most commonly used and effective method for solving linear algebraic equations is the direct method based on Gaussian elimination. The algorithms for parallel solving of large-scale sparse linear algebraic equations mainly include: triangular factorization; forward and backward substitution calculations.

[0028] Assuming state estimation, the generalized power flow matrix is a symmetric matrix, and the symmetric matrix of the Jacobian matrix A is stored only for the upper triangular part to save storage space; As Figure 5 shown, perform LDU factorization on the Jacobian matrix A, that is: A = LDU, where D is a diagonal matrix, L is a unit lower triangular matrix (diagonal elements are 1, upper triangular elements are all 0), and U is a unit upper triangular matrix, with lower triangular elements all 0; The solution process for the linear equation of each node includes: Decompose the linear equation using the triangular factorization algorithm based on Gaussian elimination to obtain a triangular matrix; As Figure 3 and Figure 4 shown, the triangular factorization algorithm based on Gaussian elimination is described as: when eliminating the p-th row and p-th column currently; For the update and correction of the elements of the diagonal matrix, the elimination operation should use the following formula: ; In a weighted directed graph, it is to correct the self-edge weights on the receiving nodes j, k, l of the edges emitted by node p, and the edge weight is reduced ; Among them, is the element in the p-th row and p-th column of the diagonal matrix; is the element in the p-th row and i-th column of the upper triangular matrix.

[0029] As Figure 4 shown, for the non-zero elements in the upper triangular part, there are three elements that need to be corrected, and they are , , .

[0030] Among them, is the element in the j-th row and k-th column of the upper triangular matrix; is the element in the k-th row and l-th column of the upper triangular matrix; is the element in the j-th row and l-th column of the upper triangular matrix; The formula for the elimination operation is: ; Among them, is the element in the i-th row and m-th column of the upper triangular matrix; is the element in the i-th row and p-th column of the lower triangular matrix; is the element in the p-th row and m-th column of the upper triangular matrix; Because only the upper triangular part of the matrix is stored, so should be replaced by , so there is: ; After the node optimization numbering for power flow calculation, the factorization process of the factor table matrix is to execute the following steps in ascending order of node numbers on the weighted directed factor graph. Taking the processing of node p as an example: Step (1) Divide the edge weight of the mutual edges sent out by node p by the self-edge weight of node p; Step (2) For the receiving nodes at the opposite ends of the mutual edges sent out by node p, subtract the product of the square of the edge weight of this mutual edge and the self-edge weight of node p from the self-edge weight at this point; ; Step (3) For all the mutual edges sent out by node p, the edge weights between these mutual edges in pairs should be subtracted by the product of the edge weights of the two sandwiching edges and the self-edge weight of node p. The situation where there is no edge between the sandwiched node pairs before the operation can be regarded as having a zero-weight edge.

[0031] Step (4) Cover all the edges connected to p, select the next node, and return to Step (1).

[0032] Finally, the decomposed triangular matrix is used for forward substitution and back substitution to solve the linear equations for the unknown state variable x.

[0033] During the forward substitution process, the forward substitution calculations are performed in parallel according to the different levels to which the nodes belong; The triangular decomposition and forward substitution calculations start from layer i = 0 and end at layer i = n - 1, and the back substitution calculation starts from layer i = n - 1 and ends at layer i = 0; in each layer, the number of GPU threads is allocated according to the number of sibling node sets at the current level.

[0034] The triangular decomposition algorithm includes the following steps: The Jacobi matrix formed by the linear equations is decomposed using the triangular decomposition algorithm into an upper triangular matrix and a lower triangular matrix; the unknown variables of the nodes are solved, and the unknown variables include the voltage amplitude and voltage phase angle of the nodes; The diagonal elements of the upper triangular matrix are updated and corrected; at the same time, the non-zero elements of the upper triangular matrix are corrected.

[0035] The operation processes of the forward substitution and back substitution are both: ; where, is the i-th element of the unknown variable; is the unbalance; is the element in the i-th row and j-th column of the upper triangular matrix; Through the forward substitution and back substitution calculations, the unknown variables of the nodes of the weighted directed factor graph are updated.

[0036] If the points on the weighted directed factor graph are represented by and the edge weight of the mutual edge on the weighted directed factor graph is , then the above program can be written as: ; where, represents the node point of node j, and the condition means is the edge weight of the edge emitted from node i, which implies the condition .

[0037] Normalization process: Divide the point of node i after the forward substitution by the self-edge weight of node i on the weighted directed factor graph, that is ; is the result of normalization.

[0038] Update the point positions on the weighted directed factor graph to the values after the previous generation and normalization. On this graph, the node number j starts from n and decreases from large to small. For all the edges pointing to j, correct the point positions of the originating nodes i. The correction formula is: ; The steps of the back substitution calculation on the directed factor graph are as follows: Assign the non-zero elements of the independent vector b to the point positions on the weighted directed factor graph; Scan i from 1 to n - 1 and use the formula to correct the point positions of the opposite-end nodes j of the edges emitted by node i; For all nodes, use the formula to normalize the point positions; Scan j from n to 2. For all the originating nodes i of the edges pointing to node j, use the formula to correct their point positions.

[0039] From the sparse matrix factorization described by graph theory and the back substitution calculation, it can be seen that there is a certain forward and backward dependence among the nodes in the matrix calculation, and there is also a large amount of parallelism among a large number of nodes. On the directed factor graph, select the edge with the smallest receiving node number from the edges emitted by each node as the path on the sparse decomposition path graph. The analysis of the current node only affects the amplitude and phase angle of the nodes on its own sparse decomposition path graph. Therefore, during the previous generation process, according to the different levels to which the nodes belong, the previous generation calculations can be performed in parallel without affecting each other. The sparse matrix calculations of the same-level nodes can be performed simultaneously in parallel because these nodes have an upward dependence and depend on the results of the nodes in the upper layer.

[0040] As Figure 2 shown, the triangular decomposition process is: The processing process for node 1 is: ; In the formula, represents the element in the 1st row and 3rd column of matrix A, represents the element in the 3rd row and 3rd column of matrix A, represents the element in the 1st row and 1st column of matrix A, represents the element in the 1st row and 6th column of matrix A, represents the element in the 6th row and 6th column of matrix A, represents the element in the 6th row and 1st column of matrix A, represents the element in the 3rd row and 6th column of matrix A, represents the element in the 3rd row and 1st column of matrix A; The processing process for node 2 is: ; In the formula, Represents the element in the 2nd row and 3rd column of matrix A, Represents the element in the 2nd row and 2nd column of matrix A, Represents the element in the 2nd row and 8th column of matrix A, Represents the element in the 8th row and 8th column of matrix A, Represents the element in the 3rd row and 3rd column of matrix A, Represents the element in the 8th row and 2nd column of matrix A, Represents the element in the 3rd row and 8th column of matrix A, Represents the element in the 3rd row and 2nd column of matrix A; Predecessor calculation for node 1: ; In the formula, Represents the element in the 1st row and 1st column of matrix A, Represents the element in the 1st row of the right - hand side term, 3 , Represents the element in the 6th row of the right - hand side term, Represents the element in the 1st row and 3rd column of the upper triangular matrix U, Represents the element in the 1st row and 6th column of the upper triangular matrix U; Predecessor calculation for node 2: ; In the formula, Represents the element in the 1st row and 1st column of the matrix, Represents the element in the 2nd row of the right - hand side term, Represents the element in the 8th row of the right - hand side term, Represents the element in the 2nd row and 3rd column of the upper triangular matrix U, Represents the element in the 2nd row and 8th column of the upper triangular matrix U; The calculations for subsequent nodes are similar to the above and will not be elaborated.

[0041] From the above calculation process, it can be seen that when executed sequentially, when processing node 1, node 1 is connected to node 3, as shown in formulas (1b) and (3b), 、 There are updated calculations for the diagonal element of node 3 and the right - hand side term element When processing node 2, node 2 is connected to node 3, as shown in formulas (2b) and (4b), 、 There are updated calculations for the diagonal element of node 3 and the right - hand side term element and there is an accumulative update result with the processing of node 1, that is: ; 。

[0042] Example 2 This example is described in combination with Example 1. The parallel solution system for implementing the parallel solution method of the linear equations of a power system disclosed in this example includes a sparse decomposition thread tree establishment module, a node definition module, a thread allocation module, and a GPU; The sparse decomposition thread tree establishment module is used to generate a sparse decomposition thread tree for all nodes in the power system; The node definition module is used to define the nodes of the sparse decomposition thread tree to form a sibling node set; The thread allocation module is used to match threads according to the sibling node set of the sparse decomposition thread tree; The GPU performs parallel computing according to threads at the same level.

[0043] The GPU includes a triangular decomposition module and a forward and backward substitution calculation module; The triangular decomposition module is used to decompose the linear equation based on the triangular decomposition algorithm of Gaussian elimination to obtain a triangular matrix; The forward and backward substitution calculation module is used to perform forward and backward substitution calculations using the decomposed triangular matrix to solve the unknown variables of the linear equation.

[0044] Example 3 This example is described in combination with Example 1 and Example 2. A method for parallel solution of linear equations of a power system grid flow based on a sparse decomposition path thread tree of BroNode disclosed in this example is as follows Figure 2 shown. The steps for parallel solution of the linear equations of the grid flow based on the sparse decomposition path thread tree of BroNode are as follows: 1. Initialize the levels of all nodes: ; 2. Then traverse all nodes, and assign nodes with an in-degree of 0 to layer 0; ; For nodes with an in-degree > 0, define the level of node i as: ; where, is the column number of the non-zero non-diagonal element in the i-th row of the lower triangular matrix L; 3. Establish a sparse decomposition path thread tree according to the above method; 4. Traverse the nodes at the same level, and put the nodes with the same parent node into the sibling node set BroNodeSet; 5. Starting from i = 0 to the i = n - 1 layer (decomposition and previous generation calculation) or from the i = n - 1 layer to the i = 0 layer (back substitution calculation), allocate the number of GPU threads according to the number of sibling node sets at the current layer at each layer; 6. Perform parallel calculation at the current layer.

[0045] Example 1: As shown in Table 1: Taking a large power grid in a certain area as an example, the number of computing nodes of this power grid is 3956. Through the above-mentioned sparse decomposition path thread layering, the number of nodes in the 0th layer is 1586. The calculations between these nodes are independent of each other. Therefore, theoretically, 1586 threads can be allocated to the matrix decomposition calculation kernel function. However, through analysis, it is found that among these 1586 nodes, 79 nodes have the situation of having the same parent node. If 1586 threads are allocated for parallel calculation, then the calculation of the parent nodes of these 79 nodes will be incorrect.

[0046] As shown in Table 2, the number of computing nodes in the second test example is 2027, and the number of computing nodes in the 0th layer is 995, among which 53 have the situation of sharing a common parent node. If the thread mutex implemented by the atomic CAS function in CUDA is used, although it can ensure the independence of the access to a certain piece of memory during multi-threaded parallelism, after the cuda mutex locks a certain piece of memory, all threads executing the current statement have to wait for this piece of memory to be unlocked before they can continue to execute, greatly reducing the parallel computing efficiency.

[0047] In Example 1, the 79 mutex threads are serial. After they finish execution, other threads can be parallel. When performing matrix decomposition in the above example, the average out-degree of each node is 3. Therefore, on average, there are 8 multiplication and addition calculations for processing one node in matrix decomposition. For a Linux server with a main frequency of 2.1 GHz, 8G of memory, and an nvidia A40 GPU running the program, the calculation time of the above 0th layer is compared as follows, and the average value is taken after multiple calculations.

[0048] Table 1: Example 1

[0049] Table 2: Example 2

[0050] From the comparison time of Table 1 and Table 2 above, the method of the present invention greatly improves the computing efficiency while ensuring the calculation accuracy.

Claims

1. A GPU parallel solution method for linear equations in a power system, characterized in that, The steps are as follows: Define an initial node among all the nodes of the power system, and divide the remaining nodes into different levels according to the distance from the initial node, so as to establish a sparse decomposition thread tree; The child nodes with the same parent node in the same level are called sibling nodes, and a sibling node set is established; Allocate threads in the GPU according to the sibling node set; Perform parallel solution calculations for the sparse matrix of each node in the same level.

2. The method for parallel solving of the linear equation set of the power system by GPU according to claim 1, characterized in that The steps of establishing the sparse decomposition thread tree include: Describe all the nodes of the power system by using a weighted directed factor graph; p, j, k, and l are all nodes for power flow calculation, where p is the power generation point, and j, k, and l are the power receiving points; On the weighted directed factor graph, among the edges emitted from each node, select the edge with the smallest receiving point number as the path on the sparse decomposition path graph.

3. The method for parallel solving of the linear equations of the power system by GPU according to claim 2, wherein, In the parallel solution calculation of the sparse matrix for each node in the same level, the sparse matrix calculations for the same-level nodes are calculated in parallel at the same time. The specific process includes: On the sparse decomposition path graph of node j, the definition of the parent node number parent(j) of node j is: ; In the formula, is the element in the \(j\)-th column of the \(p\)-th row of the lower triangular matrix \(L\) in the Jacobian matrix \(A\). It is said that node \(p\) is the parent node of node \(j\), and node \(j\) is the child node of node \(p\); The in-degree In_Degree(i) of node i is defined as the sum of the number of nodes connected to node i and with node numbers less than i, that is: ; wherein, is the edge weight of each node j pointing to node i, where i = p, j, k or l; The out-degree Out_Degree(i) of node i is defined as the sum of the number of nodes connected to node i and with node numbers greater than i, that is: ; Among them, is the edge weight of node j starting from node i.

4. The method for parallel solving of the linear equation system of the power system by GPU according to claim 3, wherein In the parallel solution calculation of the sparse matrix for each node in the same level, only when the child node calculation is completed, the parent node starts to calculate, so as to realize the stratification of the sparse decomposition path graph. The specific stratification process includes: Initialize the levels of all nodes; ; In the formula, level(i) represents the i-th level; Traverse all nodes, and assign the nodes with in-degree 0 to level 0; ; For the nodes with in-degree greater than 0, the level of node i is defined as: ; wherein, is the column number of the non-zero non-diagonal element in the i-th row of the lower triangular matrix L.

5. The GPU parallel solution method for the linear equation system of the power system according to claim 3 or 4, characterized in that The solution process for the linear equation of each node includes: Decompose the linear equation by using the triangular decomposition algorithm based on Gaussian elimination to obtain a triangular matrix; Use the decomposed triangular matrix for forward substitution and back substitution calculations to solve the unknown variables of the linear equation.

6. The method for parallel solving of the linear equation system of the power system by GPU according to claim 5, characterized in that In the forward substitution process, perform forward substitution calculations in parallel according to the different layer sets to which the nodes belong; The triangular decomposition algorithm and the forward substitution calculation start from layer i = 0 to layer i = n - 1, and the back substitution calculation ends from layer i = n - 1 to layer i = 0; In each layer, allocate the number of GPU threads according to the number of sibling node sets in the current layer.

7. The method for parallel solving of the linear equation system of the power system by GPU according to claim 5, characterized in that, The triangular decomposition algorithm includes the following steps: Decompose the Jacobian matrix formed by the linear equations by using the triangular decomposition algorithm into an upper triangular matrix and a lower triangular matrix; solve the unknown variables of the nodes, and the unknown variables include the voltage amplitude and voltage phase angle of the nodes; Update and correct the diagonal elements of the upper triangular matrix; at the same time, correct the non-zero elements of the upper triangular matrix.

8. The method for parallel solving of linear equations of a power system by GPU according to claim 5, characterized in that, The operation processes of the forward substitution and the back substitution are both: ; Among them, is the i-th element of the unknown variable; is the unbalance; is the element in the i-th row and j-th column of the upper triangular matrix; Update the unknown variables of the nodes of the weighted directed factor graph through forward substitution and back substitution calculations.

9. A parallel solution system for implementing the GPU parallel solution method of the linear equation set of the power system according to any one of claims 1-8, characterized in that, It includes a sparse decomposition thread tree establishment module, a node definition module, a thread allocation module, and a GPU; The sparse decomposition thread tree establishment module is used to generate a sparse decomposition thread tree for all nodes in the power system; The node definition module is used to define the nodes of the sparse decomposition thread tree to form a sibling node set; The thread allocation module is used to match threads according to the sibling node set of the sparse decomposition thread tree; The GPU performs parallel calculations according to threads at the same level.

10. The parallel solution system of the GPU parallel solution method for the power system linear equation set according to claim 9, characterized in that, The GPU includes a triangular decomposition module and a forward and backward substitution calculation module; The triangular decomposition module is used to decompose a linear equation based on the triangular decomposition algorithm of Gaussian elimination to obtain a triangular matrix; The forward and backward substitution calculation module is used to perform forward and backward substitution calculations using the decomposed triangular matrix to solve the unknown variables of the linear equation.

Citation Information

Patent Citations

  • Large-scale power grid flow correction equation parallel solving method

    CN104158182A

  • Load flow parallel computing method for power system

    CN110649624A

  • Electric power system load flow calculation method based on operation tree GPU parallel acceleration model

    CN111740424A

  • Electricity-gas comprehensive energy system collaborative optimization method

    CN112330020A

  • Method and device for solving sparse triangular matrix of power system based on double-layer division

    CN116307113A

Cited By

  • Layered iterative parallel computing method and device for load flow sparse equation of multiple leading edge surfaces

    CN120613739A