Static security analysis parallelization method and device based on GPU acceleration

By adopting the parallelization method of static safety analysis based on GPU acceleration in the power system, the problem of excessive time-consuming calculation of large-scale power grid systems is solved, efficient static safety analysis is achieved, real-time analysis needs are met, and the stability of the power system is improved.

CN119988070APending Publication Date: 2025-05-13STATE GRID LIAONING ELECTRIC POWER CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411995801.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When handling large-scale power grid systems, the calculation volume is large and the scene combination explosion is expected, resulting in too long calculation time and it is difficult to meet real-time requirements. Data communication between distributed computing platforms has become a bottleneck, which seriously restricts computing efficiency.

Method used

The static security analysis parallelization method based on GPU acceleration is adopted. By obtaining system information on the CPU and transmitting it to the GPU, the admission matrix batch processing and parallel modification of the induction matrix in multiple different fault scenarios is carried out to form an extended correction equation system, solve it in parallel, and finally perform trend calculation and cross-limit judgment.

Benefits of technology

It significantly improves the computing efficiency of static safety analysis, can meet the real-time analysis needs of the evolution of large-scale power grid chain fault paths, and improves the stability of the power system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988070A_ABST
    Figure CN119988070A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of electrical engineering, and particularly relates to a static safety analysis parallelization method and device based on GPU acceleration. The method comprises the following steps: acquiring system information after disturbance on a CPU (Central Processing Unit), and transmitting the system information to a GPU (Graphics Processing Unit); according to the system information, batch processing and parallel modification are carried out on the admittance matrixes under the multiple different fault scenes on the GPU, and a corrected admittance matrix is obtained; obtaining an expansion correction equation according to the corrected admittance matrix; analyzing expansion and correction equation sets under different fault scenes according to static safety, and performing parallel solution on the expansion and correction equation sets; and according to a load flow calculation result, load flow distribution out-of-limit judgment and screening under different fault scenes are carried out, and a fault set and a sorting result are transmitted back to the CPU. According to the method, the calculation speed and efficiency of static security analysis are remarkably improved, the number of calculation scenes is reduced, and the system security and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of electrical engineering technology, and in particular relates to a static safety analysis parallelization method and device based on GPU acceleration. Background Art

[0002] In recent years, the frequent occurrence of extreme disasters has increased the risk of large-scale power outages. The expansion of the scale and complexity of the power system has put forward higher requirements for disaster warning and the study of the evolution of cascading failure paths. In the analysis of the evolution of cascading failure paths in power systems, static safety analysis is a key step in determining the set of high-risk fault components after the initial disturbance. Static safety analysis is an analytical method for evaluating whether the power system is safe under normal and abnormal operating conditions. GPU is the abbreviation of graphics processing unit. Compared with CPU, which is good at processing logic-type instructions, it can show better performance in processing data-intensive calculations. A cascading failure of a power system refers to a new failure that occurs when the failure of the previous fault component ends but the system state is not healthy enough, and there is a causal relationship between the successive failures of the components.

[0003] Traditional methods mostly use serial computing or parallel computing based on multi-core CPUs. However, when dealing with large-scale power grid systems, these methods take too long to calculate due to the large amount of calculation and the explosion of scenario combinations, making it difficult to meet real-time requirements. In particular, when the system scale increases, data communication between distributed computing platforms becomes a bottleneck, which seriously restricts computing efficiency.

[0004] Although some studies have attempted to achieve parallel acceleration through database distributed computing, multi-wavefront algorithms, and other methods, these existing methods either lack parallelism or suffer from reduced algorithm effectiveness and insufficient computational accuracy due to matrix splitting, making it difficult to fully utilize the parallel computing advantages of GPUs.

[0005] In view of the problems that the above-mentioned existing technologies fail to give full play to the parallel computing advantages of GPU, have limited acceleration effects, and are difficult to meet real-time requirements, further technical updates and improvements are urgently needed. Summary of the invention

[0006] In view of the shortcomings of the above-mentioned prior art, the present invention provides a method and device for parallelizing static safety analysis based on GPU acceleration. The purpose is to achieve efficient use of GPU parallel computing capabilities, significantly improve the efficiency of static safety analysis calculations, and meet the real-time analysis needs of the evolution of large-scale power grid cascading fault paths.

[0007] The technical solution adopted by the present invention to achieve the above-mentioned purpose is:

[0008] The parallelization method of static security analysis based on GPU acceleration includes:

[0009] Obtain system information after the disturbance on the CPU, and transmit the system information to the GPU;

[0010] Based on the obtained system information after the disturbance, the admittance matrices under multiple different fault scenarios in CUDA are modified in batches and in parallel on the GPU to obtain a corrected admittance matrix;

[0011] According to the corrected admittance matrix, the extended correction equation group under different fault scenarios is obtained;

[0012] According to the extended correction equation group under different fault scenarios, the extended correction equation group is solved in parallel to obtain the power flow calculation results under different fault scenarios;

[0013] According to the power flow calculation results, the power flow distribution under different fault scenarios is judged to be out of limit, the fault scenarios that exceed the limit are screened out, and the fault set and sorting results are transmitted back to the CPU.

[0014] Furthermore, the system information after the disturbance includes: the admittance matrix after the disturbance, the power of each node and the initial value of the voltage vector.

[0015] Furthermore, the method of performing batch processing and parallel modification on the GPU for the admittance matrices under multiple different fault scenarios according to the obtained system information after the disturbance occurs, to obtain a corrected admittance matrix, includes:

[0016] Create a kernel function, index the mutual admittance element of the admittance matrix according to the thread index, and calculate the branch number corresponding to the thread number index in parallel. The k-th thread calculates the mutual admittance of the k-th branch. The parallel design of self-admittance is the same as the parallel design of mutual admittance.

[0017] Using the unified computing device architecture CUDA, the initialized admittance matrix is ​​copied to N fault , N fault is the number of fault scenarios;

[0018] Create multiple threads, index the fault scenario by thread, and calculate the correction elements of the admittance matrix of the fault scenario; for the case of 1 component failure, correct 4 elements; for the case of N component failures, correct 4*N elements;

[0019] The batch processing of AC power flow calculation and parallel modification of the admittance matrix under multiple different fault scenarios includes: forming and modifying the batch processing admittance matrix, forming the batch processing Jacobian matrix and calculating the batch processing node injection power;

[0020] The formation and modification of the batch admittance matrix includes:

[0021] (1) Create a kernel function on the GPU to perform coarse-grained parallel calculations on the mutual admittance and self-admittance elements of the admittance matrix;

[0022] (2) For different fault scenarios, the batch admittance matrix is ​​modified by copying the initial admittance matrix and modifying the admittance matrix elements corresponding to the faulty component positions.

[0023] Furthermore, the extended correction equation is obtained according to the corrected admittance matrix, including:

[0024] Create a kernel function on the GPU and use an N fault ×(n+m-1) length vector, where n is the number of system nodes minus one and m is the number of PQ nodes, storing the right-hand side vectors under different fault scenarios;

[0025] Calculate the node injection power:

[0026]

[0027] In the above formula: P k represents the node injection power, k represents the thread index number of the kernel function SKernel, NYseq[k] represents the starting index of the non-zero elements in the kth row of the admittance matrix, NYseq[k+1]-1 represents the ending index of the non-zero elements in the kth row of the admittance matrix, V j represents the voltage value corresponding to node j, θ k represents the voltage phase angle corresponding to node k, θ j represents the voltage phase angle corresponding to the j node, G represents the conductance element corresponding to the admittance matrix, and B represents the susceptance element corresponding to the admittance matrix;

[0028] Calculate the injection power of node k with thread index k=threadIdx.x, and index fault scenario i with thread coordinate i=threadIdx.y, fetch the admittance matrix element of the i-th fault scenario from the memory corresponding to the GPU, calculate the corresponding node injection power, and store the calculation result at the i×(n+m-1) position of the vector in parallel;

[0029] At the fine-grained parallel level, batch parallel formation and modification of the admittance matrix, node injection power, and Jacobian matrix;

[0030] The batch processing parallel design of the Jacobian matrix refers to the parallel design of the admittance matrix elements, including:

[0031] A single Jacobian matrix is ​​designed in parallel, including four sub-matrices Hn×n, Nn×m, Mm×n and Lm×m; during the parallel design, the influence of the node type is considered, and different index vectors "H2Y, N2Y, M2Y, L2Y" are created to form a corresponding relationship between the Jacobian sub-matrix and the admittance matrix, which is used as the subscript index during parallel calculation; different index vectors "H2Ja, N2Ja, M2Ja, L2Ja" are created to form an element correspondence between the Jacobian matrix and its sub-matrix; to form a corresponding relationship, kernel functions "HKernel", "NKernel", "MKernel" and "LKernel" are created to parallelly calculate the Jacobian sub-matrix, forming different dimensions of different sub-matrices; the elements of the sub-matrix are calculated by thread in the GPU, and stored in the corresponding position of the Jacobian matrix according to the formed corresponding relationship;

[0032] Create the kernel function YaKernel and copy the obtained Jacobian matrix to the GPU fault The copies are stored in a sparse format, and a thread i = threadIdx.x is created to index the fault scenario number. The admittance matrix element corresponding to the memory is indexed according to the fault number. The Jacobian matrix under different fault scenarios is modified in batches in parallel to realize the parallel design of the extended correction equation and obtain the extended correction equation.

[0033] Furthermore, the extended correction equations under different fault scenarios are solved in parallel, and the power flow calculation results under different fault scenarios are calculated; including:

[0034] Perform reverse Cuthill-McKee reordering on the extended correction equation to reduce the condition number of the matrix, which is implemented on the GPU through cusolverSpXcsrsymrcmHost in cusolver;

[0035] The function cusolverSpXcsrluAnalysisHost is used to perform symbolic analysis on the coefficient matrix, laying the foundation for subsequent batch solution. The correction equation under a single fault scenario is solved in parallel by cusolverSpDcsrluSolveHost. The error residual is calculated in parallel by cusparseDcsrmv. The necessary parameters for batch parallel calculation are set by the function cusolverRfBatchSetupHost. The extended correction equation is solved in batches by the function cusolverRfBatchSolve, and the power flow calculation results under different fault scenarios are obtained.

[0036] Furthermore, the above-mentioned method of judging the over-limit of the power flow distribution under different fault scenarios according to the power flow calculation result, screening out the over-limit fault scenarios, and transmitting the fault set and the sorting result back to the CPU includes:

[0037] Perform batch power flow over-limit judgment in GPU;

[0038] The judgment index includes the power limit index of each branch and the voltage limit index of each node. The formula is as follows:

[0039] |P ij |≤P ij max

[0040] V ij min ≤V ij ≤V ij max

[0041] In the above formula: P ij and P ij max They represent the power and upper power limit of branch j under fault scenario i respectively; V ij 、V ij min and V ij max They represent the voltage, voltage lower limit and voltage upper limit of node j under fault scenario i respectively;

[0042] The parallel design of over-limit judgment adopts the creation of multiple threads, indexing the fault scenario of static safety analysis by thread i = threadIdx.y, indexing the voltage of each node and the power of each branch under fault scenario i by k = threadIdx.x, and comparing the values ​​of i×(n+m-1) position in the voltage vector and the power vector;

[0043] When the power or voltage exceeds the limit, the fault scenario i corresponding to the limit is recorded in the fault set, the limit percentage is calculated to sort the faults, and finally the fault set and fault sorting results are output to the CPU.

[0044] The parallel design device for static security analysis based on GPU acceleration includes:

[0045] An acquisition module, used for acquiring system information after disturbance on the CPU and transmitting the system information to the GPU;

[0046] A parallel modification module is used to perform batch processing and parallel modification on the admittance matrices under multiple different fault scenarios on the GPU according to the obtained system information after the disturbance occurs, so as to obtain a corrected admittance matrix;

[0047] A calculation module, used for obtaining an extended correction equation according to the corrected admittance matrix;

[0048] A parallel solution module is used to perform parallel solutions of the extended correction equations according to the extended correction equations under different fault scenarios of static safety analysis;

[0049] The judgment module is used to judge the over-limit of the power flow distribution under different fault scenarios according to the power flow calculation results, filter out the over-limit fault scenarios, and transmit the fault set and sorting results back to the CPU.

[0050] Furthermore, the GPU-accelerated static security analysis parallelization device is used to implement the steps of any one of the GPU-accelerated static security analysis parallelization methods.

[0051] A computer device comprises a storage medium, a processor and a computer program stored on the storage medium and executable on the processor, wherein when the processor executes the computer program, the steps of any one of the methods for parallelizing static security analysis based on GPU acceleration are implemented.

[0052] A computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of any one of the methods for parallelizing static security analysis based on GPU acceleration are implemented.

[0053] The present invention has the following beneficial effects and advantages:

[0054] The present invention proposes a parallel static safety analysis method based on GPU acceleration, which makes full use of the parallel computing capability of GPU, significantly improves the computational efficiency of static safety analysis, effectively addresses the problems of combinatorial explosion and exponential growth of the number of scenarios in the analysis of the evolution of cascading fault paths, provides strong support for the stable operation of the power system, and meets the real-time analysis requirements of the evolution of cascading fault paths in large-scale power grids.

[0055] The method of the present invention realizes parallel processing of power flow calculation under various fault scenarios by exploiting the coarse-grained parallelism and fine-grained parallelism in static safety analysis. In terms of coarse-grained parallelism, the multi-threaded characteristics of GPU are utilized, and each thread independently calculates the AC power flow under a fault scenario, thereby doubling the calculation speed. In terms of fine-grained parallelism, batch parallel design is performed for the formation of the admittance matrix, node injection power, Jacobian matrix and the solution of the extended correction equation under each fault scenario, further optimizing memory access and reducing unnecessary calculation steps.

[0056] The present invention realizes batch parallel formation of admittance matrix, node injection power and Jacobian matrix. Through CUDA programming, the corresponding kernel function is created to realize parallel modification of admittance matrix elements and batch parallel calculation of node injection power and Jacobian matrix under different fault scenarios. This method not only improves the calculation efficiency, but also reduces the time loss of data communication.

[0057] The present invention proposes a parallel solution strategy for the extended correction equation. In view of the large scale and high condition number of the extended correction equation, the reverse Cuthill-McKee reordering method is used to reduce the matrix condition number, and the cusolver API of CUDA is used to realize batch parallel solution of the extended correction equation. This method effectively improves the convergence speed and stability of the solution.

[0058] The present invention realizes the parallelization of over-limit judgment and fault set screening. In the GPU, it is judged in parallel whether the power flow distribution of each fault scenario is over-limit, and the over-limit fault scenarios and the ranking results of their severity are output to the CPU. This method reduces the transmission time of a large amount of data between the GPU and the CPU, and improves the efficiency of fault set screening. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The above and / or additional aspects and advantages of the present invention will become apparent and easily understood from the description of the embodiments in conjunction with the following drawings, in which:

[0060] Figure 1 It is a flow chart of parallel design of static security analysis based on GPU of the present invention;

[0061] Figure 2 It is a calculation flow chart of the GPU-based static security analysis parallel method of the present invention. DETAILED DESCRIPTION

[0062] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present invention and the features in the embodiments can be combined with each other without conflict.

[0063] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0064] Refer to the following Figure 1-Figure 2 The technical solutions of some embodiments of the present invention are described.

[0065] Example 1

[0066] The present invention provides an embodiment, which is a parallel static security analysis method based on GPU acceleration. Figure 1 As shown, Figure 1 This is a flow chart of the parallel design of static security analysis based on GPU in the present invention. It specifically includes the following steps:

[0067] Step 1. Obtain system information after the disturbance on the CPU and transmit the system information to the GPU.

[0068] The system information after the disturbance includes: the admittance matrix after the disturbance, the power of each node, the initial value of the voltage vector, etc.

[0069] Step 2. Based on the system information after the disturbance obtained in step 1, batch processing and parallel modification of the admittance matrices under multiple different fault scenarios are performed on the GPU. The final modified admittance matrix specifically includes:

[0070] Create a kernel function "adMatrixKernel", index the mutual admittance elements of the admittance matrix according to the thread index, and calculate the mutual admittance parallelly by the branch number threadIdx.x corresponding to the thread number index. The k-th thread calculates the mutual admittance of the k-th branch. The parallel design of self-admittance is the same as the parallel design of mutual admittance.

[0071] The initialized admittance matrix is ​​copied to N using the memory copy Cudamemcopy in the unified computing device architecture CUDA (Compute Unified Device Architecture). fault , N fault is the number of fault scenarios;

[0072] Create multiple threads, index the fault scenario by thread, and calculate the correction elements of the admittance matrix of the corresponding fault scenario; if one component fails, 4 elements need to be corrected; if N components fail, 4*N elements need to be corrected.

[0073] The batch processing and parallel modification of the admittance matrix under multiple different fault scenarios includes: forming and modifying the batch processing admittance matrix, forming the batch processing Jacobian matrix, and calculating the batch processing node injection power.

[0074] The formation and modification of the batch admittance matrix specifically includes the following steps:

[0075] (1) On the GPU, a kernel function is created to perform coarse-grained parallel calculations on the mutual admittance and self-admittance elements of the admittance matrix;

[0076] (2) For different fault scenarios, the batch admittance matrix is ​​modified by copying the initial admittance matrix and modifying the admittance matrix elements corresponding to the faulty component positions.

[0077] Step 3. According to the modified admittance matrix obtained in step 2, a Jacobian matrix is ​​formed, the node injection power is calculated, and an extended correction equation is obtained; specifically, the steps include:

[0078] Create a kernel function on the GPU and use an Nfault ×(n+m-1) length vector, where n is the number of system nodes minus one and m is the number of PQ nodes, storing the right-hand side vectors under different fault scenarios;

[0079] The node injection power is calculated according to the following formula:

[0080]

[0081] In the above formula: P k represents the node injection power; k represents the thread index number of the kernel function SKernel; NYseq[k] represents the starting index of the non-zero elements in the kth row of the admittance matrix; NYseq[k+1]-1 represents the ending index of the non-zero elements in the kth row of the admittance matrix; V j represents the voltage value corresponding to node j; θ k represents the voltage phase angle corresponding to node k; θ j represents the voltage phase angle corresponding to the j node; G represents the conductance element corresponding to the admittance matrix; B represents the susceptance element corresponding to the admittance matrix.

[0082] Calculate the injection power of node k with thread index k=threadIdx.x, and index fault scenario i with thread coordinate i=threadIdx.y, fetch the admittance matrix element of the i-th fault scenario from the memory corresponding to the GPU, calculate the corresponding node injection power, and store the calculation result at the i×(n+m-1) position of the vector in parallel;

[0083] At the fine-grained parallel level, batch parallel formation and modification of the admittance matrix, node injection power, and Jacobian matrix;

[0084] The batch parallel design of Jacobian matrix elements can refer to the parallel design of admittance matrix elements:

[0085] First, a single Jacobian matrix is ​​designed in parallel, including four sub-matrices Hn×n, Nn×m, Mm×n and Lm×m. When designing in parallel, it is necessary to consider the influence of the node type and create different index vectors "H2Y, N2Y, M2Y, L2Y" to form the correspondence between the Jacobian sub-matrix and the admittance matrix, which is used as the subscript index during parallel calculation; it is also necessary to create different index vectors "H2Ja, N2Ja, M2Ja, L2Ja" to form the element correspondence between the Jacobian matrix and its sub-matrix; then, based on the corresponding relationship, the kernel functions "HKernel", "NKernel", "MKernel" and "LKernel" are created to parallelly calculate the Jacobian sub-matrix, where it is necessary to pay attention to the different dimensions of different sub-matrices. The elements of the sub-matrix are calculated by thread in the GPU, and stored in the corresponding position of the Jacobian matrix according to the corresponding relationship formed above.

[0086] On this basis, we create the kernel function YaKernel and copy the obtained Jacobian matrix to N on the GPU. fault The copies are stored in a sparse format, and then a thread i = threadIdx.x is created to index the fault scenario number, and the admittance matrix element of the corresponding memory is indexed according to the fault number. The Jacobian matrix under different fault scenarios is modified in batches in parallel, and finally the purpose of parallel design of the extended correction equation is achieved, and the extended correction equation is obtained.

[0087] Step 4. According to step 3, the extended correction equations under different fault scenarios are obtained, and the extended correction equations are solved in parallel to obtain the power flow calculation results under different fault scenarios. Figure 2 As shown, Figure 2 It is a calculation flow chart of the GPU-based static security analysis parallel method of the present invention, which specifically includes the following steps:

[0088] Perform a reverse Cuthill-McKee rearrangement of the extended correction equation.

[0089] Cuthill-McKee is a reordering algorithm used to reduce the bandwidth of sparse matrices. It can reduce the condition number of the matrix, thereby improving numerical stability and solution efficiency. It can reduce the condition number of the matrix. On the GPU, it can be implemented through cusolverSpXcsrsymrcmHost in the CUDA solver cusolver; a function in the NVIDIA cuSolver library is used to perform reverse Cuthill-McKee reordering of symmetric sparse matrices on the host Host to reduce the bandwidth of the matrix.

[0090] The function cusolverSpXcsrluAnalysisHost is used to perform symbolic analysis on the coefficient matrix, laying the foundation for subsequent batch processing. Then, cusolverSpDcsrluSolveHost is used to solve the correction equations in a single fault scenario in parallel, and cusparseDcsrmv is used to perform parallel calculations on the error residuals. The function cusolverRfBatchSetupHost is used to set the necessary parameters for batch parallel computing, and finally, the matrix batch solver function cusolverRfBatchSolve is used to perform batch processing on the extended correction equations to obtain the power flow calculation results under different fault scenarios.

[0091] The cusolverSpXcsrluAnalysisHost is used for symbolic analysis of LU decomposition of a sparse matrix on the host side, and is used to determine the filling in the decomposition process, that is, the increase of non-zero elements.

[0092] The cusolverSpDcsrluSolveHost is a function for solving the sparse matrix LU decomposition performed on the host side.

[0093] The cusparseDcsrmv is a function in the NVIDIA cuSparse library, which is used to perform matrix-vector multiplication operations on sparse matrices and vectors.

[0094] The matrix batch solver function cusolverRfBatchSetupHost is a necessary parameter for setting batch parallel computing.

[0095] Step 5. According to the power flow calculation results obtained in step 4, the power flow distribution under different fault scenarios is judged to be out of limit, the fault scenarios that exceed the limit are screened out, and the fault set and sorting results are transmitted back to the CPU. The specific steps include the following:

[0096] In order to avoid wasting a lot of data transmission time, batch processing power flow over-limit judgment is performed in the GPU. The judgment indicators include the power over-limit index of each branch and the voltage over-limit index of each node. The formula is as follows:

[0097] |P ij |≤P ij max

[0098] V ij min ≤V ij ≤V ij max

[0099] In the above formula: P ij and P ij max They represent the power and upper power limit of branch j under fault scenario i respectively; V ij 、V ij min and V ij max They represent the voltage, voltage lower limit and voltage upper limit of node j under fault scenario i respectively.

[0100] The parallel design of over-limit judgment also uses the creation of multiple threads. The fault scenario of static safety analysis is indexed by thread i = threadIdx.y, and the voltage of each node and the power of each branch under fault scenario i are indexed by k = threadIdx.x. The values ​​of the i×(n+m-1) position in the voltage vector and the power vector are compared.

[0101] If the power or voltage exceeds the limit, the fault scenario i corresponding to the limit is recorded in the fault set, and then the limit percentage is calculated to sort the faults, and finally the fault set and fault sorting results are output to the CPU.

[0102] Example 2

[0103] The present invention further provides an embodiment, which is a static security analysis parallelization method based on GPU acceleration.

[0104] The present invention uses the Polish high-voltage transmission network modified by IEEE 118, IEEE 300, case2383, and the European high-voltage transmission network system of case9241 in Matpower as examples to verify the feasibility of the proposed static safety analysis parallelization design method based on GPU acceleration.

[0105] For the formation of the admittance matrix, taking the 2383-node Polish high-voltage transmission network system as an example, the serial calculation takes 5.3ms, while the GPU batch calculation takes only 1.3ms. The acceleration ratio is defined to reflect the acceleration effect of the method, and its value is the ratio of the operation time of the serial algorithm to the parallel acceleration algorithm of the present invention, and its acceleration ratio is 4.08. Since the time required to calculate the admittance matrix is ​​short and there is no need to repeat the calculation during the batch solution process, the calculation of the batch admittance matrix has little effect on the acceleration effect of the overall calculation. Therefore, no further analysis is performed on the examples of other node scales.

[0106] During the solution process, the coefficient matrix of the modified equation needs to be continuously updated, so the parallel acceleration design of this process has a great impact on the acceleration effect of the overall calculation. In the test cases of four nodes of different scales, the number of nodes in the IEEE 118 and IEEE300 node systems is small, and the batch Jacobian matrix formation takes less time. The case2383 and case9241 node systems are tested and analyzed below. As shown in Table 1, Table 1 is a time comparison of the serial calculation and batch calculation of the Jacobian matrix. The test data shows that the acceleration ratios of the case2383 node system and the case9241 node system are 8.12 and 6.63 respectively, indicating that the calculation of the GPU batch Jacobian matrix formation can achieve a good acceleration effect. However, the speedup ratio of the 9241-scale system is slightly lower than that of the 2383-node system. The main reason is that the processing efficiency of the data transmission link in the program is not ideal. When the system scale increases, the data transmission between CPUs becomes more frequent and the transmission takes longer, which eventually leads to a slight decrease in the speedup ratio. However, the method of the present invention has achieved considerable acceleration effect as a whole, and greatly reduces the calculation time compared to serial calculation. It can be expected that the parallel batch calculation of the Jacobian matrix will also produce good effects on larger-scale systems.

[0107]

[0108] Table 1 Comparison of the time consumption of Jacobian matrix serial calculation and batch calculation

[0109] The acceleration effect of the overall calculation of the GPU parallel static safety analysis module is analyzed below, as shown in Table 2, which is a comparison of the overall time consumption of GPU parallel static safety analysis calculations for systems of different scales. From the calculation results of the IEEE 118-node and IEEE 300-node systems, it can be seen that due to the small node scale, the batch scale and the system scale are close, and the speedup ratio is less than 2, so GPU parallel static safety analysis has a certain acceleration effect for small-scale systems, but the acceleration effect of parallel computing is limited.

[0110] From the calculation results of the 2383-node and 9241-node systems, good acceleration effects can be obtained for large systems. The acceleration ratio in the case2383 system is 5.29, and the acceleration ratio in the case9241 system is 5.08. Theoretically, as the system scale increases, the GPU parallelism is greater, and there will be a greater acceleration ratio, but the acceleration ratio of the case9241 system is slightly reduced. The reason is that on the one hand, the time loss of frequent data transmission between the CPU and the GPU is too large, resulting in an acceleration effect lower than the theoretical situation; on the other hand, the scale of the batch processing is too large, and the scale of the extended correction equation will increase. In addition, the Jacobian matrix in the iteration process needs to be updated each time. Compared with the fixed Jacobian matrix, there is a certain time loss in the calculation process. The above factors combined lead to a slight reduction in the final acceleration ratio.

[0111] Therefore, the batch size of GPU parallel static security analysis should not be set too large, nor too small. If the batch size is too small, it will lead to low parallelism, reduced GPU acceleration efficiency, and reduced speedup ratio.

[0112]

[0113]

[0114] Table 2 Comparison of GPU parallel SSA calculation time of different scale systems

[0115] Compared with traditional methods, the GPU parallel static safety analysis algorithm proposed in this invention achieves better acceleration effect. Taking the Polish high-voltage transmission network and the European high-voltage transmission network system modified by IEEE 118, IEEE 300, case2383 in Matpower as examples, it is first explained that the acceleration effect of batch calculation of the coefficient matrix of the modified equation can achieve a 5 to 8 times acceleration ratio for large-scale systems. The batch parallel execution of the Jacobian matrix modification process will greatly save the calculation time. Then the overall acceleration effect of batch parallel static safety analysis is analyzed. The algorithm has a general acceleration effect on small-scale systems, but a significant acceleration effect on large-scale systems, and the reasons for the decrease in the acceleration ratio are analyzed.

[0116] The results show that the parallel design method of static safety analysis based on GPU acceleration can greatly improve the computing efficiency, meet the real-time analysis requirements of the evolution of large-scale power grid cascading fault paths, and improve the stability of the power system.

[0117] Example 3

[0118] The present invention further provides an embodiment, which is a static security analysis parallelization device based on GPU acceleration, including: a static security analysis parallelization device based on GPU acceleration, characterized in that: it includes:

[0119] An acquisition module, used for acquiring system information after disturbance on the CPU and transmitting the system information to the GPU;

[0120] A parallel modification module is used to perform batch processing and parallel modification on the admittance matrices in CUDA under multiple different fault scenarios on the GPU according to the obtained system information after the disturbance occurs, and finally obtain the corrected admittance matrix;

[0121] A calculation module, used for obtaining an extended correction equation group under different fault scenarios according to the corrected admittance matrix;

[0122] A parallel solution module is used to perform parallel solutions of the extended correction equations according to the extended correction equations under different fault scenarios to obtain the power flow calculation results under different fault scenarios;

[0123] The judgment module is used to judge the over-limit of the power flow distribution under different fault scenarios according to the power flow calculation results, filter out the over-limit fault scenarios, and transmit the fault set and sorting results back to the CPU.

[0124] The GPU-accelerated static security analysis parallelization design device is used to implement the steps of the GPU-accelerated static security analysis parallelization method as described in any one of Embodiments 1 or 2.

[0125] Example 4

[0126] Based on the same inventive concept, an embodiment of the present invention further provides a computer device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor. When the processor executes the computer program, the steps of any one of the GPU-accelerated static security analysis parallelization methods described in Embodiment 1 or 2 are implemented.

[0127] Example 5

[0128] Based on the same inventive concept, an embodiment of the present invention further provides a computer storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the GPU-accelerated static security analysis parallelization methods described in Embodiment 1 or 2 are implemented.

[0129] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0130] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0131] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0133] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A parallel static security analysis method based on GPU acceleration, characterized by: include: Obtain system information after the disturbance on the CPU, and transmit the system information to the GPU; Based on the obtained system information after the disturbance, the admittance matrices under multiple different fault scenarios in CUDA are modified in batches and in parallel on the GPU to obtain a corrected admittance matrix; According to the corrected admittance matrix, the extended correction equation group under different fault scenarios is obtained; According to the extended correction equation group under different fault scenarios, the extended correction equation group is solved in parallel to obtain the power flow calculation results under different fault scenarios; According to the power flow calculation results, the power flow distribution under different fault scenarios is judged to be out of limit, the fault scenarios that exceed the limit are screened out, and the fault set and sorting results are transmitted back to the CPU.

2. The method for parallelizing static security analysis based on GPU acceleration according to claim 1 is characterized in that: The system information after the disturbance includes: the admittance matrix after the disturbance, the power of each node and the initial value of the voltage vector.

3. The method for parallelizing static security analysis based on GPU acceleration according to claim 1 is characterized in that: The method comprises: performing batch processing and parallel modification on the GPU on the admittance matrices in CUDA under multiple different fault scenarios according to the obtained system information after the disturbance occurs, so as to obtain a corrected admittance matrix; comprising: Create a kernel function, index the mutual admittance element of the admittance matrix according to the thread index, and calculate the branch number corresponding to the thread number index in parallel. The k-th thread calculates the mutual admittance of the k-th branch. The parallel design of self-admittance is the same as the parallel design of mutual admittance. Using the unified computing device architecture CUDA, the initialized admittance matrix is ​​copied to N fault , N fault is the number of fault scenarios; Create multiple threads, index the fault scenario by thread, and calculate the correction elements of the admittance matrix of the fault scenario; for the case of 1 component failure, correct 4 elements; for the case of N component failures, correct 4*N elements; The batch processing of AC power flow calculation and parallel modification of the admittance matrix under multiple different fault scenarios includes: forming and modifying the batch processing admittance matrix, forming the batch processing Jacobian matrix and calculating the batch processing node injection power; The formation and modification of the batch admittance matrix includes: (1) Create a kernel function on the GPU to perform coarse-grained parallel calculations on the mutual admittance and self-admittance elements of the admittance matrix; (2) For different fault scenarios, the batch admittance matrix is ​​modified by copying the initial admittance matrix and modifying the admittance matrix elements corresponding to the faulty component positions.

4. The method for parallelizing static security analysis based on GPU acceleration according to claim 1 is characterized in that: According to the modified admittance matrix, an extended correction equation group under different fault scenarios is obtained, including: Create a kernel function on the GPU and use an N fault ×(n+m-1) length vector, where n is the number of system nodes minus one and m is the number of PQ nodes, storing the right-hand side vectors under different fault scenarios; Calculate the node injection power: In the above formula: P k represents the node injection power, k represents the thread index number of the kernel function SKernel, NYseq[k] represents the starting index of the non-zero elements in the kth row of the admittance matrix, NYseq[k+1]-1 represents the ending index of the non-zero elements in the kth row of the admittance matrix, V j represents the voltage value corresponding to node j, θ k represents the voltage phase angle corresponding to node k, θ j represents the voltage phase angle corresponding to the j node, G represents the conductance element corresponding to the admittance matrix, and B represents the susceptance element corresponding to the admittance matrix; Calculate the injection power of node k with thread index k=threadIdx.x, and index fault scenario i with thread coordinate i=threadIdx.y, fetch the admittance matrix element of the i-th fault scenario from the memory corresponding to the GPU, calculate the corresponding node injection power, and store the calculation result at the i×(n+m-1) position of the vector in parallel; At the fine-grained parallel level, batch parallel formation and modification of the admittance matrix, node injection power, and Jacobian matrix; The batch processing parallel design of the Jacobian matrix refers to the parallel design of the admittance matrix elements, including: parallel design of a single Jacobian matrix, including four sub-matrices Hn×n, Nn×m, Mm×n and Lm×m; during the parallel design, considering the influence of the node type, creating different index vectors "H2Y, N2Y, M2Y, L2Y", forming a corresponding relationship between the Jacobian sub-matrix and the admittance matrix, and using it as the subscript index during parallel calculation; creating different index vectors "H2Ja, N2Ja, M2Ja, L2Ja", forming an element corresponding relationship between the Jacobian matrix and its sub-matrix; forming the corresponding relationship, creating kernel functions "HKernel", "NKernel", "MKernel", "LKernel" to parallelly calculate the Jacobian sub-matrix, forming different dimensions of different sub-matrices; calculating the elements of the sub-matrix by thread in the GPU, and storing them in the corresponding position of the Jacobian matrix according to the formed corresponding relationship; Create the kernel function YaKernel and copy the obtained Jacobian matrix to the GPU fault The copies are stored in a sparse format, and a thread i = threadIdx.x is created to index the fault scenario number. The admittance matrix element corresponding to the memory is indexed according to the fault number. The Jacobian matrix under different fault scenarios is modified in batches in parallel to realize the parallel design of the extended correction equation and obtain the extended correction equation.

5. The method for parallelizing static security analysis based on GPU acceleration according to claim 1 is characterized in that: The method comprises: performing parallel solutions to the extended correction equations under different fault scenarios to obtain power flow calculation results under different fault scenarios; Perform reverse Cuthill-McKee reordering on the extended correction equation to reduce the condition number of the matrix, which is implemented on the GPU through cusolverSpXcsrsymrcmHost in cusolver; The function cusolverSpXcsrluAnalysisHost is used to perform symbolic analysis on the coefficient matrix, laying the foundation for subsequent batch solution. The correction equation under a single fault scenario is solved in parallel by cusolverSpDcsrluSolveHost. The error residual is calculated in parallel by cusparseDcsrmv. The necessary parameters for batch parallel calculation are set by the function cusolverRfBatchSetupHost. The matrix batch solution function cusolverRfBatchSolve is used to perform batch solution on the extended correction equation to obtain the power flow calculation results under different fault scenarios.

6. The method for parallelizing static security analysis based on GPU acceleration according to claim 1 is characterized in that: According to the power flow calculation results, the power flow distribution under different fault scenarios is judged to be out of limit, the fault scenarios that exceed the limit are screened out, and the fault set and the sorting results are transmitted back to the CPU, including: Perform batch power flow over-limit judgment in GPU; The judgment index includes the power limit index of each branch and the voltage limit index of each node. The formula is as follows: |P ij |≤P ijmax In ijmin ≤V ij ≤V ijmax In the above formula: P ij and P ijmax They represent the power and upper power limit of branch j under fault scenario i respectively; V ij 、V ijmin and V ijmax They represent the voltage, voltage lower limit and voltage upper limit of node j under fault scenario i respectively; The parallel design of over-limit judgment adopts the creation of multiple threads, indexing the fault scenario of static safety analysis by thread i = threadIdx.y, indexing the voltage of each node and the power of each branch under fault scenario i by k = threadIdx.x, and comparing the values ​​of i×(n+m-1) position in the voltage vector and the power vector; When the power or voltage exceeds the limit, the fault scenario i corresponding to the limit is recorded in the fault set, the limit percentage is calculated to sort the faults, and finally the fault set and fault sorting results are output to the CPU.

7. A parallel static security analysis device based on GPU acceleration, characterized by: include: An acquisition module, used for acquiring system information after disturbance on the CPU and transmitting the system information to the GPU; A parallel modification module is used to perform batch processing and parallel modification on the admittance matrices in CUDA under multiple different fault scenarios on the GPU according to the obtained system information after the disturbance occurs, so as to obtain a corrected admittance matrix; A calculation module, used for obtaining an extended correction equation group under different fault scenarios according to the corrected admittance matrix; A parallel solution module is used to perform parallel solutions of the extended correction equations under different fault scenarios according to the extended correction equations under static safety analysis to obtain the power flow calculation results under different fault scenarios; The judgment module is used to judge the over-limit of the power flow distribution under different fault scenarios according to the power flow calculation results, filter out the over-limit fault scenarios, and transmit the fault set and sorting results back to the CPU.

8. The GPU-accelerated static security analysis parallelization device according to claim 7, characterized in that: The device is used to implement the steps of the GPU-accelerated static security analysis parallelization method described in any one of claims 2 to 6.

9. A computer device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the static security analysis parallelization method based on GPU acceleration described in any one of claims 1 to 6 are implemented.

10. A computer storage medium, characterized in that: The computer storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the static security analysis parallelization method based on GPU acceleration described in any one of claims 1 to 6 are implemented.