Multi-head attention dynamic optimization method oriented to power grid topology self-adaption

Through the adaptive multi-head attention dynamic optimization method for grid topology, the problem that traditional methods are difficult to capture real-time state changes in the power grid are solved, and efficient and accurate grid topology analysis is achieved, meeting the needs of real-time grid operation.

CN120146321AActive Publication Date: 2025-06-13SHANDONG ZHIHECHUANG INFORMATION TECH CO LTD +1
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510616844.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

Traditional grid topology analysis methods are difficult to quickly and accurately capture real-time state changes in the power grid, and traditional multi-head attention mechanisms are difficult to adaptively select effective features, which cannot meet the needs of real-time operation of the power grid.

Method used

A dynamic optimization method for adaptive multi-head attention is proposed for power grid topology. Through data acquisition and preprocessing, a multi-head attention structure is constructed, including input data processing, adaptive selection of multi-head division, cross-head information interaction and attention calculation, and the focus score is adjusted using scaling dot product attention calculation and grid context information.

Benefits of technology

It realizes adaptive processing of high-dimensional and complex power grid data, accurately captures real-time state changes in the power grid, improves the accuracy and real-time nature of power grid topology analysis, and meets the needs of real-time operation of the power grid.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146321A_ABST
    Figure CN120146321A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of deep learning, and particularly relates to a multi-head attention dynamic optimization method oriented to power grid topology self-adaption. The method comprises the following steps: firstly, collecting and preprocessing a plurality of data such as a power grid node connection relationship and line transmission parameters; then, a multi-head attention structure is constructed, operations such as input data processing, multi-head division adaptive selection and attention calculation are carried out, according to adaptive selection, a neural network is used for extracting features to calculate and select scores to screen heads, and the attention calculation is combined with context information to adjust the scores; and multi-head output is spliced, and final output is obtained through learning weight matrix transformation. Compared with a traditional method, the method has the advantages that high-dimensional complex power grid data can be processed, real-time state changes can be accurately captured, the accuracy and real-time performance of power grid topology analysis are improved, the real-time operation requirement of a power grid is met, and power grid operation is effectively optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of deep learning, and particularly relates to a dynamic optimization method for multi-head attention for power grid topology adaptation. Background Art

[0002] Under the background of the current rapid development and digital transformation of the power grid, the scale of the power grid is constantly expanding, and the topological structure is becoming increasingly complex, which puts forward higher requirements for the accuracy and real-time performance of power grid topology analysis and operation optimization. Traditional power grid topology analysis methods, such as rule-based and model-based methods, are difficult to handle large-scale and highly dynamic power grid data, and have limitations in dealing with complex topological structures, and cannot quickly and accurately capture the real-time state changes of the power grid. With the development of deep learning technology, the attention mechanism has achieved remarkable results in many fields, but there are still challenges in directly applying the traditional attention mechanism in power grid topology analysis. Power grid data has the characteristics of high-dimensionality, non-linearity, and strong coupling. The traditional multi-head attention mechanism is difficult to adaptively select effective features and lacks sufficient consideration of the dynamic characteristics of power grid operation, resulting in inaccurate analysis results and unable to meet the needs of real-time power grid operation. Summary of the Invention

[0003] The present invention aims at the technical problems existing in the above background art, and proposes a dynamic optimization method for multi-head attention for power grid topology adaptation.

[0004] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0005] S1. Data collection, collecting various data related to the power grid topology, including the connection relationship data of power grid nodes, the transmission parameter data of lines, the power distribution data of each node, and the real-time state data of power grid operation;

[0006] S2. Data processing, preprocessing the collected data;

[0007] S3. Constructing a multi-head attention structure, including input data processing, multi-head division adaptive selection, cross-head information interaction, and attention calculation; the input data processing: performing a linear transformation on the preprocessed data, and respectively obtaining a query vector Q, a key vector K, and a value vector V through learnable weight matrices , , ; multi-head division adaptive selection: dividing the obtained query vector Q, key vector K, and value vector V into multiple heads, introducing adaptive head selection, and dynamically determining whether each head participates in subsequent attention calculation; attention calculation: performing scaled dot-product attention calculation on each head participating in the calculation;

[0008] S4. Concatenating the outputs of all heads participating in the calculation;

[0009] S5. Obtain the final output through the linear transformation of the final learnable weight matrix .

[0010] Preferably, in the multi - head partition adaptive selection in step S3, the query vector Q, key vector K, and value vector V can be partitioned into multiple heads as follows: ; ; ; where represents the feature dimension of each head, h represents the total number of heads, and the query vector Q, key vector K, and value vector V are partitioned into h heads.

[0011] Preferably, after obtaining the weight matrices of h heads, the specific operation of adaptive head selection is as follows: Use a convolutional neural network to extract the features of the power grid topology data; for each head, calculate a selection score; set a selection threshold. If the selection score is greater than the selection threshold, this head participates in the subsequent attention calculation, otherwise this head does not participate in the calculation.

[0012] Preferably, the selection score is calculated using the method of information gain. For each head h, calculate the information gain of its corresponding query vector , key vector and value vector with the input feature X;

[0013] The information gain is obtained by calculating the change in entropy. For the query vector, the information gain , and similarly, can be obtained, , and then the selection score is obtained: , where is the weight coefficient.

[0014] Preferably, between calculating the selection score and setting the selection threshold, it is also necessary to judge the correlation between heads to adjust the calculation score. Mutual information is used to measure the correlation between heads. For each head h, the calculation of its selection score can be adjusted to: , where function is the mutual information function.

[0015] Preferably, in step S3, the specific steps of attention calculation are as follows: Calculate the original attention score; introduce the context information of the power grid operation state, and the context information is obtained through a pre - trained power - grid - related model; according to the context information and the features of the power grid topology data, calculate the adjustment factor for each attention score; the calculation of the adjustment factor is: , where C is the context information, is an adjustment factor for the obtained attention score, is average pooling; multiply the original attention scores element-wise by the adjustment factor to obtain adjusted attention scores, then apply the softmax function to the adjusted attention scores to obtain the final attention weights; finally, multiply the obtained final attention weights of each head by the value vector to obtain the output of each head.

[0016] Compared with the prior art, the advantages and positive effects of the present invention are as follows. In data processing, multi-source power grid data is collected and finely preprocessed to ensure data quality. The multi-head attention structure is adopted to obtain multi-vectors through linear transformation, adaptively divide and screen the heads for participation in calculations, and improve the calculation pertinence and efficiency. The scaled dot-product attention calculation is used to adjust the attention scores in combination with the power grid context information to accurately focus on key features. In the output stage, the multi-head outputs are concatenated and transformed by a learnable weight matrix to enhance the result accuracy. Overall, the present invention can adaptively process high-dimensional and complex power grid data, break through the limitations of traditional methods, accurately capture the real-time state changes of the power grid, effectively improve the accuracy and real-time performance of power grid topology analysis, and meet the real-time operation requirements of the power grid. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following-described drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0018] Figure 1 is the overall structure flowchart of the present invention; Figure 2 is the sub-flowchart for constructing the multi-head attention structure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] In order to be able to more clearly understand the above objects, features, and advantages of the present invention, the following will further illustrate the present invention with reference to the drawings and embodiments. It should be noted that, without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other.

[0020] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Therefore, the present invention is not limited by the specific embodiments disclosed in the following specification.

[0021] In an embodiment, with the continuous expansion of the power grid scale and the increasing complexity of the topological structure, the security and stability of power grid operation face severe challenges. Existing power grid topology analysis methods have problems in quickly and accurately capturing the real-time state changes of the power grid when dealing with large-scale and highly dynamic power grid data. At the same time, directly applying traditional attention mechanisms cannot fully adapt to the high-dimensional, non-linear, and strongly coupled characteristics of power grid data, making it difficult to adaptively select effective features, resulting in the analysis results not meeting the requirements of real-time power grid operation. To achieve precise analysis of the power grid topology, a multi-head attention dynamic optimization method for power grid topology adaptation is proposed.

[0022] First, data collection is carried out, collecting various data related to the power grid topology, including the connection relationship data of power grid nodes, the transmission parameter data of lines, the power distribution data of each node, and the real-time state data of power grid operation.

[0023] Then, the collected data is preprocessed. Using data cleaning techniques, noise and outliers in the data are removed. By setting reasonable upper and lower limit thresholds, data points that deviate significantly from the normal range are identified and excluded to avoid interference from these abnormal data in subsequent analysis. For data with missing values, the KNN algorithm based on machine learning is used for processing, estimating the missing values according to the characteristics of similar data points to ensure the integrity of the data. After that, the data is standardized, using the Z-Score standardization method to unify data with different dimensions to the same scale.

[0024] Next, a multi-head attention structure is constructed, including input data processing, multi-head division adaptive selection, cross-head information interaction, and attention calculation.

[0025] Input data processing is to perform a linear transformation on the preprocessed data, obtaining the query vector Q, key vector K, and value vector V through learnable weight matrices , , respectively. Specifically, through matrix multiplication operations , , , where X is the preprocessed input data. After calculation, the query vector Q, key vector K, and value vector V are obtained.

[0026] Multi-head division adaptive selection: The obtained query vector Q, key vector K, and value vector V are divided into multiple heads, and adaptive head selection is introduced to dynamically determine whether each head participates in subsequent attention calculations. Specifically, in multi-head division adaptive selection, the division of the query vector Q, key vector K, and value vector V into multiple heads can be expressed as: ; ; ; where Denote the feature dimension of each head as \(d\), and \(h\) as the total number of heads. The query vector \(Q\), key vector \(K\), and value vector \(V\) are partitioned into \(h\) heads.

[0027] Then, after obtaining the \(h\) head weight matrices as above, the specific operation of adaptive head selection is as follows: Use a neural network to extract the features of the power grid topology data; for each head, calculate a selection score; set a selection threshold. If the selection score is greater than the selection threshold, this head participates in the subsequent attention calculation, otherwise it does not participate in the calculation.

[0028] Specifically, use a convolutional neural network (CNN) to extract the features of the power grid topology data. First, organize the power grid topology data into a multi-dimensional matrix. The convolutional layer of the CNN slides a learnable convolutional kernel over the data for convolution. Convolutional kernels of different sizes capture local details and macroscopic features respectively. For example, small convolutional kernels focus on the connection characteristics near nodes, and large convolutional kernels obtain the connection patterns of regional power grids. The pooling layer reduces the dimension of the output of the convolutional layer, retains the main features, and reduces the computational amount. Finally, the fully connected layer integrates the features extracted previously and outputs the power grid topology data features for subsequent calculation of the head selection score, which helps to achieve adaptive head selection. Then, for each head \(h\), calculate its corresponding query vector , key vector and value vector with the information gain of the input feature \(X\); the information gain is obtained by calculating the change in entropy. For the query vector, the information gain , and similarly, can be obtained, , and then the selection score is obtained: , where is the weight coefficient. Finally, between calculating the selection score and setting the selection threshold, it is also necessary to judge the correlation between heads to adjust the calculation score. Use mutual information to measure the correlation between heads. For each head \(h\), the calculation of its selection score can be adjusted to: , where function is the mutual information function.

[0029] In addition, between calculating the selection score and setting the selection threshold, it is also necessary to judge the correlation between heads to adjust the calculation score. Use mutual information to measure the correlation between heads. For each head \(h\), the calculation of its selection score can be adjusted to: , where function is the mutual information function. It is used to control the influence degree of head correlation on the selection score. In this way, the higher the correlation between heads, the greater the influence on the selection score, and it can more reasonably screen out the heads participating in the subsequent attention calculation, improving the accuracy and effectiveness of the model for power grid topology analysis.

[0030] Attention calculation: Perform scaled dot-product attention calculation on each head participating in the calculation.

[0031] Specifically, the calculation of the original attention score is as follows. First, calculate and the transposed dot product of, that is . Since the numerical size of the dot product result will affect the performance of the subsequent softmax function, to avoid it being too large and causing gradient disappearance or instability, divide the dot product result by Finally, obtain the original attention score.

[0032] The context information is obtained through a pre-trained power grid-related model. The pre-trained power grid-related model can be an LSTM trained based on historical power grid data. Collect a large amount of historical power grid operation data. Organize these data into sequences in chronological order and input them into the LSTM model. The model captures the laws and trends of power grid operation by continuously learning the time series features and dependencies in historical data. During the training process, adjust the model parameters to minimize the error between the predicted value and the actual value. After training, when new power grid data is input, the model outputs context information containing key information such as power grid operation trends and potential abnormal risks based on the learned laws and in combination with the current input data.

[0033] According to the context information and the characteristics of the power grid topology data, calculate the adjustment factor for each attention score. The calculation of the adjustment factor is as follows: , where C is the context information, is the adjustment factor for the obtained attention score, is average pooling.

[0034] Finally, multiply each element of the original attention score by the adjustment factor to obtain the adjusted attention score. Through this element-wise multiplication, the attention score can fully consider the real-time state and topological characteristics of power grid operation. Then apply the softmax function to the adjusted attention score to obtain the final attention weights. The softmax function converts the scores into a probability distribution, ensuring that the sum of all attention weights is 1, so that the relative importance of different features in the whole can be clearly reflected. Multiply the final attention weight of each head obtained by the value vector to get the output of each head. Each head focuses on data features from different angles. By multiplying with the value vector, the attention weight is converted into the result of information screening and integration, so that the output of each head contains valuable information related to its own focus, providing rich and targeted data for subsequent multi-head splicing.

[0035] Next, the results of all multi-head attention calculations are concatenated and output. Specifically, after completing the attention calculation for each head and obtaining the corresponding output, these outputs are concatenated in a specific dimension in sequence, and then the concatenated vector is input into the final learnable weight matrix for linear transformation. The parameters of the learnable weight matrix are continuously optimized and adjusted during the model training process. The purpose is to further fuse and transform the concatenated feature vectors to better meet the actual requirements of power grid topology analysis. Through this linear transformation operation, the final output is obtained.

[0036] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A multi-head attention dynamic optimization method for power grid topology adaptation, characterized in that: The following steps are involved: S1. Data collection: collecting various data related to the power grid topology, including connection relationship data of power grid nodes, transmission parameter data of lines, power distribution data of each node, and real-time status data of power grid operation; S2, data processing, pre-processing the collected data; S3, building a multi-head attention structure, including input data processing, multi-head partitioning adaptive selection, cross-head information interaction, and attention calculation; Input data processing: linear transformation of preprocessed data through a learnable weight matrix , , Get the query vector Q, key vector K and value vector V respectively; Multi-head partitioning adaptive selection: The query vector Q, key vector K and value vector V are divided into multiple heads, and adaptive head selection is introduced to dynamically determine whether each head participates in subsequent attention calculations; Attention calculation: perform scaled dot product attention calculation on each head involved in the calculation; S4, concatenate the outputs of all heads involved in the calculation; S5, through the final learnable weight matrix The final output is obtained by linear transformation.

2. According to claim 1, a multi-head attention dynamic optimization method for power grid topology adaptation is characterized in that: The division of the query vector Q, the key vector K and the value vector V into multiple heads in the multi-head partition adaptive selection in step S3 can be expressed as: ; ; ; in represents the feature dimension of each head, h represents the total number of heads, and the query vector Q, key vector K, and value vector V are divided into h heads.

3. The multi-head attention dynamic optimization method for power grid topology adaptation according to claim 2 is characterized in that: The specific operation of performing adaptive head selection after obtaining h head weight matrices is as follows: Convolutional neural networks are used to extract features of power grid topology data; For each head, a selection score is calculated; A selection threshold is set. If the selection score is greater than the selection threshold, the head participates in the subsequent attention calculation, otherwise the head does not participate in the calculation.

4. The multi-head attention dynamic optimization method for power grid topology adaptation according to claim 3 is characterized in that: The selection score calculation adopts the information gain method. For each head h, the corresponding query vector is calculated. , key vector Sum value vector Information gain with input feature X; The information gain is obtained by calculating the change in entropy. For the query vector, the information gain is , similarly we can get , , and then get the selection score : ,in is the weight coefficient.

5. The multi-head attention dynamic optimization method for power grid topology adaptation according to claim 3 is characterized in that: Before calculating the selection score and setting the selection threshold, it is also necessary to determine the correlation between the heads to adjust the calculation score. The mutual information is used to measure the correlation between the heads. For each head h, the calculation of its selection score can be adjusted as follows: ,in The function is the mutual information function.

6. The multi-head attention dynamic optimization method for power grid topology adaptation according to claim 1 is characterized in that: The specific steps of attention calculation in step S3 are as follows: Calculate the raw attention score; Introducing context information of the power grid operation status, wherein the context information is obtained through a pre-trained model related to the power grid; According to the context information and the characteristics of the power grid topology data, the adjustment factor of each attention score is calculated; the adjustment factor is calculated as follows: , where C is the context information, is the adjustment factor of the obtained attention score, is average pooling; Multiply the original attention score by the adjustment factor element by element to obtain an adjusted attention score, and then apply the softmax function to the adjusted attention score to obtain the final attention weight; Finally, the final attention weight of each head is multiplied by the value vector Get the output for each head.

Citation Information

Patent Citations

  • Power load prediction method and system based on modal decomposition and batch decomposition

    CN118095529A

  • Multi-node short-term power load prediction method based on MST-GCN and Transform fusion

    CN118572664A

  • Power grid load prediction system and method based on machine learning

    CN119398274A

  • Low-voltage transformer area topology identification method based on graph attention network

    CN119518720A

  • Power load prediction method based on dual-channel cross attention network

    CN119669732A