Intelligent Substation Secondary Equipment Fault Location Method Based on Improved Transformer

By adopting the improved Transformer model in the secondary equipment fault location of intelligent substations, the THA and ICB modules are used to improve feature interaction and multi-scale extraction capabilities, solving the inefficiency problem of existing models when processing dynamic topology structures and high-dimensional complex data, achieving more efficient fault location and real-time requirements.

CN119807982BActive Publication Date: 2025-06-27SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510298129.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-27
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The existing secondary equipment fault positioning model of smart substations is difficult to cope with the dynamic changes in the substation topology, it is difficult to fully capture global features, and has low computational efficiency when processing high-dimensional complex data.

Method used

Using the improved Transformer model, the feature interaction capability and multi-scale feature extraction capability are improved by introducing THA module and ICB module, and combining residual connection and layer normalization processing, the model's processing capability of high-dimensional complex data is enhanced.

Benefits of technology

It significantly improves the overall efficiency of fault location, reduces the need for manual intervention, and meets the real-time requirements in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807982B_ABST
    Figure CN119807982B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for fault location of secondary equipment in intelligent substations based on an improved Transformer. The steps include: constructing an improved Transformer model; training the constructed improved Transformer model; determining whether the fault online diagnosis threshold is reached, and when the fault online diagnosis threshold is reached, inputting unknown fault category samples F test into the improved Transformer model; calculating the probability distribution of the fault samples belonging to each fault category by the improved Transformer model, so as to realize the fault location of the fault samples. This fault location method can handle high-dimensional complex features, meet the real-time requirements in practical applications, significantly improve the overall efficiency of fault handling, and reduce the need for manual intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for fault location of secondary equipment, in particular to a method for fault location of secondary equipment in an intelligent substation based on an improved Transformer. Background Art

[0002] Since the release of the national standard "Technical Guidelines for Intelligent Substations" in 2013, intelligent substations have been widely used across the country, marking the official entry of China's power system into a new era of intelligent development. Intelligent substations have become an indispensable part of the power system. In an intelligent substation, the stability of secondary equipment has a direct impact on the safe and stable operation of the substation and even the entire power grid. Since secondary equipment failures are usually accompanied by a large number of device and alarm signal interactions, accurately locating secondary equipment failures has become a key technical challenge for ensuring the safe and stable operation of the power system.

[0003] In the early days, the fault location of secondary equipment in intelligent substations relied on regular inspections by staff and judgment based on historical experience, resulting in long response times, high work intensity, and the accuracy of location being extremely vulnerable to human factors. To improve the accuracy of fault location and reduce the workload of staff, the Apriori association rule mining algorithm is used in the prior art to screen evaluation indicators, the attribute hierarchy process and the anti-entropy weight method are used to calculate subjective and objective weights, and the combined weight is obtained based on the cooperative game model. Finally, a risk assessment model for secondary equipment in the power grid based on the cloud model is constructed. Although certain results have been achieved in the fault location of secondary equipment, these methods not only highly rely on prior knowledge but also are difficult to adapt to the complexity and dynamic variability of the system.

[0004] In recent years, due to their advantages in dealing with non-linear mapping problems, artificial intelligence algorithms have gradually been applied to fault location research. A fault location method based on Probabilistic Neural Networks (PNN) has been proposed in the prior art. However, the performance of PNN is greatly affected by the selection of the smoothing factor, which easily leads to unstable model performance during the training process. The prior art has also proposed using GRU to achieve fault location of secondary equipment, effectively alleviating the problems of gradient disappearance and explosion in RNN. However, in an intelligent substation, due to the variety of secondary equipment, the fault characteristics show diversity, and GRU cannot fully capture the global features in the data.

[0005] Based on the current situation of the above-mentioned existing technologies, it can be seen that the current fault location models of secondary equipment in intelligent substations implemented based on artificial intelligence algorithms have the following problems: 1) Existing models are difficult to cope with the dynamic changes of substation topological structures, which makes it difficult to ensure the adaptability and accuracy of the original models when configurations or topologies are adjusted; 2) With the diversification of fault characteristics, existing models are difficult to fully capture global characteristics; 3) Existing models may face the problem of low computational efficiency when dealing with high-dimensional and complex data. Summary of the Invention

[0006] The object of the invention is to provide a method for fault location of secondary equipment in intelligent substations based on an improved Transformer, which can handle high-dimensional complex features, meet the real-time requirements in practical applications, significantly improve the overall efficiency of fault handling, and reduce the need for manual intervention.

[0007] Technical solution: The method for fault location of secondary equipment in intelligent substations based on the improved Transformer of the present invention includes the following steps:

[0008] Step 1, construct an improved Transformer model for fault location of secondary equipment in the substation;

[0009] Step 2, use the fault sample set to train the constructed improved Transformer model;

[0010] Step 3, continuously judge whether the online fault diagnosis threshold is reached. If the online fault diagnosis threshold is not reached, notify the staff to perform manual reasoning to obtain the fault source. If the online fault diagnosis threshold is reached, the formed fault feature set F i is used as an unknown fault category sample F test and input it into the trained improved Transformer model;

[0011] Step 4, calculate the probability distribution of the fault sample belonging to each fault category by the trained improved Transformer model, and select the category with the largest probability as the fault category, so as to realize the fault location of the fault sample.

[0012] Further, in step 1, the constructed improved Transformer model includes a THA module, an ICB module, a first RC&LN module, an FFN module, a second RC&LN module, a linear layer, and a Softmax module; the input feature data is processed by the THA module, and the processed feature data is respectively sent to the ICB module and the first RC&LN module. The ICB module extracts important features and sends the extracted important features to the first RC&LN module. The first RC&LN module performs residual connection and layer normalization processing on the two input data streams and then sends them to the FFN module. The FFN module extracts complex features, and then the second RC&LN module performs residual connection and layer normalization processing on the extracted complex features and sends them to the linear layer. The linear layer maps the hidden layer representation of the input data to the output space of the fault location task, and finally, the Softmax module generates the probability distribution of each fault location and outputs the final fault location.

[0013] Further, in step 2, the specific steps for training the constructed improved Transformer model using the fault sample set are as follows:

[0014] Step 2.1, obtain the fault sample set as , where F i is the fault feature set of the i th fault sample, N is the total number of fault samples, represents the class label to which the i th fault sample belongs, C is the total number of fault classes;

[0015] Step 2.2, input the fault feature set into the constructed improved Transformer model to generate the probability distribution belonging to each fault class as , where is the fault feature set F i belonging to the fault class j predicted probability;

[0016] Step 2.3, use the cross-entropy loss function to train the improved Transformer model. The cross-entropy loss function is: , where B is the number of fault samples in the fault sample set, δ ( Y i ,j ) is an indicator function. When Y i = j , δ (Y i ,j ) = 1, otherwise 0, where log represents the natural logarithm.

[0017] Furthermore, in step 3, the online fault diagnosis threshold is obtained by evaluating historical fault feature information. The specific steps are as follows:

[0018] First, obtain the historical fault feature information of each group of the substation to be evaluated within the evaluation period. Each group of historical fault feature information is a string of feature information data composed of 0 and / or 1. Each digit in the feature information data represents the situation information of each monitoring point of the substation. Data 1 indicates that the corresponding abnormal situation appears at the monitoring point, and data 0 indicates that the corresponding abnormal situation does not appear at the monitoring point;

[0019] Then, perform statistical analysis on the obtained historical fault feature information to obtain the number of record groups of the historical fault feature information within the evaluation period and the number of 1s in each group of historical fault feature information Z j , and then find the Z j maximum value Zmax and the minimum value Zmin ;

[0020] Then, calculate the feature information record frequency as P times / year, and then find the average value of 1s in each group of historical fault feature information Zx ;

[0021] Finally, establish the evaluation formula for the online fault diagnosis threshold as: , where K 1, K 2, and K 3 are all adjustment coefficients used to fine-tune the output online fault diagnosis threshold, R is the threshold value used to adjust the intervention threshold of online diagnosis.

[0022] Furthermore, in step 3, the formed fault feature set F i is: , where is the self-checking feature subset, is the fault warning feature subset, is the process layer status feature subset, is the communication layer status feature subset;

[0023] The self-checking feature subset integrates the self-checking information of the hardware module , the memory exception information , and the self-checking information of the communication module , the formula is: , In the formula, are respectively the total number of devices of the measurement and control device, protection device, merging unit, and intelligent terminal, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's hardware self-check alarm, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's memory abnormality alarm, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's communication module self-check alarm;

[0024] Fault alarm feature subset Integrates sampling fault alarm , trip fault alarm , communication alarm signal , body alarm signal , the formula is: , In the formula, are respectively the total number of devices of the measurement and control device, protection device, merging unit, intelligent terminal, and switch, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's sampling fault alarm, is the l th protection device's protection action fault alarm, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, the n th intelligent terminal device, and the o th switch's communication fault alarm, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the nOntology fault alarm of an intelligent terminal device;

[0025] State feature subset Including current sampling values , voltage sampling values And digital input abnormal status , the formula is: , In the formula, Are the current values of phase A, phase B, and phase C in channel 1 respectively, Are the current values of phase A, phase B, and phase C in channel 2 respectively, Are the voltage values of phase A, phase B, and phase C in channel 1 respectively, Are the voltage values of phase A, phase B, and phase C in channel 2 respectively, and are the total numbers of devices of the measuring and control device, protection device, and merging unit respectively, Is the k th measuring and control device, the l th protection device, and the m th merging unit's switch quantity signal related to the state;

[0026] Communication layer state feature subset Used to comprehensively reflect various state information of the communication layer, including SV message anomaly , GOOSE message anomaly And possible switch status anomaly , the formula is: , In the formula, Are the total numbers of devices of the measuring and control device, protection device, merging unit, intelligent terminal, and switch respectively, Are the l th protection device, the m th merging unit, the n th intelligent terminal device's SV message anomaly alarm, Are the k th measuring and control device, the l th protection device, the m th merging unit, the n th intelligent terminal device's GOOSE message anomaly alarm, Are the o th switch's port status, network traffic status, and link detection status.

[0027] Furthermore, in step 4, the calculation formula for selecting the category with the highest probability is: , in the formula, Is the test sample F testThe predicted category, is the test sample F test belongs to the category j The predicted probability, is to solve for the category index that maximizes j .

[0028] Furthermore, the calculation formula of the THA module is: , where Q, K, and V are the query vector, key-value vector, and value vector respectively, W o is a learnable weight matrix, h is the attention head head The total number of head The calculation formula for each attention head is: , where head i is the output of each attention head, P w and P l are two different linear projection matrices respectively, and are the query transformation matrix and key transformation matrix of the i th head respectively, is the query factor, softmax The function is used to transform the similarity into a probability distribution.

[0029] Furthermore, the calculation formula of the ICB module is: , where A 1 and A 2 are the first interaction feature and the second interaction feature respectively, Conv 3 is the third convolutional layer, and the first interaction feature A 1 and the second interaction feature A 2 The calculation formulas are: , , where is the temporal feature output by the THA module, Conv 1 is the small kernel convolution for extracting fine-grained local features, Conv 2 is the large kernel convolution for capturing global dependencies, is the GeLU activation function, represents the element-wise multiplication operation.

[0030] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) By introducing the THA module into the improved Transformer model, the feature interaction ability of the improved Transformer model is effectively enhanced, enabling it to maintain high adaptability and accuracy even under the condition of dynamic changes in the substation topology structure; (2) Utilizing the ICB model to extract local fine-grained features through small kernel convolution and obtain global context information through large kernel convolution, and adding it to the output part of the THA module due to its characteristic of being able to fully capture feature information at different scales under diverse fault feature conditions, significantly improves the comprehensiveness and accuracy of feature expression. The ICB module dynamically modulates the convolution path, promotes the complementary fusion of multi-scale features, and effectively enhances the processing ability of the improved Transformer model for high-dimensional complex data; (3) The improved Transformer model can not only handle high-dimensional complex feature processing but also meet the real-time requirements in practical applications, and significantly improves the overall efficiency of fault handling, reducing the need for manual intervention. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is the flowchart of the fault location method of the present invention;

[0032] Figure 2 is the structure diagram of the improved Transformer model of the present invention;

[0033] Figure 3 is the structure diagram of the ICB module of the present invention;

[0034] Figure 4 is the positioning effect diagram of the improved Transformer of the present invention;

[0035] Figure 5 is the line interval topology diagram of the 220KV intelligent substation of the present invention;

[0036] Figure 6 is the curve of the accuracy rate and loss function value of the present invention;

[0037] Figure 7 is the structure diagram of the traditional Transformer model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings, but the protection scope of the present invention is not limited to the described embodiments.

[0039] As Figure 1 shown, the intelligent substation secondary equipment fault location method based on the improved Transformer disclosed by the present invention includes the following steps:

[0040] Step 1, construct an improved Transformer model for secondary equipment fault location in a substation;

[0041] Step 2, use the fault sample set to train the constructed improved Transformer model;

[0042] Step 3, judge in real time whether the online fault diagnosis threshold is reached. If the online fault diagnosis threshold is not reached, notify the staff to perform manual reasoning to obtain the fault source. When performing manual reasoning, the fault source is inferred according to the existing secondary equipment fault knowledge base. If the online fault diagnosis threshold is reached, the formed fault feature set F i is used as an unknown fault class sample F test and input it into the trained improved Transformer model;

[0043] Step 4, calculate the probability distribution of the fault sample belonging to each fault class by the trained improved Transformer model, and select the class with the largest probability as the fault class, so as to realize the fault location of the fault sample.

[0044] Furthermore, in Step 1, the constructed improved Transformer model includes a THA module, an ICB module, a first RC&LN module, an FFN module, a second RC&LN module, a linear layer, and a Softmax module. As Figure 2 shown, the THA module processes the input feature data, and the processed feature data are respectively sent to the ICB module and the first RC&LN module. The ICB module extracts important features and sends the extracted important features to the first RC&LN module. The first RC&LN module performs residual connection and layer normalization processing on the two input data and then sends them to the FFN module. The FFN module extracts complex features, and then the second RC&LN module performs residual connection and layer normalization processing on the extracted complex features and sends them to the linear layer. The linear layer maps the hidden layer representation of the input data to the output space of the fault location task, and finally generates the probability distribution of each fault location through the Softmax module and outputs the final fault location.

[0045] The THA module replaces the original MHA of the traditional Transformer, enhances the feature interaction ability of the model, and improves the adaptability and accuracy of the model.

[0046] The ICB module can extract local fine-grained features and global context information respectively through the combination of small-kernel convolution and large-kernel convolution, thereby effectively capturing features of different scales, enhancing the model's ability to express complex features. At the same time, ICB promotes the complementary fusion between multi-scale features by dynamically modulating the convolution path, improving the comprehensive utilization efficiency of the model for high-dimensional features.

[0047] Both the first RC&LN module and the second RC&LN module are used to solve the problems of gradient vanishing and explosion during model training, accelerate the training of the model, and improve the performance of the model.

[0048] The FFN module is used to perform non-linear transformation on the features after attention calculation, enabling the model to learn more complex feature representations.

[0049] The linear layer converts the features extracted in the model into the final prediction result, that is, the probability distribution of the fault location. In the fault location task, the output of the linear layer is the probability distribution of each fault location.

[0050] The Softmax operation maps the score of each location to between 0 and 1 and ensures that the sum of the probabilities of all locations is 1, so that the model can output the prediction probability of each fault location.

[0051] Furthermore, in step 2, the specific steps for training the constructed improved Transformer model using the fault sample set are as follows:

[0052] Step 2.1, obtain the fault sample set as , where F i is the fault feature set of the i th fault sample, N is the total number of fault samples, represents the class label to which the i th fault sample belongs, C is the total number of fault classes;

[0053] Step 2.2, input the fault feature set into the constructed improved Transformer model to generate the probability distribution belonging to each fault class as , where is the fault feature set F i belonging to the fault class j of the prediction probability;

[0054] Step 2.3, use the cross-entropy loss function to train the improved Transformer model, and the cross-entropy loss function is: , where B is the number of fault samples in the fault sample set,δ ( Y i ,j ) is an indicator function. When Y i = j , δ ( Y i ,j ) = 1, otherwise it is 0. log represents the natural logarithm.

[0055] Furthermore, in step 3, the online fault diagnosis threshold is obtained by evaluating historical fault feature information. The specific steps are as follows:

[0056] First, obtain each group of historical fault feature information of the substation to be evaluated within the evaluation period. Each group of historical fault feature information is a string of feature information data composed of 0 and / or 1. Each digit in the feature information data represents the situation information of each monitoring point of the substation. Data 1 indicates that the corresponding abnormal situation appears at the monitoring point, and data 0 indicates that the corresponding abnormal situation does not appear at the monitoring point. The number of digits of the feature information data is the same as the number of monitoring points of the current substation. The evaluation period is set to 1 year, that is, obtain all historical fault feature information within 1 year before the current moment. If the recording duration of the historical fault feature information is less than 1 year, then count all historical fault feature information;

[0057] The monitoring point situation information includes protection device locking, protection device action, intelligent terminal feedback, circuit breaker action, merging unit device abnormality, merging unit self-check alarm, merging unit SV data abnormality, merging unit input self-check loop error, merging unit dual-position input inconsistency, merging unit sampling abnormality, merging unit synchronization signal interruption, etc. The monitoring point situation information varies according to the settings of the substation, and there are differences in the monitoring points of substations of different scales. Therefore, the online fault diagnosis threshold of each substation needs to be customized;

[0058] Then, statistically analyze the obtained historical fault feature information to obtain the number of recorded groups of historical fault feature information within the evaluation period and the number of data 1 in each group of historical fault feature information Z j , and then find the maximum value Z j in Zmax and the minimum value Zmin ;

[0059] Then, calculate the feature information recording frequency as P times / year according to the number of recorded groups, and then find the average value Zx of data 1 in each group of historical fault feature information;

[0060] Finally, establish the evaluation formula for the online fault diagnosis threshold as: , where K 1,[[]] K 2 and K 3 are all adjustment proportionality coefficients, used to finely adjust the online fault diagnosis threshold of the output, so that between the stage output values, R is the threshold value, used to adjust the intervention threshold of online diagnosis.

[0061] By setting the online fault diagnosis threshold, it is possible to avoid frequent intervention of the positioning method for automatic calculation caused by a small number of false alarm messages; by setting the threshold value R it is possible to overall increase the threshold value and avoid frequent intervention of online diagnosis, so as to distinguish from manual diagnosis.

[0062] Furthermore, in step 3, the formed fault feature set F i is: , where is the self-check feature subset, is the fault alarm feature subset, is the process layer status feature subset, is the communication layer status feature subset;

[0063] The self-check feature subset combines the self-check information of the hardware module , memory exception information and the self-check information of the communication module , and the formula is: , where are the total numbers of devices of the measurement and control device, protection device, merging unit, and intelligent terminal respectively, are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's hardware self-check alarm, such as I / O board hardware failure, power supply anomaly alarm, etc., are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the n th intelligent terminal device's memory anomaly alarm, such as FLASH error, internal memory anomaly, etc., are respectively the k th measurement and control device, the l th protection device, the m th merging unit, and the nCommunication module self-check alarm of an intelligent terminal device, such as abnormal X network of GOOSE board, communication interruption of SV board, etc.;

[0064] Fault alarm feature subset Integrates sampling fault alarms and trip fault alarms and communication alarm signals and body alarm signals , the formula is: , In the formula, are the total numbers of equipment of the measurement and control device, protection device, merging unit, intelligent terminal, and switch respectively, are respectively the k rd measurement and control device, the l th protection device, the m th merging unit, and the n th sampling fault alarms of intelligent terminal devices, such as abnormal SV data, invalid SV transmission quality, etc., is the protection action fault alarm of the l th protection device, such as protection device locking, differential protection locking, etc., are respectively the k rd measurement and control device, the l th protection device, the m th merging unit, the n th intelligent terminal device, and the o th communication fault alarms of switches, such as communication abnormality, VLAN allocation error, etc., are respectively the k rd measurement and control device, the l th protection device, the m th merging unit, and the n th body fault alarms of intelligent terminal devices, such as self-check alarm, device operation abnormality, etc.;

[0065] Status feature subset Includes current sampling values and voltage sampling values and digital input abnormal states , the formula is: , in the formula, are the current values of phase A, phase B, and phase C in channel 1 respectively, are the current values of phase A, phase B, and phase C in channel 2 respectively, are the voltage values of phase A, phase B, and phase C in channel 1 respectively, are the voltage values of phase A, phase B, and phase C in channel 2 respectively, and are the total numbers of equipment of the measurement and control device, protection device, and merging unit respectively, For the k th measurement and control device, the l th protection device, and the m th merging unit, the status-related digital signals, such as abnormal input digital signals, abnormal output signals, etc.;

[0066] Communication layer status feature subset used to comprehensively reflect various status information of the communication layer, including abnormal SV messages , abnormal GOOSE messages and possible abnormal switch status , the formula is: , In the formula, are the total numbers of the measurement and control device, protection device, merging unit, intelligent terminal, and switch respectively, are respectively the l th protection device, the m th merging unit, the n th intelligent terminal device's SV message abnormal alarms, such as SV total alarm, SV board X network abnormal, etc., are respectively the k th measurement and control device, the l th protection device, the m th merging unit, the n th intelligent terminal device's GOOSE message abnormal alarms, such as GOOSE data abnormal, GOOSE communication abnormal, etc., are respectively the o th switch's port status, network traffic status, and link detection status, such as broadcast storm alarm, traffic abnormal, etc.

[0067] Furthermore, in step 4, the formula for selecting the category with the highest probability is: , in the formula, is the test sample F test 's predicted category, is the predicted probability that the test sample F test belongs to the category j , is to solve the category index that makes the largest j .

[0068] Furthermore, the formula of the THA module is: , in the formula, Q, K, and V are the query vector, key vector, and value vector respectively, W o is the learnable weight matrix, h is the attention headhead The total number, for each attention head head The calculation formula is: , where head i is the output of each attention head, P w and P l are two different linear projection matrices respectively, and are respectively i the query transformation matrix and the key transformation matrix of the th head, softmax The function is used to convert the similarity into a probability distribution.

[0069] Furthermore, the structure of the ICB module is as shown in Figure 3 The calculation formula of the ICB module is: , where A 1 and A 2 are the first interaction feature and the second interaction feature respectively, Conv 3 is the third convolutional layer, and the first interaction feature A 1 and the second interaction feature A 2 The calculation formulas are: , , where is the temporal feature output by the THA module, Conv 1 is the small kernel convolution, which is used to extract fine-grained local features, Conv 2 is the large kernel convolution, which is used to capture global dependencies, is the GeLU activation function, represents the element-wise multiplication operation.

[0070] In order to verify the location effect of the fault location method of the present invention, taking the line interval of a typical 220KV intelligent substation as the research object, the "direct sampling and direct tripping" mode is adopted to analyze secondary equipment such as protection and measurement and control. The involved topological structure is as shown in Figure 5 . Combining the actual test cases of the Electric Power Research Institute, 2730 fault samples are collected, covering 28 fault situations. Each fault type is classified according to the number shown in Table 1. In addition, the fault feature set F i contains a total of 204 secondary equipment fault feature information. In order to ensure the effectiveness of the experiment, the present invention randomly divides these fault samples into a training set and a test set according to a ratio of 8:2 for model training and evaluation.

[0071] Table 1 is a statistical table of secondary equipment failure types and numbers

[0072]

[0073] Due to the dimensional differences in the characteristic information, before model training, it is necessary to perform Max-Min normalization on all input quantities. The normalization range is [0,1], and the normalization formula is: , where is the result after data normalization, x is the original input data, and are the maximum and minimum values of the original data respectively.

[0074] The present invention uses the commonly used model performance evaluation indicators in classification tasks to verify the performance of the proposed model, specifically as follows:

[0075] (1) Accuracy , which represents the proportion of samples correctly predicted by the model in all predictions. The calculation formula is: ;

[0076] (2) Precision P , which represents the proportion of samples actually belonging to the fault category among the samples predicted as faults by the model. The calculation formula is: ;

[0077] (3) Recall , which represents the proportion of samples actually belonging to the fault category that are correctly identified as faults by the model. The calculation formula is: ;

[0078] (4) F1 score, which is used to comprehensively measure the balance between the accuracy rate and the coverage rate. The calculation formula is: ;

[0079] In the above four formulas, TP represents the number of samples correctly predicted as the positive class, TN represents the number of samples correctly predicted as the negative class, FP represents the number of samples actually being the negative class but mispredicted as the positive class (i.e., false alarm), and FN represents the number of samples actually being the positive class but mispredicted as the negative class (i.e., missed alarm).

[0080] The experimental environment of the present invention uses the Windows11 operating system and the NVIDIA RTX3060 graphics card. The algorithm model uses Python 3.6.13 as the programming language, and the model is built based on the deep learning framework Pytorch.

[0081] To avoid overtraining or under-training of the model, experiments were conducted on the impact of the number of iterations on the fault location performance. As Figure 6As shown, as the number of iterations increases, the accuracy and loss value of the model show a continuously changing trend. Especially when the number of iterations approaches 300, the accuracy and loss function value tend to stabilize, and the change amplitude decreases significantly, indicating that the model has basically converged at this time. To ensure that the model can fully converge, the maximum number of iterations is set to 600 in the subsequent experiments.

[0082] Since the performance of the optimizer in gradient descent may significantly affect the training efficiency and final performance of the model, the present invention conducts a comparative analysis of the performance of multiple typical optimizers through experiments, so as to screen out the optimizer most suitable for the subsequent experiments to ensure the efficiency and reliability of model training. The experimental results are shown in Table 2. From the perspective of the accuracy of the comprehensive model positioning and the training time consumption of the model in Table 2, the SGD optimizer performs relatively balanced. Therefore, the SGD optimizer is selected for use in the subsequent experiments of the present invention.

[0083] Table 2 is a statistical table of the fault location accuracy when using different optimizers

[0084]

[0085] To explore the performance of hyperparameters on the model, the present invention selects and experiments on the hyperparameters number of attention heads Head, learning rate lr, and batch_size with the traditional Transformer model as the baseline. The experimental results are shown in Table 3.

[0086] As can be seen from Table 3, when Head = 4 and batch_size = 16, when lr = 0.1, the performance indicators of the model reach the best level, indicating that this learning rate can effectively balance the optimization speed and stability. When the learning rate increases to lr = 0.5, the model accuracy drops to 85.37%, which may be because too large a learning rate will cause the model gradient to be unstable or gradient explosion. When the learning rate decreases, that is, lr = 0.01 and lr = 0.05, although the model can converge stably, the model performance is slightly lower than that of lr = 0.1. Therefore, an appropriate learning rate can better improve the model performance.

[0087] When Head = 4 and lr = 0.1, the batch size has an obvious impact on the model performance and training time consumption. Experiments show that when the batch size is 16, the model performance is the best and the training time consumption is moderate, meeting the actual task requirements. Reducing the batch size to 8 will significantly increase the training time consumption, while increasing it to 32 shortens the training time, but the model performance decreases, indicating that too large a batch size may lead to inaccurate gradient estimation and affect the model effect.

[0088] When batch_size = 16 and lr = 0.1, the model performs best with Head = 4, and the accuracy reaches 97.56%. When the number of heads is reduced to Head = 2, the model performance drops to 96.22%, indicating that too few attention heads limit the feature interaction ability. After increasing to Head = 6, the accuracy only increases to 97.07%, but the training time increases significantly, indicating that too many attention heads will increase the computational complexity rather than significantly improve the performance.

[0089] Table 3 is the statistical table of hyperparameter selection experiments

[0090]

[0091] In summary, the traditional Transformer model performs best when Head = 4, batch_size = 16, and lr = 0.1. Subsequent experiments on improving the model will be carried out based on this parameter setting to evaluate the effectiveness of the improvement measures on the model performance.

[0092] To verify the effectiveness of the positioning method proposed in the present invention, the model parameters when the above model performs best are adopted in the present invention, and experimental verification is carried out. The positioning effect is as Figure 4 shown. In addition, to verify the performance of the improved Transformer model constructed in the present invention, the performance index accuracy A cc , precision P, recall R c and F1 score are further calculated, and comparative experiments are carried out with models commonly used in the field, such as the fully convolutional network (FCN), deep neural network (DNN), GRU (gated recurrent unit), Transformer, etc. The experimental results are shown in Table 4. It can be seen that all performance indicators of the improved Transformer model constructed in the present invention are higher than those of FCN, DNN, GRU, and Transformer. This result indicates that the improved Transformer model constructed in the present invention can accurately classify and identify fault samples in the secondary equipment fault location task, thereby reducing false alarms and missed detections. Although the processing time of the improved Transformer model constructed in the present invention increases compared with other models, the time gap is small, and it can still meet the real-time requirements in practical applications and significantly improve the overall efficiency of fault handling, reducing the need for manual intervention.

[0093] Table 4 is the statistical table of comparative experiment results

[0094]

[0095] The present invention also verifies the effectiveness of each optimization module for the constructed improved Transformer model through ablation experiments, and specifically sets the following experiments:

[0096] Experiment 1: The unimproved traditional Transformer is used as a baseline comparison to evaluate the performance of the model without any optimization. The traditional Transformer is as Figure 7 shown.

[0097] Experiment 2: Replace the traditional MHA with the THA module of the present invention to verify the effectiveness of the THA module in enhancing feature interaction ability and global feature expression.

[0098] Experiment 3: Add the ICB module after the MHA of the traditional Transformer to verify the role of the ICB module in multi-scale feature extraction, feature fusion, and improvement of computational efficiency.

[0099] Experiment 4: Replace the traditional MHA with the THA module and combine it with the ICB module for joint application to comprehensively evaluate the collaborative optimization effect of the THA module and the ICB module, especially the enhancement in the accuracy of fault feature extraction and the overall model performance.

[0100] Table 5 is the statistical table of the comparative experiments

[0101]

[0102] As shown in Table 5, by comparing Experiments 1-4, it can be seen that Experiment 1 performed the worst in all indicators, indicating that the traditional Transformer model has obvious deficiencies in feature extraction and global information expression capabilities; by comparing the results of Experiments 1 and 2, it can be seen that the addition of the THA module significantly improved the performance of the traditional Transformer model, especially the accuracy and recall rates increased by 0.64% and 0.66% respectively, verifying that the THA module enhanced the Head feature interaction ability while also enhancing the comprehensive utilization effect of the model on the feature subspace, and improving the model's ability to extract key features in complex classification tasks; by comparing Experiments 1 and 3, it can be seen that adding the ICB module after the MHA of the traditional Transformer further improved the multi-scale feature extraction ability of the Transformer model, increasing the accuracy and recall rates of the Transformer model by 0.37% and 0.27% respectively; by comparing Experiments 1 and 4, it can be seen that the improved Transformer model of the present invention's comprehensive utilization ability for high-dimensional features significantly improved the fault location performance of the Transformer model, and compared with the traditional Transformer, the accuracy, precision, recall rate, and F1 score were increased by 1.34%, 1.25%, 1.49, and 1.23% respectively.

[0103] In summary, the present invention addresses the problems that the current intelligent substation secondary equipment fault location model has difficulty in coping with the dynamic changes of the substation topology structure, difficulty in fully capturing global features, and may face low computational efficiency when dealing with high-dimensional and complex data. An improved Transformer model is disclosed, and its effectiveness is verified through experiments. The specific conclusions are as follows:

[0104] (1) By introducing the THA module to replace the MHA of the traditional model, the present invention effectively improves the feature interaction ability of the Transformer model, enabling it to still maintain high adaptability and accuracy in the case of dynamic changes in the substation topology structure. The experimental results show that compared with the traditional Transformer model, the THA module solves the cross-head feature interaction problem, making the feature information obtained by different Heads no longer independent of each other, and enhancing the comprehensive utilization effect of the feature subspace.

[0105] (2) By using the ICB model, local fine-grained features are extracted through small kernel convolution, and global context information is obtained through large kernel convolution. Considering the characteristic of being able to fully capture feature information at different scales under diverse fault features, it is added to the output part of the THA module, significantly enhancing the comprehensiveness and accuracy of feature representation. At the same time, the ICB module dynamically modulates the convolution path, promoting the complementary fusion of multi-scale features and effectively enhancing the processing ability of the Transformer model for high-dimensional complex data.

[0106] (3) Compared with models such as Transformer, FCN, GRU, and DNN commonly used in current research in this field. The improved Transformer model proposed by the present invention can not only handle high-dimensional complex features but also meet the real-time requirements in practical applications, and significantly improve the overall efficiency of fault handling, reducing the need for manual intervention.

[0107] Experimental results show that the improved Transformer model proposed by the present invention in the secondary equipment fault location task of intelligent substations is not only significantly superior to traditional methods in terms of performance but also meets the real-time requirements in practical applications in terms of efficiency. This provides reliable support for the safe operation and efficient management of smart grids.

[0108] As described above, although the present invention has been shown and described with reference to specific preferred embodiments, it should not be construed as a limitation of the present invention itself. Various changes in form and detail may be made without departing from the spirit and scope of the present invention defined by the appended claims.

Claims

1. A method for locating faults of secondary equipment in smart substation based on improved Transformer, characterized in that: The steps include: Step 1, constructing an improved Transformer model for substation secondary equipment fault location; Step 2: Use the fault sample set to train the constructed improved Transformer model; Step 3: Determine in real time whether the fault online diagnosis threshold is reached. If the fault online diagnosis threshold is not reached, notify the staff to perform manual reasoning to obtain the fault source. If the fault online diagnosis threshold is reached, the fault feature set F i As the unknown fault class sample F test Input into the trained improved Transformer model; Step 4: The trained improved Transformer model is used to calculate the probability distribution of the fault sample belonging to each fault category, and the category with the highest probability is selected as the fault category, thereby realizing the fault location of the fault sample; In step 1, the constructed improved Transformer model includes a THA module, an ICB module, a first RC&LN module, an FFN module, a second RC&LN module, a linear layer and a Softmax module; the THA module processes the input feature data, and the processed feature data is sent to the ICB module and the first RC&LN module respectively, the ICB module extracts important features, and sends the extracted important features to the first RC&LN module, the first RC&LN module performs residual connection and layer normalization processing on the two input data and then sends them to the FFN module, the FFN module extracts complex features, and then the second RC&LN module performs residual connection and layer normalization processing on the extracted complex features and then sends them to the linear layer, the linear layer maps the hidden layer representation of the input data to the output space of the fault location task, and finally generates the probability distribution of each fault location through the Softmax module, and outputs the final fault location; The calculation formula of ICB module is: ICB =Conv3(A1+A2), where A1 and A2 are the first interactive feature and the second interactive feature respectively, Conv3 is the third convolutional layer, and the calculation formula of the first interactive feature A1 and the second interactive feature A2 is: A1=φ(Conv1(F'))☉Conv2(F'), A2=φ(Conv2(F'))☉Conv1(F'), where F' is the temporal feature output by the THA module, Conv1 is a small kernel convolution for extracting fine-grained local features, Conv2 is a large kernel convolution for capturing global dependencies, φ(·) is the GeLU activation function, and ⊙ represents the element-wise multiplication operation; The calculation formula of the THA module is: THA(Q,K,V)=[head1, head2,...,head h ]W o , where Q, K and V are query vector, key value vector and value vector respectively, and W o is a learnable weight matrix, h is the total number of attention heads, and the calculation formula for each attention head is: In the formula, head i For the output of each attention head, P w and P l They are two different linear projection moments. Array, and are the query transformation matrix, key transformation matrix and value projection matrix of the i-th head respectively, For the query factor, the softmax function is used to transform the similarity into a probability distribution.

2. The method for locating faults of secondary equipment in smart substation based on improved Transformer according to claim 1 is characterized in that: In step 2, the specific steps of training the constructed improved Transformer model using the fault sample set are: Step 2.1, obtain the fault sample set as In the formula, F i is the fault feature set of the i-th fault sample, N is the total number of fault samples, Y i ∈{1,2,...,C} represents the category label to which the i-th fault sample belongs, and C is the total number of fault categories; Step 2.2: The fault feature set F i Input into the constructed improved Transformer model to generate the probability distribution of each fault category: In the formula, F is the fault feature set F i The predicted probability of belonging to fault category j; Step 2.3, use the cross entropy loss function to train the improved Transformer model. The cross entropy loss function is: Where B is the number of fault samples in the fault sample set, δ(Y i , j) is the indicator function, when Y i =j,δ(Y i , j) = 1, otherwise it is 0, log represents the natural logarithm.

3. The method for locating faults of secondary equipment in smart substation based on improved Transformer according to claim 1 is characterized in that: In step 3, the fault online diagnosis threshold is obtained by evaluating the historical fault feature information. The specific steps are as follows: First, each group of historical fault feature information of the substation to be evaluated within the evaluation period is obtained. Each group of historical fault feature information is a string of feature information data composed of 0 and / or 1. Each digit in the feature information data represents the situation information of each monitoring point of the substation. Data 1 indicates that the corresponding abnormal situation occurs at the monitoring point, and data 0 indicates that the corresponding abnormal situation does not occur at the monitoring point. Then, the obtained historical fault feature information is statistically analyzed to obtain the number of record groups of historical fault feature information within the evaluation period and the number of data 1 in each group of historical fault feature information Z j , and then find the quantity Z j The maximum value Zmax and the minimum value Zmin in the characteristic information are calculated; then the characteristic information recording frequency is calculated as P times / year according to the number of record groups, and then the average value Zx of data 1 in each group of historical fault characteristic information is calculated; Finally, the evaluation formula for the fault online diagnosis threshold is established as follows: Wherein, K1, K2 and K3 are adjustment coefficients used to fine-tune the output fault online diagnosis threshold, and R is the threshold value used to adjust the intervention threshold of online diagnosis.

4. The method for locating faults of secondary equipment in smart substation based on improved Transformer according to claim 1 is characterized in that: In step 3, the fault feature set F i For: F i =[F Si , F Ei ,F Oi ,F Ci ], i = 1, 2, 3, ..., N, where F Si is the self-check feature subset, F Ei is the fault alarm feature subset, F Oi is the process layer state feature subset, F Ci is a subset of communication layer state features; Self-check feature subset F Si Integrated hardware module self-test information S H , memory abnormal information S M And the communication module self-test information S C , the formula is: In the formula, a, b, c, and d are the total number of measurement and control devices, protection devices, merging units, and intelligent terminals, respectively. They are the hardware self-check alarms of the kth measurement and control device, the lth protection device, the mth merging unit, and the nth intelligent terminal device. They are the abnormal alarms of the storage devices of the kth measurement and control device, the lth protection device, the mth merging unit and the nth intelligent terminal device. They are the self-test alarms of the communication modules of the kth measurement and control device, the lth protection device, the mth merging unit and the nth intelligent terminal device; Fault alarm feature subset F Ei Integrated sampling fault alarm E SV , trip fault alarm E T 、Communication alarm signal E CA 、Main body alarm signal E EA , the formula is: In the formula, a, b, c, d, and e are the total number of measurement and control devices, protection devices, merging units, intelligent terminals, and switches, respectively. They are the sampling fault alarms of the kth measurement and control device, the lth protection device, the mth merging unit and the nth intelligent terminal device. This is the protection action failure alarm of the lth protection device. They are the communication fault alarms of the kth measurement and control device, the lth protection device, the mth merging unit, the nth intelligent terminal device, and the oth switch. are the fault alarms of the kth measurement and control device, the lth protection device, the mth merging unit, and the nth intelligent terminal device; the state feature subset F Oi Contains current sampling value O I , voltage sampling value O U And the abnormal state of the switch quantity O SW , the formula is: In the formula, They are the current values ​​of phase A, phase B, and phase C in channel 1, They are the current values ​​of phase A, phase B, and phase C in channel 2, They are the voltage values ​​of phase A, phase B, and phase C in channel 1, are the voltage values ​​of phase A, phase B and phase C in channel 2, respectively, and are the total number of devices of the measurement and control device, protection device and merging unit, respectively. The switch quantity signal related to the status of the kth measurement and control device, the lth protection device and the mth merging unit; Communication layer state feature subset F Ci Used to comprehensively reflect various status information of the communication layer, including SV message abnormality C SVM 、GOOSE message abnormality C GM And possible switch status abnormality C SW , the formula is: In the formula, a, b, c, d, and e are the total number of measurement and control devices, protection devices, merging units, intelligent terminals, and switches, respectively. They are the SV message abnormal alarms of the lth protection device, the mth merging unit, and the nth intelligent terminal device. They are the GOOSE message abnormal alarms of the kth measurement and control device, the lth protection device, the mth merging unit, and the nth intelligent terminal device. They are the port status, network traffic status and link detection status of the oth switch respectively.

5. The method for locating faults of secondary equipment in smart substation based on improved Transformer according to claim 1 is characterized in that: In step 4, the formula for selecting the category with the highest probability is: In the formula, For the test sample F test The predicted category, For the test sample F test The predicted probability of belonging to category j, To solve The largest category index j.

Citation Information

Patent Citations

  • Secondary system fault positioning method based on improved IDNN model

    CN113283462A

  • Intelligent substation secondary system fault positioning method based on Transform-GRU

    CN114492662A