An Automatic Risk Rating Method and Device for Grid On-site Operations Based on BERT
By enhancing text and improving the quality of the grading table and history library text of the power grid field operations, and training using the BERT model, the problem of automatic rating of the risk level of the power grid field operations is solved, achieving higher rating accuracy and lower manual workload.
Patent Information
- Application Number
- CN202310217606.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-08
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2043-03-08
AI Technical Summary
The prior art is difficult to effectively and automatically rating the risk level of power grid on-site operations, resulting in errors in manual ratings, affecting the safety and economics of operations.
Using a deep learning method based on BERT, we use text enhancement of the hierarchical table text and redundant text deletion and risk level correction of the history library text, a relatively complete sample set is built, and the BERT model is trained to achieve automatic rating of the risk level of the power grid on-site operation.
It effectively reduces manual workload, reduces data quality requirements, improves the accuracy of automatic rating results of on-site operation risk levels, and this method is versatile and can be applied to other deep learning models.
Smart Images

Figure CN116341902B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the power system, and specifically relates to a method and device for automatically rating the risks of on-site power grid operations based on BERT. Background Art
[0002] Production safety is the basic guarantee for the operation and development of the power system, which is related to China's economic development and national life safety. To ensure that no personal injury accidents, malicious misoperation incidents, or equipment fault tripping or temporary stoppage incidents caused by operation and maintenance management responsibilities occur at the production operation site, the State Grid Corporation of China issued a notice on further strengthening the risk control of on-site production operations in January 2022. Based on the original regulations on operation safety risk control, it was revised and supplemented, and implementation rules for various on-site operation risk controls were given. An operation risk grading table and corresponding risk grading basis were formulated.
[0003] The operation risk grading table (hereinafter referred to as the grading table) divides the operation risk level from high to low into levels I to V considering factors such as personal risk, equipment importance, operation and maintenance operation risk, operation control difficulty, and process technology difficulty. However, the content of actual on-site operations (hereinafter referred to as the historical database) is rich, and there are often operation contents not covered or partially covered in the above documents. At this time, it is necessary to manually refer to the similar operation scope and operation content at the same voltage level to determine the risk level.
[0004] Limited by the differences in experience and professional knowledge, it is difficult to completely avoid the situation where the actual operation risk level is misclassified, which affects the safety and economy of on-site operations. Therefore, there is a need for intelligent rating of the actual operation risk level. If text mining technology can be explored to meet this need, it can not only assist front-line staff in quickly determining the risk level of the current actual operation and avoiding misrating, but also improve the professional level of business personnel. Therefore, it is of great significance to explore text mining technology for on-site power grid operations.
[0005] At present, some scholars have tried to conduct data mining through other data such as images to complete the identification of power safety risks, but there is no risk research targeting text data yet. In addition, although the research on text mining in the power field has received attention, there is still room for improvement. From the perspective of the text content mined, the current research mainly focuses on texts such as power equipment defect texts, power dispatching plans, and power dispatching information, lacking research on on-site operation texts. From the perspective of the text objects mined, the current research related to power text classification often only focuses on historical texts and discards standard and specification documents with guiding and normative natures, that is, only uses historical texts as the sample set to train the classification model, and then uses the trained model to classify newly input texts. This research method will be extremely dependent on the quality of historical texts. When there are problems such as insufficient coverage and low accuracy in historical texts, the classification effect of the model will drop significantly. Therefore, how to combine historical texts with standard guidelines should also be one of the important factors to be considered in the research on power text classification. Summary of the Invention
[0006] The technical problem to be solved by the present invention is the problem of automatic rating of on-site operation risks in the power grid, that is, how to use deep learning methods to utilize on-site operation text data in the power grid to achieve automatic rating of on-site operation risk levels.
[0007] Based on this, the present invention proposes a method and device for automatic rating of on-site operation risks in the power grid based on BERT.
[0008] The first aspect of the present invention provides a method for automatic rating of on-site operation risks in the power grid based on BERT.
[0009] First, perform text enhancement on the grading table text and convert it into text similar to the content of the historical database text.
[0010] Next, delete redundant texts and correct risk levels in the historical database text to improve its text quality.
[0011] Then, use the text-enhanced grading table text and the historical database text with improved text quality to construct a relatively complete sample set, and use this sample set to train the BERT model.
[0012] Finally, use the trained BERT model to achieve automatic rating of on-site operation risk levels in the power grid.
[0013] The second aspect of the present invention provides a device for automatic rating of on-site operation risks in the power grid based on BERT.
[0014] A grading table text enhancement module, which is used to perform text enhancement on the grading table text and convert it into text similar to the content of the historical database text.
[0015] The historical database text quality improvement module is used to delete redundant text and correct the risk level of the historical database text in the historical database to improve its text quality.
[0016] The BERT model training module uses the graded table text after text enhancement and the historical database text after text quality improvement to construct a relatively complete sample set, and uses this sample set to train the BERT model.
[0017] The rating module uses the trained BERT model to automatically rate the risk level of on-site power grid operations.
[0018] Advantages of the present invention: The present invention can make full use of the graded table and the historical database text and improve their quality. Therefore, it can effectively reduce the manual workload, lower the data quality requirements, and improve the accuracy of the automatic rating result of the on-site operation risk level. At the same time, the methods of text enhancement and risk level correction proposed by the present invention are universal and can be applied to other deep learning models. Brief Description of the Drawings
[0019] Figure 1 is the BERT single-text topic classification model;
[0020] Figure 2 is the model training flow chart;
[0021] Figure 3 is the schematic structural diagram of the device of the present invention. Detailed Embodiment
[0022] The core technical concept of the present invention: According to the characteristics of two types of on-site power grid operation texts, namely the graded table and the historical database, the text quality is improved through processing methods such as text enhancement, redundant text deletion, and risk level correction, and a better classification effect is achieved on the BERT (Bidirectional Encoder Representations from Transformers) single-text topic classification model.
[0023] Among them, the BERT model is a natural language processing model proposed by Google in 2018, which has achieved the best results in tasks such as sentiment analysis, text classification, and entity information recognition. It is a milestone model achievement in the history of natural language processing development. Therefore, the present invention selects the single-text topic classification model of BERT as the model choice for deep learning, and the structure of this model is as Figure 1 shown. The model takes a single text T toki(i = 1, 2, …, n) is used as the input, adding a specific token [CLS] at the beginning of the sentence and a specific delimiter [SEP] at the end. The input in natural language form needs to be encoded, and the vectorized representation E of the input is obtained through token embedding, segment embedding, and position embedding. CLS and E i (i = 1, 2, …, n), E SEP . The vectorized input is passed into a bidirectional Transformer structure, and the hidden vector T corresponding to each word is obtained by means of a multi-head self-attention mechanism. CLS and T i (i = 1, 2, …, n), T SEP . The input of the model is comprehensively represented by the hidden vector T of the specific token. CLS The calculation formula is as follows:
[0024] P sen = softmax(T CLS W T ) (1)
[0025] where P sen represents the predicted value of the risk level of the input operation text; T CLS is the final hidden state of the specific token [CLS] that comprehensively represents the semantics of the input text in the BERT model; W represents the weight coefficient matrix of the fully connected layer.
[0026] The risk level assessment of on-site power grid operations is actually a multi-classification problem. Therefore, the output P sen is a 5-dimensional vector, and its elements respectively represent the probability values of the risk level assessment of the input text being rated from I to V. The level corresponding to the maximum value is selected as the predicted level.
[0027] The technical solution of the present invention is as follows: First, perform text enhancement on the classification table text to convert it into text similar to the historical database text; then, perform redundant text deletion and risk level error correction on the historical database text to improve its text quality; then, use the text-enhanced classification table text and the historical database text with improved text quality to construct a relatively complete sample set, and use this sample set to train the BERT model; finally, use the trained BERT model to achieve automatic rating of the risk level of on-site power grid operations, and the process is as Figure 2 shown.
[0028] At the same time, according to the above technical solution, the present invention also gives the corresponding device structure diagram, as Figure 3 shown, where:
[0029] The classification table text enhancement module is used to perform text enhancement on the classification table text to convert it into text similar to the historical database text.
[0030] A historical database text quality improvement module for deleting redundant texts and correcting risk levels in the historical database text to improve its text quality.
[0031] A BERT model training module that constructs a relatively complete sample set using the text-enhanced grading table text and the historical database text with improved text quality, and uses this sample set to train the BERT model.
[0032] A rating module that uses the trained BERT model to automatically rate the risk levels of on-site power grid operations.
[0033] The present invention adopts the following specific steps:
[0034] Step 1: Read in the grading table text. The read grading table text mainly includes content such as operation categories, voltage levels, operation scopes, operation contents, and operation risk levels.
[0035] Step 2: Enhance the grading table text. Enhancing the grading table text means converting general text into content closer to actual on-site operation text. Specifically, two methods are adopted:
[0036] The first method is to split the operation content in the risk table using semicolons and " / " as markers, and convert text involving multiple operation objects or operation contents into multiple single operation contents involving a single operation object.
[0037] The second method is to convert Class A, B, C, D, and E maintenance in the operation content into specific operation contents.
[0038] Step 3: Output the enhanced grading table to the sample set. Supplement the content of the text-enhanced grading table to the sample set, aiming to alleviate problems such as incorrect risk level content and overly large gaps in the proportions of different risk levels existing in the historical database text, and achieve the comprehensive utilization of both. Among them, the incorrect risk level content refers to the incorrect ratings of on-site operations by on-site operation personnel; the overly large gap in the proportions of different risk levels means that in on-site operation texts, the number of high-risk operation texts is far less than the number of low-risk operation texts.
[0039] Step 4: Read in the historical database text. The read historical database text mainly includes content such as voltage level, operation start time, operation end time, operation content, operation category, power outage status, and operation risk level.
[0040] Step 5: Remove redundant texts from the historical database. Redundant texts refer to on-site operation texts that are reported multiple times due to the operation duration reaching several days. The specific method for removing redundancy is as follows: First, divide the texts into several groups according to the start time and end time of the operations in the historical database; then, use jieba to segment the operation content texts within each group, splitting the complete sentences into several words; finally, calculate the word overlap rate between the texts within each group respectively, and consider the operation content with an overlap rate of over 95% as the same operation and delete it. Among them, jieba is a commonly used Chinese text segmentation tool.
[0041] Step 6: Correct the risk levels in the historical database. To reduce the manual workload, the present invention proposes a semi-supervised error correction method, which is divided into two major steps: rough error correction and fine error correction. The specific implementation is as follows.
[0042] 1) Rough error correction. First step, divide the on-site operation content into several parts according to the voltage level and professional type; second step, use the TF-IDF algorithm to screen out the keywords in each part, and complete clustering using the k-means clustering algorithm based on the screened keywords; third step, count the risk levels in each cluster result, take the risk level with the largest quantity as the cluster label, and correct the sample risk levels that are inconsistent with the cluster label. Among them, the TF-IDF algorithm is a commonly used keyword screening algorithm, and the k-means clustering algorithm is a commonly used clustering algorithm.
[0043] 2) Fine error correction. First step, randomly divide the sample set after rough error correction into two parts in a ratio of 9:1, which are used as the training set and the validation set respectively. Use the training set and the validation set to train the BERT model, and use the trained model to predict the complete sample set; second step, list the texts with inconsistent risk level labels in the prediction results and the labels in the historical database as suspicious data, manually verify this part of the data, and replace the original data with the verified and corrected data.
[0044] Step 6: Output the content of the historical database to the sample set.
[0045] Step 7: Conduct rating training and testing using the BERT model.
[0046] Step 8: Output the trained rating model. Thereafter, the trained rating model can be used to automatically rate the risk levels of on-site operation texts of the power grid.
[0047] Application example:
[0048] Use the grading table and 45,662 on-site operation historical databases of a provincial company in a week as the original data set for application verification.
[0049] Step 1: Read in the classification table text. Since the operation categories in the historical database mainly include power transmission, substation, and power distribution, a total of 250 pieces of relevant operation category classification table text are read in. Table 1 shows the content of the on-site operation classification table text for some substation operation categories.
[0050] Table 1 Example of the on-site operation classification table for some substations
[0051]
[0052] Step 2: Enhance the classification table text. Enhance the content of the classification table. The original 250 pieces of text are enhanced to 694 pieces of text. Table 2 shows the result after enhancing the content related to Class B maintenance of transformers in Table 1.
[0053] Table 2 Example of text enhancement
[0054]
[0055]
[0056] Step 3: Output the enhanced classification table to the sample set.
[0057] Step 4: Read in the historical database text. Read in 45,662 pieces of on-site operation historical database text of a provincial company in one week. Table 3 shows some examples of the historical database text.
[0058] Table 3 Example of the historical database
[0059]
[0060] Step 5: Remove redundancy from the historical database text. Delete 8,244 pieces of redundant text. After deletion, the number of historical database text becomes 37,418 pieces.
[0061] Step 6: Correct the risk levels in the historical database. During rough correction, 3,000 cluster results are obtained through the cooperation and clustering of the TF-IDF algorithm and the k-means algorithm. A total of 4,878 pieces of data with inconsistent risk levels are corrected by comparing the risk levels that account for the main components in each cluster. During fine correction, 2,835 pieces of suspicious data are discovered using the BERT model, and a total of 1,122 pieces of data are corrected through manual verification.
[0062] Step 7: Output the content of the historical database to the sample set. Combining the content of the classification table after text enhancement in Step 3, a sample set with improved quality such as text enhancement, redundancy deletion, and level correction is obtained at this time, totaling 38,112 pieces.
[0063] Step 8: Use the BERT model for rating training and testing. To verify the classification effect of the BERT model, the common TextCNN model is selected as the comparison model of the BERT model. Among them, the key parameters of the BERT model are set as follows: the number of hidden layers is set to 12, the number of taps in the multi-tap self-attention mechanism is 12, the word vector dimension is set to 768, the maximum length of the input sentence is set to 256 words, the batch number is set to 16, and the learning rate is set to 1×10 -5 The key parameters of the TextCNN model are set as follows: the convolution kernel size is (2, 3, 4), the number of convolution kernels is set to 256, the word vector dimension is set to 192, the maximum length of the input sentence is set to 256, the batch size is set to 128, and the learning rate is set to 1×10 -3 .
[0064] In order to verify the effectiveness of the quality improvement method used in the sample set construction process, the original historical database, the original grading table, and the grading table and historical database formed by different processing methods were used as sample sets to conduct experiments on the BERT model and the comparison model. The experimental process is as follows: (1) All the relatively small number of level I and II were selected from the historical texts after redundant text deletion and risk level correction to be included in the test set, and 10% of the level III, IV, and V were randomly selected to be included in the test set, resulting in a test set of 3786 items; (2) When training each model, the samples of level III, IV, and V in the test set were removed from their respective sample sets, and then divided into training set and validation set at a ratio of 9:1 for model training, and the classification model corresponding to each sample set was obtained; (3) The rating effect of each model was tested using the test set.
[0065] The experimental environment used is as follows: Core i9-10900K processor with a main frequency of 3.70GHz, NVIDIA GeForce RTX 3090 graphics card with a memory size of 24GB, memory size of 64GB, Python version 3.8.10, and PyTorch version 1.12.1. The F1 value is used as the evaluation indicator for the automatic rating of the risk level of power grid field operations. Table 4 reflects the rating effects of the BERT model and the TextCNN model on different sample sets.
[0066] Table 4 Rating effects of BERT model and TextCNN model on different sample sets
[0067]
[0068]
[0069] As can be seen from Table 4, the BERT model has better rating effect compared with the commonly used TextCNN model. The text quality improvement methods such as text augmentation and risk level error correction proposed in this embodiment can effectively improve the rating effect of the model. Among them, the BERT model is improved by 0.063, and the TextCNN model is improved by 0.073. After redundant text is deleted, the rating performance of each model decreases slightly. On the one hand, it shows that redundant text has an impact on the performance of the rating model; on the other hand, since there are 8,244 redundant texts, accounting for about 18% of the original historical database, the number of training samples becomes much smaller after deletion, indicating that the number of samples has an impact on the rating performance of the BERT model. Thus, it can be seen that when the number of texts in the historical database is insufficient, retaining a certain proportion of redundant texts is beneficial to automatic rating; when the number of historical texts is sufficient, redundant texts can be deleted to improve text quality. The above verifies the effectiveness and generality of this embodiment.
Claims
1. A method for automatically rating risks of on-site power grid operations based on BERT, characterized in that the method comprises the following steps: First, perform text enhancement on the grading table text to convert it into text with content similar to that in the historical database; Next, delete redundant text from the historical database text and correct the risk levels in the historical database to improve its text quality; Then, use the text-enhanced grading table text and the historical database text with improved text quality to construct a relatively complete sample set, and use this sample set to train the BERT model; Finally, use the trained BERT model to achieve automatic rating of the risk levels of on-site power grid operations; The correction of the risk levels in the historical database is specifically as follows: 1) Coarse correction; The first step is to divide the on-site operation content into several parts according to the voltage level and professional type; The second step is to use the TF-IDF algorithm to screen out the keywords in each part, and complete clustering using the k-means clustering algorithm based on the screened keywords; The third step is to count the risk levels in each cluster result, take the risk level with the largest quantity as the cluster label, and correct the risk levels of the samples inconsistent with the cluster label; 2) Fine correction; The first step is to randomly divide the sample set after coarse correction into two parts with a ratio of 9:1, which are used as the training set and the validation set respectively. Use the training set and the validation set to train the BERT model, and use the trained model to predict the complete sample set; The second step is to list the texts with inconsistent risk level labels in the prediction results and the labels in the historical database as suspicious data, manually verify some data, and replace the original data with the verified and corrected data.
2. A method for automatically rating risks of on-site power grid operations based on BERT according to claim 1, characterized in that: It further includes reading in the grading table text, and the grading table text mainly includes operation categories, voltage levels, operation scopes, operation contents, and operation risk levels.
3. A method for automatically rating risks of on-site power grid operations based on BERT according to claim 2, characterized in that: The enhancement of the grading table text refers to converting general text into content close to the actual on-site operation text; The following two methods are used for enhancement: Method 1: Split the operation content in the risk table with semicolons and " / " as markers, and convert the text involving multiple operation objects or operation contents into multiple single operation contents involving a single operation object; Method 2: Convert the maintenance of categories A, B, C, D, and E in the operation content into specific operation contents.
4. A method for automatically rating risks of on-site power grid operations based on BERT according to claim 1, characterized in that: It further includes reading in the historical database text, and the historical database text mainly includes voltage levels, operation start times, operation end times, operation contents, operation categories, whether power is cut off, and operation risk levels.
5. A method for automatically rating risks of on-site power grid operations based on BERT according to claim 4, characterized in that: The specific deletion of redundant text is: First, divide the text into several groups according to the operation start time and end time in the historical database; Then, use a Chinese text tokenization tool to tokenize the operation content text within each group, splitting the complete sentences into several words; Finally, calculate the word coincidence rate between the texts within each group respectively, and regard the operation content with a coincidence rate of over 95% as the same operation and delete it.
6. A BERT-based automatic risk rating device for on-site power grid operations, used to implement the method described in claim 1, characterized in that, it includes: a grading table text enhancement module, used to enhance the grading table text and convert it into text similar to the historical database text content; a historical database text quality improvement module, used to delete redundant text from the historical database text and correct the risk levels in the historical database to improve its text quality; a BERT model training module, which constructs a relatively complete sample set using the enhanced grading table text and the historical database text with improved text quality, and uses this sample set to train the BERT model; a rating module, which uses the trained BERT model to achieve automatic rating of the on-site power grid operation risk level.