Method, system and electronic equipment for predicting operation risks of power Internet of Things

By combining adaptive comprehensive oversampling and Catboost ensemble learning model with Bayesian optimization algorithm, the problems of data imbalance and blind model parameter tuning in power Internet of Things risk analysis are solved, achieving more accurate and stable risk prediction.

CN114511194BActive Publication Date: 2025-09-19NORTHEAST DIANLI UNIVERSITY +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210015149.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-07
Publication Date
2025-09-19
Estimated Expiration
2042-01-07

AI Technical Summary

Technical Problem

Existing technologies fail to comprehensively consider information, physical, and social risks in the risk analysis of the power Internet of Things, resulting in data imbalance and decreased model prediction accuracy. In addition, the blind tuning of traditional model parameters leads to unstable prediction accuracy.

Method used

An adaptive comprehensive oversampling method is used to balance multi-source data, and combined with the Catboost ensemble learning model and Bayesian optimization algorithm, a power Internet of Things operation risk prediction model is constructed. By fusing information, physical and social measurement data through time series, pseudo samples are generated to improve model accuracy.

Benefits of technology

The prediction accuracy and stability of the power Internet of Things operation risk prediction model have been improved, the false alarm rate has been reduced, and the accuracy of risk prediction has been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114511194B_ABST
    Figure CN114511194B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electric power Internet of Things, and in particular to an operational risk prediction method, system, and electronic device for an electric power Internet of Things. In this method, based on a time series as a benchmark, multi-source data within a preset historical time period are fused to obtain a complete data set. The complete data set is subjected to data balancing processing based on an adaptive comprehensive oversampling method to obtain a balanced data set. An electric power Internet of Things operational risk prediction model is trained based on the balanced data set. Based on the current multi-source data of the electric power Internet of Things to be tested and the electric power Internet of Things operational risk prediction model, an operational risk prediction result of the electric power Internet of Things to be tested is obtained. By fusing measurement data from the information side, the physical side, and the social side and performing data balancing processing on the fused data based on the adaptive comprehensive oversampling method, the prediction accuracy of the trained electric power Internet of Things operational risk prediction model can be improved, thereby improving the accuracy of the operational risk prediction result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power Internet of Things, and in particular to a method, system and electronic equipment for predicting operation risks of the power Internet of Things. Background Art

[0002] The risks faced by the Power Internet of Things (PoI) during operation are diverse and expanding in scope. Risks such as equipment failure, cyberattacks, and human error all have significant impacts on the PoI. If these risks are not addressed promptly, they could lead to a series of cross-domain cascading failures, or even catastrophic power outages in severe cases. The earlier risks are identified and the more timely measures are taken, the lower the cost of risk control. Therefore, conducting research on the operational risk prediction of the PoI, integrating multi-source data from the information, physical, and social aspects of the PoI, and fully mining the hidden information within this data, is crucial for ensuring the safe and stable operation of the PoI. This research aims to predict various security risks facing the PoI before failures occur, thereby identifying weak links and addressing them.

[0003] There are three deficiencies in the current risk analysis of the power Internet of Things:

[0004] 1) Traditional power Internet of Things operational risk predictions mostly study the risks of the power information domain and physical domain in isolation, rarely considering the impact of social risks, and failing to conduct a comprehensive analysis of the power Internet of Things operational risks from the perspective of information, physics, and society. The power Internet of Things operational risks are essentially determined by the risks in the information, physical, and social spaces. Therefore, risk predictions must comprehensively consider the measurement data from the information, physical, and social sides.

[0005] 2) The small proportion of risk samples in power IoT operational data can cause data imbalance, leading to a bias in trained classifiers towards the majority class and reduced performance. This poses challenges to subsequent model training accuracy and can lead to false positives when predicting risks. Therefore, effective data processing of multi-source data from information, physics, and society is essential before model training. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to address the deficiencies of the existing technology and provide a method, system and electronic equipment for predicting the operation risks of the power Internet of Things.

[0007] The technical solution of the method for predicting the operation risk of the power Internet of Things of the present invention is as follows:

[0008] Based on the time series, multi-source data within a preset historical time period is integrated to obtain a complete data set. The multi-source data includes: measurement data from the information side, physical side, and social side of the power Internet of Things;

[0009] When the data in the complete data set is unbalanced, performing data balancing processing on the complete data set based on an adaptive comprehensive oversampling method to obtain a balanced data set;

[0010] A power Internet of Things operation risk prediction model is obtained based on the balanced data set training;

[0011] According to the current multi-source data of the electric power Internet of Things to be tested and the electric power Internet of Things operation risk prediction model, an operation risk prediction result of the electric power Internet of Things to be tested is obtained.

[0012] The beneficial effects of the operation risk prediction method of the power Internet of Things of the present invention are as follows:

[0013] On the one hand, by introducing measurement data from the information side, physical side, and social side that affect the security of the power Internet of Things, data fusion is performed based on time series, and a complete data set that integrates the measurement data from the information side, the physical side, and the social side is constructed based on random matrix theory. On the other hand, based on the adaptive synthetic oversampling (ADASYN) method, data balancing is performed on the fused data, which can generate pseudo samples that are highly similar to real samples, assist in constructing a balanced data set, and overcome the disadvantages of low training accuracy caused by the low number of samples in certain categories, resulting in unstable performance of the risk prediction model. Based on the above two aspects, the prediction accuracy of the trained power Internet of Things operation risk prediction model can be improved, and the accuracy of the operation risk prediction results can be improved.

[0014] Based on the above solution, the operation risk prediction method of the power Internet of Things of the present invention can also be improved as follows.

[0015] Furthermore, the power Internet of Things operation risk prediction model obtained by training based on the balance data set includes:

[0016] A Catboost ensemble learning model is constructed using a symmetric decision tree as a base classifier, and is trained based on the balanced dataset to obtain a Catboost ensemble classifier;

[0017] Using the Bayesian optimization method to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier;

[0018] All optimal parameters are passed to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model.

[0019] The beneficial effect of adopting the above further scheme is that the traditional Catboost model can improve the classification performance by merging multiple classifiers, but the model performance will be affected by key parameters, and manual parameter adjustment has a certain degree of blindness, which is easy to lose the optimal solution of the parameters, and it takes too long, which will affect the accuracy of the risk prediction model. In this application, the modeling process includes two model training and learning stages. In the first stage, a Catboost ensemble learning model is constructed with a symmetric decision tree as the base classifier, and a Catboost ensemble classifier is obtained by training; in the second stage, the Bayesian optimization algorithm is introduced to optimize the parameters of the Catboost model, so that the obtained power Internet of Things operation risk prediction model has higher prediction accuracy.

[0020] Furthermore, the multi-source data within a preset historical time period is fused based on the time series to obtain a complete data set, including:

[0021] Generate an original dataset Dataset based on the multi-source data within the preset historical time period, Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1 ,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D s Represents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iNRepresents: N measurement data collected from the physical side of the Power Internet of Things at the i-th moment in the preset historical time period, where i, n, and N are all positive integers;

[0022] Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, the complete dataset D is constructed:

[0023] The technical solution of the operation risk prediction system of the power Internet of Things of the present invention is as follows:

[0024] Includes fusion module, balancing module, training module and prediction module;

[0025] The fusion module is used to fuse multi-source data within a preset historical time period based on time series to obtain a complete data set, wherein the multi-source data includes measurement data on the information side, measurement data on the physical side, and measurement data on the social side of the power Internet of Things;

[0026] The balancing module is configured to: when the data in the complete data set is unbalanced, perform data balancing processing on the complete data set based on an adaptive comprehensive oversampling method to obtain a balanced data set;

[0027] The training module is used to: obtain a power Internet of Things operation risk prediction model based on the balance data set training;

[0028] The prediction module is used to obtain an operation risk prediction result of the power Internet of Things to be tested based on the current multi-source data of the power Internet of Things to be tested and the power Internet of Things operation risk prediction model.

[0029] The beneficial effects of the operation risk prediction system of the power Internet of Things of the present invention are as follows:

[0030] On the one hand, by introducing measurement data from the information side, physical side, and social side that affect the security of the power Internet of Things, data fusion is performed based on time series, and a complete data set that integrates the measurement data from the information side, the physical side, and the social side is constructed based on random matrix theory. On the other hand, based on the adaptive synthetic oversampling (ADASYN) method, data balancing is performed on the fused data, which can generate pseudo samples that are highly similar to real samples, assist in constructing a balanced data set, and overcome the disadvantages of low training accuracy caused by the low number of samples in certain categories, resulting in unstable performance of the risk prediction model. Based on the above two aspects, the prediction accuracy of the trained power Internet of Things operation risk prediction model can be improved, and the accuracy of the operation risk prediction results can be improved.

[0031] Based on the above solution, the operation risk prediction system of the power Internet of Things of the present invention can also be improved as follows.

[0032] Furthermore, the training module is specifically used for:

[0033] A Catboost ensemble learning model is constructed using a symmetric decision tree as a base classifier, and is trained based on the balanced dataset to obtain a Catboost ensemble classifier;

[0034] Using the Bayesian optimization method to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier;

[0035] All optimal parameters are passed to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model.

[0036] The beneficial effect of adopting the above further scheme is that the traditional Catboost model can improve the classification performance by merging multiple classifiers, but the model performance will be affected by key parameters, and manual parameter adjustment has a certain degree of blindness, which is easy to lose the optimal solution of the parameters, and it takes too long, which will affect the accuracy of the risk prediction model. In this application, the modeling process includes two model training and learning stages. In the first stage, a Catboost ensemble learning model is constructed with a symmetric decision tree as the base classifier, and a Catboost ensemble classifier is obtained by training; in the second stage, the Bayesian optimization algorithm is introduced to optimize the parameters of the Catboost model, so that the obtained power Internet of Things operation risk prediction model has higher prediction accuracy.

[0037] Furthermore, the fusion module is specifically used to:

[0038] Generate an original dataset Dataset based on the multi-source data within the preset historical time period, Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1 ,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D sRepresents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iN Represents: N measurement data collected from the physical side of the Power Internet of Things at the i-th moment in the preset historical time period, where i, n, and N are all positive integers;

[0039] Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, the complete dataset D is constructed:

[0040] A storage medium of the present invention stores instructions. When a computer reads the instructions, the computer executes any one of the above-mentioned methods for predicting the operation risks of the electric power Internet of Things.

[0041] An electronic device of the present invention includes a processor and the above-mentioned storage medium, wherein the processor executes instructions in the storage medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a flow chart of a method for predicting operational risks of an electric power Internet of Things according to an embodiment of the present invention;

[0043] Figure 2 Schematic diagram of the process of obtaining a balanced dataset;

[0044] Figure 3 A flowchart for training the risk prediction model for the power Internet of Things.

[0045] Figure 4 Schematic diagram of the GBDT training process;

[0046] Figure 5 It is a schematic diagram of the topological structure;

[0047] Figure 6 This is the confusion matrix of the risk prediction results before data balancing;

[0048] Figure 7 This is the confusion matrix of the risk prediction results after data balancing;

[0049] Figure 8 Schematic diagram of the ROC curve;

[0050] Figure 9 Schematic diagram of the precision-recall curve;

[0051] Figure 10 A schematic diagram of the confusion matrix;

[0052] Figure 11 This is the ROC curve after parameter optimization;

[0053] Figure 12 This is the precision-recall curve after parameter optimization;

[0054] Figure 13 The confusion matrix after parameter optimization;

[0055] Figure 14 This is a schematic diagram of the structure of an operation risk prediction system for the power Internet of Things according to an embodiment of the present invention; DETAILED DESCRIPTION

[0056] like Figure 1 As shown, a method for predicting the operation risk of the power Internet of Things according to an embodiment of the present invention includes the following steps:

[0057] S1. Based on the time series, multi-source data within a preset historical time period are integrated to obtain a complete data set. The multi-source data includes: measurement data from the information side, measurement data from the physical side, and measurement data from the social side of the power Internet of Things;

[0058] Among them, the preset historical time period can be set according to actual conditions, and multi-source data can be obtained by collecting measurement data from the information side, physical side and social side of multiple power Internet of Things.

[0059] Among them, the measurement data on the information side includes: attack signals on the power Internet of Things, and network traffic of the power Internet of Things, etc.; the measurement data on the physical side includes: three-phase voltage and three-phase current of the power Internet of Things, etc.; the measurement data on the social side includes meteorological data such as humidity, temperature, precipitation, etc.

[0060] S2. When the data in the complete dataset is unbalanced, the complete dataset is balanced based on the adaptive synthetic oversampling method to obtain a balanced dataset. Since the complete dataset may have data imbalance, which may lead to a decrease in subsequent model performance, the minority class samples are oversampled based on the adaptive synthetic oversampling (ADASYN) algorithm before model training, such as Figure 2 The specific steps are as follows:

[0061] S20. Calculate whether the data is balanced, that is, determine whether the data in the complete data set is balanced. Specifically, the complete data set D includes a risk data set and a normal operation data set. The risk data refers to the data generated by network attacks, system failures, system disturbances, etc. on the power Internet of Things, which are distributed on the information side, the physical side, and the social side. Based on these risk data, all risk samples in the complete data set D can be determined, such as "x 11 x 21 …x n1 y 11 y 21 …y n1 z 11 z 21 …z n1 ", etc., constitute the risk data set. The data in the normal operation data set is: the remaining samples in the complete data set D except the risk data set.

[0062] That is to say, each row of data in the complete data set D is a sample, for example, "x 11 x 21 …x n1 y 11 y 21 …y n1 z 11 z 21 …z n1 " is a sample, and the samples including risk data are determined as risk samples to form a risk data set, and the remaining samples are formed into a normal operation data set.

[0063] The number of all samples in the risk dataset is m s , the number of all samples of the normal running data set is m l , the imbalance degree is calculated by the following formula: When d is less than a preset threshold, it is determined that the data in the complete data set D is unbalanced. The preset threshold is 1% or 1‰, etc., and can also be set and adjusted according to actual conditions.

[0064] Determine whether the data in the complete data set is balanced based on the above content. If so, determine the current complete data set as a balanced data set. If not, execute S21.

[0065] S21, oversampling, i.e., performing data balancing on the complete data set, and determining the oversampled complete data set as the balanced data set. The oversampling process is as follows, specifically:

[0066] S210, calculate the number of samples G to be synthesized, G=(m l -m s)*b, where b∈[0,1], and the specific value of b can be set according to the actual situation. For example, when b=1, the sum of the number of synthesized samples and the number of risk samples in the risk dataset is equal to the number of samples in the normal operating dataset. In addition, b can also be set to different values ​​such as 0.5;

[0067] S211. For each risk sample, find its K nearest neighbor samples and calculate: Where Δ is the number of samples in the K nearest neighbors that are in the normal running dataset, and Z is a normalization factor to ensure that r forms a distribution. Thus, the more samples in the normal running dataset surrounding a risky sample, the higher its r. Both Δ and K are integers.

[0068] S212, through the formula: g j =r j ×G, calculate the number of synthetic samples required for each risk sample, r j Indicates r, g corresponding to the j-th risk sample j It represents the number of samples to be synthesized corresponding to the j-th risk sample, where j is an integer.

[0069] S213. Synthesize a synthetic sample corresponding to the j-th risk sample. Specifically, the risk sample serves as the minority class sample in the ADASYN algorithm, and the samples in the normal operating dataset serve as the majority class samples in the ADASYN algorithm. Synthesize a synthetic sample corresponding to each risk sample. The ADASYN algorithm automatically determines the number of synthetic samples required for each minority class sample to obtain sufficient pseudo data that is highly similar to the original data. This effectively eliminates the impact of data imbalance on the training accuracy of machine learning algorithms.

[0070] S3. Develop a power Internet of Things operation risk prediction model based on the balanced data set training;

[0071] S4. Obtain an operation risk prediction result of the power Internet of Things to be tested based on the current multi-source data of the power Internet of Things to be tested and the power Internet of Things operation risk prediction model.

[0072] On the one hand, by introducing measurement data from the information side, physical side, and social side that affect the security of the power Internet of Things, data fusion is performed based on time series, and a complete data set that integrates the measurement data from the information side, the physical side, and the social side is constructed based on random matrix theory. On the other hand, based on the adaptive synthetic oversampling (ADASYN) method, data balancing is performed on the fused data, which can generate pseudo samples that are highly similar to real samples, assist in constructing a balanced data set, and overcome the disadvantages of low training accuracy caused by the low number of samples in certain categories, resulting in unstable performance of the risk prediction model. Based on the above two aspects, the prediction accuracy of the trained power Internet of Things operation risk prediction model can be improved, and the accuracy of the operation risk prediction results can be improved.

[0073] Optionally, in the above technical solution, the power Internet of Things operation risk prediction model is obtained based on the balanced data set training, including:

[0074] S30, using the symmetric decision tree as the base classifier, constructing a Catboost ensemble learning model, and training it based on the balanced data set to obtain a Catboost ensemble classifier; Figure 3 As shown, specifically:

[0075] The balanced data set is divided into a training set and a test set. Then, a Catboost ensemble learning model is constructed using a symmetric decision tree as the base classifier. Specifically, multiple symmetric decision trees can be constructed. Each symmetric decision tree pair is trained for classification based on the training set and the test set to obtain a Catboost ensemble classifier.

[0076] S31, using the Bayesian optimization method to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier; Figure 3 As shown, specifically:

[0077] S310, determining a parameter initialization population, specifically: obtaining each parameter of the Catboost ensemble classifier and establishing a parameter initialization population;

[0078] S311. Establish a substitution probability model and bring the parameter-initialized population into the substitution probability model;

[0079] S312, calculating the objective function;

[0080] S313, building a Bayesian network;

[0081] S314. The Bayesian network samples and determines whether the maximum number of iterations has been reached. If so, the optimal parameter combination is output. If not, the probability model is modified and the process returns to S311 until the optimal parameter combination is output. The optimal parameter combination includes the optimal parameters corresponding to each parameter of the Catboost ensemble classifier.

[0082] Among them, the specific technical details of S310 to S314 are well known to those skilled in the art and are not elaborated here.

[0083] S32. All optimal parameters are passed to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model. Specifically:

[0084] Taking the symmetric decision tree as the base classifier, a Catboost ensemble learning model is constructed. In order to further improve the model performance, the Bayesian optimization algorithm is used to find the optimal parameters of the model. First, a Catboost risk prediction model is constructed with the symmetric decision tree as the base classifier; then, Bayesian optimization is used to update the posterior distribution of the objective function according to the given objective function by continuously adding sample points, so as to obtain the optimal parameters of the model, thereby improving the classification accuracy of the model for samples. A scheme for constructing a power Internet of Things operation risk prediction model based on BO-Catboost. Specifically: Gradient boosting decision tree (GBDT) is an ensemble learning framework based on decision trees. It uses an additive model (i.e., a linear combination of basis functions) and an algorithm that continuously reduces the residuals generated by the training process to achieve data classification or regression. During the training process, the accuracy of the final classifier is continuously improved by reducing the deviation. The weak classifiers obtained in each round of training are weighted and summed to obtain the final total classifier. The training process of GBDT is as follows: Figure 4 shown.

[0085] With the exponential growth of data volumes, the GDBT algorithm suffers from the common drawbacks of overfitting and slow training. Catboost improves upon GBDT by introducing a ranking boosting strategy to address the gradient bias and prediction offset issues inherent in the standard GBDT model. It also employs a fully symmetric decision tree to improve the model's generalization and prediction speed, while ensuring both training and prediction accuracy.

[0086] 1) Quick scoring

[0087] Catboost uses fully symmetric decision trees (ODTs) as its base learner. Unlike typical decision trees, ODTs use identical features and thresholds for splitting internal nodes of the same depth. Therefore, ODTs can be transformed into decision tables with 2D entries, where d represents the number of layers in the decision tree. This structure is more balanced and processes features much faster than a typical decision tree.

[0088] 2) Sorting-type boosting algorithm

[0089] Catboost uses the ranking promotion method to reduce the gradient deviation and solve the prediction offset problem. k, use all sample data except this sample to train the corresponding model M k , and continuously train the weak learner by calculating the gradient estimate of the sample data to obtain an optimized model, thereby improving the generalization ability of the model. The algorithm processing flow is as follows:

[0090] Input: training set The number of iterations T;

[0091] Output: Model M1, M2, ..., M Q ;

[0092] ① Randomly generate a sequence σ, sort the sample data in the training set W according to the sequence value, and calculate the corresponding models M1, M2, ..., M respectively. Q ;

[0093] ② Each model M k , M k ∈(M1,M2,...,M Q ) are obtained by training with the first k randomly arranged samples;

[0094] ③In the process of iterative update, model M k-1 It is an unbiased estimation of the gradient based on the kth sample;

[0095] ④ Based on the sample gradient, the weak learner is continuously trained until the number of iterations reaches the maximum, the final model is output, and training stops. Where k and Q are both positive integers.

[0096] The Catboost model can improve classification performance by combining multiple classifiers. However, model performance is affected by key parameters. Manual parameter tuning requires considerable effort and is inherently blind, making it easy to miss the optimal parameter solution, thus impacting the accuracy of the risk prediction model. Compared to other hyperparameter optimization algorithms, such as grid search, random search, and genetic algorithms, Bayesian optimization requires fewer initial sample points and offers high optimization efficiency, making it more suitable for model hyperparameter tuning scenarios.

[0097] The Bayesian optimization algorithm replaces crossover and mutation in the genetic algorithm with Bayesian network sampling. First, after sampling, the joint probability distribution of the better solution is obtained and a Bayesian network model is generated. The network model is then sampled to generate new candidate solutions for the next iteration. By looping this process, the optimal solution is finally obtained. The specific algorithm flow is as follows: Figure 3 shown.

[0098] In order to find a suitable set of hyperparameters for the model and improve the prediction accuracy of the model, this paper constructs a power Internet of Things operation risk prediction model based on BO-Catboost. The specific implementation process is as follows:

[0099] ① Set the optimization range of Catboost algorithm parameters. Each parameter can take any value within the range.

[0100] ② Initialize the regression prediction model of the Catboost algorithm, where the Catboost algorithm is selected as the training target and the model performance index is used as the evaluation standard.

[0101] ③ Initialize the Bayesian optimization algorithm and establish an alternative probability model. Then, select a parameter combination from the parameter set as the initial parameters for the Catboost algorithm model and train it. After training, test the model using the test set. The test set and the result set are used as inputs to the evaluation function for evaluation. The result is the performance evaluation value of the Catboost algorithm model using this parameter combination.

[0102] ④ According to the effect evaluation value obtained in ③, search for the best parameters on the proxy model and output the corresponding evaluation value and parameters at the same time.

[0103] ⑤ When the number of iterations to find the parameters reaches the maximum, the optimization is stopped and the parameter combination that maximizes the evaluation value is found from the surrogate model. The optimal parameter combination is the parameter of the Catboost algorithm, and the final prediction model is obtained through training.

[0104] The traditional Catboost model can improve classification performance by combining multiple classifiers, but the model performance will be affected by key parameters, and manual parameter adjustment has a certain degree of blindness, which is easy to lose the optimal solution of parameters, and takes too long, which will affect the accuracy of the risk prediction model. In this application, the modeling process includes two model training and learning stages. In the first stage, a Catboost ensemble learning model is constructed with a symmetric decision tree as the base classifier, and a Catboost ensemble classifier is obtained by training; in the second stage, the Bayesian optimization algorithm is introduced to optimize the parameters of the Catboost model, so that the obtained power Internet of Things operation risk prediction model has higher prediction accuracy.

[0105] Optionally, in the above technical solution, based on the time series, multi-source data within a preset historical time period is fused to obtain a complete data set, including:

[0106] S10, generate the original dataset Dataset based on the multi-source data within the preset historical time period, Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D s Represents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period; where i, n and N are all positive integers.

[0107] Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, we construct the complete dataset D: Specifically:

[0108] In order to conduct a comprehensive analysis of the operational risks of the power Internet of Things from the perspectives of information, physical, and social, a random matrix method is introduced. Within any period of time, the measurement data collected from any attribute of any node on the information, physical, and social sides can form a column vector. Among them, any attribute of any node, such as the attack signal of the second node + voltage of item A + temperature, is extracted from the measurement data on the information, physical, and social sides to form the original dataset. Using time series as a benchmark, data from different spaces at the same time are fused. When fusion occurs, the time series in one of the data files is selected as the benchmark. This file is called the benchmark file, and the parameters of the other data stream files must be unified to this time benchmark. Data from the information side, physical side, and social side at each sampling moment are imported according to the time series to construct a high-dimensional random matrix. This high-dimensional random matrix is ​​the complete dataset D:

[0109]

[0110] Among them, the measurement data on the information side includes: attack signals on the power Internet of Things, and network traffic of the power Internet of Things, etc.; the measurement data on the physical side includes: three-phase voltage and three-phase current of the power Internet of Things, etc.; the measurement data on the social side includes meteorological data such as humidity, temperature, precipitation, etc.

[0111] The following is an example to illustrate the technical effect of the method for predicting the operation risk of the power Internet of Things in this application. Specifically:

[0112] RT-LAB and OPNET are used to jointly simulate and build a 16-node topology structure to simulate the measurement data of the information side and the physical side. Figure 5 As shown, the element selected by the circular frame represents a circuit breaker, the element selected by the solid rectangular frame represents a transmission line, the element selected by the dotted rectangular frame represents a transformer, and the element selected by the octagonal frame represents a power supply.

[0113] In the established 16-node topology structure, more than 20,000 data items were collected over 200 seconds at an interval of 0.01 seconds. Among them, the single-phase short circuit risk was set during 15 to 16.5 seconds, the two-phase short circuit risk was set during 45 to 45.5 seconds, the two-phase grounding risk was set during 75 to 76.5 seconds, the three-phase short circuit risk was set during 120 to 120.5 seconds, the false command attack injection risk was set during 150 to 151 seconds, and the human error risk was set between 195 and 200 seconds. Measurement data on the information and physical sides were collected.

[0114] At the same time, social meteorological data for the same period was downloaded from the China Meteorological Network, primarily including humidity, temperature, and precipitation. After obtaining the information, physical, and social measurement data, the collected time series was used as a benchmark to fuse these information and social measurement data into the simulated physical dataset, resulting in a complete dataset, as shown in Table 1, where "node" represents each node.

[0115] Table 1:

[0116]

[0117] In order to analyze the degree of improvement of oversampling on the risk prediction model, the Catboost algorithm was used to train the data sets before and after balancing, and the risk prediction performance of the model was compared and analyzed. The confusion matrix of the risk prediction results before and after data balancing is shown in the figure below. Figure 6 and Figure 7 shown.

[0118] Depend on Figure 6 It can be seen that the risk categories 1, 2, 3, and 4 in the original data have low prediction accuracy and high false positives due to the small number of samples. Figure 7Analysis shows that after data balancing, the prediction accuracy for risk categories 1, 2, 3, and 4 increased by 10%, 2%, 9%, and 10%, respectively, and the false positive rate also decreased significantly. Therefore, oversampling minority class samples to obtain a balanced dataset plays an important role in reducing the false positive rate of model risk prediction and improving model stability.

[0119] Precision, recall, and F1-Score are used as performance metrics for risk prediction models. Precision measures the proportion of true positive samples among all samples predicted as positive. Recall measures the proportion of positive samples among all correctly classified samples. F1-Score is the harmonic mean of precision and recall. The calculation formulas for each metric are as follows:

[0120]

[0121] Among them, TP (True Positive) means that the positive sample is predicted as a positive example; TN (True Negative) means that the positive sample is predicted as a negative example; FP (False Positive) means that the negative sample is predicted as a positive example; FN (False Negative) means that the negative sample is predicted as a negative example.

[0122] The oversampled balanced data was put into the model for training, and the training set and test set were divided into two sets in a ratio of 7:3. The average precision, average recall rate and average F1-Score of the Catboost algorithm for risk prediction were 99.76%, 98.09% and 98.9% respectively. The ROC curve, precision-recall curve and confusion matrix of the Catboost algorithm are shown in Figure 2. Figure 8-10 shown.

[0123] Depend on Figure 8 The analysis shows that the inflection point of the ROC curve of the Catboost model is close to (0, 1), indicating that the model can achieve high prediction accuracy under the condition of low false alarm rate. Figure 9 From the analysis, we can see that the inflection point of the curve is close to (1, 1), which means that the model can achieve high precision under the condition of high recall rate. Figure 10 The analysis shows that the overall classification accuracy is good, with only categories 1, 3, and 4 requiring slight improvement. The above analysis demonstrates that the Catboost model has a certain applicability for handling operational risk prediction problems in the power Internet of Things.

[0124] Catboost model performance is affected by several key parameters, listed in Table 2. The maximum number of trees affects the model's computational cost and can lead to overfitting. The learning rate affects the total training time. The maximum tree depth also significantly affects model performance and overfitting.

[0125] Table 2:

[0126] parameter meaning default value iterations Maximum number of trees 500 learning_rate Learning rate 0.03 depth Maximum depth of the tree 6

[0127] In order to further improve the performance of the model, the Bayesian optimization method is used to find the optimal parameters of the model. In the experiment, the parameter interval settings of the Bayesian optimization are shown in Table 3.

[0128] Table 3:

[0129] parameter meaning Optimization interval iterations Maximum number of trees [100,1000] learning_rate Learning rate [0.01,0.3] depth Maximum depth of the tree [1,10]

[0130] The parameter optimization process is as follows. To avoid the randomness of the training results, five-fold cross-validation is used for training. The mean AUC of the model cross-validation 5 times is used as the objective function. Finally, the optimal parameter set of the risk prediction model is {iterations = 883, learning_rate = 0.13, depth = 9}.

[0131] By setting the optimal parameter combination in the model, the ROC curve, precision-recall curve and confusion matrix are obtained as follows: Figure 11-13 shown.

[0132] Depend on Figure 11-13 Analysis shows that the BO-Catboost algorithm achieves average precision, recall, and F1-Score for risk prediction of 99.77%, 98.8%, and 99.07%, respectively, representing improvements of 0.01%, 0.71%, and 0.17%, respectively, compared to the pre-parameter optimization period. The optimized model demonstrates excellent overall performance, with an overall false positive rate of only 1%, demonstrating the robustness of this risk prediction model.

[0133] (1) The operational risks of the power Internet of Things are essentially determined by the risks in the information, physical, and social spaces. Therefore, risk prediction should comprehensively consider the measurement data of the information, physical, and social spaces. Most current risk prediction methods only consider the measurement data of the information and physical spaces, while ignoring the measurement data of the social space.

[0134] (2) However, the proportion of risk samples in the power Internet of Things operation data is very small, which will cause data imbalance, making the trained classifier more biased towards the majority class, resulting in a decrease in the performance of the classifier, bringing challenges to the accuracy of subsequent model training, and causing the model to produce false positives when predicting risks. Therefore, it is necessary to effectively process multi-source data from information, physics, and society before model training. The traditional SMOTE algorithm randomly selects a minority class sample as a main sample, randomly selects one from its K nearest neighbor minority class samples, and uses the convex combination of the two as a synthetic sample. It does not simply copy the minority class, which reduces the impact of overfitting. However, it is sensitive to noise samples. When the main sample is a noise sample, the newly synthesized sample may also be a noise sample.

[0135] (3) Traditional risk analysis in the power Internet of Things mainly focuses on post-event control, that is, quickly and accurately solving problems that have occurred to reduce actual losses, such as risk assessment and fault location. The emergency response method is too passive and is not conducive to maintaining the safe operation of the power Internet of Things. Accurately and timely predicting the risk categories faced by the power Internet of Things during operation can help power grid personnel to isolate risks and troubleshoot faults in a timely manner. Therefore, in model design, the main focus should be on the accuracy of the model's risk prediction and the performance of risk prediction in unknown data. The traditional Catboost model can improve classification performance by merging multiple classifiers, but the model performance will be affected by key parameters, and manual parameter adjustment has a certain degree of blindness, which is easy to lose the optimal solution of parameters and takes too long, which will affect the accuracy of the risk prediction model.

[0136] Explanation of terms:

[0137] Power Internet of Things: Power Internet of Things is the application of Internet of Things in smart grids. It is the result of the development of information and communication technology to a certain stage. It will effectively integrate communication infrastructure resources and power system infrastructure resources, improve the level of informatization of the power system, improve the utilization efficiency of the existing infrastructure of the power system, and provide important technical support for the power grid's generation, transmission, transformation, distribution, and consumption.

[0138] Power cyber-physical system (CPS): The information side and the physical side of the power system are increasingly interactively coupled. A large number of electrical equipment, data acquisition devices, and computing terminals are interconnected through two physical networks, the power grid and the communication network, gradually forming a power cyber-physical system that integrates the computing system, communication network, and physical environment, namely the power CPS.

[0139] In the above embodiments, although the steps are numbered S1, S2, etc., these are only specific embodiments given in this application. Those skilled in the art can adjust the execution order of S1, S2, etc. according to actual conditions, which is also within the scope of protection of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0140] like Figure 14 As shown, an operation risk prediction system 200 of the power Internet of Things according to an embodiment of the present invention includes a fusion module 210, a balancing module 220, a training module 230 and a prediction module 240;

[0141] The fusion module 210 is used to fuse multi-source data within a preset historical time period based on time series to obtain a complete data set, where the multi-source data includes measurement data from the information side, the physical side, and the social side of the power Internet of Things.

[0142] The balancing module 220 is configured to: when the data in the complete data set is unbalanced, perform data balancing processing on the complete data set based on an adaptive comprehensive oversampling method to obtain a balanced data set;

[0143] The training module 230 is used to: obtain a power Internet of Things operation risk prediction model based on the balanced data set training;

[0144] The prediction module 240 is used to obtain an operation risk prediction result of the power Internet of Things to be tested based on the current multi-source data of the power Internet of Things to be tested and the power Internet of Things operation risk prediction model.

[0145] On the one hand, by introducing measurement data from the information side, physical side, and social side that affect the security of the power Internet of Things, data fusion is performed based on time series, and a complete data set that integrates the measurement data from the information side, the physical side, and the social side is constructed based on random matrix theory. On the other hand, based on the adaptive synthetic oversampling (ADASYN) method, data balancing is performed on the fused data, which can generate pseudo samples that are highly similar to real samples, assist in constructing a balanced data set, and overcome the disadvantages of low training accuracy caused by the low number of samples in certain categories, resulting in unstable performance of the risk prediction model. Based on the above two aspects, the prediction accuracy of the trained power Internet of Things operation risk prediction model can be improved, and the accuracy of the operation risk prediction results can be improved.

[0146] Optionally, in the above technical solution, the training module 230 is specifically used to:

[0147] A Catboost ensemble learning model is constructed using a symmetric decision tree as the base classifier, and trained on a balanced dataset to obtain a Catboost ensemble classifier.

[0148] The Bayesian optimization method is used to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier;

[0149] All optimal parameters are passed to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model.

[0150] The traditional Catboost model can improve classification performance by combining multiple classifiers, but the model performance will be affected by key parameters, and manual parameter adjustment has a certain degree of blindness, which is easy to lose the optimal solution of parameters, and takes too long, which will affect the accuracy of the risk prediction model. In this application, the modeling process includes two model training and learning stages. In the first stage, a Catboost ensemble learning model is constructed with a symmetric decision tree as the base classifier, and a Catboost ensemble classifier is obtained by training; in the second stage, the Bayesian optimization algorithm is introduced to optimize the parameters of the Catboost model, so that the obtained power Internet of Things operation risk prediction model has higher prediction accuracy.

[0151] Optionally, in the above technical solution, the fusion module 210 is specifically configured to:

[0152] Generate the original dataset Dataset based on multi-source data within the preset historical time period. Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1 ,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D s Represents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iNRepresents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iN Represents: N measurement data collected from the physical side of the Power Internet of Things at the i-th moment in the preset historical time period, where i, n, and N are all positive integers;

[0153] Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, we construct the complete dataset D:

[0154] The above-mentioned parameters and steps for each unit module to implement corresponding functions in the operation risk prediction system 200 of an electric power Internet of Things of the present invention can refer to the parameters and steps in the embodiment of the operation risk prediction method of an electric power Internet of Things above, and will not be repeated here.

[0155] A storage medium in an embodiment of the present invention stores instructions. When a computer reads the instructions, the computer executes any one of the above-mentioned methods for predicting the operation risks of the power Internet of Things.

[0156] An electronic device according to an embodiment of the present invention includes a processor and the above-mentioned storage medium. The processor executes instructions in the storage medium. The electronic device can be a computer, a mobile phone, etc.

[0157] Those skilled in the art will appreciate that the present invention may be implemented as a system, method or computer program product.

[0158] Therefore, the present disclosure may be embodied in the following forms: entirely in hardware, entirely in software (including firmware, resident software, microcode, etc.), or in a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present disclosure may be embodied in the form of a computer program product embodied in one or more computer-readable media, wherein the computer-readable media contains computer-readable program code.

[0159] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device.

[0160] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for predicting operational risks of an electric power Internet of Things, characterized in that: include: Based on the time series, multi-source data within a preset historical time period is integrated to obtain a complete data set. The multi-source data includes: measurement data from the information side, physical side, and social side of the power Internet of Things; When the data in the complete data set is unbalanced, performing data balancing processing on the complete data set based on an adaptive comprehensive oversampling method to obtain a balanced data set specifically includes: S210, calculate the number of samples G to be synthesized, G=(m l -m s )*b, where b∈[0,1], the complete dataset includes the risk dataset and the normal operation dataset, m s Represents: the number of all samples in the risk dataset, m l Indicates: the number of all samples in the dataset that runs normally; S211. For each risk sample’s K nearest neighbor samples, calculate: Where Δ is the number of samples in the K nearest neighbors that belong to the normal operation data set, Z is the normalization factor to ensure that r forms a distribution, and Δ and K are both integers; S212, through the formula: g j =r j ×G, calculate the number of synthetic samples required for each risk sample, r j Indicates r, g corresponding to the j-th risk sample j Indicates the number of samples to be synthesized corresponding to the j-th risk sample, where j is an integer; S213, synthesize the synthetic sample corresponding to the j-th risk sample; A power Internet of Things operation risk prediction model is obtained based on the balanced data set training; Obtaining an operation risk prediction result of the power Internet of Things to be tested based on current multi-source data of the power Internet of Things to be tested and the power Internet of Things operation risk prediction model; The power Internet of Things operation risk prediction model obtained by training based on the balance data set includes: A Catboost ensemble learning model is constructed using a symmetric decision tree as a base classifier, and is trained based on the balanced dataset to obtain a Catboost ensemble classifier; Using the Bayesian optimization method to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier; All optimal parameters are passed to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model.

2. The method for predicting the operation risk of the electric power Internet of Things according to claim 1 is characterized in that: The above method uses time series as a benchmark to fuse multi-source data within a preset historical time period to obtain a complete data set, including: Generate an original dataset Dataset based on the multi-source data within the preset historical time period, Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1 ,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D s Represents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iN Represents: N measurement data collected from the social side of the Power Internet of Things at the i-th moment in the preset historical time period, where i, n, and N are all positive integers; Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, the complete dataset D is constructed:

3. An operation risk prediction system for the power Internet of Things, characterized by: Includes fusion module, balancing module, training module and prediction module; The fusion module is used to fuse multi-source data within a preset historical time period based on time series to obtain a complete data set, wherein the multi-source data includes measurement data on the information side, measurement data on the physical side, and measurement data on the social side of the power Internet of Things; The balancing module is configured to: when the data in the complete data set is unbalanced, perform data balancing processing on the complete data set based on an adaptive comprehensive oversampling method to obtain a balanced data set, specifically comprising: S210, calculate the number of samples G to be synthesized, G=(m l -m s )*b, where b∈[0,1], the complete dataset includes the risk dataset and the normal operation dataset, m s Represents: the number of all samples in the risk dataset, m l Indicates: the number of all samples in the dataset that runs normally; S211. For each risk sample’s K nearest neighbor samples, calculate: Where Δ is the number of samples in the K nearest neighbors that belong to the normal operation data set, Z is the normalization factor to ensure that r forms a distribution, and Δ and K are both integers; S212, through the formula: g j =r j ×G, calculate the number of synthetic samples required for each risk sample, r j Indicates r, g corresponding to the j-th risk sample j Indicates the number of samples to be synthesized corresponding to the j-th risk sample, where j is an integer; S213, synthesize the synthetic sample corresponding to the j-th risk sample; The training module is used to: obtain a power Internet of Things operation risk prediction model based on the balance data set training; The prediction module is used to obtain an operation risk prediction result of the power Internet of Things to be tested based on the current multi-source data of the power Internet of Things to be tested and the power Internet of Things operation risk prediction model; The training module is specifically used for: A Catboost ensemble learning model is constructed using a symmetric decision tree as a base classifier, and is trained based on the balanced dataset to obtain a Catboost ensemble classifier; Using the Bayesian optimization method to obtain the optimal parameters corresponding to each parameter of the Catboost ensemble classifier; Passing all optimal parameters to the Catboost ensemble classifier to obtain the power Internet of Things operation risk prediction model; The fusion module is specifically used for: Generate an original dataset Dataset based on the multi-source data within the preset historical time period, Among them, x i =(x i1 ,x i2 ,...x iN ) T ,y i =(y i1 ,y i2 ,...y iN ) T z i =(z i1 ,z i2 ,...z iN ) T , D c Indicates: the measurement data of the information side of the power Internet of Things within the preset historical time period, D p Indicates: the measurement data of the physical side of the power Internet of Things within a preset historical period, D s Represents the measurement data of the social side of the power Internet of Things within a preset historical period, x i1 ,x i2 ,...x iN Represents: N measurement data collected from the information side of the power Internet of Things at the i-th moment in the preset historical time period, y i1 ,y i2 ,...y iN Represents: N measurement data collected from the physical side of the power Internet of Things at the i-th moment in the preset historical time period, z i1 ,z i2 ,...z iN Represents: N measurement data collected from the social side of the Power Internet of Things at the i-th moment in the preset historical time period, where i, n, and N are all positive integers; Based on the original dataset Dataset, taking the time series as the benchmark and using random matrix theory, the complete dataset D is constructed:

4. A storage medium, characterized in that The storage medium stores instructions, and when a computer reads the instructions, the computer executes an operation risk prediction method for an electric power Internet of Things according to any one of claims 1 to 2.

5. An electronic device, characterized in that: The device comprises a processor and the storage medium according to claim 4, wherein the processor executes instructions in the storage medium.

Citation Information

Patent Citations

  • State grid mechanical external damage prediction method

    CN111310785A