Wideband Oscillation Classification Method Based on Deep Reinforcement Learning and Pattern Mining

By applying deep reinforcement learning and mode mining methods in power systems, the oscillation classification model is trained, and the problems of low efficiency and poor accuracy of broadband oscillation classification in the existing technology are solved, and more efficient and accurate oscillation classification is achieved.

CN119513670BActive Publication Date: 2025-05-30HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510072918.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-30
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing control science analysis and ordinary artificial intelligence algorithms are difficult to accurately capture the internal laws of broadband oscillation of the power system, resulting in low classification efficiency and poor accuracy, affecting equipment safety and power consumption quality.

Method used

Using a method based on deep reinforcement learning and pattern mining, the oscillation classification model is trained through the Markov algorithm, and the multi-modal time series is used to input and output the predicted oscillation category, and combining the reinforcement learning module to train the discriminant mode and classification model.

Benefits of technology

It improves the accuracy and efficiency of broadband oscillation classification, can more accurately capture the internal laws of broadband oscillation in the power grid, and provides a more reliable, scalable and generalizable oscillation classification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119513670B_ABST
    Figure CN119513670B_ABST
Patent Text Reader

Abstract

This application relates to the field of power transmission technology, in particular to a broadband oscillation classification method based on deep reinforcement learning and pattern mining. The broadband oscillation classification method based on deep reinforcement learning and pattern mining proposed in this application encodes multivariate time series data such as voltage and current in the power grid line into a univariate clustering sequence, then learns multi-pattern time series using candidate patterns extracted from the clustering sequence, trains discriminant patterns through a reinforcement learning module, and trains a classification model at the same time. This application maps the category of the sample from the multi-pattern time series using a neural network, and updates the multi-pattern time series using a reinforcement learning method, which can take advantage of the high fitting ability of the neural network while ensuring that the classification method has a certain degree of interpretability. At the same time, it can effectively explore the temporal relationship between data existing in the broadband oscillation phenomenon while implementing the classification algorithm; it overcomes the problems of low accuracy and low efficiency of existing oscillation classification methods.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of power transmission, and in particular, to a broadband oscillation classification method based on deep reinforcement learning and pattern mining. Background Art

[0002] In a "dual-high" power system, the interaction between power generation equipment, transmission networks, power loads, etc. will cause instability oscillations in the frequency range from a few hertz to several thousand hertz. Oscillation refers to the dynamic process in which electrical quantities (such as voltage, current, power, etc.) of a power system fluctuate periodically with time due to the influence of its own or external factors, and the interaction between power electronic devices and the power grid, and the oscillation frequency changes within a relatively wide range is called broadband oscillation of the power system. At present, due to the complex internal mechanism of broadband oscillation formation, it is difficult for existing control theory analysis and ordinary artificial intelligence algorithms to efficiently classify it while capturing its internal complex features for further suppression.

[0003] The problem of broadband oscillation in the power system seriously affects equipment safety and power quality, restricts the efficient consumption of new energy, and threatens the safety and stability of the power grid, which has attracted extensive attention from academia and industry. However, at present, people have not yet formed a unified understanding of the problem of broadband oscillation, and the physical mechanism cannot be accurately revealed. How to classify broadband oscillation often can only rely on traditional control theory analysis or ordinary artificial intelligence methods, and it is difficult to accurately capture its internal laws for broadband oscillation phenomena, which brings certain difficulties to subsequent suppression. In view of this, it is necessary to provide a method that can accurately capture the internal laws of broadband oscillation and efficiently classify the broadband oscillation phenomena that occur in the power grid. Summary of the Invention

[0004] In order to overcome the deficiencies of low accuracy and inefficiency in broadband oscillation classification in the above-mentioned prior art, this application proposes a broadband oscillation classification method based on deep reinforcement learning and pattern mining. The broadband oscillation classification of the power grid transmission line based on deep reinforcement learning and pattern mining improves the accuracy and efficiency of broadband oscillation classification.

[0005] A broadband oscillation classification method based on deep reinforcement learning and pattern mining proposed by the present invention first trains an oscillation classification model through the Markov algorithm. The input of the oscillation classification model is a multi-pattern time series obtained by encoding multivariate time series representing the operating conditions of the power system, and the output is the predicted oscillation category.

[0006] The method for encoding multivariate time series data to obtain multi - mode time series is as follows: First, discretize the multivariate time series data collected over consecutive time periods into single - time - point data, then cluster the single - time - point data, and replace each single - time - point data in the multivariate time series data with the corresponding cluster number to form a clustering sequence; Obtain the common subsequences of the clustering sequences corresponding to the multivariate time series data numbers under different oscillation categories, and extract the common subsequences in the clustering sequences corresponding to the multivariate time series data to form multi - mode time series;

[0007] Obtain multivariate time series data as the test object, combine the known clustering results, obtain the clustering sequence of the test object, and extract the common subsequence to form the multi - mode time series of the test object; Input the multi - mode time series of the test object into the oscillation classification model to obtain the predicted value of the oscillation category of the test object.

[0008] Preferably, the training method of the oscillation classification model includes the following steps:

[0009] St1. Obtain multivariate time series data from the historical operation data of the power system, convert it into multi - mode time series as learning samples and store them in the experience pool. The learning samples are labeled with oscillation categories; Construct a basic model; The basic model includes a feature extraction network and an SAC network; The SAC network includes an actor network, a first state evaluation network, a second state evaluation network, a first action evaluation network, and a second action evaluation network;

[0010] The feature extraction network generates a hidden state h and a predicted category y' for the multi - mode time series; The actor network generates an adjustment strategy based on the multi - mode time series and h; After the multi - mode time series executes the adjustment strategy, a new multi - mode time series is formed; The inputs of the first state evaluation network, the second state evaluation network, the first action evaluation network, and the second action evaluation network are all the multi - mode time series, h, and a reward r, and the outputs are the first state evaluation value, the second state evaluation value, the first action evaluation value, and the second action evaluation value respectively;

[0011] The calculation formula for the reward r is:

[0012] ;

[0013] ;

[0014] where m is the number of oscillation categories, f i is the representation of sample s in the i - th oscillation category; is the value of the i - th oscillation in the predicted category obtained by sample s through the feature extraction module; y i is the value of the i - th oscillation in the oscillation category labeled by sample s; ε is a set smoothing parameter;

[0015] St2. Randomly select multiple samples from the experience pool and input them into the basic model. Calculate the loss function based on the predicted category and the true oscillation category, and update the feature extraction network; calculate the reward r, as well as the action loss and state loss; update the first action evaluation network and the second action evaluation network according to the action loss, update the first state evaluation network according to the state loss, and copy the first state evaluation network to the second state evaluation network;

[0016] St3. Calculate the policy loss by combining the updated first action evaluation network, and update the Actor network according to the policy loss;

[0017] St4. Input the sample s into the updated basic model. The Actor network obtains the action policy and the new multi-mode time series, and associate the new multi-mode time series with the original oscillation category and put it into the experience pool as a sample;

[0018] St5. Determine whether the number of updates of the basic model reaches the set number; if not, return to step St2; if so, fix the feature extraction network as the oscillation classification model, and the oscillation classification model predicts the oscillation category according to the input multi-mode time series.

[0019] Preferably, the action loss function is:

[0020] Lq(i') = ∑ s∈B [q(s, a; wi') - U(q)] 2 / |B|, i' = 1 or 2;

[0021] U(q) = r + γv(s')

[0022] Wherein, s represents the multi-mode time series input as a sample by the basic model; B is the training batch, |B| is the training batch size; wi' represents the approximate parameter of the i'-th action evaluation network, a represents the policy action generated by the actor network when the input sample of the basic model is s; q(s, a; wi') represents the action evaluation value output by the i'-th action evaluation network when the input sample of the basic model is s and the action policy is a; U(q) represents the true value estimation of the sample; s' is the new multi-mode time series formed by adjusting s in combination with the action policy a; v(s') is the first state evaluation value output by the first state evaluation network when the input of the basic model is s'; γ is the set reward discount coefficient;

[0023] In step St3, update the first action evaluation network according to the first action loss Lq(1), and update the second action evaluation network according to the second action loss Lq(2).

[0024] Preferably, the state loss function is:

[0025] Lv = ∑ s∈B[q(s,a)-U(v)] 2 / |B|;

[0026] U(v)=E a~π(a|s;θ) (s,a;wi') - αlnπ(a|s;θ)]

[0027] Where s represents the multi - mode time series with the base model as the sample input; B is the training batch, and |B| represents the training batch size; π(a|s;θ) is the set of probabilities output by the actor network; θ is the approximate parameter of the first state - value network; q(s,a) is the probability value corresponding to the action policy a in π(a|s;θ), U(v) is the true value estimate of the sample s; E a~π(a|s;θ) represents the expectation of the distribution where π(a|s;θ) is located; (s,a;wi') represents taking the smaller value between q(s,a;w1) and q(s,a;w2), where q(s,a;w1) and q(s,a;w2) respectively represent the action evaluation values output by the first action evaluation network and the second action evaluation network when the input sample of the base model is s and the action policy is a; α is the entropy parameter in the maximum entropy exploration of the SAC algorithm.

[0028] Preferably, the policy loss function is:

[0029] La = ∑ s∈B E a~π(a|s;θ) [q 0 (s,a) - αlnπ(a|s;θ)] / |B|

[0030] Where q 0 (s,a) represents the first action evaluation value generated by the updated first action evaluation network when the input of the base model is s.

[0031] Preferably, the acquisition of the learning samples in St1 includes the following sub - steps:

[0032] S1. Obtain a set of multivariate time series data labeled with oscillation categories, and discretize the multivariate time series data in the set with time points as units to obtain a set of single - time - point data;

[0033] S2. Cluster the data in the set of single - time - point data to obtain the mapping relationship between each data in the multivariate time series data and the clustering clusters;

[0034] S3. Replace the data in the multivariate time series data with the corresponding cluster numbers to form a clustering sequence;

[0035] ​S4. Extract common subsequences from the clustering sequences under the same oscillation category to form continuous subsequences, and use the continuous subsequences as the pattern candidates for the corresponding categories; obtain the union L of the pattern candidates for all oscillation categories;

[0036] S5. For the clustering sequences of multivariate time series data, replace the continuous cluster number arrays with continuous subsequences, and extract the continuous subsequences to form the multi-pattern time series output of the multivariate time series data.

[0037] Preferably, in S2, the Toeplitz inverse covariance clustering method is used for clustering;

[0038] Step S5 specifically includes the following sub-steps:

[0039] S51. Construct a feature alternative set for the multivariate time series data, and the initial state of the feature alternative set is an empty set;

[0040] S52. Search for continuous numerical combinations in the clustering sequences of the multivariate time series data that form any continuous subsequence in the union L;

[0041] S53. Form continuous subsequences from the continuous numerical combinations and store them in the feature alternative set, and then delete the continuous numerical combinations from the clustering sequences;

[0042] S54. Determine whether there are continuous numerical combinations in the clustering sequences that can form any continuous subsequence in the union L;

[0043] If yes, return to step S53;

[0044] If no, arrange the continuous subsequences in the feature alternative set in order to form the multi-pattern time series output.

[0045] Preferably, the feature extraction network uses an RNN network.

[0046] A broadband oscillation classification system based on deep reinforcement learning and pattern mining proposed by this application includes a memory and a processor. A computer program is stored in the memory, and the processor is connected to the memory. The processor is used to execute the computer program to implement the broadband oscillation classification method based on deep reinforcement learning and pattern mining.

[0047] A storage medium proposed by this application stores a computer program, and the computer program is used to implement the broadband oscillation classification method based on deep reinforcement learning and pattern mining when executed.

[0048] The advantages of this application are as follows:

[0049] (1) The broadband oscillation classification method based on deep reinforcement learning and pattern mining proposed in this application encodes multivariate time series data such as voltage and current in the power grid line into a univariate clustering sequence, then learns a multi-pattern time series from the candidate patterns extracted from the clustering sequence, trains a discriminant pattern through a reinforcement learning module, and trains a classification model at the same time.

[0050] (2) In this application, a clustering method is used to classify the discrete time series data to obtain a univariate clustering sequence of the time series data, extract the multi-pattern time series of each sample from it, map the category of the sample from the multi-pattern time series with a neural network, and use the reinforcement learning method to update the multi-pattern time series, which can take advantage of the high fitting ability of the neural network while ensuring a certain degree of interpretability of the classification method, and can effectively explore the temporal relationship between data in the broadband oscillation phenomenon while implementing the classification algorithm.

[0051] (3) This application uses a method of extracting the multi-pattern time series of discriminant patterns and a framework optimized by reinforcement learning. Therefore, the oscillation classification model proposed in this application has good reliability, scalability, and generalization ability, and can discriminate time series data faster when dealing with similar data or additional tasks.

[0052] (4) This application uses the Soft Actor-Critic method to establish a reinforcement learning framework. This method adopts the idea of offline reinforcement learning and the maximum entropy action selection strategy, which can ensure the stability of network training and the exploration of action selection, make the network training not easily fall into local optimum, and improve the speed and accuracy of network training.

[0053] (5) For the reinforcement learning method adopted in this application, the input of the SAC algorithm when updating the Critic network (i.e., the state evaluation network and the action evaluation network) is the multi-pattern time series of the current sample and the hidden state of its mapping space, which keeps the update of the Critic network synchronized with that of the deep neural network, speeds up the fitting speed, and improves the efficiency during training. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a flowchart of a feature extraction method for multivariate sequence data;

[0055] Figure 2 It is a flowchart of the oscillation classification model training;

[0056] Figure 3It is a diagram of a deep reinforcement learning framework. Among them, M is a multi-mode time series, h is the hidden state extracted by the feature extraction module, S represents the Critic network including the first state evaluation network, the second state evaluation network, the first action evaluation network and the second action evaluation network, r represents the reward, and π θ represents the actor network; is the predicted oscillation category vector, and y is the true oscillation category vector;

[0057] Figure 4 It is a flowchart of a broadband oscillation classification method based on deep reinforcement learning and pattern mining;

[0058] Figure 5 They are the ROC curves of different models. Specific implementation manners

[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0060] Referring to Figure 1 , a feature extraction method for multi-variable sequence data proposed in this implementation manner includes the following steps:

[0061] S1. Obtain a sample set X of multivariate time series data Xn labeled with oscillation categories, and convert X into a multi-variable sequence X'. The essence of X' is a set of single-time-point data;

[0062] X = {X1, X2, …, Xn, …, XN}

[0063] Xn = {x(n,1), x(n,2), …, x(n,j), …, x(n,t n )}

[0064] X' = {x(n,j)|1 ≤ j ≤ t n , 1 ≤ n ≤ N}

[0065] Among them, X1, X2, Xn, and XN respectively represent the 1st, 2nd, nth, and Nth multivariate time series data samples;

[0066] x(n,1), x(n,2), x(n,j), x(n,t n ) respectively represent the acquisition data of the time series data sample Xn at the 1st, 2nd, jth, and t n th time steps; N is the number of samples, and t n is the number of time slots included in Xn; 1 ≤ j ≤ tn , where \(1\leq n\leq N\);

[0067] It should be noted that each sample \(X_n\) may have one or more oscillations. The oscillation categories of the sample can be represented by a vector of length \(m\), where \(m\) is the number of oscillation categories. For example, when \(m = 8\), the oscillation category vector is denoted as \(G\);

[0068] \(G=\{G_1;G_2;G_3;G_4;G_5;G_6;G_7;G_8\}\)

[0069] \(G_1\) represents the oscillation caused by poor matching between the generator and the load, \(G_2\) represents the oscillation caused by the dynamic characteristics of power electronic devices, \(G_3\) represents the oscillation caused by the reaction delay of the generator control system, \(G_4\) represents the oscillation caused by load fluctuations, \(G_5\) represents the oscillation caused by the high-frequency response of power electronic devices, \(G_6\) represents the non-linear oscillation caused by the non-linear characteristics of power electronic devices, \(G_7\) represents the oscillation caused by harmonics, and \(G_8\) represents the oscillation caused by system short circuit or system disturbance;

[0070] When the sample has the oscillation \(G_g\), \(G_g\) represents the amplitude of this type of oscillation, with the unit of \(KHz\); when the sample does not have the oscillation \(G_g\), \(G_g\) represents \(0\).

[0071] For example, assume that the oscillation category vector of sample \(X_1\) is denoted as \((10;5;0;0;0;0;0;0)\), which means that the amplitude of the oscillation \(G_1\) caused by poor matching between the generator and the load in \(X_1\) is \(10KHz\), the amplitude of the oscillation \(G_2\) caused by the dynamic characteristics of power electronic devices is \(5\), and \(X_1\) does not have other types of oscillations.

[0072] S2. Cluster the data \(x(n,j)\) in the multivariate sequence \(X'\) to obtain \(C\) clusters, and obtain the mapping relationship between the data \(x(n,j)\) and the clusters; define \(K(n,j)\) as the cluster to which the data \(x(n,j)\) belongs, that is, \(K(n,j)\in K\);

[0073] \(K = \{K_1,K_2,\cdots,K_c,\cdots,K_C\}\)

[0074] where \(K\) represents the cluster set, \(K_1\), \(K_2\), \(K_c\), \(K_C\) respectively represent the 1st, 2nd, \(c\)th, and \(C\)th clusters obtained by clustering, and \(C\) is the total number of clusters.

[0075] Specifically, the Toeplitz inverse covariance clustering method can be used to cluster the data \(x(n,j)\) in \(X'\) to obtain the cluster set \(K\).

[0076] S3. Obtain the clustering sequence corresponding to \(X_n\), denoted as \(U(n)\);

[0077] \(U(n)=\{K(n,1),K(n,2),\cdots,K(n,j),\cdots,K(n,t n )\}\)

[0078] S4. For the clustering sequence U(n) corresponding to the sample Xn under the same oscillation category, extract the common subsequence to form a continuous subsequence, and use the continuous subsequence as the pattern candidate for the corresponding category; obtain the union of the pattern candidates for all oscillation categories and denote it as L;

[0079] L = {L1, L2, …, Lp, …, LP}

[0080] Among them, P is the number of continuous subsequences in the union L, and L1, L2, Lp, LP are the 1st, 2nd, pth, and Pth continuous subsequences in the union L respectively; obviously, each continuous subsequence Lp is the set of the numbers of at least two clusters in the cluster set K;

[0081] S5. For U(n) = {K(n, 1), K(n, 2), …, K(n, j), …, K(n, t n )}, replace the continuous cluster number array that constitutes the continuous subsequence in the union L with the continuous subsequence, extract the continuous subsequence to form the multi-pattern time series of Xn and output it.

[0082] Specifically, step S5 specifically includes the following sub-steps: S51. Construct the feature candidate set of Xn and clear it;

[0083] S52. Search in U(n) for the continuous numerical combination that can form any continuous subsequence in the union L;

[0084] S53. Store the continuous subsequence formed by the continuous numerical combination into the feature candidate set, and then delete the continuous numerical combination from U(n);

[0085] Specifically in implementation, in step S53, always search for the continuous numerical combination that can form a continuous subsequence from the first numerical value.

[0086] S54. Determine whether there is a continuous numerical combination in U(n) that can form any continuous subsequence in the union L;

[0087] If yes, return to step S53;

[0088] If no, arrange the continuous subsequences in the feature candidate set in order to form a multi-pattern time series and output it.

[0089] Referring to Figure 2 , Figure 3 , this application also proposes a method for training an oscillation classification model, including the following steps:

[0090] St1. Using the above steps S1 - S5, convert the obtained multi - variable time series data Xn into a multi - mode time series and store it in the experience pool as a learning sample. The learning sample is labeled with oscillation categories; construct a basic model; the basic model includes a feature extraction network and an SAC network; the SAC network includes an actor network, a first state evaluation network, a second state evaluation network, a first action evaluation network, and a second action evaluation network;

[0091] The feature extraction network generates a hidden state h and a predicted category y' for the multi - mode time series; the actor network generates an adjustment strategy based on the multi - mode time series and h; after the multi - mode time series executes the adjustment strategy, a new multi - mode time series is formed; the first state evaluation network generates a first state evaluation value based on the multi - mode time series, h, and a reward r; the second state evaluation network generates a second state evaluation value based on the multi - mode time series, h, and a reward r; the first action evaluation network generates a first action evaluation value based on the multi - mode time series, h, and a reward r; the second action evaluation network generates a second action evaluation value based on the multi - mode time series, h, and a reward r;

[0092] St2. Randomly select multiple samples from the experience pool and input them into the basic model. Calculate the loss function according to the predicted category and the true oscillation category and update the feature extraction network. Calculate the reward r, the action losses Lq(1), Lq(2), and the state loss Lv. Update the first action evaluation network according to Lq(1), update the second action evaluation network according to Lq(2), update the first state evaluation network according to Lv, and copy the first state evaluation network to the second state evaluation network;

[0093] ;

[0094] ;

[0095] where m is the number of oscillation categories, f i is the representation of sample s in the i - th oscillation category; is the value of the i - th oscillation in the predicted category obtained by sample s through the feature extraction module; y i is the value of the i - th oscillation in the oscillation category labeled for sample s; ε is a smoothing parameter;

[0096] Lq(1)=∑ s∈B [q(s,a;w1)-U(q)] 2 / |B|

[0097] Lq(2)=∑ s∈B [q(s,a;w2)-U(q)] 2 / |B|

[0098] Among them, s represents the multi-modal time series with the base model as the sample input; B is the training batch, that is, a batch of samples are input into the base model; |B| is the training batch size, that is, the number of samples in a batch; w1 represents the approximate parameter of the first action evaluation network, a represents the policy action generated by the actor network when the sample input to the base model is s; q(s, a; w1) represents the first action evaluation value output by the first action evaluation network when the sample input to the base model is s and the action policy is a; w2 represents the approximate parameter of the second action evaluation network, and q(s, a; w2) represents the second action evaluation value output by the second action evaluation network when the sample input to the base model is s and the action policy is a; U(q) represents the true value estimate of the sample;

[0099] U(q)=r + γv(s')

[0100] s' is the new multi-modal time series formed after s is adjusted by combining the action policy a when the input of the base model is s; v(s') is the first state evaluation value output by the first state evaluation network when the input of the base model is s'; γ is the set reward discount coefficient, which can be specifically set to 0.05; r is the reward function corresponding to s;

[0101] Lv = ∑ s∈B [q(s, a) - U(v)] 2 / |B|

[0102] U(v)=E a~π(a|s;θ) (s, a; wi') - αlnπ(a|s; θ)]

[0103] Among them, π(a|s; θ) is the set of probabilities output by the actor network; θ is the approximate parameter of the first state value network; q(s, a) is the probability value corresponding to the action policy a in π(a|s; θ), and U(v) is the true value estimate of the sample s; E a~π(a|s;θ) represents the expectation of the distribution where π(a|s; θ) is located; (s, a; wi') represents taking the smaller value between q(s, a; w1) and q(s, a; w2), where q(s, a; w1) and q(s, a; w2) respectively represent the action evaluation values output by the first action evaluation network and the second action evaluation network when the sample input to the base model is s and the action policy is a; α is the entropy parameter in the maximum entropy exploration of the SAC algorithm, and its size represents the exploration enthusiasm for non-optimal paths; ln represents the natural logarithm;

[0104] St3, calculate the policy loss La by combining the updated first action evaluation network, and update the Actor network according to La;

[0105] La = ∑ s∈B E​a~π(a|s;θ) [q 0 (s,a) - αlnπ(a|s;θ)] / |B|

[0106] where q 0 (s,a) represents the first action evaluation value generated by the updated first action evaluation network when the input of the base model is s;

[0107] St4. Input s into the updated base model, the Actor network obtains the action policy and the new multi-modal time series, and associates the new multi-modal time series with the original oscillation category as a sample and puts it into the experience pool;

[0108] St5. Determine whether the number of updates of the base model reaches the set number; if not, return to step St2; if yes, fix the feature extraction network as the oscillation classification model, and the oscillation classification model predicts the oscillation category according to the input multi-modal time series.

[0109] Referring to Figure 4 , a broadband oscillation classification method based on deep reinforcement learning and pattern mining proposed in this application includes the following steps:

[0110] SA1. Collect the power system operating condition data, denoted as X0 = {x(0,1), x(0,2), …, x(0,z), …, x(0,t 0 )}; x(0,1), x(0,2), x(0,z), x(0,t 0 ) respectively represent the collected data at the 1st, 2nd, zth, and tth 0 time steps in X0, where 1 ≤ z ≤ t 0 ;

[0111] SA2. Obtain the clustering results obtained in step S2 of the above feature extraction method for multi-variable sequence data, and then classify x(0,z) (1 ≤ z ≤ t 0 ) to obtain the cluster number to which each collected data x(0,z) belongs, and construct a clustering sequence; then execute steps S3 - S5 to obtain the multi-modal time series corresponding to X0 as the observation data;

[0112] Specifically, the method for classifying x(0,z) is: calculate the distance between x(0,z) and the centroids of each cluster in S2, and select the cluster corresponding to the minimum distance as the cluster to which x(0,z) belongs.

[0113] SA3. Input the observation data into the oscillation classification model to obtain the predicted category.

[0114] The following combines specific embodiments to elaborate and verify the above broadband oscillation classification method based on deep reinforcement learning and pattern mining and the oscillation classification model.

[0115] In this embodiment, about 15,000 pieces of multivariate time series data Xn are extracted from the historical data of the power system. Each piece of multivariate time series data contains 20,000 time slots, and the time slot length is set to 0.02 seconds. The data elements of the multivariate time series data Xn include three-phase voltage, three-phase current, and power.

[0116] In this embodiment, the oscillation categories are set in combination with the causes of oscillations, specifically including eight categories, namely: oscillations caused by poor matching between generators and loads, oscillations caused by the dynamic characteristics of power electronic devices, oscillations caused by the reaction delay of generator control systems, oscillations caused by load fluctuations, oscillations caused by the high-frequency response of power electronic devices, nonlinear oscillations caused by the nonlinear characteristics of power electronic devices, oscillations caused by harmonics, and oscillations caused by system short circuits or system disturbances.

[0117] During the calculation of this embodiment, the smoothing parameter ε = 0.1 is set.

[0118] In this embodiment, the extracted multivariate time series data Xn is divided into a training set and a test set.

[0119] When implementing the broadband oscillation classification method based on deep reinforcement learning and pattern mining proposed in this application, the oscillation classification model proposed in this application is constructed by using the above steps St1 - St5. The feature extraction network uses an RNN network, and the data in the training set is processed into training samples by using the above steps S1 - S5 for training the oscillation classification model.

[0120] After the oscillation classification model converges, it is used to classify the oscillations in the test set to verify the model performance. First, the data in the test set is clustered in combination with the clustering clusters of the training samples, and a multi-pattern time series of each data in the test set is constructed in combination with the clustering results as test samples, and then the test samples are input into the trained oscillation classification model for oscillation category prediction.

[0121] To verify the effectiveness of the oscillation classification model proposed in this application, the oscillation classification model UMR proposed in this application is compared with a comparison model through simulation. The comparison models include the traditional convolutional neural network RNN, the classification model Transformer with a multi-head attention mechanism, and the global feature algorithm DTWD for simulation comparison. The input of the comparison model is the multivariate time series data Xn, and the output is the oscillation category.

[0122] In this embodiment, the three comparison models are all trained using machine learning algorithms on the training set and then tested on the test set. Table 1 shows the comparison of the accuracies of the four models on the test set. It can be seen that the UMR model leads in accuracy compared with other models, and its leading advantage in AUC (Area Under Curve, the area enclosed by the ROC curve and the coordinate axis) is relatively large.

[0123] Table 1 Experimental Results of Each Algorithm

[0124] ;

[0125] Figure 5 are the ROC curves of four algorithms. Combining Table 1, Figure 5 it can be seen that the UMR model provided by this application not only has better accuracy and precision compared to other deep learning models, but also the mechanism of its algorithm design ensures that it has interpretability lacking in other deep learning models. Because while the UMR model provided by this application completes classification, it refines the temporal relationship in the data when broadband oscillation occurs, which is impossible for general deep learning models to achieve.

[0126] Of course, for those skilled in the art, this application is not limited to the details of the above exemplary embodiments, but also includes the same or similar structures that can be implemented in other specific forms without departing from the spirit or basic characteristics of this application. Therefore, from any perspective, the embodiments should be regarded as exemplary and non-limiting. The scope of this application is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed within this application. Any reference signs in the claims should not be construed as limiting the claimed rights.

[0127] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

[0128] The technologies, shapes, and structures not detailed in this application are all well-known technologies.

Claims

1. A broadband oscillation classification method based on deep reinforcement learning and pattern mining, characterized in that: Firstly, the oscillation classification model is trained by using the Markov algorithm. The input of the oscillation classification model is the multi-mode time series obtained by encoding the multivariate time series data representing the working conditions of the power system, and the output is the predicted oscillation category. The method of encoding multivariate time series data to obtain multimodal time series is as follows: first, the multivariate time series data collected in a continuous time period is discretized into single time point data, and then the single time point data is clustered, and each single time point data in the multivariate time series data is replaced with the corresponding cluster cluster number to form a cluster sequence; the common subsequence of the cluster sequence corresponding to the multivariate time series data number under different oscillation categories is obtained, and the common subsequence in the cluster sequence corresponding to the multivariate time series data is extracted to form a multimodal time series; Obtain multivariate time series data as the test object, combine it with the known clustering results, obtain the clustering sequence of the test object, and extract the common subsequence to form the multi-mode time series of the test object; input the multi-mode time series of the test object into the shock classification model to obtain the shock category prediction value of the test object; The training method of the oscillation classification model includes the following steps: St1. Obtain multivariate time series data from the historical operation data of the power system and convert it into a multi-mode time series as a learning sample and store it in the experience pool. The learning sample is marked with the oscillation category; build a basic model; the basic model includes a feature extraction network and a SAC network; the SAC network includes an actor network, a first state evaluation network, a second state evaluation network, a first action evaluation network and a second action evaluation network; The feature extraction network generates a hidden state h and a predicted category y' for the multimodal time series; the actor network generates an adjustment strategy based on the multimodal time series and h; the multimodal time series forms a new multimodal time series after executing the adjustment strategy; the inputs of the first state evaluation network, the second state evaluation network, the first action evaluation network, and the second action evaluation network are all the multimodal time series, h, and the reward r, and the outputs are the first state evaluation value, the second state evaluation value, the first action evaluation value, and the second action evaluation value, respectively; The calculation formula of reward r is: ; ; Among them, m is the number of shock categories, f i is the representation of sample s in the i-th oscillation category; is the value of the ith oscillation in the predicted category obtained by the feature extraction module for sample s; y i is the value of the ith shock in the shock category marked by sample s; ε is the set smoothing parameter; St2, randomly select multiple samples from the experience pool and input them into the basic model, calculate the loss function and update the feature extraction network according to the predicted category and the real shock category; calculate the reward r as well as the action loss and state loss; update the first action evaluation network and the second action evaluation network according to the action loss, update the first state evaluation network according to the state loss, and copy the first state evaluation network to the second state evaluation network; St3, calculate the strategy loss based on the updated first action evaluation network, and update the Actor network based on the strategy loss; St4, input sample s into the updated basic model, the Actor network obtains the action strategy and the new multi-mode time series, and associates the new multi-mode time series with the original oscillation category as a sample and puts it into the experience pool; St5, determine whether the number of basic model updates reaches the set number; if not, return to step St2; if yes, fix the feature extraction network as the oscillation classification model, and the oscillation classification model predicts the oscillation category according to the input multi-modal time series.

2. The broadband oscillation classification method based on deep reinforcement learning and pattern mining according to claim 1, characterized in that: The action loss function is: Lq(i')=∑ s∈B [q(s,a;wi') - U(q)] 2 / |B|, i' = 1 or 2; U(q)=r+γv(s'); Where s represents the multimodal time series of the basic model as sample input; B is the training batch, |B| is the training batch size; wi' represents the approximate parameters of the i'th action evaluation network, a represents the policy action generated by the actor network when the basic model input sample is s; q(s,a;wi') represents the action evaluation value output by the i'th action evaluation network when the basic model input sample is s and the action strategy is a; U(q) represents the true value estimate of the sample; s' is the new multimodal time series formed by adjusting s in combination with the action strategy a; v(s') is the first state evaluation value output by the first state evaluation network when the basic model input is s'; γ is the set reward discount coefficient; In step St3, the first action evaluation network is updated according to the first action loss Lq(1), and the second action evaluation network is updated according to the second action loss Lq(2).

3. The broadband oscillation classification method based on deep reinforcement learning and pattern mining as claimed in claim 1, characterized in that: The state loss function is: Lv=∑ s∈B [q(s,a)-U(v)] 2 / |B|; U(v)=E a~π(a|s;θ) [ (s,a;wi')-αlnπ(a|s;θ)]; Where s represents the multimodal time series of the basic model as sample input; B is the training batch, |B| is the training batch size; π(a|s;θ) is the probability set of the actor network output; θ is the approximate parameter of the first state value network; q(s,a) is the probability value corresponding to the action strategy a in π(a|s;θ), and U(v) is the true value estimate of sample s; E a~π(a|s;θ) represents the expectation of the distribution of π(a|s;θ); (s,a;wi') means taking the smaller value between q(s,a;w1) and q(s,a;w2), q(s,a;w1) and q(s,a;w2) respectively represent the action evaluation values ​​output by the first action evaluation network and the second action evaluation network when the basic model input sample is s and the action strategy is a; α is the entropy parameter in the maximum entropy exploration in the SAC algorithm.

4. The broadband oscillation classification method based on deep reinforcement learning and pattern mining as claimed in claim 3, characterized in that: The policy loss function is: La=∑ s∈B E a~π(a|s;θ) [q0(s,a)-αlnπ(a|s;θ)] / |B|; Among them, q0(s,a) represents the first action evaluation value generated by the updated first action evaluation network when the basic model input is s.

5. The broadband oscillation classification method based on deep reinforcement learning and pattern mining as claimed in claim 1, characterized in that: The acquisition of learning samples in St1 includes the following steps: S1. Obtain a set of multivariate time series data marked with shock categories, and discretize the multivariate time series data in the set in units of time points to obtain a set of single time point data; S2, clustering the data in the set of single time point data to obtain the mapping relationship between each data in the multivariate time series data and the clustering clusters; S3, replacing the data in the multivariate time series data with the serial numbers of the corresponding clusters to form a cluster sequence; S4, extracting common subsequences from clustered sequences under the same oscillation category to form continuous subsequences, and using the continuous subsequences as pattern candidates for the corresponding category; Get the union L of pattern candidates of all oscillation categories; S5. For the clustering sequence of the multivariate time series data, replace the continuous cluster number array with a continuous subsequence, and extract the continuous subsequence to form a multi-mode time series output of the multivariate time series data.

6. The broadband oscillation classification method based on deep reinforcement learning and pattern mining according to claim 5, characterized in that: In S2, the Toeplitz inverse covariance clustering method is used for clustering; step S5 specifically includes the following sub-steps: S51, constructing a feature candidate set for multivariate time series data, where the feature candidate set is initially an empty set; S52, searching for a continuous numerical combination constituting any continuous subsequence in the union L in the cluster sequence of the multivariate time series data; S53, combining continuous numerical values ​​to form a continuous subsequence and storing it in a feature candidate set, and then deleting the continuous numerical combination from the clustering sequence; S54, judging whether there is a continuous value combination in the cluster sequence that can constitute any continuous subsequence in the union set L; If yes, return to step S53; If not, the continuous subsequences in the feature candidate set are arranged in order to form a multi-modal time series output.

7. The broadband oscillation classification method based on deep reinforcement learning and pattern mining as claimed in claim 1, characterized in that: The feature extraction network uses RNN network.

8. A broadband oscillation classification system based on deep reinforcement learning and pattern mining, characterized in that: It includes a memory and a processor, the memory stores a computer program, the processor is connected to the memory, and the processor is used to execute the computer program to implement the broadband oscillation classification method based on deep reinforcement learning and pattern mining as described in any one of claims 1 to 7.

9. A storage medium, characterized in that: A computer program is stored, and when the computer program is executed, it is used to implement the broadband oscillation classification method based on deep reinforcement learning and pattern mining as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Generation method, device and equipment of robot automation process and storage medium

    CN115953123A

  • Broadband oscillation mode recognition method, device and equipment and storage medium

    CN119226949A