Substation GIS equipment LCC prediction method based on improved Attention-LSTM algorithm
By improving the Attention-LSTM algorithm, the problem of abnormal data and missing data in the life cycle cost calculation of GIS equipment in the substation is solved, more accurate cost prediction is achieved, and the asset cost management level of power grid companies is improved.
Patent Information
- Application Number
- CN202510036829.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to accurately calculate the life cycle cost of GIS equipment in substations, mainly due to large prediction errors caused by abnormal data and missing data.
Using a method based on the improved Attention-LSTM algorithm, an accurate LCC data prediction model is constructed through the introduction of data collection, screening and correction, dimensionless processing and Attention mechanism.
It effectively overcomes the problems of data abnormalities and missing, improves the accuracy of the cost calculation of the GIS equipment in the substation, and thus improves the level of asset expense management of power grid companies.
Smart Images

Figure CN120086756A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power systems, in particular to the technology for calculating the life cycle cost of substation GIS equipment, and specifically relates to a method for predicting the life cycle cost data of substation GIS equipment based on an improved Attention-LSTM algorithm. Background Technique
[0002] With the improvement of the asset cost management level of power grid companies, the life cycle cost (LCC) of substation equipment, which is an important reference index for asset cost management, becomes particularly important. As an important power distribution device in a substation, the accurate calculation of the LCC data of GIS equipment has become an urgent problem to be solved. Considering the difficulty in collecting the early LCC cost data of substation GIS equipment and the short operation time of some GIS equipment, there are problems of missing and insufficient cost data of GIS equipment within the whole life cycle. In addition, due to possible errors in entering cost materials during the data collection process, there are also data anomaly problems in the LCC data. These LCC data problems are transformed into the problem of large prediction errors in LCC data during the LCC calculation of substation GIS equipment, making it difficult to accurately calculate the life cycle cost of substation GIS equipment.
[0003] Therefore, a method is needed to process the LCC data of substation GIS equipment, overcome the influence of LCC data problems on the data calculation results, and achieve accurate calculation of the life cycle cost of substation GIS equipment, thereby effectively improving the asset cost management level of power grid companies. Summary of the Invention
[0004] The purpose of the present invention is to solve the technical problem that the existing technology for obtaining the life cycle data of GIS equipment is prone to large prediction errors in LCC data due to data anomalies, and to propose a method for predicting the LCC of substation GIS equipment based on an improved Attention-LSTM algorithm.
[0005] To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0006] A method for predicting the LCC of substation GIS equipment based on an improved Attention-LSTM algorithm includes the following steps:
[0007] Step S1: Carry out the work of collecting the LCC data of substation GIS equipment, and collect and statistically analyze the life cycle cost data of in-service substation GIS equipment, including: the initial investment cost of GIS equipment, the daily operation, maintenance and fault costs, and the retirement and disposal costs;
[0008] Step S2: Establish a data screening and correction model according to the data characteristics of the whole life cycle cost data of substation GIS equipment, identify abnormal data, and perform data screening and repair;
[0009] Step S3: For the convenience of predicting LCC data, perform dimensionless and uniform processing on the processed LCC data of GIS equipment according to types;
[0010] Step S4: Considering the influence of various data types on the LCC measurement results, based on the LSTM algorithm, introduce the Attention algorithm to improve the prediction algorithm, and construct an accurate prediction model for LCC data based on the improved Attention-LSTM;
[0011] Step S5: Divide the predicted LCC data of substation GIS equipment into a prediction set and a verification set, and substitute the prediction set to start data prediction;
[0012] Step S6: Import the predicted LCC data of substation GIS equipment into Step S2 for validity evaluation. If the requirements are met, the prediction is completed; otherwise, re-prediction is required.
[0013] The whole life cycle cost data of the substation GIS equipment includes the initial investment cost, daily maintenance and repair cost, and retirement and disposal cost of the GIS equipment. The initial investment cost mainly includes construction engineering cost, equipment purchase cost, installation engineering cost, and other initial investment costs; the daily maintenance and repair cost mainly includes operation inspection cost, equipment maintenance cost, maintenance material cost, maintenance service cost, fault repair cost, and other operation and maintenance fault costs; the retirement and disposal cost is mainly calculated at 5% of the initial investment cost.
[0014] The specific form of the data screening and correction model in Step S2 is as follows:
[0015] First, establish a data screening and correction model according to the data characteristics of the collected dataset of substation GIS equipment, judge the validity of all the original LCC data of GIS equipment in the substation, and correct the abnormal data. The specific steps are as follows:
[0016] S2-1: Let the discrete observation value of the whole life cycle cost data of substation GIS equipment be l k (k = 1, 2,..., m), and its corresponding time series be t k (k = 1, 2,...., m). Starting from the k-th data of the discrete observation value, use the multi-level interpolation method to screen and repair abnormal data;
[0017] In the formula, l k is the LCC data of the GIS equipment for the k-th data, and its corresponding k-th time series is t k, the total number of discrete observed values of the LCC data of the GIS device is m;
[0018] S2-2: Use the multi-level difference method to perform the three-level discriminant difference of the original effective fitting data. Select the k-th data and the next 3 consecutive data as l k , l k+1 , l k+2 , …, l k+3 , then its hierarchical difference calculation is as follows formula (1):
[0019] Δ k = l k - 3l k+1 + 3l k+2 - l k+3 (1)
[0020] In the formula, l k , l k+1 , l k+2 , l k+3 are the LCC data of the k-th GIS device, the LCC data of the k+1-th GIS device, the LCC data of the k+2-th GIS device, and the LCC data of the k+3-th GIS device respectively. Δ k is the multi-level difference of the LCC data of these four GIS devices;
[0021] S2-3: Use the residual formula ξ k ~(μ, σ 2 ) to set the threshold to judge whether the difference of the original effective fitting data meets the requirements. μ is the mean of the residuals, σ is the standard deviation of the residuals. According to the 3σ principle of the normal distribution, the threshold ε is set as the following formula (2):
[0022]
[0023] In the formula, m is the total number of discrete observed values of the LCC data of the GIS device, in which y k is the smoothed value of the LCC data of the GIS device, is the average value of the smoothed value y k of the LCC data of the GIS device, and ξ k refers to the residual;
[0024] S2-4: Compare the multi-level difference result with the set threshold ε to judge the validity of the original fitting data. If the multi-level difference result meets the following formula, it is determined that the original fitting data is valid. The discrimination is as follows formula (3):
[0025] Δ k <ε (3)
[0026] where ε is a set threshold value used to judge the validity of the original fitting data, Δ k refers to the multi-layer difference result of the LCC data;
[0027] S2-5: If it is determined that the original fitting data is valid, a promotion model for the whole life cycle cost data of the substation GIS equipment is constructed with this data; if it is determined that the original fitting data is invalid, multiple consecutive data are reselected for data validity judgment until multiple consecutive valid data are found;
[0028] S2-6: Using the characteristic point T i and the characteristic basis function J i,p (a) Establish the total equation of the characteristic curve. The recursive equation of the characteristic basis function is shown in formula (5), and the total equation of the characteristic curve Q(a) is shown in formula (5):
[0029]
[0030] where T i is the characteristic point of the fitting characteristic curve, J i,P (u) is the characteristic basis function of the fitting characteristic curve, p is the degree of the fitting curve, T i,0 (a) is the initial recurrence basis function of the characteristic basis function of the fitting characteristic curve, a i is the i-th node of the knot vector U = {a 0 , a 1 ,..., a i+P+1 ,}, a i+1 is the (i + 1)-th node of the knot vector U, a i+p is the (i + p)-th node of the knot vector U, a i+p+1 is the (i + p + 1)-th node of the knot vector U, a p+1 is the (p + 1)-th node of the knot vector U, J i,P-1 (a) is the previous characteristic basis function of the characteristic basis function J i,P (a) of the fitting characteristic curve; Q(a) refers to the total equation of the characteristic curve, P i refers to the characteristic point of the control curve, J i,p (a) refers to the characteristic basis function, and m is the total number of discrete observation values of the LCC data of the GIS equipment;
[0031] S2-7: Using the established total equation of the characteristic curve Q(a) to establish a forward promotion model for the characteristics of the whole life cycle cost data of the substation GIS equipment. The established model is shown in formula (6):
[0032]
[0033] where p a is the degree of fitting for establishing the model, is the predicted value of the LCC data for the k-th GIS device, for the fitting order of p a, The knot vector is a k when the eigenbasis function, and the value of k at this time is determined by the i value of the i-th knot, where i represents the order selected for fitting the eigencurve; y 预测 refers to the number of prediction points of the LCC data;
[0034] S2-8: Perform validity judgment on all data of the whole life cycle cost of the forward substation GIS device according to the established forward promotion feature model, and the judgment method is shown in formula (7):
[0035]
[0036] In the formula is the difference between the data and the standard value, is the j-th l k predicted value of the LCC data of the GIS device. If l k is abnormal, continue to judge the next data until the normal data is judged until, at this time the fitting order L at the minimum a (min);
[0037] In the formula is the judged normal and valid data, and L a (min) is to record the normal and valid data the absolute value of the difference from the k-th original data the minimum forward fitting order;
[0038] S2-9: Use the established total eigencurve equation Q(a) to establish a reverse promotion feature model for the whole life cycle cost data of the reverse substation GIS device, and the established model is as follows in formula (8):
[0039]
[0040] for the fitting order of p b, The knot vector is a k when the eigenbasis function, and the value of k at this time is determined by the i' value of the i'-th knot;
[0041] S2-10: Perform validity judgment on all data of the whole life cycle cost of the reverse substation GIS device according to the established reverse promotion feature model, and the judgment method is shown in formula (7);
[0042]
[0043] S2-11: If l k is abnormal, continue to judge the next data until normal data is judged and record the fitting times L at this moment b (min);
[0044] In the formula is the normal and valid data for judgment, and L b (min) is the minimum reverse fitting times of the absolute value of the difference between the recorded normal and valid data and the k-th original data; ;
[0045] S2-12: Screen and repair the valid data according to the forward and reverse validity judgment data results of the whole life cycle cost of the substation GIS equipment in S2-8 and S2-10;
[0046] (1) The forward and reverse validity judgment results are both abnormal data
[0047] ① If L a (min) ≠ L b (min), determine that this point is abnormal data to be deleted and delete it;
[0048] ② If L a (min) = L b (min) = p, determine that this point is abnormal data that can be repaired, and repair it using formula (9), that is, continue to promote a set of original valid fitting data in the order of a k :
[0049]
[0050] (2) One of the forward or reverse validity judgment results is abnormal data
[0051] Use a time series a k in the reverse time sequence direction of a k" data fitting promotion model to judge the valid data again, as shown in formula (10):
[0052]
[0053] is the characteristic basis function when the fitting times is p a, and the knot vector is a k” ;
[0054] If there exists any layer Δ k that satisfies formula (7), then determine a k as LCC valid data;
[0055] (3) The positive and reverse validity judgment results are both valid data
[0056] If the positive and reverse validity judgment results of the LCC data are both valid, it can be considered that the data is high-quality data meeting the requirements.
[0057] In step S3, the dimensionless and consistency processing steps are as follows:
[0058] S3-1. Perform consistency processing on all LCC valid data for the index type:
[0059] Perform min-type data consistency on all LCC valid data using formula (11):
[0060]
[0061] In the formula: l ij is the min-type LCC valid data in the i-th row and j-th column, and L is the maximum value in the min-type LCC valid data l ij ; is the result of min-type data consistency;
[0062] Perform interval-type data consistency processing on all LCC valid data using formula (12):
[0063]
[0064] In the formula, [a, b] is the optimal interval of the LCC valid data l ij ; is the result after interval-type data consistency. The max(a - m, L - b) in the formula refers to the maximum value of (a - m, L - b);
[0065] S3-2. Perform dimensionless and normalization processing on all LCC valid data for the index data:
[0066] Perform standard 0-1 transformation processing on all LCC valid data using formula (13):
[0067]
[0068] In the formula, m is the minimum value in the min-type LCC valid data l ij ; is the result of dimensionless and normalization processing of the LCC valid data for the index data, and l ij is the min-type LCC valid data in the i-th row and j-th column;
[0069] Perform linear proportional transformation processing on all LCC valid data using formula (14):
[0070]
[0071] Wherein: is the maximum value of l ij ; is the result of linearly scaling the LCC effective data, and l ij is the minimum-type LCC effective data in the i-th row and j-th column;
[0072] Normalize all LCC effective data using formula (15):
[0073]
[0074] Wherein: is the result of normalizing all LCC effective data, and l ij is the minimum-type LCC effective data in the i-th row and j-th column.
[0075] In step S4, the calculation of the LCC data accurate prediction model using the improved Attention-LSTM is as follows:
[0076] S4-1. Divide all the LCC data that have been dimensionless and normalized into a prediction set and a validation set; then use the LSTM prediction model for calculation, and the modeling process of the calculation is as follows:
[0077]
[0078] Wherein, tanh and σ are activation functions, and X t is the data input at time t, and H t is the hidden state at time t. The subscripts i, f, and o of z represent the input gate, forget gate, and output gate respectively, and c t is the candidate memory unit at time t; W f , W i , W c , W o are the weight matrices corresponding to each module; Y t is the predicted value output;
[0079] Considering that the dataset of substation GIS equipment comes from multiple different GIS equipment, the Attention mechanism is used to identify the influence degree of the LCC data of different GIS equipment on the prediction result during the LSTM prediction process, and weight assignment is carried out according to the influence degree. The weight assignment formula of the Attention mechanism is as follows:
[0080]
[0081] Wherein, A is the input data after weight assignment, and X iis the i-th dimensionless and normalized LCC data, i∈[1,n], n is the number of input LCC data, s() is the scoring function for scoring based on correlation;
[0082] Since there are many input LCC time series, in order to reduce the complexity of prediction and facilitate calculation, a clustering algorithm is introduced to screen the multiple input time series to improve the prediction model. The improvements are as follows:
[0083] (1) For x* ij Two different time series x′=[x′ 1 ,x′ 2 ,…,x′ t ] and x″=[x″ 1 ,x″ 2 ,…,x″ t ], calculate the similarity of these two time series:
[0084]
[0085] In the formula, G s-t (x′, x″) is the correlation coefficient, s∈{1,2,…,2t-1} represents a sequence of length 2t-1;
[0086] (2) Find the reference sequence x* with the greatest square similarity to all time series:
[0087]
[0088] In the formula, is the initial reference sequence, arg is the mean function, and max is the maximum function;
[0089] (3) Determination of similarity judgment value D:
[0090] D(x′,x″)=1-max(G s-t (x′,x″)) (20)
[0091] The larger the value of D, the higher the similarity;
[0092] By calculating the sum of squared errors (SSE) under different numbers of clusters, the point with the largest SSE decrease rate is selected as the optimal number of clusters, that is, the number of sequences that remain at the end;
[0093] S4-2, use the improved Attention-LSTM LCC data accurate prediction model established in S4 to predict future data;
[0094] S4-3. Import the data predicted in step S4-2 into the data screening and correction model established in step S2, and use formula (7) in step S2-8 to evaluate the validity of the data;
[0095] If all the predicted data are valid data of the life cycle cost data (LCC) of the substation GIS equipment, it indicates that the prediction model has a good effect;
[0096] If some of the data predictions do not meet the requirements, the threshold ε needs to be adjusted again, the data needs to be processed again, and the model for predicting the life cycle cost data (LCC) of the substation GIS equipment needs to be rebuilt.
[0097] Compared with the prior art, the present invention has the following technical effects:
[0098] 1) The present invention can perform validity screening and repair on all calculation data in advance, and then perform preprocessing such as unifying the types of cost data and making the index data dimensionless and normalized. It solves the problem of large fluctuations in LCC data, performs data preprocessing in advance for realizing accurate calculation of the life cycle cost of substation GIS equipment, and thus effectively improves the level of asset cost management of power grid companies;
[0099] 2) The present invention collects and statistically analyzes the existing life cycle cost (LCC) data of GIS equipment in substations, performs validity screening and repair on all LCC data. The repaired data is processed for unifying the index types, normalizing, making dimensionless, and establishing and correcting a prediction model. Finally, the established model is used to predict the LCC data, solves the problem of large fluctuations in LCC data during LCC calculation, and realizes accurate calculation of the life cycle cost (LCC) of substation GIS equipment;
[0100] 3) The purpose of the present invention is to solve the problem of accurate calculation of the life cycle cost (LCC) data of GIS equipment in substations and improve the level of asset cost management of power grid companies. Aiming at the difficulties in collecting the early LCC cost data of substation GIS equipment, and the fact that the operation time of some GIS equipment is relatively short, there are deficiencies in the collection of basic LCC data within the life cycle, and the defects in the early LCC data collection method, resulting in abnormal situations in some LCC data. This leads to the problem of large fluctuations in LCC data during LCC calculation, making it difficult to accurately calculate the life cycle cost of substation GIS equipment. Therefore, the present invention provides a method to process the LCC data, overcome the influence of data fluctuations during LCC calculation on the data calculation results, realize accurate calculation of the life cycle cost of substation GIS equipment, and thus effectively improve the level of asset cost management of power grid companies. BRIEF DESCRIPTION OF THE DRAWINGS
[0101] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0102] Figure 1 is the flow schematic diagram of the present invention;
[0103] Figure 2 is the flow chart of LCC data screening and filling provided by the present invention;
[0104] Figure 3 is the comparison chart of the predicted value and the true value obtained by implementing the present invention in the example. Specific embodiments
[0105] A method for predicting the LCC of substation GIS equipment based on an improved Attention-LSTM algorithm includes the following steps:
[0106] Step S1: Carry out the work of collecting LCC data of substation GIS equipment, collect and count the whole life cycle cost data of in-service substation GIS equipment, including: the initial investment cost of GIS equipment, the daily operation, maintenance and fault costs, and the retirement and disposal costs;
[0107] Step S2: Establish a data screening and correction model according to the data characteristics of the whole life cycle cost data of substation GIS equipment, identify abnormal data and perform data screening and repair;
[0108] Step S3: For the convenience of predicting LCC data, perform dimensionless and uniform processing on the processed GIS equipment LCC data according to types;
[0109] Step S4: Considering the influence of various data types on the LCC calculation result, based on the LSTM algorithm, introduce the Attention algorithm to improve the prediction algorithm, and construct an accurate prediction model of LCC data based on the improved Attention-LSTM;
[0110] Step S5: Divide the predicted LCC data of substation GIS equipment into a prediction set and a verification set, and substitute the prediction set to start the prediction of data;
[0111] Step S6: Import the predicted LCC data of substation GIS equipment into step S2 for effectiveness evaluation. If the requirements are met, the prediction is completed; otherwise, re-prediction is required.
[0112] The specific form of the data screening and correction model in step S2 is:
[0113] First, establish a data screening and correction model according to the data characteristics of the collected dataset of substation GIS equipment, identify the effectiveness of all the original LCC data of GIS equipment in the substation, and correct the abnormal data. The specific steps are as follows:
[0114] S2-1: Let the discrete observable of the life cycle cost data of the substation GIS equipment be \(l\) k \((k = 1, 2, \cdots, m)\), and its corresponding time series be \(t\) k \((k = 1, 2, \cdots, m)\). Starting from the \(k\)th data of the discrete observable, the multi-level interpolation method is used to screen and repair abnormal data;
[0115] In the formula, \(l\) k is the LCC data of the GIS equipment for the \(k\)th data, and its corresponding \(k\)th time series is \(t\) k , and the total number of discrete observables of the LCC data of the GIS equipment is \(m\);
[0116] S2-2: Use the multi-level difference method to perform three-level discriminant differences on the original effective fitting data. Select the \(k\)th data and the next 3 consecutive data behind it as \(l\) k , \(l\) k+1 , \(l\) k+2 , \(\cdots\), \(l\) k+3 , then its hierarchical difference calculation is as follows formula (1):
[0117] \(\Delta\) k \(= l\) k \(- 3l\) k+1 \(+ 3l\) k+2 \(- l\) k+3 (1)
[0118] In the formula, \(l\) k , \(l\) k+1 , \(l\) k+2 , \(l\) k+3 are the LCC data of the \(k\)th GIS equipment, the LCC data of the \((k + 1)\)th GIS equipment, the LCC data of the \((k + 2)\)th GIS equipment, and the LCC data of the \((k + 3)\)th GIS equipment respectively. \(\Delta\) k is the multi-level difference of the LCC data of these four GIS equipments;
[0119] S2-3: Use the residual formula \(\xi\) k \(\sim (\mu, \sigma\) 2 ) to set the threshold to judge whether the original effective fitting data difference meets the requirements. \(\mu\) is the mean of the residuals, \(\sigma\) is the standard deviation of the residuals. According to the 3\(\sigma\) principle of the normal distribution, the threshold \(\varepsilon\) is set as the following formula (2):
[0120]
[0121] In the formula, \(m\) is the total number of discrete observables of the LCC data of the GIS equipment, In \(y\) k is the smoothed value of the LCC data of the GIS equipment, For the smoothed value y of the LCC data of the GIS device k is the average value, ξ k refers to the residual;
[0122] S2-4: Compare the multi-level difference result with the set threshold ε to determine the validity of the original fitting data. If the multi-level difference result satisfies the following formula, it is determined that the original fitting data is valid, and the discrimination is as shown in formula (3):
[0123] Δ k <ε (3)
[0124] In the formula, ε is the set threshold used to judge the validity of the original fitting data, and Δ k refers to the multi-level difference result of the LCC data;
[0125] S2-5: If it is determined that the original fitting data is valid, use this data to construct a promotion model for the life cycle cost data of the substation GIS device; if it is determined that the original fitting data is invalid, reselect multiple consecutive data for data validity judgment until multiple consecutive valid data are found;
[0126] S2-6: Use the characteristic point T i and the characteristic basis function J i,p (a) to establish the total equation of the characteristic curve. The recursive equation of the characteristic basis function is as shown in formula (5), and the total equation of the characteristic curve Q(a) is as shown in formula (5):
[0127]
[0128] In the formula, T i is the characteristic point of the fitting characteristic curve, J i,P (u) is the characteristic basis function of the fitting characteristic curve, p is the degree of the fitting curve, T i,0 (a) is the initial recurrence basis function of the characteristic basis function of the fitting characteristic curve, a i is the i-th node of the knot vector U = {a 0 ,a 1 ,...,a i+P+1 ,}, a i+1 is the (i + 1)-th node of the knot vector U, a i+p is the (i + p)-th node of the knot vector U, a i+p+1 is the (i + p + 1)-th node of the knot vector U, a p+1 is the (p + 1)-th node of the knot vector U, J i,P-1 (a) is the previous characteristic basis function of the characteristic basis function J i,P (a) of the fitting characteristic curve; Q(a) refers to the total equation of the characteristic curve, P i refers to the characteristic point of the control curve, J i,p(a) refers to the characteristic basis function, and m is the total number of discrete observed values of the LCC data of the GIS device;
[0129] S2-7: Use the established total characteristic curve equation Q(a) to establish a forward promotion model for the characteristics of the whole life cycle cost data of the substation GIS device. The established model is shown in formula (6):
[0130]
[0131] In the formula, p a is the number of fittings for establishing the model, is the predicted value of the LCC data of the kth GIS device, is the characteristic basis function when the number of fittings is p a, and the knot vector is a k At this time, the value of k is determined by the value of i of the ith knot. i represents the number of times of selecting the fitting characteristic curve; y 预测 refers to the number of prediction points of the LCC data;
[0132] S2-8: According to the established forward promotion characteristic model, judge the validity of all the data of the whole life cycle cost of the substation GIS device. The judgment method is shown in formula (7):
[0133]
[0134] In the formula is the difference between the data and the standard value, is the predicted value of the LCC data of the jth l k GIS device. If l k is abnormal, continue to judge the next data until the normal data is judged. At this time the minimum number of fitting times L a (min);
[0135] In the formula is the judged normal and valid data, and L a (min) is the minimum positive fitting number of times for recording the absolute value of the difference between the normal and valid data and the kth original data; ;
[0136] S2-9: Use the established total characteristic curve equation Q(a) to establish a reverse promotion model for the characteristics of the whole life cycle cost data of the substation GIS device. The established model is as follows in formula (8):
[0137]
[0138] is the number of fittings for pb, The characteristic basis function when the knot vector is a k At this time, the value of k is determined by the value of i' of the i'-th knot;
[0139] S2-10: According to the established reverse promotion characteristic model, judge the validity of all data of the whole life cycle cost of the reverse substation GIS equipment, and the judgment method is shown in formula (7);
[0140]
[0141] S2-11: If l k is abnormal, continue to judge the next data until the normal data is judged up to, and record the fitting times L at this moment when it is the smallest b (min);
[0142] In the formula is the judged normal and valid data, and L b (min) is to record the normal and valid data and the absolute value of the difference from the k-th original data of the minimum reverse fitting times;
[0143] S2-12: According to the forward and reverse validity judgment data results of all the whole life cycle costs of the substation GIS equipment in S2-8 and S2-10, screen and repair the valid data;
[0144] (1) The forward and reverse validity judgment results are both abnormal data
[0145] ① If L a (min) ≠ L b (min), determine that this point is an abnormal data to be deleted and delete it;
[0146] ② If L a (min) = L b (min) = p, determine that this point is an abnormal data that can be repaired, and use formula (9) to repair it, that is, continue to promote a set of original valid fitting data in the order of a k :
[0147]
[0148] (2) One of the forward or reverse validity judgment results is abnormal data
[0149] Adopt a time series a k in the reverse time sequence direction of a k" data fitting promotion model to judge the valid data again, as shown in formula (10):
[0150]
[0151] For the fitting order of p a, The knot vector is a k” when the eigenbasis function;
[0152] If there exists any layer Δ k satisfying formula (7), then a is determined k as the valid data of LCC;
[0153] (3) The forward and reverse validity judgment results are both valid data
[0154] The LCC data with both forward and reverse validity judgment results being valid can be considered as high-quality data meeting the requirements.
[0155] In step S3, the dimensionless and unification processing steps are as follows:
[0156] S3-1. Perform index type unification processing on all LCC valid data:
[0157] Perform extremely small data unification on all LCC valid data using formula (11):
[0158]
[0159] In the formula: l ij is the extremely small LCC valid data in the i-th row and j-th column, and L is the maximum value in the extremely small LCC valid data l ij ; is the result of extremely small data unification;
[0160] Perform interval type data unification processing on all LCC valid data using formula (12):
[0161]
[0162] In the formula, [a, b] is the best interval of the LCC valid data l ij ; is the result after interval type data unification, and max(a - m, L - b) in the formula refers to the maximum value of (a - m, L - b);
[0163] S3-2. Perform index data dimensionless and normalization processing on all LCC valid data:
[0164] Perform standard 0-1 transformation processing on all LCC valid data using formula (13):
[0165]
[0166] Where m is the minimum value of the effective data l of the ultra-small LCC ij in it, is the result of dimensionless and normalization processing of the index data for the LCC effective data, l ij is the ultra-small LCC effective data in the i-th row and j-th column;
[0167] Perform linear proportional transformation processing on all LCC effective data using formula (14):
[0168]
[0169] Where: is l ij the maximum value of is the result of linear proportional transformation processing for the LCC effective data, l ij is the ultra-small LCC effective data in the i-th row and j-th column;
[0170] Perform normalization processing on all LCC effective data using formula (15):
[0171]
[0172] Where: is the result of normalization processing for all LCC effective data, l ij is the ultra-small LCC effective data in the i-th row and j-th column.
[0173] In step S4, the calculation steps of the LCC data accurate prediction model using the improved Attention-LSTM are as follows:
[0174] S4-1. Divide all LCC data that have completed dimensionless and normalization processing of the index data into a prediction set and a validation set; then use the LSTM prediction model for calculation, and the modeling process of the calculation is as follows:
[0175]
[0176] Where, tanh, σ are activation functions, X t is the data input at time t, H t is the hidden state at time t, and the subscripts i, f, o of z represent the input gate, forget gate, and output gate respectively, c t is the candidate memory cell at time t; W f 、W i 、W c 、W o are the weight matrices corresponding to each module; Y t is the predicted value output;
[0177] Considering that the data set of the substation GIS equipment comes from multiple different GIS equipment, the Attention mechanism is used to identify the influence degree of the LCC data of different GIS equipment on the prediction result during the LSTM prediction process, and weight assignment is carried out according to the influence degree. The weight assignment formula of the Attention mechanism is as follows:
[0178]
[0179] In the formula, A is the input data after weight assignment, Xi is the i-th dimensionless and normalized LCC data, i ∈ [1, n], n is the number of input LCC data, s() is the scoring function for scoring based on and related to;
[0180] Since the number of input LCC time series is large, in order to reduce the complexity of prediction and facilitate calculation, a clustering algorithm is introduced to screen the input multiple time series to improve the prediction model. The improvement is as follows:
[0181] (1) For two different time series x′ = [x′1, x'2, …, x't] and x" = [x"1, x"2, …, x"t] in x*ij, calculate the similarity of these two time series:
[0182]
[0183] In the formula, G s-t (x′, x″) is the correlation coefficient, s ∈ {1, 2, …, 2t - 1} represents a sequence of length 2t - 1;
[0184] (2) Find the reference sequence x* with the largest squared similarity to all time series:
[0185]
[0186] In the formula, is the initial reference sequence, arg is the mean function, and max is the maximum value function;
[0187] (3) Determination of the similarity judgment value D:
[0188] D(x′, x″) = 1 - max(G s-t (x′, x″)) (20)
[0189] The larger the value of D, the higher the similarity;
[0190] By calculating the sum of squared errors SSE under different clustering numbers, select the point with the largest SSE decline rate as the optimal clustering number, that is, the number of sequences finally retained;
[0191] S4-2. Use the improved Attention-LSTM-based LCC data accurate prediction model established in S4 to predict future data;
[0192] S4-3. Import the data predicted in step S4-2 into the data screening and correction model established in step S2, and use formula (7) in step S2-8 to evaluate the validity of the data;
[0193] If all the predicted data are valid data of the life cycle cost data LCC of the substation GIS equipment, it indicates that the prediction model has a good effect;
[0194] If some of the data predictions do not meet the requirements, the threshold ε needs to be adjusted again, the data needs to be processed again, and the life cycle cost data LCC data of the substation GIS equipment needs to be re-modeled and predicted.
[0195] Example:
[0196] The present invention was tested on the prediction of the life cycle cost of the 110 kV main transformer equipment of a substation in Henan to verify the accuracy of a GIS equipment life cycle cost prediction method based on an improved Attention-LSTM algorithm proposed by the present invention. First, calculate the LCC cost of each year within the life cycle of the main transformer from the obtained data such as the initial investment cost, operation and maintenance cost, and retirement cost of the main transformer. Then, use the data of the previous 30 years as the test set and the data of the next 10 years as the validation set. The data of the test set are as follows:
[0197] Table 1
[0198]
[0199]
[0200] For the data of the test set, use the data screening and correction model proposed by the present invention to identify abnormal data, and perform data screening and repair. After discrimination and repair, the obtained data are as follows:
[0201] Table 2
[0202]
[0203] Substitute the repaired data into the improved Attention-LSTM algorithm proposed by the present invention, and the prediction results of the LCC cost for the next 10 years can be obtained. The actual LCC cost and the predicted cost are shown in Table 2.
[0204] Table 3
[0205]
[0206] Make a line chart of the actual LCC cost data and the predicted LCC data as shown in the appendix. Figure 3 As can be seen from the comparison effect of the true value and the predicted value in the figure, the algorithm proposed by the present invention has a good effect.
[0207] To evaluate the prediction effect of the algorithm proposed by the present invention, the root mean square error (RMSE) and the coefficient of determination (R 2 ) are introduced here, and their calculation methods are as follows:
[0208] (1) Root mean square error (RMSE)
[0209]
[0210] where L i refers to the actual LCC cost in the i-th operation year, refers to the predicted LCC cost in the i-th operation year.
[0211] Substitute the actual LCC cost data and the predicted LCC data into this formula, and the RMSE can be obtained as 709.81. Compared with the magnitude of the original cost, this root mean square error is relatively small, indicating that the prediction effect is good.
[0212] (2) Coefficient of determination (R 2 )
[0213]
[0214] where H 残差和 refers to the sum of squared residuals of the LCC cost data, and H 总平方和 refers to the total sum of squares of the LCC cost data. The coefficient of determination can reflect the goodness of the fitted data, and its value will be between 0 and 1. The closer it is to 1, the better the prediction effect of the model. Substitute the data of the main transformer into this calculation formula, and the calculated coefficient of determination is 0.988, indicating that the prediction of the algorithm proposed by the present invention is relatively accurate.
Claims
1. A method for LCC prediction of substation GIS equipment based on improved Attention-LSTM algorithm, characterized in that: The following steps are involved: Step S1: Carry out LCC data collection of substation GIS equipment, collect and count the full life cycle cost data of substation GIS equipment in operation, including: initial investment cost of GIS equipment, daily operation and maintenance failure cost and decommissioning disposal cost; Step S2: according to the data characteristics of the full life cycle cost data of the substation GIS equipment, a data screening and correction model is established to identify abnormal data and perform data screening and repair; Step S3: To facilitate the prediction of LCC data, dimensionless and uniform processing is performed on the processed GIS equipment LCC data according to type; Step S4: Considering the impact of various data types on the LCC calculation results, based on the LSTM algorithm, the Attention algorithm is introduced to improve the prediction algorithm, and an accurate LCC data prediction model based on the improved Attention-LSTM is constructed; Step S5: Divide the LCC data predicted by the substation GIS equipment into a prediction set and a verification set, and substitute the prediction set to start data prediction; Step S6: Import the LCC data predicted by the substation GIS equipment into step S2 for effectiveness evaluation. If the requirements are met, the prediction is completed, otherwise, re-prediction is required.
2. The method according to claim 1, characterized in that The specific form of the data screening and correction model in step S2 is: First, a data screening and correction model is established based on the data characteristics of the collected substation GIS equipment data set, the validity of all GIS equipment LCC raw data in the substation is judged, and abnormal data is corrected. The specific steps are as follows: S2-1: Assume that the discrete observation of the life cycle cost data of substation GIS equipment is l k (k=1,2,...,m), the corresponding time series is t k (k=1,2,....,m), multi-level interpolation method is used to screen and repair abnormal data starting from the kth data of discrete observations; Where l k is the LCC data of the kth data GIS device, and its corresponding kth time series is t k , the total number of discrete observations of LCC data of GIS equipment is m; S2-2: Use the multi-level difference method to perform the three-level discriminant difference of the original effective fitting data, select the kth data and the following three consecutive data as l k ,l k+1 ,l k+2 ,…,l k+3 , then its level difference is calculated as follows: D k =l k -3l k+1 +3l k+2 -l k+3 (1) Where l k , l k+1 , l k+2 , l k+3 are the LCC data of the kth GIS device, the LCC data of the k+1th GIS device, the LCC data of the k+2th GIS device, and the LCC data of the k+3th GIS device, Δ k It is the multi-layer difference of LCC data of these four GIS devices; S2-3: Using the residual formula ξ k ~(μ,σ 2 ) Set the threshold to judge whether the difference of the original effective fitting data meets the requirements, μ is the mean of the residual, σ is the standard deviation of the residual, and according to the 3σ principle of normal distribution, the threshold ε is set as follows (2): Where m is the total number of discrete observations of LCC data of GIS equipment, middle y k is the smoothed value of LCC data of GIS equipment, is the smoothed value y of LCC data of GIS equipment k The average value of k refers to the residual; S2-4: The multi-level difference result and the set threshold ε are used to judge the validity of the original fitting data. If the multi-level difference result satisfies the following formula, the original fitting data is judged to be valid. The judgment is as follows: D k <e (3) Where ε is the set threshold value, which is used to judge the validity of the original fitting data, Δ k Refers to the multi-layer difference results of LCC data; S2-5: If the original fitting data is determined to be valid, the data promotion model of the life cycle cost of substation GIS equipment is constructed using this data; if the original fitting data is determined to be invalid, multiple continuous data are reselected to determine the data validity until multiple continuous valid data are found; S2-6: Using feature point T i and the characteristic basis function J i,p (a) Establish the total equation of the characteristic curve. The characteristic basis function recursive equation is shown in formula (5). The total equation of the characteristic curve Q(a) is shown in formula (5): Where T i is the characteristic point of the fitting characteristic curve, J i,P (u) is the characteristic basis function of the fitting characteristic curve, p is the degree of the fitting curve, T i,0 (a) is the initial recursive basis function of the characteristic basis function of the fitting characteristic curve, a i The node vector U = {a0, a1, ..., a i+P+1 ,} the i-th node, a i+1 is the i+1th node of the node vector U, a i+p is the i+pth node of the node vector U, a i+p+1 is the i+p+1th node of the node vector U, a p+1 is the p+1th node of the node vector U, J i,P-1 (a) is the characteristic basis function J of the fitting characteristic curve i,P (a) is the previous characteristic basis function; Q(a) refers to the total equation of the characteristic curve, P i Refers to the characteristic points of the control curve, J i,p (a) refers to the characteristic basis function, and m is the total number of discrete observations of the LCC data of the GIS device; S2-7: Use the established characteristic curve total equation Q(a) to establish a forward substation GIS equipment full life cycle cost data characteristic promotion model, and establish the model as shown in formula (6): Where p a is the number of times the model is fitted. is the predicted value of the LCC data of the kth GIS device, The number of fitting times is p a, The node vector is a k The characteristic basis function at this time, the value of k is determined by the value of i at the i-th node, i represents the number of times the fitting characteristic curve is selected; y 预测 Refers to the number of prediction points of LCC data; S2-8: Based on the established forward promotion characteristic model, the validity of all data of the full life cycle cost of the forward substation GIS equipment is judged. The judgment method is shown in formula (7): In the formula is the difference between the data and the standard value, For the jth l k The predicted value of LCC data of GIS equipment, if l k If abnormal, continue to judge the next data until normal data is judged So far, at this time The minimum number of fitting times L a (min); In the formula For normal and effective data, L a (min) is to record normal valid data The absolute value of the difference with the kth original data The minimum number of forward fits; S2-9: Use the established characteristic curve total equation Q(a) to establish a reverse substation GIS equipment full life cycle cost data characteristic promotion model, and establish the model as follows: The number of fitting times is p b, The node vector is a k The characteristic basis function at this time, the value of k is determined by the value of i' of the i'th node; S2-10: Based on the established reverse promotion characteristic model, the validity of all data of the full life cycle cost of the reverse substation GIS equipment is judged. The judgment method is shown in formula (7); S2-11: If l k If abnormal, continue to judge the next data until normal data is judged Until now, and record this moment The minimum number of fitting times L b (min); In the formula For normal and effective data, L b (min) is to record normal valid data The absolute value of the difference with the kth original data The minimum number of inverse fits; S2-12: Screen and repair valid data based on all the forward and reverse validity judgment data results of the full life cycle cost of substation GIS equipment in S2-8 and S2-10; (1) Both the positive and negative validity judgment results are abnormal data ①If L a (min)≠L b (min), determine that the point is abnormal data that needs to be deleted and delete it; ②If L a (min) = L b (min) = p, the point is determined to be repairable abnormal data, and is repaired using formula (9), that is, according to a k Continue to generalize a set of original valid fitting data backward in the order of: (2) One of the positive or negative validity judgment results is abnormal data Use a k A timing sequence in the reverse direction k" The data fitting promotion model determines the effective data again, as shown in formula (10): The number of fitting times is p a, The node vector is a k” The characteristic basis function when ; If there is any layer Δ k If formula (7) is satisfied, then a is determined k It is the LCC valid data; (3) Both the positive and negative validity judgment results are valid data LCC data with both positive and reverse validity judgment results being valid can be considered as high-quality data that meets the requirements.
3. The method according to claim 1, characterized in that In step S3, the dimensionless and uniform processing steps are as follows: S3-1. All LCC valid data shall be processed in accordance with the indicator type: Formula (11) is used to perform minimal data consistency on all LCC valid data: Where: l ij is the effective data of the very small LCC in the i-th row and j-th column, and L is the effective data of the very small LCC l ij The maximum value in This is the result of the unification of very small data; Formula (12) is used to perform interval data consistency processing on all LCC valid data: Where [a, b] is the LCC effective data l ij The best interval of It is the result after the interval data is unified. The max(am,Lb) in the formula refers to the maximum value of (am,Lb); S3-2. All LCC valid data are dimensionless and normalized: Formula (13) is used to perform standard 0-1 transformation on all LCC valid data: Where m is the effective data of the very small LCC ij The minimum value in is the result of dimensionless and normalized processing of LCC effective data. ij The valid data of the very small LCC in the i-th row and j-th column; Formula (14) is used to perform linear proportional transformation on all LCC valid data: Where: for l ij The maximum value of is the result of linear proportional transformation of LCC valid data, l ij The valid data of the very small LCC in the i-th row and j-th column; All LCC valid data are normalized using formula (15): Where: is the result of normalizing all LCC valid data, l ij It is the valid data of the very small LCC in the i-th row and j-th column.
4. The method according to claim 1, characterized in that In step S4, the calculation steps of the LCC data accurate prediction model using the improved Attention-LSTM are as follows: S4-1. Divide all LCC data that have completed dimensionless and normalized indicator data into a prediction set and a validation set; then use the LSTM prediction model for calculation. The calculation modeling process is as follows: In the formula, tanh, σ is the activation function, X t is the data input at time t, H t is the hidden state at time t. The subscripts i, f, and o of z represent the input gate, forget gate, and output gate, respectively. t is the candidate memory unit at time t; W f , W i , W c , W o is the weight matrix corresponding to each module; Y t Output is the predicted value; Considering that the data set of substation GIS equipment comes from multiple different GIS devices, the Attention mechanism is used to identify the influence of LCC data of different GIS devices on the prediction results during the LSTM prediction process, and weights are assigned according to the influence. The weight assignment formula of the Attention mechanism is as follows: Where A is the input data after weight assignment, Xi is the i-th dimensionless and normalized LCC data, i∈[1,n], n is the number of input LCC data, and s() is the scoring function based on and related scoring; Since there are many input LCC time series, in order to reduce the complexity of prediction and facilitate calculation, a clustering algorithm is introduced to screen the multiple input time series to improve the prediction model. The improvements are as follows: (1) For x* ij Two different time series x'=[x′1,x′2,…,x′ t ] and x″=[x″1,x″2,…,x″ t ], calculate the similarity of these two time series: In the formula, G s-t (x′, x″) is the correlation coefficient, s∈{1,2,…,2t-1} represents a sequence of length 2t-1; (2) Find the reference sequence x* with the greatest square similarity to all time series: In the formula, is the initial reference sequence, arg is the mean function, and max is the maximum function; (3) Determination of similarity judgment value D: D(x′,x″)=1-max(G s-t (x′,x″)) (20) The larger the value of D, the higher the similarity; By calculating the sum of squared errors (SSE) under different numbers of clusters, the point with the largest SSE decrease rate is selected as the optimal number of clusters, that is, the number of sequences remaining at the end; S4-2, use the improved Attention-LSTM LCC data accurate prediction model established in S4 to predict future data; S4-3, importing the data predicted in step S4-2 into the data screening and correction model established in step S2, and using formula (7) in step S2-8 to evaluate the validity of the data; If all the prediction data are LCC valid data of the life cycle cost of substation GIS equipment, it means that the prediction model is effective; If some data predictions do not meet the requirements, it is necessary to readjust the threshold ε, reprocess the data, and remodel and predict the LCC data of the life cycle cost of substation GIS equipment.
Citation Information
Cited By
Energy data synchronous processing method, system, equipment and medium
CN121327326A
An energy data synchronization processing method, system, device and medium
CN121327326B