A method and system for predicting transmission line loss based on multi-dimensional influencing factors
Through a multi-dimensional influence method, the gradient lifting tree model is used to predict the line loss of transmission lines, which solves the problem of low line loss prediction accuracy in the existing technology and achieves more accurate line loss management.
Patent Information
- Application Number
- CN201911213450.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-12-02
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2039-12-02
AI Technical Summary
The line loss prediction accuracy of existing transmission lines is not high, making it difficult to achieve accurate concurrent line loss calculations, affecting the accuracy of line loss management.
Using a multi-dimensional influence amount based method, the multi-dimensional influence amount data of the transmission line is collected, data preprocessed, the training set, verification set and test set are divided, the linear loss is predicted using the gradient lift tree model, and the weight of the multi-dimensional influence amount is analyzed.
The accuracy of line loss prediction of 500kV overhead transmission lines has been improved, the problem of low line loss prediction accuracy has been solved, and more accurate line loss management reference has been provided.
Smart Images

Figure CN110956330B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electrical energy metering in electrical engineering, and more specifically, to a method and system for predicting transmission line loss based on multi-dimensional influencing factors. Background Art
[0002] Line loss reflects the planning, production, and management levels of the power grid, and is an important criterion for evaluating power departments. However, the error in theoretical line loss calculation will cause the report to not accurately reflect the actual line loss situation, bringing great obstacles to line loss management. With the advancement of refined line loss management work, there is an urgent need for an accurate synchronous line loss calculation method.
[0003] Therefore, a technology is needed to realize the technology of predicting transmission line loss based on multi-dimensional influencing factors. Summary of the Invention
[0004] The technical solution of the present invention provides a method and system for predicting transmission line loss based on multi-dimensional influencing factors to solve the problem of low accuracy in predicting existing transmission line loss.
[0005] To solve the above problems, the present invention provides a method for predicting transmission line loss based on multi-dimensional influencing factors, and the method includes:
[0006] Collect multi-dimensional influencing factors of transmission line loss;
[0007] Perform data preprocessing on the non-temporal data in the line body data and the pole tower data at all levels of the transmission line, so that the non-temporal data forms fixed parameters of the transmission line, and incorporate the fixed parameters into the temporal data table of the transmission line to generate a multi-dimensional influencing factor data table;
[0008] Divide the data in the multi-dimensional influencing factor data table into training set data, validation set data, and test set data; train the gradient boosting tree model through the training set data and output a training result;
[0009] Verify the training result through the validation set data. When the training result passes the verification, generate a trained gradient boosting tree model;
[0010] Predict the test set data through the trained gradient boosting tree model and output the line loss prediction result of the transmission line.
[0011] Preferably, the step of training the gradient boosting tree model through the training set data and outputting a training result further includes:
[0012] Analyze the weights of the multi-dimensional influencing factors through the gradient boosting tree model.
[0013] Preferably, the analysis of the weights of the multi-dimensional influencing quantities by the gradient boosting tree model further includes:
[0014] Training the gradient boosting tree model with the training set data to obtain a trained gradient boosting tree model;
[0015] Extracting all split node data n of the trained gradient boosting tree model;
[0016] Calculating the number of split nodes nk generated by each dimension of the k-dimensional influencing quantity k ;
[0017] Calculating the weight Q of the k-dimensional influencing quantity according to all the split node data n and the number of split nodes nk generated by each dimension of the influencing quantity k , Q k = n k / n × 100%.
[0018] Preferably, when verifying the training result with the validation set data and generating a trained gradient boosting tree model when the training result passes the verification, it further includes:
[0019] Initializing the gradient boosting tree model, estimating the model parameters that minimize the loss function, and creating an initial gradient boosting tree model based on the model parameters;
[0020] Within the expected maximum number of iterations, sequentially iterate through steps S1, S2, S3, and S4:
[0021] S1: Calculating the loss function and negative gradient of the initial gradient boosting tree model;
[0022] S2: Using the negative gradient as new training set data and fitting a regression tree model according to the new training set data;
[0023] S3: Calculating the best fitting value for each leaf node of the regression tree model;
[0024] S4: Incorporating the best fitting value into the regression tree model;
[0025] When the maximum number of iterations is reached, outputting the trained gradient boosting tree model.
[0026] Preferably, the multi-dimensional influencing quantities include: the electricity sales volume of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, the commissioning time of the line body data of the transmission line, the overhead line length, conductor type, conductor cross-section, number of splits, span of each level of tower of the line, call height, pole height, number of circuits erected on the same pole, phase sequence, terrain and geology, tower type, and turning direction.
[0027] Based on another aspect of the present invention, there is provided a system for predicting transmission line loss based on multi-dimensional influencing factors, the system comprising:
[0028] An acquisition unit for acquiring multi-dimensional influencing factors of transmission line loss;
[0029] A generation unit for preprocessing non-temporal data in the line body data and tower data at all levels of the transmission line to form fixed parameters of the transmission line, and incorporating the fixed parameters into the temporal data table of the transmission line to generate a multi-dimensional influencing factor data table;
[0030] A training unit for dividing the data in the multi-dimensional influencing factor data table into training set data, validation set data, and test set data; training a gradient boosting tree model with the training set data and outputting a training result;
[0031] A validation unit for validating the training result with the validation set data, and generating a trained gradient boosting tree model when the training result passes the validation;
[0032] A prediction unit for predicting the test set data with the trained gradient boosting tree model and outputting a line loss prediction result of the transmission line.
[0033] Preferably, the training unit is used to train a gradient boosting tree model with the training set data and output a training result, and is further used for:
[0034] Analyzing the weights of the multi-dimensional influencing factors through the gradient boosting tree model.
[0035] Preferably, analyzing the weights of the multi-dimensional influencing factors through the gradient boosting tree model further includes:
[0036] Training the gradient boosting tree model with the training set data to obtain a trained gradient boosting tree model;
[0037] Extracting all split node data n of the trained gradient boosting tree model;
[0038] Calculating the number of split nodes nk generated by each dimension of the k-dimensional influencing factor k ;
[0039] Calculating the weight Q of the k-dimensional influencing factor according to all the split node data n and the number of split nodes nk generated by each dimension of the influencing factor k , Q k =n k / n×100%.
[0040] Preferably, the verification unit is used to verify the training result with the verification set data. When the training result passes the verification, a trained gradient boosting tree model is generated. The verification unit is further used to:
[0041] Initialize the gradient boosting tree model, estimate the model parameters that minimize the loss function, and create an initial gradient boosting tree model based on the model parameters;
[0042] Within the expected maximum number of iterations, sequentially iterate through steps S1, S2, S3, and S4:
[0043] S1: Calculate the loss function and negative gradient of the initial gradient boosting tree model;
[0044] S2: Use the negative gradient as the new training set data, and fit a regression tree model according to the new training set data;
[0045] S3: Calculate and obtain the best fitting value for each leaf node of the regression tree model;
[0046] S4: Incorporate the best fitting value into the regression tree model;
[0047] When the maximum number of iterations is reached, output the trained gradient boosting tree model.
[0048] Preferably, the multi-dimensional influencing quantities include: the electricity sales volume of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, the commissioning time of the line body data of the transmission line, the overhead line length, conductor type, conductor cross-section, number of splits, span of each level of pole tower of the line, calling height, pole height, number of circuits erected on the same pole, phase sequence, terrain and geology, pole tower property, and corner direction.
[0049] The technical solution of the present invention provides a method and system for predicting the line loss of a transmission line based on multi-dimensional influencing factors. The method includes: collecting multi-dimensional influencing factors of the line loss of the transmission line; performing data preprocessing on the non-temporal data in the line body data and the pole tower data at all levels of the transmission line to form fixed parameters of the transmission line from the non-temporal data, and incorporating the fixed parameters into the temporal data table of the transmission line to generate a multi-dimensional influencing factor data table; dividing the data in the multi-dimensional influencing factor data table into training set data, validation set data, and test set data; training a gradient boosting tree model with the training set data and outputting a training result; verifying the training result with the validation set data, and when the training result passes the verification, generating a trained gradient boosting tree model; predicting the test set data with the trained gradient boosting tree model and outputting the line loss prediction result of the transmission line. To improve the accuracy of the line loss prediction of the 500 kV overhead transmission line, the technical solution of the present invention proposes a method for predicting the line loss of the 500 kV overhead transmission line based on multi-dimensional influencing factors, solves the problem of low accuracy of the current line loss prediction of the 500 kV overhead transmission line, and provides a certain reference for the direction of line loss management. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The exemplary embodiments of the present invention can be more fully understood by reference to the following drawings:
[0051] Figure 1 FIG. is a flowchart of a method for predicting the line loss of a transmission line based on multi-dimensional influencing factors according to a preferred embodiment of the present invention;
[0052] Figure 2 FIG. is a flowchart for training a gradient boosting tree model according to a preferred embodiment of the present invention;
[0053] Figure 3 FIG. is a flowchart for training a gradient boosting tree model according to a preferred embodiment of the present invention; and
[0054] Figure 4 FIG. is a structural diagram of a system for predicting the line loss of a transmission line based on multi-dimensional influencing factors according to a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0055] Now, the exemplary embodiments of the present invention will be described with reference to the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not intended to limit the present invention. In the drawings, the same unit / element is denoted by the same reference numeral.
[0056] Unless otherwise specified, the terms used herein (including scientific and technical terms) have the ordinary meaning as understood by those skilled in the relevant technical field. Additionally, it can be understood that terms defined in commonly used dictionaries should be construed to have a meaning consistent with the context of their relevant fields, and should not be construed as having an idealized or overly formal meaning.
[0057] Figure 1 This is a flowchart of a method for predicting transmission line loss based on multi-dimensional influencing factors according to a preferred embodiment of the present invention. This application provides a method for predicting the line loss of 500 kV overhead transmission lines based on multi-dimensional influencing factors, including: collecting multi-dimensional influencing factor data of 500 kV overhead transmission lines; preprocessing the line tower feature data; dividing the preprocessed line data, training a gradient boosting tree model to predict the line loss of the line, and calculating the weights of multi-dimensional influencing factors according to the prediction results. This application monitors the power metering data, environmental data set, and line body data of 500 kV overhead transmission lines; and predicts the line loss of 500 kV overhead transmission lines based on the monitored data, guides the line loss control direction through the calculation results of the weights of multi-dimensional influencing factors in the prediction model, effectively avoids economic losses caused by abnormal line losses, and significantly improves the line loss management level of 500 kV overhead transmission lines. As Figure 1 shown, a method for predicting transmission line loss based on multi-dimensional influencing factors, the method includes:
[0058] Preferably, in step 101: collect multi-dimensional influencing factors of the transmission line loss. Preferably, the multi-dimensional influencing factors include: the sold electricity of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, the commissioning time of the line body data of the transmission line, the length of the overhead line, the conductor type, the conductor cross-section, the number of splits, the span of each level of line towers, the call height, the pole height, the number of circuits erected on the same pole, the phase sequence, the terrain and geology, the pole tower property, and the corner direction. This application collects multi-dimensional influencing factors related to the line loss of 500 kV overhead transmission lines.
[0059] Preferably, in step 102: perform data preprocessing on the line body data in the transmission line route and the non-temporal data in the data of each level of transmission line towers, so that the non-temporal data forms fixed parameters of the transmission line, and incorporate the fixed parameters into the temporal data table of the transmission line route to generate a multi-dimensional impact quantity data table. In this application, data preprocessing is performed on two types of non-temporal data, namely line body data and data of each level of transmission line towers, so that they form fixed parameters of the line and are incorporated into the temporal data table of the line to form a multi-dimensional impact quantity data table. In the non-temporal data preprocessing method of this application, numerical data such as overhead line length, conductor cross-section, span, and nominal height are averaged, and categorical data such as phase sequence and terrain geology are encoded using ONE-HOT, that is, each possible category of categorical data is used as a single fixed parameter of the line, and the value of the fixed parameter of the line is the number of towers of the category to which the line belongs.
[0060] Preferably, in step 103: divide the data in the multi-dimensional impact quantity data table into training set data, validation set data, and test set data; train the gradient boosting tree model with the training set data and output the training result. Preferably, train the gradient boosting tree model with the training set data and output the training result, and it also includes: analyzing the weights of multi-dimensional impact quantities through the gradient boosting tree model. Preferably, when the gradient boosting tree model analyzes the weights of multi-dimensional impact quantities, it also includes: training the gradient boosting tree model with the training set data to obtain the trained gradient boosting tree model; extracting all split node data n of the trained gradient boosting tree model; calculating the number of split nodes n generated by each dimension of the k-dimensional impact quantity k ; according to all split node data n and the number of split nodes n generated by each dimension of the impact quantity k , calculate the weight Q of the k-dimensional impact quantity k , Q k = n k / n×100%. This application divides the preprocessed data. Based on the data recording date, the multi-dimensional impact quantity data table is divided into 6 equal parts, among which 5 parts are used as the training and validation data of the model, and 1 part is used as the test data for training the model. Using the 5-fold cross-validation method, the 5 parts of training and validation data are divided into 5 combinations of training sets and validation sets. Each time, 4 out of the 5 parts with incomplete sameness are taken as the training set, and 1 part of data is taken as the validation set, and the training is repeated 5 times. This application divides the preprocessed data to form training set data, validation set data, and test set data. This application uses the training set data to train the gradient boosting tree model, and screens according to the performance of the validation set data. The formed model set predicts the test set data, and the results are averaged and output as the line loss prediction result to obtain the 500kV overhead transmission line line loss prediction model. Finally, the weights of the multi-dimensional impact quantities in the model are analyzed. The input of the 500kV overhead transmission line line loss prediction model of this application is the multi-dimensional impact quantity data table in the training set data, and the output is the line loss value of the 500kV overhead transmission line.
[0061] This application extracts the training set data and trains the gradient boosting tree model; among them, a finite data set D = {Z, y} is given, where Z = [Z1,..., Z i ,..., Z N is the input, and y = [y1,..., y N is the output.
[0062] As Figure 2 shown, extract all the splitting node numbers n of the trained gradient boosting tree model;
[0063] Calculate the number of splitting nodes n k generated by each dimension of the k-dimensional impact quantity;
[0064] Calculate the impact quantity weight Q k of the k-dimensional impact quantity, Q k = n k / n×100%.
[0065] Preferably, in step 104: the training result is verified by the validation set data. When the training result passes the verification, a trained gradient boosting tree model is generated. Preferably, verifying the training result by the validation set data and generating a trained gradient boosting tree model when the training result passes the verification further includes: initializing the gradient boosting tree model, estimating the model parameters that minimize the loss function, and creating an initial gradient boosting tree model based on the model parameters; iteratively executing steps S1, S2, S3, and S4 according to the expected maximum number of iterations: S1: calculating the loss function and negative gradient of the initial gradient boosting tree model; S2: using the negative gradient as the new training set data and fitting a regression tree model according to the new training set data; S3: calculating and obtaining the best fitting value for each leaf node of the regression tree model; S4: fusing the best fitting value into the regression tree model; when the maximum number of iterations is reached, outputting the trained gradient boosting tree model.
[0066] As Figure 3 shown, the present application first initializes the gradient boosting tree model. For a given finite data set D = {Z, y}, where Z = [Z1,…,Z i ,…Z N is the input and y = [y1,…y N is the output. Estimate the model parameters γ that minimize the loss function L(y,γ), N is the number of data samples, and use it as the initial model f0(Z i ), that is:
[0067]
[0068] Let C be the number of iterations. For the c-th iteration, c = 1, 2,…C, execute steps S1 - S4;
[0069] S1 calculates the current model loss function and the negative gradient r ic , that is, the residual:
[0070]
[0071] S2 uses r ic as the new label of the sample Z i . In the formula, f(Z i ) is the model obtained in the (c - 1)-th iteration, and a new sample data set [(Z i , r ic ), i = 1, 2,…N] is obtained. Use it as the new training data and fit to obtain the next regression tree model. The new tree model consists of leaf nodes R jc (j = 1, 2,…J). J is the number of leaf nodes of the regression tree model.
[0072] S3 for each leaf node R jc, calculate the best - fit value γ of the sample jc .
[0073]
[0074] where f c-1 (Z i ) is the model obtained from the (c - 1) - th iteration;
[0075] S4 Update the model of the m - th iteration:
[0076]
[0077] I(Z i ∈R jc ) is an indicator function. When the sample Z i belongs to the leaf node R jc , the function value is 1, otherwise it is 0;
[0078] Output the final gradient - boosted tree model f C (Z i ).
[0079]
[0080] f o (Z i ) is the initial model; I(Z i ∈R jc ) is an indicator function;
[0081] Preferably, in step 105: Use the trained gradient - boosted tree model to predict the test - set data and output the line - loss prediction result of the transmission line.
[0082] Figure 4 is a system structure diagram for predicting the line loss of a transmission line based on multi - dimensional influencing factors according to a preferred embodiment of the present invention. As Figure 4 shown, the present application provides a system for predicting the line loss of a transmission line based on multi - dimensional influencing factors. The system includes:
[0083] A collection unit 401, which is used to collect the multi - dimensional influencing factors of the line loss of the transmission line. Preferably, the multi - dimensional influencing factors include: the sold electricity of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, the commissioning time, overhead line length, conductor type, conductor cross - section, number of splits, span, call height, pole height, number of circuits erected on the same pole, phase sequence, terrain and geology, pole tower property, and corner direction of the line body data of the transmission line. The present application collects the multi - dimensional influencing factors related to the line loss of 500kV overhead transmission lines.
[0084] A generating unit 402 is configured to perform data preprocessing on the line body data in the transmission line route and the non-temporal data in the tower data at all levels of the line, so that the non-temporal data forms fixed parameters of the transmission line, and incorporate the fixed parameters into the temporal data table of the transmission line route to generate a multi-dimensional influence quantity data table. In this application, data preprocessing is performed on two types of non-temporal data, namely, line body data and tower data at all levels of the line, so that they form fixed parameters of the line and are incorporated into the temporal data table of the line to form a multi-dimensional influence quantity data table. In the non-temporal data preprocessing method of this application, numerical data such as overhead line length, conductor cross-section, span, and nominal height are averaged, and categorical data such as phase sequence and terrain geology are encoded using ONE-HOT, that is, each possible category of the categorical data is used as a single fixed parameter of the line, and the value of the fixed parameter of the line is the number of towers of the category to which it belongs in the line.
[0085] A training unit 403 is configured to divide the data in the multi-dimensional influence quantity data table into training set data, validation set data, and test set data; train the gradient boosting tree model with the training set data and output a training result. Preferably, the training unit is configured to train the gradient boosting tree model with the training set data and output a training result, and is further configured to: analyze the weights of the multi-dimensional influence quantities through the gradient boosting tree model. Preferably, analyzing the weights of the multi-dimensional influence quantities through the gradient boosting tree model further includes: training the gradient boosting tree model with the training set data to obtain the trained gradient boosting tree model; extracting all split node data n of the trained gradient boosting tree model; calculating the number of split nodes n generated by each dimension of the k-dimensional influence quantity k ; according to all split node data n and the number of split nodes n generated by each dimension of the influence quantity k calculate the weight Q of the k-dimensional influence quantity k , Q k = n k / n×100%. This application divides the preprocessed data. Based on the data recording date, the multi-dimensional impact quantity data table is divided into 6 equal parts, among which 5 parts are used as the training and validation data of the model, and 1 part is used as the test data for training the model. Using the 5-fold cross-validation method, the 5 parts of training and validation data are divided into 5 combinations of training sets and validation sets. Each time, 4 out of the 5 parts that are not exactly the same are taken as the training set, and 1 part of data is taken as the validation set, and the training is repeated 5 times. This application divides the preprocessed data into training set data, validation set data, and test set data. This application uses the training set data to train the gradient boosting tree model, and filters according to the performance of the validation set data. The formed model set predicts the test set data, and the results are averaged and output as the line loss prediction result to obtain the 500kV overhead transmission line line loss prediction model. Finally, the weights of the multi-dimensional impact quantities in the model are analyzed. The input of the 500kV overhead transmission line line loss prediction model of this application is the multi-dimensional impact quantity data table in the training set data, and the output is the line loss value of the 500kV overhead transmission line.
[0086] This application extracts the training set data and trains the gradient boosting tree model; among them, a finite data set D = {Z, y} is given, where Z = [Z1,..., Z i ,…Z N is the input, and y = [y1,..., y N is the output.
[0087] As Figure 2 shown, extract all the split node numbers n of the trained gradient boosting tree model;
[0088] Calculate the number of split nodes n k generated by each dimension of the k-dimensional impact quantity;
[0089] Calculate the impact quantity weight Q k of the k-dimensional impact quantity, Q k = n k / n×100%.
[0090] A verification unit 404 is used to verify the training result with the verification set data. When the training result passes the verification, a trained gradient boosting tree model is generated. Preferably, the verification unit 404 is used to verify the training result with the verification set data. When the training result passes the verification, a trained gradient boosting tree model is generated, and is further used to: initialize the gradient boosting tree model, estimate the model parameters that minimize the loss function, and create an initial gradient boosting tree model based on the model parameters; within the expected maximum number of iterations, sequentially iterate through steps S1, S2, S3, and S4: S1: Calculate the loss function and negative gradient of the initial gradient boosting tree model; S2: Use the negative gradient as the new training set data, and fit a regression tree model according to the new training set data; S3: Calculate and obtain the best fit value for each leaf node of the regression tree model; S4: Incorporate the best fit value into the regression tree model; when the maximum number of iterations is reached, output the trained gradient boosting tree model.
[0091] As Figure 3 shown, the present application first initializes the gradient boosting tree model, estimates the model parameters γ that minimize the loss function L(y,γ), where N is the number of data samples, and uses it as the initial model f0(Z i ),, that is:
[0092]
[0093] Let C be the number of iterations. For the c-th iteration, c = 1, 2,... C, execute steps S1 - S4;
[0094] S1 calculates the current model loss function and the negative gradient r of the model according to the following formula ic , that is, the residual:
[0095]
[0096] S2 uses r ic as the new label for the sample Z i , where f(Z i ) is the model obtained from the (c - 1)-th iteration, and a new sample data set [(Z i , r ic ), i = 1, 2,... N] is obtained, which is used as the new training data to fit the next regression tree model. The new tree model consists of leaf nodes R jc (j = 1, 2,... J). J is the number of leaf nodes of the regression tree model.
[0097] S3 calculates the best fit value γ jc for each leaf node R jc .
[0098]
[0099] where f c-1 (Z i ) is the model obtained from the (c - 1)-th iteration;
[0100] S4 Updates the model for the m-th iteration:
[0101]
[0102]
[0103] I(Z i ∈R jc ) is an indicator function, which has a value of 1 when the sample Z i belongs to the leaf node R jc and 0 otherwise;
[0104] Output the final gradient boosting tree model f C (Z i ).
[0105]
[0106] f o (Z i ) is the initial model; I(Z i ∈R jc ) is an indicator function;
[0107] The prediction unit 405 is configured to predict the test set data through the trained gradient boosting tree model and output the line loss prediction result of the output circuit line.
[0108] A system 400 for predicting the line loss of a transmission line based on multi-dimensional influencing factors according to a preferred embodiment of the present invention corresponds to a method 100 for predicting the line loss of a transmission line based on multi-dimensional influencing factors according to a preferred embodiment of the present invention, and will not be elaborated herein.
[0109] The present invention has been described by referring to a few embodiments. However, as is well known to those skilled in the art, other embodiments equivalent to those disclosed above of the present invention equally fall within the scope of the present invention as defined by the appended patent claims.
[0110] Generally, all terms used in the claims are construed according to their ordinary meanings in the technical field, unless otherwise clearly defined therein. All references to "a / the [device, component, etc.]" are to be construed openly as at least one instance of the device, component, etc., unless otherwise clearly stated. The steps of any method disclosed herein need not be performed in the exact order disclosed, unless clearly stated.
Claims
1. A method for predicting transmission line loss based on multi-dimensional influencing factors, the method comprising: Collecting multi-dimensional influencing factors of transmission line loss; The multi-dimensional influencing factors include: electricity sales volume of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, operation time of the line body data of the transmission line, overhead line length, conductor type, conductor cross-section, number of splits, span of each level of poles and towers of the line, nominal height, pole height, number of circuits erected on the same pole, phase sequence, terrain and geology, pole and tower property, corner direction; Performing data preprocessing on the non-temporal data in the line body data and the data of each level of poles and towers in the transmission line, so that the non-temporal data forms fixed parameters of the transmission line, and incorporating the fixed parameters into the temporal data table of the transmission line to generate a multi-dimensional influencing factor data table; wherein, performing data preprocessing on the non-temporal data in the line body data and the data of each level of poles and towers includes: performing averaging processing on numerical data, and using the averaged value after processing as a fixed parameter; for categorical data, using the number of categories of the categorical data as a fixed parameter; Dividing the data in the multi-dimensional influencing factor data table into training set data, validation set data and test set data; training a gradient boosting tree model with the training set data and outputting a training result; Validating the training result with the validation set data, and when the training result passes the validation, generating a trained gradient boosting tree model, further comprising: Initializing the gradient boosting tree model, estimating model parameters that minimize the loss function, and creating an initial gradient boosting tree model based on the model parameters; Within the expected maximum number of iterations, sequentially iterating through steps S1, S2, S3, and S4: S1: Calculating the loss function and negative gradient of the initial gradient boosting tree model; S2: Using the negative gradient as new training set data and fitting a regression tree model according to the new training set data; S3: Calculating and obtaining the best fitting value for each leaf node of the regression tree model; S4: Incorporating the best fitting value into the regression tree model; When the maximum number of iterations is reached, outputting the trained gradient boosting tree model; Predicting the test set data with the trained gradient boosting tree model and outputting the line loss prediction result of the transmission line.
2. The method according to claim 1, wherein the training the gradient boosting tree model with the training set data and outputting a training result further comprises: Analyzing the weights of the multi-dimensional influencing factors through the gradient boosting tree model.
3. The method according to claim 2, wherein the analyzing the weights of the multi-dimensional influencing factors through the gradient boosting tree model further comprises: Training the gradient boosting tree model with the training set data to obtain a trained gradient boosting tree model; Extracting all split node data n of the trained gradient boosting tree model; Calculate the number of split nodes n generated by each dimension of the k-dimensional influence quantity k ; The number of splitting nodes $n$ generated according to all the splitting node data $n$ and the influence amount per dimension k , calculate the weight $Q$ of the $k$-dimensional influence amount k , $Q$ k = $n$ k / $n$ × 100%.
4. A system for predicting transmission line loss based on multi-dimensional influencing factors, the system comprising: A collection unit for collecting multi-dimensional influencing factors of transmission line loss; The multi-dimensional influencing quantities include: the electricity sales volume of the transmission line, voltage, current, temperature, humidity, air pressure, wind speed and direction of the environment where the transmission line is located, the commissioning time of the line body data of the transmission line, the overhead line length, conductor type, conductor cross-section, number of splits, span of each level of pole towers of the line, nominal height, pole height, number of circuits erected on the same pole, phase sequence, terrain and geology, pole tower property, and corner direction; The generation unit is used to perform data preprocessing on the non-temporal data in the line body data and the data of each level of pole towers in the transmission line, so that the non-temporal data forms the fixed parameters of the transmission line, and incorporate the fixed parameters into the temporal data table of the transmission line to generate a multi-dimensional influencing quantity data table; wherein, performing data preprocessing on the non-temporal data in the line body data and the data of each level of pole towers includes: performing averaging processing on the numerical data, and using the processed average value as the fixed parameter; for categorical data, using the number of categories of the categorical data as the fixed parameter; The training unit is used to divide the data in the multi-dimensional influencing quantity data table into training set data, validation set data and test set data; train the gradient boosting tree model through the training set data, and output the training result; The validation unit is used to validate the training result through the validation set data. When the training result passes the validation, generate the trained gradient boosting tree model, and is also used for: Initialize the gradient boosting tree model, estimate the model parameters that minimize the loss function, and create an initial gradient boosting tree model based on the model parameters; Within the expected maximum number of iterations, sequentially execute steps S1, S2, S3 and S4: S1: Calculate the loss function and negative gradient of the initial gradient boosting tree model; S2: Use the negative gradient as the new training set data, and fit a regression tree model according to the new training set data; S3: Calculate and obtain the best fitting value for each leaf node of the regression tree model; S4: Incorporate the best fitting value into the regression tree model; When the maximum number of iterations is reached, output the trained gradient boosting tree model; The prediction unit is used to predict the test set data through the trained gradient boosting tree model, and output the line loss prediction result of the transmission line.
5. The system according to claim 4, wherein the training unit is used to train the gradient boosting tree model through the training set data, output the training result, and is also used for: Analyze the weights of the multi-dimensional influencing quantities through the gradient boosting tree model.
6. The system according to claim 5, wherein analyzing the weights of the multi-dimensional influencing quantities through the gradient boosting tree model further includes: Train the gradient boosting tree model through the training set data to obtain the trained gradient boosting tree model; Extract all split node data n of the trained gradient boosting tree model; Calculate the number of split nodes n generated by each dimension of influence in the k-dimensional influence quantity k ; The number of split nodes n generated based on all the split node data n and the influence amount per dimension k , calculate the weight Q of the k-dimensional influence amount k , Q k = n k / n × 100%.
Citation Information
Patent Citations
Cleaning method for distribution feeder statistical line loss rate data based on AMI data
CN107301499A
Intelligent electric energy meter time series data processing method and processing device
CN110502518A