Method for predicting ground synthetic electric field based on improved Stacking algorithm

Through the improved Stacking algorithm combined with multiple machine learning models, the problem of unconsidered environmental factors in the prediction of ground synthesis electric field of DC line is solved, and the prediction effect with higher accuracy and wider range is achieved.

CN115544879BActive Publication Date: 2025-07-08TAIAN XINHENG TECH SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202211214405.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2025-07-08
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

The existing DC line ground synthesis electric field prediction methods fail to fully consider complex environmental factors such as temperature, humidity, wind direction, air particles of different particle sizes, wind speed, etc., resulting in low prediction accuracy and limited use range.

Method used

The improved Stacking algorithm is adopted, combining five machine learning models (stochastic forest, gradient enhancement decision tree, lightweight gradient lifter, extreme gradient lift tree, and K nearest neighbor algorithm) as the basis learner, and the first layer model is trained through five-fold cross-validation, and the multivariate linear regression model is used as the second layer meta learner to build a ground synthetic electric field prediction model to avoid overfitting the prediction results.

Benefits of technology

The accuracy of ground synthesis electric field prediction is improved and the scope of use is expanded. The mean square error, root mean square error and average absolute error are all lower than that of a single machine learning model, and the prediction effect is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115544879B_ABST
    Figure CN115544879B_ABST
Patent Text Reader

Abstract

Method for predicting ground synthetic electric field based on improved Stacking algorithm, comprising the following steps: obtaining sample data, including ground synthetic electric field data and meteorological data; preprocessing the obtained sample data, including outlier detection and normalization processing; using five-fold cross-validation method to perform prediction modeling on the preprocessed sample data, and using the improved Stacking algorithm prediction model to predict the synthetic electric field. Using mean square error MSE, root mean square error RMSE, and mean absolute error MAE to evaluate the performance of the improved Stacking algorithm prediction model. The method of the present invention uses a machine learning algorithm with higher prediction accuracy as the base learner. Compared with the traditional synthetic electric field prediction method, machine learning prediction can fully consider the influence of complex environment on the ground synthetic electric field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of calculating the ground synthetic electric field of ultra / extra high voltage DC lines, and specifically relates to a method for predicting the ground synthetic electric field based on an improved Stacking algorithm. Background Technique

[0002] UHV transmission lines are an important part of China's power grid, and the surrounding electromagnetic environment is complex and changeable. The ground synthetic electric field is one of the main electromagnetic environment indicators of UHV transmission lines. Accurate prediction and long-term monitoring are of great significance for the safe operation of the power grid. UHVDC transmission lines have the advantages of long transmission distance and large power compared with general lines, which can greatly facilitate residential and industrial electricity use, but at the same time will cause electromagnetic environment impacts, such as synthetic electric fields and audible noises. The ground electric field of DC transmission lines is called the DC synthetic electric field. Since the transmission frequency is 0Hz, the space charge generated by corona on the conductor after the line operates will form an ion current in space. After the ion current field is superimposed with the nominal electric field of the conductor itself, the ground synthetic electric field intensity will be significantly increased.

[0003] At present, the methods for predicting the ground synthetic electric field of DC lines mainly include calculation methods based on Deutsch hypothesis, upwind finite element method, flux line method, etc. However, these numerical calculation methods fail to fully reflect the influence mechanism of complex environments on the synthetic electric field, such as temperature, humidity, wind direction, air particles with different particle sizes, wind speed, etc., and cannot effectively predict the distribution of the synthetic electric field. Summary of the Invention

[0004] In view of the above technical problems, the present invention provides a method for predicting the ground synthetic electric field based on an improved Stacking algorithm. Compared with the traditional synthetic electric field calculation method, this prediction method can fully consider the influence of complex environmental factors on the ground synthetic electric field, and overcomes the problem that it is difficult to obtain good prediction results due to the deficiencies of a single machine learning model in some aspects, improving the prediction accuracy and expanding the scope of use; compared with the general Stacking algorithm, this method effectively avoids overfitting of the prediction results and does not reduce the prediction accuracy.

[0005] The technical solution adopted by the present invention is as follows:

[0006] A method for predicting the ground synthetic electric field based on an improved Stacking algorithm, comprising the following steps:

[0007] Step 1: Obtain sample data, including ground synthetic electric field data and meteorological data;

[0008] Step 2: Preprocess the sample data obtained in Step 1, including outlier detection and normalization processing;

[0009] Step 3: Use the five-fold cross-validation method to perform predictive modeling on the sample data that has been processed in Step 2 to obtain an improved Stacking algorithm prediction model.

[0010] Step 4: Use the improved Stacking algorithm prediction model in Step 3 to predict the synthetic electric field.

[0011] Step 5: Use the mean squared error MSE, root mean squared error RMSE, and mean absolute error MAE to evaluate the performance of the improved Stacking algorithm prediction model.

[0012] In the above Step 1, according to the national measurement standard DL / T 1089-2008, use the corresponding equipment to collect synthetic electric field data and meteorological data. The meteorological data includes temperature, humidity, air pressure, wind speed, PM1.0, PM2.5, and PM10.

[0013] In the above Step 2, the outlier detection method used is the Local Outlier Factor (LOF) algorithm. The LOF algorithm is an unsupervised algorithm for determining outliers in multivariate data. It quantifies the degree of abnormality of each point in the dataset by calculating distance and density.

[0014] In the above Step 2, the normalization method used is

[0015] where: S is the result of normalizing each feature; s is the original data of each feature; S max and S min are the maximum and minimum values of each feature.

[0016] The above Step 3 includes the following steps:

[0017] S3.1: After normalizing the initial data, obtain the dataset G = {(y i , x i ); i = 1, 2, … M}, where x i is the i-th sample feature vector and y i is the output value corresponding to the i-th sample; M is the number of collected synthetic electric field samples.

[0018] S3.2: Divide the dataset G into a training set G train and a test set G test , where the number of training set samples is N and the number of test set samples is M - N. Use the five-fold cross-validation method to train the base learners in the first layer for the training set G train : Divide G train into five equal parts to obtain G train1 , G train2 , G train3 , Gtrain4 , G train5 , which are used as the test set and training set of different base learners respectively.

[0019] Select one fold of them as the test set in turn, and the remaining four folds as the training set. During the five training processes, the prediction results of the base learner H k , and the prediction results of z kn , where k represents the serial number of the first-layer base learner, and K represents the number of the first-layer base learners.

[0020] Finally, the output results of the K base learners and G train The y in n constitute a new training set used by the second-layer meta-learner: G train-new ={(y n , z 1n ,…, z kn ); n = 1,…, N, k = 1,…, K};

[0021] Among them: n represents the sample serial number in the training set, and N represents the number of training set samples; y n represents the predicted value corresponding to n samples; z 1n ,…, z kn represents the new synthetic electric field data predicted by the K base learners according to the N samples.

[0022] S3.3: During the five-fold cross-validation process in S3.2, each fold of cross-validation will train the base learner H k . Through the well-trained base learner H k , the y in the test set G test is predicted. After the five-fold cross-validation is finally completed, the predicted value P i output by each base learner for G test is: P kn =(P kn +P kn1 +P kn2 +P kn3 +P kn4 +P kn5 ) / 5;

[0023] Among them, P kn1 , P kn2 , P kn3 , P kn4 , P kn5 are the predicted values output by the k-th base learner for the test set G test after five different trainings respectively. P kn is the average value of the predicted values output by the k-th base learner for the test set G test . n is the test set G testNumber of samples.

[0024] The results output by the final K base learners and G test in y i constitute a new test set used by the second-layer meta-learner:

[0025] G test-new ={(y n , P 1n ,…, P kn ); n = 1,…, M - N, k = 1,…, K};

[0026] P 1n ,…, P kn represent the new synthetic electric field data predicted by the K base learners based on m - N samples.

[0027] S3.4: Train the second-layer meta-learner with the new training set G train-new and use the new test set G test-new to detect the accuracy of the model.

[0028] In step 4, for the first-layer base learners of the improved Stacking algorithm prediction model, five prediction algorithms with strong learning ability and large structural differences are selected: Random Forest (RF), Gradient Boosting Decision Tree (GBDT), LightGBM, Extreme Gradient Boosting (XGBoost), and K-Nearest Neighbor (KNN). For the second-layer meta-learner, a linear regression model is selected for the final prediction.

[0029] In step 5, the mean square error formula is:

[0030] The root mean square error formula is:

[0031] The mean absolute error formula is:

[0032] Where: y i and represent the true value and the predicted value; n represents the number of predicted values and true values; the smaller the mean square error, root mean square error, and mean absolute error, the better the prediction effect of the algorithm on the ground synthetic electric field.

[0033] The method for predicting the ground synthetic electric field based on the improved Stacking algorithm of the present invention has the following technical effects:

[0034] 1) The method of the present invention uses machine learning algorithms with high prediction accuracy as base learners. Compared with traditional synthetic electric field prediction methods, machine learning prediction can fully consider the influence of complex environments on the ground synthetic electric field.

[0035] 2) The present invention uses a new combined prediction method: the Stacking method, which utilizes five single machine learning models: Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), and K-Nearest Neighbor Algorithm (KNN). The single prediction results of these models are used as the training set to train the second-layer meta-learner for predicting the ground synthetic electric field. This overcomes the problem that it is difficult to obtain good prediction results due to the deficiencies of a single model in certain aspects. The prediction results show that the mean squared error (MSE) of the improved Stacking method is 0.1228, the root mean squared error (RMSE) is 0.3235, and the mean absolute error (MAE) is 0.2396, all of which are lower than those of the five single machine learning models. The improved Stacking algorithm improves the prediction accuracy and expands the scope of use.

[0036] 3) In the training stage of a general Stacking model, the training set and test set of the second-layer meta-learner are only composed of the new data sets generated by the first-layer base learners, which inevitably leads to overfitting of the meta-learner's prediction results. To address this issue, the present invention proposes an improved Stacking algorithm: after completing the training of the first-layer base learners, the prediction results obtained based on the initial training set G train and the initial training set G train are combined as the training set G train-new for the meta-learner; the same applies to the test set of the meta-learner. The prediction results obtained by the trained base learners based on the initial test set G test and the initial test set G test are combined as the test set of the meta-learner. This improved Stacking algorithm can effectively avoid overfitting of the ground synthetic electric field prediction results. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is a flowchart of the method of the present invention.

[0038] Figure 2(a) shows the results of univariate outlier detection;

[0039] Figure 2(b) shows the results of bivariate outlier detection;

[0040] Figure 2(c) shows the results of multivariate outlier detection.

[0041] Figure 3 is a schematic diagram of five-fold cross-validation.

[0042] Figure 4 is a framework diagram of the Stacking ensemble learning model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The present invention will be further described below with reference to the accompanying drawings:

[0044] The Stacking framework is an ensemble learning model with a serial structure. Different from ensemble learning methods such as Bagging, Boosting, and Voting, the meta-learner in the second layer of Stacking uses the output results of multiple base learners in the first layer for stacking to obtain the result. As Figure 1 shown, the present invention provides a method for predicting the ground synthetic electric field based on an improved Stacking algorithm, which specifically includes the following steps:

[0045] Step S1: Obtain sample data, including ground synthetic electric field data and meteorological data;

[0046] Step S2: Perform data preprocessing on the sample data obtained in Step S1, including outlier detection and normalization processing;

[0047] Step S3: Use the five-fold cross-validation method to perform predictive modeling on the sample data that has been processed in Step S2 to obtain an improved Stacking algorithm prediction model;

[0048] Step S4: Use the model based on the improved Stacking algorithm to predict the synthetic electric field;

[0049] Step S5: Use the mean square error MSE, root mean square error RMSE, and mean absolute error MAE to evaluate the performance of the synthetic electric field prediction model.

[0050] In Step S1, the obtained sample data should be collected according to the national measurement standard DL / T 1089-2008, and the corresponding equipment should be used to collect the synthetic electric field data and meteorological data. The synthetic electric field data is the ground synthetic field strength under the UHV transmission line collected by the TFMS field mill probe, unit: kV / m; the meteorological data includes temperature, humidity, air pressure, wind speed, PM1.0, PM2.5, and PM10.

[0051] In Step S2, due to equipment anomalies and external environmental interferences and other emergencies during the collection process of the synthetic electric field, there is a great possibility of randomly generating abnormal data in the dataset, which will affect the final prediction result. The determination of single-variable outliers is shown in Figure 2(a). There are some points with insignificant fluctuations in value, and it is difficult to determine whether they are outliers; the determination of multi-variable outliers is shown in Figures 2(b) and 2(c). The more variables there are, the more obvious the outliers are. In this application, the multi-variable outlier detection algorithm - Local Outlier Factor (LOF) algorithm is selected.

[0052] In the present invention, there is generally a strong linear relationship between the synthetic electric fields collected by each probe. Therefore, the Euclidean distance is used for distance calculation between points, which specifically includes the following steps:

[0053] 1) Calculate the k-th distance of a point p in the synthetic electric field dataset as: d k (p) = d(p, o)

[0054] Where: d(p, o) is the distance between point p and point o. In the dataset, there are at least k points other than p whose distances to point p are less than or equal to d(p, o), and at most k - 1 points other than p whose distances to point p are less than or equal to d(p, o).

[0055] 2) Calculate the k-th distance neighborhood of a point p in the synthetic electric field dataset as: N k (p) = {q ∈ D\{p}|d(p, q) ≤ d k (p)}

[0056] Where: point q is in the neighborhood of point p; D\{p} is all points except point p.

[0057] 3) Calculate the local reachability density of a point p in the synthetic electric field dataset as:

[0058] Where: the reachability distance from point o to point p in the dataset is: Rd k (p, o) = max{N k (o), d(p, o)}, and max{N k (o), d(p, o)} represents the maximum value between the k-th distance and the straight-line distance, where o ∈ N k (o).

[0059] 4) The local outlier factor of a point p in the dataset is:

[0060] The LOF algorithm is an unsupervised algorithm for determining outliers in multivariate data. It quantifies the degree of abnormality of each point in the dataset by calculating distance and density. This algorithm calculates an outlier factor LOF for each point in the dataset, and determines whether it is an outlier factor by judging whether LOF is close to 1. If LOF is much greater than 1, it is considered an outlier factor; if it is close to 1, it is a normal point.

[0061] This application uses 7 features including temperature, humidity, air pressure, wind speed, PM10, PM2.5, and PM1.0 to train an algorithm model to predict the ground synthetic electric field of ±800 kV transmission lines. Since the 7 meteorological data are in different dimensions, normalization is required to avoid the influence of dimensional differences on the prediction accuracy, and the values are converted to the range of 0 to 1. The normalization method used is Where: S is the result of normalizing each feature; s is the original data of each feature; S max and S min are the maximum and minimum values of each feature respectively.

[0062] In step S3, it specifically includes the following steps:

[0063] Step 3.1: After normalizing the initial data, the dataset G = {(y i , x i ); i = 1, 2, … m} is obtained, where x i is the i-th sample feature vector and y i is the predicted value corresponding to the i-th sample.

[0064] Step 3.2: Divide the data G into a training set G train and a test set G test . Use 5-fold cross-validation to train the base learners of the first layer for G train : Divide G train into five equal parts to obtain G train1 , G train2 , G train3 , G train4 , G train5 . Select one fold as the test set in turn, and the remaining four folds as the training set.

[0065] As Figure 3 shown, it specifically includes the following steps:

[0066] 1) Use G train1 , G train2 , G train3 , G train4 as the training set and G train5 as the test set to train the base learner H1, and the prediction result is z 1n1 ; Use G train1 , G train2 , G train3 , G train5 as the training set and G train4 as the test set to train the base learner H1, and the prediction result is z 1n2 ; Use G train1 , G train2 , G train4 , G train5 as the training set and G train3 as the test set to train the base learner H1, and the prediction result is z 1n3 ; Use G train1 , G train3 , G train4 , G train5 as the training set and G train2 as the test set to train the base learner H1, and the prediction result is z 1n4 ; Use G train2 , G train3 , G train4 , Gtrain5 As the training set, G train1 Use G as the test set to train the base learner H1, and the prediction result is z 1n5 . Then the prediction result of the base learner H1 is z 1n = z 1n1 + z 1n2 + z 1n3 + z 1n4 + z 1n5 , where n is the number of samples N in the test set.

[0067] 2) Train the base learners H2, H3, H4, and H5 successively as in step 1), and the prediction results z 2n , z 3n , z 4n , z 5n will be obtained successively.

[0068] 3) During the five training processes, the prediction results of the base learner H K , where k = 1, …, K, can be expressed as z kn . Finally, the results output by the K base learners and the y train in G n constitute a new training set used by the second-layer meta-learner:

[0069] G train-new = {(y n , z 1n , …, z kn ); n = 1, …, N}

[0070] Step 3.3: During the five-fold cross-validation process in step 3.2, each fold of cross-validation will train the base learner H k . Through the trained H k each time, predict the y test in G i . Finally, after the five-fold cross-validation ends, the predicted value P test output by each base learner is: P kn = (P kn + P kn1 + P kn2 + P kn3 + P kn4 + P kn5 ) / 5, which specifically includes the following steps:

[0071] 1) Use G train1 , G train2 , G train3 , G train4 as the training set, and G train5 as the test set to train the base learner H1. After training, output the prediction result P for the test set G test ​1n1 ; G train1 、G train2 、G train3 、G train5 as the training set, G train4 as the test set to train the base learner H1. After training is completed, the prediction result P is output for the test set G test ; G 1n2 ; G train1 、G train2 、G train4 、G train5 as the training set, G train3 as the test set to train the base learner H1. After training is completed, the prediction result P is output for the test set G test ; G 1n3 ; G train1 、G train3 、G train4 、G train5 as the training set, G train2 as the test set to train the base learner H1. After training is completed, the prediction result P is output for the test set G test ; G 1n4 ; G train2 、G train3 、G train4 、G train5 as the training set, G train1 as the test set to train the base learner H1. After training is completed, the prediction result P is output for the test set G test ; G 1n5 .

[0072] After the final five-fold cross-validation is completed, the predicted value P of G output by the first base learner test is: P 1n = (P 1n + P 1n1 + P 1n2 + P 1n3 + P 1n4 + P 1n5 ) / 5, where n is the number of samples in the test set G test i.e., M - N.

[0073] 2) Train the base learners H2, H3, H4, H5 in sequence as in step 1), and the prediction results P 2n 、P 3n 、P 4n 、P 5n .

[0074] 3) The results output by the final K base learners and y test in G i constitute the new test set used by the second-level meta-learner:

[0075] G test-new = {(y n , P 1n , …, P kn ); n = 1, …, M - N}

[0076] Step 3.4: Train the meta - learner of the second layer through G train-new and use G test-new to detect the accuracy of the model. In the training stage of the general Stacking model, the training set and test set of the meta - learner of the second layer are only composed of the new data sets generated by the base learners of the first layer, which will inevitably lead to over - fitting of the prediction results of the meta - learner. The new training set G train-new and the new test set G test-new composed of the improved Stacking model adopted by the present invention can effectively avoid over - fitting of the prediction results. In step S4, for the method of predicting the ground synthetic electric field based on the improved Stacking algorithm, as Figure 4 shown, in this Stacking ensemble learning model, five machine learning algorithms, namely Random Forest (RF), Gradient Boosting Decision Tree (GBDT), Light Gradient Boosting Machine (LightGBM), Extreme Gradient Boosting (XGBoost), and K - Nearest Neighbor algorithm (KNN), are selected in the base learners of the first layer. Among them: XGBoost, LightGBM, and GBDT are different implementation forms of the Boosting ensemble learning method based on decision trees; RandomForest is an implementation form of the Voting ensemble learning method based on decision trees; KNN is a machine learning method for calculating distances in the feature space, and this method is relatively mature in theory and application. These five algorithms have strong learning abilities and large differences in structure, resulting in relatively high prediction accuracy and non - overly similar prediction results.

[0077] This application predicts the ground synthetic electric field of the ±800 kV transmission line by training an algorithm model with 7 features including temperature, humidity, air pressure, wind speed, PM10, PM2.5, and PM1.0. After normalizing the meteorological data, the values are converted to between 0 and 1 to avoid the influence of the dimension differences between different features on the prediction accuracy. The base learners of the first layer make predictions based on the training set and test set divided from the measured synthetic electric field data. Specifically: 1. Combine the prediction results of the training set and the initial training set as the training set of the meta - learner of the second layer to improve the accuracy of the synthetic electric field model predicted by the improved Stacking algorithm; 2. Combine the prediction results of the test set and the initial test set as the test set of the meta - learner of the second layer to verify the effect of the synthetic electric field model predicted based on the improved Stacking algorithm.

[0078] In the improved Stacking algorithm, a multiple linear regression model is selected as the meta-learner in the second layer for the final prediction. The multiple linear regression model is usually used to describe the stochastic linear relationship between variable y and x, that is: y = β0 + β1x1 + β2x2 + … + β n x n + ξ, where x1, …, x k are non-stochastic variables, specifically the 7 meteorological factors of temperature, humidity, air pressure, wind speed, PM10, PM2.5, and PM1.0 in this application; β0, β1 … β n are regression coefficients; ξ is a stochastic error term; y is a stochastic dependent variable, specifically the synthetic electric field prediction value in this application.

[0079] In step S5, the mean square error MSE, root mean square error RMSE, and mean absolute error MAE are used to evaluate the performance of the synthetic electric field prediction model, and their formulas are respectively:

[0080]

[0081]

[0082]

[0083] where y i and represent the true value and predicted value of the synthetic electric field; n represents the number of predicted values and true values. The smaller the mean absolute error and root mean square error, the better the prediction effect of the algorithm.

[0084] Example:

[0085] Select the measured ground synthetic electric field data in May and November of a certain year for the ±800 kV UHV DC transmission project. After data cleaning, data normalization and other processing of the original data, use the present invention for training and prediction. For comparison, a single machine learning model and an improved Stacking model are used for training and prediction at the same time, and the evaluation of the ground synthetic electric field prediction results shown in Tables 1 and 2 is obtained.

[0086] Table 1 Comparison of prediction algorithm errors based on May data

[0087]

[0088] Table 2 Comparison of prediction algorithm errors based on November data

[0089]

[0090] The results in Table 1 and Table 2 show that when predicting the ground synthetic electric field, each single machine learning algorithm has its own advantages and disadvantages, while the accuracy of the Stacking prediction method is significantly improved compared with each machine learning model. The mean square error MSE, root mean square error RMSE, and mean absolute error MAE of the Stacking algorithm are all lower than those of the five single machine learning models. The performance of the Stacking prediction model is verified to be superior to that of each single machine learning model through the comparison of the prediction errors of the two groups of data. Therefore, the present invention selects the improved Stacking method for predicting the ground synthetic electric field, which has practical application value through practical tests.

Claims

1. Method for predicting ground synthetic electric field based on improved Stacking algorithm, characterized in that It includes the following steps: Step 1: Obtain sample data, including ground synthetic electric field data and meteorological data; Step 2: Preprocess the sample data obtained in Step 1, including outlier detection and normalization; Step 3: Use the five-fold cross-validation method to perform predictive modeling on the sample data that has been processed in Step 2 to obtain an improved Stacking algorithm prediction model; Step 4: Use the improved Stacking algorithm prediction model in Step 3 to predict the synthetic electric field; Step 3 includes the following steps: S3.1: After normalizing the initial data, the dataset G = {(y i , x i ); i = 1, 2, … M}, where x i is the i-th sample feature vector, and y i is the output value corresponding to the i-th sample; M is the number of collected synthetic electric field samples; S3.2: Divide the dataset G into a training set G train and a test set G test , where the number of samples in the training set is N, and the number of samples in the test set is M - N. Use five-fold cross-validation to train the base learners in the first layer on the training set G train : Divide G train into five equal parts to obtain G train1 , G train2 , G train3 , G train4 , G train5 , and use them as the test sets and training sets for different base learners respectively; Select one of the folds as the test set in turn, and the remaining four folds as the training set. During the five training processes, the base learner H k , and the prediction results of z kn , where k represents the serial number of the base learner in the first layer, and K represents the number of base learners in the first layer; The output results of the final K base learners and y in G train form a new training set used by the second-layer meta-learner: G n ={(y train-new , z n , …, z 1n , …, z kn ); n = 1, …, N, k = 1, …, K}; Where: n represents the sample serial number in the training set, and N represents the number of samples in the training set; y n represents the predicted values corresponding to the n samples; z 1n , …, z kn represents the new synthetic electric field data obtained by the K base learners according to the prediction of the N samples; S3.3: During the five-fold cross-validation process of S3.2, the base learner H will be trained once for each fold of cross-validation. k , and through the base learner H trained each time k , the y in the test set G test will be predicted. Finally, after the five-fold cross-validation is completed, the predicted value P i of G output by each base learner is: P test = (P kn + P kn + P kn1 + P kn2 + P kn3 + P kn4 + P kn5 ) / 5; Among them, P kn1 , P kn2 , P kn3 , P kn4 , P kn5 are the predicted values output by the k-th base learner after five different trainings for the test set G test , and P kn is the average value of the predicted values output by the k-th base learner for the test set G test , where n is the number of samples in the test set G test ; The results output by the final K base learners and G test in y i constitute a new test set used by the second-layer meta-learner: G test-new = {(y n , P 1n , …, P kn ); n = 1, …, M - N, k = 1, …, K}; P 1n , …, P kn represent the new synthetic electric field data obtained by K base learners based on the prediction of m - N samples; S3.4: Through the new training set G train-new Train the second-level meta-learner and use the new test set G test-new Detect the accuracy of the model.

2. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 1, wherein: It also includes Step 5: Use the mean square error MSE, root mean square error RMSE, and mean absolute error MAE to evaluate the performance of the improved Stacking algorithm prediction model.

3. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 1, wherein: In Step 1, the synthetic electric field data is the ground synthetic field strength collected under the UHV transmission line, unit: kV / m; the meteorological data includes temperature, humidity, air pressure, wind speed, PM1.0, PM2.5, and PM10.

4. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 1, characterized in that: In Step 2, the outlier detection method used is the Local Outlier Factor (LOF) algorithm. The LOF algorithm quantifies the degree of abnormality of each point in the dataset through the calculation of distance and density.

5. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 4, characterized in that: In the said step 2, the normalization method used is where: S is the result of normalizing each feature; s is the original data of each feature; S max and S min are the maximum and minimum values of each feature respectively.

6. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 1, wherein: In Step 4, for the improved Stacking algorithm prediction model, the first-layer base learners select five prediction algorithms with strong learning ability and significant structural differences, namely Random Forest (RF), Gradient Boosting Decision Tree (GBDT), LightGBM, Extreme Gradient Boosting (XGBoost), and K-Nearest Neighbor (KNN). The second-layer meta-learner selects a linear regression model for the final prediction. Specifically: The first-layer base learners make predictions based on the training set and test set divided from the measured synthetic electric field data. Specifically: The prediction results of the training set and the initial training set are combined as the training set of the second-layer meta-learner; the prediction results of the test set and the initial test set are combined as the test set of the second-layer meta-learner; The linear regression model is used to describe the stochastic linear relationship between variables y and x, i.e., y = β0 + β1x1 + β2x2 + … + β n x n + ξ, where x1, …, x k are non-stochastic variables, specifically the 7 meteorological factors of temperature, humidity, air pressure, wind speed, PM10, PM2.5, and PM1.0 here; β0, β1, …, β n are regression coefficients; ξ is a random error term; y is a random dependent variable, specifically the predicted value of the synthetic electric field here.

7. The method for predicting the ground synthetic electric field based on the improved Stacking algorithm according to claim 2, wherein: In the said step 5, the mean square error formula is as follows: The root mean square error formula is as follows: The mean absolute error formula is as follows: Where: y i and represent the true value and the predicted value; n represents the number of predicted values and true values; the smaller the mean square error, the root mean square error, and the mean absolute error are, the better the algorithm's prediction effect on the ground synthetic electric field is.