Method and System for Generating Hydroelectric Power Production Simulation Data Based on Overseas New Energy Consumption

通过训练和校验模型,生成了准确的水电生产模拟数据,解决了海外新能源项目中数据不足的问题,支持项目开发和投资决策。

CN119884760BActive Publication Date: 2025-07-11CEEC HUNAN ELECTRIC POWER DESIGN INST
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510360760.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the development of overseas new energy projects, the lack of power system data disclosure has led to the inability to accurately analyze the level of new energy consumption, especially the generation of hydropower data is relatively uncertain, which affects the calculation of installed capacity.

Method used

By collecting meteorological, load and hydropower output coefficient data, using the training set and test set to train multiple models, select the model with the smallest prediction error, calculate the Shapley value and distribution difference value of the feature, construct the data sets C and D, perform secondary training and prediction error verification, and generate hydropower production simulation data.

Benefits of technology

It provides high-accurate hydropower production simulation data, provides basic data support for overseas new energy project development and investment decisions, and solves the problems of difficult data acquisition and insufficient disclosure.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119884760B_ABST
    Figure CN119884760B_ABST
Patent Text Reader

Abstract

A method and system for generating hydropower production simulation data based on overseas new energy consumption. The method includes: 1. Collect meteorological, load, and hydropower output coefficient data of a certain place to obtain dataset A, train and test multiple models to obtain a selected model; 2. Collect meteorological data of the overseas prediction location to obtain dataset B excluding the hydropower output coefficient feature, and use the selected model to predict the hydropower output coefficient in dataset B; 3. Construct dataset C; 4. Exclude the data in dataset A that is the same as dataset C to obtain an alternative dataset, and then select a specified number of data from the alternative dataset to construct dataset D; 5. Use dataset C to perform secondary training on the selected model in 1, use the selected model after secondary training to calculate predictions for dataset D and calculate the prediction error, and then confirm whether the prediction error is correct through non-parametric hypothesis testing. The present invention provides decision-making support for the development and investment in overseas new energy markets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data generation, and particularly relates to a method and system for generating hydropower production simulation data based on overseas new energy consumption. Background Art

[0002] Domestic related enterprises that want to develop new energy projects overseas must analyze the consumption capacity of a country based on its existing power system and planned power system. The 87,600-hour data of each power source and load throughout the year will be the key data for calculating the consumption level and power adequacy. Among them, the data generation of wind power is highly correlated with wind speed, and that of photovoltaic power is highly correlated with irradiation amount. However, there are many influencing factors for hydropower data generation and the uncertainty is relatively large.

[0003] Domestic project developers who go abroad for project development need to generally know the consumption capacity of that country, and then calculate the installed capacity of the exploitable new energy in that country. However, in many cases, the disclosure degree of power system data abroad is not sufficient, so the consumption level analysis cannot be carried out. Summary of the Invention

[0004] The present invention provides a method and system for generating hydropower production simulation data based on overseas new energy consumption to solve the technical problems mentioned in the background art.

[0005] To achieve the above object, the technical solution of the present invention is realized as follows:

[0006] The present invention provides a method for generating hydropower production simulation data based on overseas new energy consumption, including the following steps:

[0007] S1. Collect meteorological, load, and hydropower output coefficient data of a certain place to obtain dataset A. Then divide dataset A into two parts and name them as the training set and the test set respectively. Use the training set and the test set to train and test multiple selected models respectively to obtain multiple test results. Evaluate the multiple test results and select the model with the smallest prediction error as the selected model;

[0008] S2. Calculate the marginal contribution of each feature in dataset A according to the predicted value of the selected model, and then calculate the Shapley value of each feature in dataset A based on the marginal contribution of each feature; collect the meteorological data of the overseas prediction location to obtain dataset B excluding the hydropower output coefficient feature, and calculate the Shapley value of each feature in dataset B; use the selected model to predict the hydropower output coefficient in dataset B, and calculate the distribution difference value of each feature in dataset B and dataset A;

[0009] S3. Select a set number of data from dataset A, and construct dataset C based on the Shapley values of each feature in dataset A and dataset B, and the distribution difference values of each feature in dataset A and dataset B;

[0010] S4. Remove the data in dataset A that is the same as dataset C to obtain an alternative dataset, then select a specified number of data from the alternative dataset, and construct dataset D based on the distribution difference values of each feature in dataset A and the Shapley values of each feature in dataset A;

[0011] S5. Use dataset C to perform secondary training on the selected model in S1 to obtain the selected model after secondary training; use the selected model after secondary training to calculate predictions for dataset D and calculate the prediction error, and then confirm whether the prediction error is correct through a non-parametric hypothesis testing method to quantitatively characterize the prediction accuracy of dataset A predicting dataset B according to the recurrence relationship of the distance.

[0012] Further, the specific steps of S1 are as follows:

[0013] S11. Collect the meteorology, load, and hydropower output coefficient of a certain place to obtain dataset A;

[0014] S12. Then divide dataset A into a training set and a test set according to a set ratio;

[0015] S13. On the basis of considering weekdays and holidays, use the training set to train multiple selected models. The selected models are linear regression models or non-linear regression models. After training, use the test set to perform tests to obtain multiple test results;

[0016] S14. Evaluate the multiple test results based on the root mean square error RMSE and the mean absolute percentage error MAPE, and then select the model with the smallest prediction error as the selected model ;

[0017] Among them, the calculation formulas for the root mean square error RMSE and the mean absolute percentage error MAPE are as follows:

[0018] ;

[0019] ;

[0020] Among them, is the number of samples, is the th observation value, is the th predicted value.

[0021] Further, S2 specifically includes the following steps:

[0022] S21. Input the dataset A into the selected model where represents a set of independent variable data in the dataset A input into the selected model ; and , represents a set of data containing k features, and the k features are respectively , , ,..., ;

[0023] S22. For the th set of data , calculate the marginal contribution , where S represents a subset of the dataset A; and the marginal contribution is specifically as follows:

[0024] ;

[0025] where represents the predicted value of the selected model after adding the feature to the subset S, represents the predicted value of the selected model when only using the features in the subset S;

[0026] S23. Obtain the Shapley value of each feature in the dataset A by weighted averaging the marginal contributions of each feature in the dataset A over all possible subsets ;

[0027] S24. Collect the meteorological data of the overseas prediction location, obtain the other data of the dataset B except for the hydropower output coefficient feature, and use the selected model to predict the hydropower output coefficient in the dataset B to obtain the dataset B containing all features;

[0028] S25. Use the same method as S21 to S23 to obtain the Shapley value of each feature in the dataset B;

[0029] S26. Bin the dataset A and the dataset B by using equal-width binning or equal-frequency binning; then calculate the proportion of each binned bin in the dataset A and the dataset B, so as to obtain the proportion of each binned bin in the dataset A and the proportion

[0030] S27. Calculate the distribution difference value of each feature in dataset A and dataset B based on the proportion of each bin in different datasets obtained from the previous calculation.

[0031] Further, the Shapley value of the i-th feature in S23 is calculated as follows:

[0032] ;

[0033] where represents the subset obtained by removing the group of data from N, N represents the set of all features, and ; represents the number of features in the subset , is the number of permutations of

[0034] Further, the calculation formula for the distribution difference value of each feature in S26 is as follows:

[0035] ;

[0036] where PSI is the distribution difference value of the feature.

[0037] Further, S3 specifically includes the following steps:

[0038] Select the feature data including all inputs to the selected model in dataset A to obtain multiple selected data;

[0039] S32. Construct constraint one and objective function one; Constraint one includes two, one of which is that the selected data in S31 are all within dataset A, and the other is that the selected data in S31 need to be greater than the set quantity;

[0040] Then, use the Shapley value of each feature in dataset A and dataset B and the distribution difference value of each feature in dataset A and dataset B to solve the distance between dataset B and dataset C;

[0041] S33. Use a heuristic algorithm to solve constraint one and objective function one, find the multiple data corresponding to the approximate optimal solution, and construct dataset C.

[0042] Further, the objective function one in S32 is specifically as follows:

[0043] ;

[0044] where Denote the th data in the pre - constructed dataset C, denote the th data in the dataset B, denote the distance between the dataset B and the dataset C.

[0045] Furthermore, the S4 specifically includes the following steps:

[0046] S41. Remove the data in dataset A that is the same as the dataset C to obtain an alternative dataset, and then select a specified number of multiple data from the alternative dataset to obtain multiple selected data;

[0047] S42. Construct the second constraint condition and the second objective function. The second constraint condition includes two parts. One is that the number of selected data needs to be greater than the specified number; the other is that the distance between the dataset composed of multiple selected data and the dataset C needs to be greater than or equal to the distance between the dataset C and the dataset B; the dataset composed of multiple selected data is the pre - constructed dataset D;

[0048] Then use the Shapley value of each feature in dataset A and the distribution difference value of each feature in dataset A to solve the distance between the dataset C and the pre - constructed dataset D;

[0049] The second objective function is specifically as follows:

[0050] ;

[0051] ;

[0052] where, denote the th data in the pre - constructed dataset D, denote the distance between the dataset C and the pre - constructed dataset D;

[0053] S43. Use a heuristic algorithm to solve the second constraint condition and the second objective function, find the multiple data corresponding to the approximate optimal solution, and construct the dataset D.

[0054] Furthermore, the S5 specifically includes the following steps:

[0055] S51. Use the dataset C to train the selected model in S1 to obtain the trained selected model;

[0056] S52. Then, under the constraint condition that the distance between the dataset C and the dataset D is greater than or equal to the distance between the dataset C and the dataset B, use the trained selected model to calculate the prediction error of the dataset D;

[0057] S53. Conduct a non-parametric hypothesis test on the prediction errors in the dataset D, and confirm whether the prediction errors in S52 are correct through the non-parametric hypothesis test, so as to quantitatively characterize the prediction accuracy of the dataset A for predicting the dataset B according to the recurrence relationship of the distances.

[0058] Among them, the recurrence relationship of the distances is the second constraint condition in S42 and the constraint condition in S52.

[0059] On the other hand, the present invention also provides a hydropower production simulation data generation system, including a computer device, which is programmed or configured to execute the above hydropower production simulation data generation method.

[0060] Advantages of the present invention:

[0061] The present invention discloses a hydropower production simulation data generation method based on overseas new energy consumption. In the case where the disclosure degree of overseas power data is insufficient and the data acquisition is difficult, the present invention can provide basic data for consumption production simulation for project developers, and provide an accuracy evaluation of the basic data, so as to provide decision-making support for the development and investment in overseas new energy markets. Description of the Drawings

[0062] Figure 1 It is a flowchart of the hydropower production simulation data generation method in the present invention. Detailed Embodiments

[0063] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Preferred embodiments of the present invention are shown in the drawings. However, the present invention can be implemented in many other different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the understanding of the disclosure of the present invention more thorough and comprehensive.

[0064] Refer to Figure 1 , the embodiment of the present application provides a hydropower production simulation data generation method based on overseas new energy consumption, including the following steps:

[0065] S1. Collect the meteorological, load, and hydropower output coefficients of a certain place to obtain the dataset A, then divide the dataset A into two parts, and name them the training set and the test set respectively. Use the training set and the test set to train and test a selected plurality of models respectively to obtain a plurality of test results, evaluate the plurality of test results, and select the model with the smallest prediction error as the selected model.

[0066] Preferably, the dataset A is the meteorological data of a certain province in the central part of the country, including rainfall, runoff, evaporation, 2m air temperature, dew point temperature, load, and 8760-time-point data of hydropower output.

[0067] S2. Calculate the marginal contribution of each feature in dataset A based on the predicted values of the selected model, and then calculate the Shapley value of each feature in dataset A according to the marginal contribution of each feature; collect the meteorological data of the overseas prediction location to obtain dataset B excluding the hydropower output coefficient feature, and calculate the Shapley value of each feature in dataset B; use the selected model to predict the hydropower output coefficient in dataset B, and calculate the distribution difference value of each feature in dataset B and dataset A.

[0068] Preferably, dataset B is the meteorological data of the overseas prediction location, including rainfall, runoff, evaporation, 2m air temperature, dew point temperature, load 8760 time point data, and the predicted hydropower output coefficient.

[0069] S3. Select a set number of data from dataset A, and construct dataset C based on the Shapley values of each feature in dataset A and dataset B and the distribution difference values of each feature in dataset A and dataset B.

[0070] S4. Eliminate the data in dataset A that is the same as dataset C to obtain an alternative dataset, then select a specified number of data from the alternative dataset, and construct dataset D based on the distribution difference value of each feature in dataset A and the Shapley value of each feature in dataset A.

[0071] S5. Use dataset C to perform secondary training on the selected model in S1 to obtain the selected model after secondary training; use the selected model after secondary training to calculate the prediction of dataset D and calculate the prediction error, and then confirm whether the prediction error is correct through the non-parametric hypothesis testing method, so as to quantitatively characterize the prediction accuracy of dataset A predicting dataset B according to the recurrence relationship of the distance.

[0072] The present invention discloses a method for generating hydropower production simulation data based on overseas new energy consumption. In the case where the disclosure degree of overseas power data is insufficient and the data acquisition difficulty is relatively large, the present invention can provide basic data for consumption production simulation for project developers and provide an accuracy evaluation of the basic data, providing decision-making support for the development and investment of overseas new energy markets.

[0073] In some embodiments, S1 specifically includes the following steps:

[0074] S11. Collect the meteorology, load, and hydropower output coefficient of a certain place to obtain dataset A.

[0075] S12. Then divide dataset A into a training set and a test set according to a set ratio.

[0076] S13. On the basis of considering working days and holidays, use the training set to train multiple selected models. The selected models are linear regression models or non-linear regression models. The linear regression models include linear neural networks or others; the non-linear regression models include decision trees, random forests, non-linear neural networks or others. During the training process, on the basis of features such as temperature, precipitation, evaporation, runoff, humidity, and load, the previous moment, the average value of the previous three moments, the average value of the previous six moments, the average value of the previous twelve moments, the value of the same moment on the past same day, the average value of the same moment on the past same day, the value of the same moment 7 days ago, and the average value of the same moment 7 days ago of each feature are used as extended features for research;

[0077] After training, use the test set for testing to obtain multiple test results;

[0078] S14. Evaluate the multiple test results according to the root mean square error RMSE and the mean absolute percentage error MAPE, and then select the model with the smallest prediction error as the selected model ;

[0079] Among them, the calculation formulas of the root mean square error RMSE and the mean absolute percentage error MAPE are as follows:

[0080] ;

[0081] ;

[0082] Among them, is the number of samples, is the th observed value, is the th predicted value.

[0083] In some embodiments, the said S2 specifically includes the following steps:

[0084] S21. Input the data set A into the selected model where represents a set of independent variable data in the data set A input into the selected model ; and , represents a set of data containing k features, and the k features are respectively , , ,..., ;

[0085] S22. For the th group of data , calculate the marginal contribution , where S represents a subset of the dataset A; among which the marginal contribution is as follows:

[0086] ;

[0087] Among which, represents the predicted value of the selected model after adding the feature to the subset S, and represents the predicted value of the selected model when only using the features in the subset S;

[0088] S23. Obtain the Shapley value of each feature in the dataset A by weighted averaging the marginal contributions of each feature in the dataset A over all possible subsets ;

[0089] S24. Collect the meteorological data of the overseas prediction locations, obtain the other data of the dataset B except for the hydropower output coefficient feature, and use the selected model to predict the hydropower output coefficient in the dataset B to obtain the dataset B containing all features;

[0090] S25. Use the same method as S21 to S23 to obtain the Shapley value of each feature in the dataset B;

[0091] S26. Bin the dataset A and the dataset B by equal-width binning or equal-frequency binning; then calculate the proportion of each binned bin in the dataset A and the dataset B, so as to obtain the proportion of each binned bin in the dataset A and the proportion in the dataset B ;

[0092] Among which, ; is the sample number of the th bin in the dataset A; N A is the total sample number of the dataset A;

[0093] And, , is the sample number of the th bin in the dataset B; is the total sample number of the dataset B;

[0094] S27. Calculate the distribution difference value of each feature in the dataset A and the dataset B according to the proportion of each bin in different datasets obtained previously; the distribution difference value can be the population stability index (PSI value), Kullback-Leibler divergence (KL divergence), earth mover's distance, etc.;

[0095] The distribution difference value of the features needs to specify the expected data set (i.e., data set B) as the benchmark, because the calculation of the Shap value is based on the expected data set, and the calculation of most distribution difference values is asymmetric.

[0096] In some embodiments, the Shapley value of each feature in the data set A in S23 is calculated as follows:

[0097] ;

[0098] where represents is the subset obtained by removing the th group of data from N, N represents the set of all features, and ; represents the number of features of the subset , is the number of all permutations of

[0099] In some embodiments, the calculation formula of the distribution difference value of each feature in S26 is specifically as follows:

[0100] ;

[0101] where PSI is the distribution difference value of the feature.

[0102] In some embodiments, S3 specifically includes the following steps:

[0103] S31. Select the feature data including all inputs to the selected model in the data set A to obtain a plurality of selected data;

[0104] S32. Construct constraint condition 1 and objective function 1; Constraint condition 1 includes two, one of which is that the selected data in S31 are all data in the data set A, and the other is that the selected data in S31 need to be greater than the set quantity, that is, the selected data in S31 need to be large enough;

[0105] Then, use the Shapley value of each feature in the data set A and the data set B and the distribution difference value of each feature in the data set A and the data set B to solve the distance between the data set B and the data set C;

[0106] S33. Use a heuristic algorithm to solve constraint condition 1 and objective function 1, optimize the solution speed, find the plurality of data corresponding to the approximate optimal solution, and construct the data set C. The heuristic algorithm here includes but is not limited to simulated annealing, Pareto, and particle swarm optimization.

[0107] In some embodiments, the first objective function in S32 is specifically as follows:

[0108] ;

[0109] Among them, represents the th data in the pre-constructed dataset C, represents the th data in the dataset B, represents the distance between the dataset B and the dataset C.

[0110] In some embodiments, S4 specifically includes the following steps:

[0111] S41. Remove the data in the dataset A that is the same as the dataset C to obtain an alternative dataset, and then select a specified number of multiple data from the alternative dataset to obtain multiple selected data;

[0112] S42. Construct the second constraint condition and the second objective function. There are two second constraint conditions. One is that the number of selected data needs to be greater than the specified number; that is, the number of selected data needs to be sufficient. The other is that the distance between the dataset composed of multiple selected data and the dataset C needs to be greater than or equal to the distance between the dataset C and the dataset B; the dataset composed of multiple selected data is the pre-constructed dataset D;

[0113] Then, use the Shapley value of each feature in the dataset A and the distribution difference value of each feature in the dataset A to solve the distance between the dataset C and the pre-constructed dataset D;

[0114] The second objective function is specifically as follows:

[0115] ;

[0116] ;

[0117] Among them, represents the th data in the pre-constructed dataset D, represents the distance between the dataset C and the pre-constructed dataset D;

[0118] S43. Use a heuristic algorithm to solve the second constraint condition and the second objective function, optimize the solution speed, find the multiple data corresponding to the approximate optimal solution, and construct the dataset D. The heuristic algorithm here includes but is not limited to simulated annealing, Pareto, particle swarm optimization, and genetic algorithm.

[0119] In some embodiments, S5 specifically includes the following steps:

[0120] S51. Train the selected model in S1 using dataset C to obtain the trained selected model;

[0121] S52. Then, under the constraint that the distance between dataset C and dataset D is greater than or equal to the distance between dataset C and dataset B, use the trained selected model to calculate the prediction error of dataset D;

[0122] S53. Conduct a non-parametric hypothesis test on the prediction error in dataset D, and confirm whether the prediction error in S52 is correct through the non-parametric hypothesis test, so as to quantitatively characterize the prediction accuracy of dataset A predicting dataset B according to the recurrence relationship of distances;

[0123] Among them, the recurrence relationship of distances is the second constraint in S42 and the constraint in S52.

[0124] The specific process of the non-parametric hypothesis test is as follows:

[0125] First, assume that the result obtained in S52 is correct and credible. Assume the significance level (α = 0.05), and then calculate the P value (probability value). The calculation of the P value adopts the form of a one-sided test:

[0126] One-sided test :

[0127] ;

[0128] ;

[0129] One-sided test :

[0130] ;

[0131] ;

[0132] Among them, is the population standard deviation, is the sample size, is the test sample mean, is the population mean, is the assumed population mean.

[0133] Compare the calculated P value with the pre-set significance level to obtain whether to accept the original hypothesis, and finally draw a conclusion.

[0134] For the sake of easy understanding, the following is a specific example for illustration:

[0135] Step S1. In this example, the dataset A uses the public data in the global climate atmospheric reanalysis dataset produced by the European Centre for Medium-Range Weather Forecasts (ECMWF), such as precipitation, evaporation, runoff, 2m maximum temperature, dew point temperature, 2m temperature, etc. In addition to the above public data, some derived features need to be calculated. For example, first, according to the above dew point temperature and 2m temperature, the Antoine equation or Wexler formula can be used to calculate the water vapor pressure and saturated water vapor pressure, so as to obtain the relative humidity. Secondly, combined with the load generation data, the above features and variables such as the known hydropower output coefficient of a certain province are used as basic features to conduct the hydropower generation simulation.

[0136] The relative humidity is calculated based on the water vapor pressure and saturated water vapor pressure, and the calculation formula is as follows:

[0137] Water vapor pressure (e):

[0138] ;

[0139] Saturated water vapor pressure ( ):

[0140] ;

[0141] Relative humidity (RH) (%):

[0142] ;

[0143] Among them, is the saturated water vapor pressure at 0°C, usually taking the value of 6.1078; T is the 2m temperature; is the dew point temperature; the values of a and b are related to the temperature:

[0144] When or , for the water surface: ;

[0145] When or , for the ice surface: .

[0146] The dependent variable and independent variable data of the known dataset A are shown in the following table.

[0147] Table 1: Simulation data of dataset A;

[0148]

[0149] In the process of model training and testing, the output coefficient is selected as the dependent variable. All feature data are divided into a training set and a testing set according to the ratio of 0.7:0.3, and models such as decision tree, random forest, neural network, and Lasso are used for training.

[0150] Predict the above model to obtain the model training and test results for a certain year. The selected model is Lasso, and the Lasso model is the selected model. Calculate its model evaluation value, and the results are shown in the following table:

[0151] Table 2: Evaluation indicators of the selected model training and test for dataset A;

[0152]

[0153] From the above table, it can be seen that the training and test model evaluation errors calculated by the selected model are relatively low, meeting the prediction accuracy requirements. The above model can be used for model migration of the prediction dataset.

[0154] Step S2: Calculate the feature contribution degree, distribution difference value, and hydropower output coefficient of the prediction dataset B;

[0155] Use the Shapley value additive explanation model (SHAP model) to calculate the Shapley value (Shapley value) of each feature in the expected dataset of the selected model obtained in step S1. The Shapley value can quantify the contribution degree of each feature to the model prediction.

[0156] Specifically, for each feature and each subset, calculate the marginal contribution . The Shapley value of the feature is obtained by weighted averaging of its marginal contributions in all possible subsets. The weights are determined according to the size of the subset and the number of permutations and combinations.

[0157] ;

[0158] where, , represents the number of elements in the subset , is the number of all permutations of

[0159] By calculation, the importance rankings of each feature are temperature, precipitation, evaporation, load, runoff, and humidity, and the Shapley contribution weights are 'precipitation': 0.0089,'maximum temperature': 0.0626, 'load': 0.000195, 'evaporation': 0.0257, 'runoff': 0.01288, 'humidity': 0.0369.

[0160] Using the equal-frequency binning method, each bin of each feature contains 200 pieces of data. The distribution difference values (PSI) of each feature between dataset A and dataset B are calculated as follows: precipitation: 13.86, maximum temperature: 1.97, load: 633, evaporation: 4.8, runoff: 9.58, humidity: 3.34.

[0161] And use the model selected by S1 to predict the hydropower output coefficient in dataset B.

[0162] The distance between datasets is the sum of the product of the distribution difference value of each feature and the Shapley value.

[0163] Step S3: Construct dataset C;

[0164] The format of dataset B is shown in the following table, which only contains independent variables and lacks the dependent variable to be predicted.

[0165] Table 3: Simulation data of dataset B;

[0166]

[0167] Use the simulated annealing algorithm to optimize the solution speed and find dataset C corresponding to the approximate optimal solution. The data in the above datasets can be regarded as a set of multiple particles.

[0168] ;

[0169] Among them, ;

[0170] ;

[0171] ;

[0172] Among them, Fitness(X) is the fitness value of each particle after each extended optimization, is the objective function value of the distance to the optimized target model (i.e., the selected model), is the penalty coefficient, T is the current annealing temperature, is the penalty function, is the piecewise function, is the conversion function of multiple constraint functions, represents the penalty exponent when exceeding the threshold.

[0173] Use the simulated annealing criterion to set the end condition of the extended optimization. The specific end condition of the extended optimization is:

[0174] ;

[0175] wherein, is the probability that the -th particle accepts the value after the n-th iteration, is the current annealing temperature, is the value after the n-th iteration, is the value after the (n - 1)-th iteration, is the fitness value corresponding to the (n - 1)-th iteration, is the fitness value corresponding to the n-th iteration.

[0176] The calculation result is that 1035 pieces of data are selected from the dataset A to form the dataset C, and the obtained dataset C is as follows:

[0177] Table 4: Dataset C;

[0178]

[0179] Step S4, construct the dataset D;

[0180] According to the optimization objective of the minimum distance between the dataset C and the dataset D, the dataset D obtained by using the simulated annealing algorithm, the optimization solution method is the same as that of the dataset C. According to the calculation result, a total of 100 pieces of data are selected, as shown in the following table.

[0181] Table 5: Dataset D;

[0182]

[0183] Step S5, model inference error, and prediction of the dataset B;

[0184] Step S51, train the dataset C using the selected model obtained in Step 1, and then predict the dataset D. Under the constraint condition that the distance between the dataset C and the dataset D is greater than or equal to the distance between the dataset C and the dataset B, calculate the prediction error of the dataset D. The RMSE and MAPE are 11.93% and 32.9% respectively. According to the fact that the distance between C and D is greater than the distance between C and B, it can be obtained that the error of C predicting B is lower than the error of C predicting D, that is, the error of A predicting B is lower than the model evaluation value of C predicting D.

[0185] Step S52, conduct a non-parametric hypothesis test on the prediction error of the dataset D. Assume that the conclusion obtained in the first point is correct and credible. Without presupposing the distribution types of the features in the dataset B and C, the calculated p-values are all greater than 0.1. Therefore, whether the assumed significance level is 0.05 or 0.01, the p-values are all greater than the significance level. Thus, it can be obtained that the hypothesis of the non-parametric hypothesis test is credible, that is, the regression model trained and optimized based on the dataset A, trained on the dataset C, and the error of predicting the dataset B is close to the error of predicting the dataset D.

[0186] On the other hand, the present invention also provides a hydropower production simulation data generation system, including a computer device programmed or configured to execute the above-mentioned hydropower production simulation data generation method.

[0187] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, and all of them should be covered by the protection scope of the present invention. Moreover, the technical solutions between various embodiments of the present invention can be combined with each other, but it must be based on the fact that those skilled in the art can implement them. When the combination of technical solutions appears to be contradictory or unable to be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for generating hydroelectric power production simulation data based on overseas new energy consumption, characterized in that The steps include: S1. Collect the weather, load and hydropower output coefficient of a certain place to obtain a data set A, then divide the data set A into two parts, and name them as a training set and a test set respectively, use the training set and the test set to train and test the selected multiple models respectively, obtain multiple test results, evaluate the multiple test results, and select the model with the smallest prediction error as the selected model; S2. Calculate the marginal contribution of each feature in data set A according to the predicted value of the selected model, and then calculate the Shapley value of each feature in data set A according to the marginal contribution of each feature; collect meteorological data of overseas prediction sites, obtain data set B except the hydropower output coefficient feature, and calculate the Shapley value of each feature in data set B; and use the selected model to predict the hydropower output coefficient in data set B, and calculate the distribution difference value of each feature in data set B and data set A; S3, selecting a set amount of data from data set A, and constructing data set C based on the Shapley value of each feature in data set A and data set B, and the distribution difference value of each feature in data set A and data set B; The construction of data set C must satisfy constraint condition 1 and objective function 1. Constraint condition 1 includes two conditions: one is that the selected data must be in data set A, and the other is that the selected data must be greater than the set number; Objective function 1 is to minimize the distance between data set C and data set B; S4. Eliminate the data identical to data set C in data set A to obtain an alternative data set, then select a specified number of data from the alternative data set, and construct data set D based on the distribution difference value of each feature in data set A and the Shapley value of each feature in data set A; the construction of data set D must satisfy constraint condition 2 and objective function 2. Constraint condition 2 includes two, one of which is that the number of selected data must be greater than the specified number; the other is that the distance between the pre-constructed data set D and data set C must be greater than or equal to the distance between data set C and data set B; objective function 2 is that the distance between data set C and data set D is minimized; S5. Use the data set C to perform secondary training on the selected model in S1 to obtain the selected model after secondary training; use the selected model after secondary training to calculate the prediction of data set D and calculate the prediction error, and then confirm whether the prediction error is correct through the non-parametric hypothesis testing method, so as to quantitatively characterize the prediction accuracy of data set A for predicting data set B according to the recursive relationship of distance.

2. The method for generating hydropower production simulation data according to claim 1, wherein The S1 specifically includes the following steps: S11, collecting the weather, load and hydropower output coefficient of a certain place to obtain a data set A; S12, then divide the data set A into a training set and a test set according to a set ratio; S13. Based on the consideration of working days and holidays, the selected multiple models are trained using the training set, where the selected model is a linear regression model or a nonlinear regression model; After training, the test set is used for testing to obtain multiple test results; S14. Evaluate multiple test results based on the root mean square error (RMSE) and the mean absolute percentage error (MAPE), and then select the model with the smallest prediction error as the selected model. ; Among them, the calculation formulas of root mean square error RMSE and mean absolute percentage error MAPE are as follows: ; ; Among them, is the number of samples, is the th observation value, is the th predicted value.

3. The method for generating hydropower production simulation data according to claim 1, characterized in that, The S2 specifically includes the following steps: S21. Input dataset A into the selected model wherein represents a set of independent variable data in dataset A input into the selected model ; and , represents a set of data containing k features, and the k features are respectively , , ,..., ; S22. For the group of data , calculate the marginal contribution , where S represents a subset of dataset A; the marginal contribution is specifically as follows: ; Among them, represents the selected model after adding the feature to the subset S of the predicted value, represents the predicted value of the selected model when only using the features in the subset S ; S23. Obtain the Shapley value of each feature in dataset A by taking the weighted average of the marginal contributions of each feature in dataset A across all possible subsets ; S24. Collect meteorological data of overseas prediction locations to obtain other data in dataset B except for the hydropower output coefficient feature, and use the selected model to predict the hydropower output coefficient in dataset B to obtain dataset B containing all features. S25. Use the same method as in S21 to S23 to obtain the Shapley value of each feature in dataset B. S26. Bin the dataset A and the dataset B by using equal-width binning or equal-frequency binning; then calculate the proportion of each binned box in the dataset A and the dataset B, so as to obtain the proportion of each binned box in the dataset A and the proportion in the dataset B ; S27. Calculate the distribution difference value of each feature in dataset A and dataset B according to the proportion of each bin obtained from the previous calculation in different datasets.

4. The method for generating hydroelectric production simulation data according to claim 3, wherein, The Shapley value of each feature in the dataset A in S23 The specific calculation formula is as follows: ; denote is the subset of N after removing the group of data , where N represents the set of all features, and ; denotes the number of features of the subset , and is the number of all permutations of elements.

5. The method for generating hydropower production simulation data according to claim 4, wherein The specific calculation formula for the distribution difference value of each feature in S26 is as follows: ; where PSI is the distribution difference value of the feature.

6. The method for generating hydroelectric production simulation data according to claim 5, wherein S3 specifically includes the following steps: S31. Select, from dataset A, the feature data that includes all the data input into the selected model to obtain multiple selected data. S32. Construct constraint condition 1 and objective function 1. Then, use the Shapley value of each feature in dataset A and dataset B and the distribution difference value of each feature in dataset A and dataset B to solve for the distance between dataset B and dataset C. S33. Use a heuristic algorithm to solve constraint condition 1 and objective function 1 to find multiple data corresponding to the approximate optimal solution and construct dataset C.

7. The method for generating hydroelectric power production simulation data according to claim 6, wherein The specific objective function 1 in S32 is as follows: ; Among them, represents the th data in the pre-built dataset C, represents the th data in the dataset B, represents the distance between the dataset B and the dataset C.

8. The method for generating hydropower production simulation data according to claim 7, wherein S4 specifically includes the following steps: S41. Remove the data in dataset A that is the same as dataset C to obtain an alternative dataset, and then select multiple data of a specified quantity from the alternative dataset to obtain multiple selected data. S42. Construct constraint condition 2 and objective function 2. The dataset composed of multiple selected data is the pre-constructed dataset D. Then, use the Shapley value of each feature in dataset A and the distribution difference value of each feature in dataset A to solve the distance between dataset C and the pre-constructed dataset D; The specific objective function 2 is as follows: ; ; Among them, represents the th data in the pre-built dataset D, represents the distance between the dataset C and the pre-built dataset D; S43. Use a heuristic algorithm to solve constraint condition 2 and objective function 2 to find multiple data corresponding to the approximate optimal solution and construct dataset D.

9. The method for generating hydropower production simulation data according to claim 1, characterized in that S5 specifically includes the following steps: S51. Use dataset C to train the selected model in S1 to obtain the trained selected model. S52. Then, under the constraint condition that the distance between dataset C and dataset D is greater than or equal to the distance between dataset C and dataset B, use the trained selected model to calculate the prediction error of dataset D. S53. Conduct a non-parametric hypothesis test on the prediction error in dataset D and confirm whether the prediction error in S52 is correct through the non-parametric hypothesis test, so as to quantitatively characterize the prediction accuracy of dataset A predicting dataset B according to the recurrence relationship of the distance. Among them, the recurrence relationship of the distance is the constraint condition 2 in S42 and the constraint condition in S52.

10. A hydroelectric power production simulation data generation system, including a computer device, characterized in that, The computer device is programmed or configured to execute the hydropower production simulation data generation method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Information importance cognitive defense method based on interpretable learning in automatic modulation identification

    CN119416030A