Transformer area load prediction method and device based on LGBM algorithm
By obtaining a multidimensional data set and using the mutual information method and LGBM algorithm to predict the load in the substation, the accuracy and efficiency problems of load prediction in the distributed photovoltaic power generation system are solved, and a more efficient load forecasting effect is achieved.
Patent Information
- Application Number
- CN202510688454.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-12
AI Technical Summary
Traditional substation load forecasting methods are difficult to accurately predict load changes in distributed photovoltaic power generation systems, especially small photovoltaic systems. Due to large data errors, the prediction accuracy decreases and cannot adapt to changes in load characteristics.
A load forecasting method based on the LGBM algorithm is adopted. By obtaining a multidimensional data set including historical load data, distributed photovoltaic power generation data and meteorological characteristic data, and using the mutual information method for feature selection, a load forecasting model based on the LGBM algorithm is constructed, and the feature subset associated with the substation load is screened out for prediction.
The accuracy and efficiency of substation load forecasting are improved, the data dimension and calculation amount are reduced, the overfitting problem is avoided, and the generalization ability and prediction performance of the model are enhanced.
Smart Images

Figure CN120634296A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of load forecasting, and in particular to a method and device for predicting load in a substation based on the LGBM algorithm. Background Art
[0002] A distributed photovoltaic system is a power generation system that installs solar photovoltaic modules in various locations, such as building rooftops, industrial and commercial plants, and agricultural facilities. Inverters convert direct current (DC) electricity into alternating current (AC) and directly connect it to the local power grid. It offers advantages such as flexible deployment and local consumption, effectively improving energy efficiency.
[0003] In the power system sector, traditional methods for predicting load at substations rely primarily on historical load data, employing statistical analysis and time series models. When load conditions are relatively stable, some traditional methods can produce reasonably reasonable forecasts. However, distributed photovoltaic power generation is decentralized, intermittent, and random. Light conditions and photovoltaic equipment installation vary widely across regions, leading to significant fluctuations in generated power and making accurate forecasts difficult. Small photovoltaic systems, in particular, suffer from significant errors and poor reliability in collected data due to equipment cost constraints, complex installation environments, and relatively weak monitoring technologies. This inaccurate data further interferes with accurate assessment of substation load changes, making them extremely complex. Traditional forecasting methods struggle to adapt to these new load characteristics, resulting in a significant decline in forecast accuracy. Summary of the Invention
[0004] In response to the above situation, the embodiments of the present application provide a substation load forecasting method and device based on the LGBM algorithm, which aims to solve the above problem or at least partially solve the above problem.
[0005] In a first aspect, an embodiment of the present application provides a method for predicting load in a substation area based on the LGBM algorithm, the method comprising:
[0006] Obtaining a target multidimensional dataset, wherein the target multidimensional dataset includes historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data;
[0007] Performing feature selection on the target multidimensional data set based on a mutual information method to determine a feature subset associated with the substation load;
[0008] The feature subset is input into a load forecasting model based on the LGBM algorithm to obtain a load forecasting result of the substation predicted by the load forecasting model.
[0009] In a second aspect, an embodiment of the present application further provides a device for predicting load in a substation area based on the LGBM algorithm, comprising:
[0010] An acquisition module is used to acquire a target multidimensional data set, wherein the target multidimensional data set includes historical load data of the substation area, distributed photovoltaic power generation data and meteorological characteristic data;
[0011] A feature selection module, configured to perform feature selection on the target multidimensional data set based on a mutual information method, and determine a feature subset associated with the substation load;
[0012] The prediction module is used to input the feature subset into the load prediction model based on the LGBM algorithm to obtain the load prediction result of the substation predicted by the load prediction model.
[0013] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, cause the processor to perform the steps of the first aspect described above.
[0014] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple applications, the electronic device executes the steps of the first aspect above.
[0015] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects: by obtaining the historical load data, distributed photovoltaic power generation data and meteorological characteristic data of the substation, and establishing a multidimensional data set, it comprehensively covers the various key factors affecting the substation load, more accurately reflects the actual situation of the substation load, provides a reliable data basis for subsequent forecast analysis, and improves the accuracy of the forecast. Further, the mutual information method is used to select features of the multidimensional data set, which can accurately filter out the feature subsets closely related to the substation load forecast from the numerous data features, retain the most valuable information for load forecasting, remove redundant features, not only reduce the data dimension, reduce the complexity and computational complexity of model training, but also effectively avoid overfitting problems, and significantly improve the prediction performance and generalization ability of the model. Finally, since LGBM is a distributed gradient boosting framework based on the decision tree algorithm, it has the advantages of high efficiency, parallelization, accuracy, etc., so the substation load forecast result obtained by using the load forecasting model based on the LGBM algorithm to predict the substation load is more accurate, and the prediction efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0017] Figure 1A schematic diagram of a flow chart of a method for predicting load in an area based on the LGBM algorithm provided in an embodiment of the present application is shown;
[0018] Figure 2 A schematic diagram showing a flow chart of a method for obtaining a multidimensional dataset provided in an embodiment of the present application is shown;
[0019] Figure 3 A schematic diagram showing a flow chart of a method for acquiring meteorological characteristic data for a weather forecast provided by an embodiment of the present application is shown;
[0020] Figure 4 A flow chart of a method for determining the credibility of a data set provided in an embodiment of the present application is shown;
[0021] Figure 5 A schematic diagram showing a flow chart of a method for determining a feature subset provided in an embodiment of the present application is shown;
[0022] Figure 6 A schematic diagram of a process for constructing a prediction result correction model provided in an embodiment of the present application is shown;
[0023] Figure 7 The structure diagram of the load prediction device based on the LGBM algorithm provided in the embodiment of the present application is shown;
[0024] Figure 8 A schematic structural diagram of an electronic device provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0026] It should be noted that the terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such usage is interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the term "including" and its variations are to be interpreted as open-ended terms meaning "including but not limited to."
[0027] Figure 1 The flow chart of the load forecasting method based on the LGBM algorithm provided in the embodiment of the present application is shown. Figure 1It can be seen that this application at least includes steps S101 to S103:
[0028] Step S101: Obtain a target multidimensional dataset, which includes historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data.
[0029] In this embodiment, historical load data, distributed photovoltaic power generation data, and meteorological characteristic data are collected to reflect the past distributed photovoltaic power generation of the load in the area, as well as the meteorological conditions that affect the load and photovoltaic power generation. These different types of data are integrated together to create a multidimensional dataset, providing a comprehensive data foundation for subsequent analysis and model training.
[0030] Step S102: performing feature selection on the target multidimensional dataset based on the mutual information method to determine a feature subset.
[0031] The mutual information method is used to select features from a multidimensional training dataset, filtering out a subset of features relevant to substation load forecasting. The mutual information method measures the degree of dependence between features and the target variable (substation load). The selected feature subset retains the most valuable information for load forecasting, helping to improve the model's predictive performance while reducing interference from redundant features.
[0032] Step S103: Input the feature subset into the load forecasting model based on the LGBM algorithm to obtain the load forecasting result of the substation predicted by the forecasting model.
[0033] A load forecasting model based on the LGBM algorithm is constructed and the feature subset obtained through feature selection is input into the model. By analyzing and processing the input feature subset, the load forecasting model based on the LGBM algorithm can generate substation load forecast results. LGBM (Light Gradient Boosting Machine) is a distributed gradient boosting framework based on the decision tree algorithm. It supports efficient parallel training and has the advantages of faster training speed, lower memory consumption, better accuracy, and distributed support for rapid processing of massive data.
[0034] from Figure 1It can be seen from the method shown that this application obtains the historical load data, distributed photovoltaic power generation data and meteorological characteristic data of the substation and establishes a multidimensional data set, which comprehensively covers various key factors affecting the substation load, more accurately reflects the actual situation of the substation load, provides a reliable data basis for subsequent forecast analysis, and improves the accuracy of the forecast. The mutual information method is further used to perform feature selection on the multidimensional data set, which can accurately filter out the feature subsets closely related to the substation load forecast from many data features, retain the most valuable information for load forecasting, remove redundant features, not only reduce the data dimension, reduce the complexity and computational complexity of model training, but also effectively avoid overfitting problems, and significantly improve the prediction performance and generalization ability of the model. Finally, since LGBM is a distributed gradient improvement framework based on the decision tree algorithm, it has the advantages of high efficiency, parallelization, accuracy, etc., so the substation load forecast results obtained by using the load forecast model based on the LGBM algorithm for substation load forecasting are more accurate, and the forecast efficiency is improved.
[0035] In some embodiments of the present application, in the above method, in step S101, the target multidimensional dataset is generated based on the initial multidimensional dataset, such as Figure 2 As shown, a method for obtaining a multidimensional dataset is provided, comprising the following steps:
[0036] Step S201: obtaining an initial multidimensional data set, which includes historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data of the distributed photovoltaic system.
[0037] Among them, meteorological characteristic data include irradiance, meteorological conditions, solar position and angle, solar altitude angle, altitude, degree of cloud cover, etc.
[0038] Step S202: performing data cleaning and preprocessing on the initial multidimensional dataset.
[0039] In some embodiments, data cleaning and preprocessing are performed on the initial multidimensional dataset, including filling missing values, correcting outliers, normalizing data, and performing feature correlation analysis on the multidimensional dataset.
[0040] Specifically, missing values for numerical data are filled using statistics such as the mean, median, and mode. For missing values in categorical data, reasonable inferences can be made based on the categories of adjacent or related data. Missing value filling ensures the integrity of the dataset and reduces the impact of missing data on subsequent analysis.
[0041] Identify and correct outliers in multidimensional data sets. Detect outliers through statistical analysis methods (such as standard deviation-based methods, which consider data that is more than a certain standard deviation away from the mean as outliers) or machine learning-based methods (such as the isolation forest algorithm). Then, correct the outliers according to the specific situation, such as replacing them with reasonable values or performing smoothing, to make the data set more robust.
[0042] Data of different dimensions may have different dimensions and value ranges, which can affect the convergence speed of machine learning algorithms and model performance. Data standardization, by converting data to the same scale and using normalization or standardization methods, ensures that different features have equal weight and influence in model training, helping to improve model training effectiveness and prediction accuracy.
[0043] By calculating the correlation coefficient between features (such as the Pearson correlation coefficient), we can understand the degree of association between different features. This helps to identify redundant features or collinearity issues in the dataset. For features with excessive correlation, we can select or process them according to the actual situation. For example, we can retain features that have a greater impact on the target variable (station load) and remove redundant features to reduce data dimensions, improve model training efficiency and stability, and avoid overfitting.
[0044] Through the above data cleaning and preprocessing steps, the quality of the multidimensional dataset is improved, providing a good data foundation for subsequent feature selection and model building.
[0045] Step S203: Determine the credibility of the initial multidimensional dataset after data cleaning and preprocessing.
[0046] Step S204: Determine the target multidimensional dataset based on the credibility of the initial multidimensional dataset.
[0047] In some embodiments, whether the initial multidimensional data has predictive value is determined by determining the credibility of the initial multidimensional data, for example, by setting a credibility threshold. When the credibility of the initial multidimensional data is greater than or equal to the credibility threshold, it indicates that the initial multidimensional data set has predictive value. When the credibility of the initial multidimensional data is less than the credibility threshold, it indicates that the initial multidimensional data set does not have predictive value.
[0048] In one embodiment, when the credibility of the initial multidimensional dataset indicates that it has predictive value, the initial multidimensional dataset is used as the target multidimensional dataset. Specifically, the target multidimensional dataset includes historical load data for the substation, distributed photovoltaic power generation data, and meteorological characteristic data for the distributed photovoltaic system. In this embodiment, if the credibility of the initial multidimensional dataset indicates that it has predictive value—that is, the data quality is good and can provide reliable support for prediction—then the multidimensional dataset is used in the subsequent feature selection steps, laying the foundation for accurate load forecasting results.
[0049] In another embodiment, when the credibility indicates that the initial multidimensional dataset does not have predictive value, meteorological characteristic data from weather forecasts are obtained and used as supplementary data for the initial multidimensional dataset to generate a target multidimensional dataset. That is, the target multidimensional dataset includes the historical load data of the substation, distributed photovoltaic power generation data, meteorological characteristic data of the distributed photovoltaic system, and meteorological characteristic data from weather forecasts. In this embodiment, when the credibility of the multidimensional dataset does not have predictive value, it indicates that the current dataset may be insufficient. In this case, meteorological characteristic data from weather forecasts are obtained as supplementary data for the multidimensional dataset. By supplementing the new data, the quality and predictive value of the dataset are improved to meet the needs of subsequent load forecasting.
[0050] In an embodiment of the present application, the meteorological characteristic data of the distributed photovoltaic system is obtained through large-scale centralized photovoltaic system layout detection equipment, which can directly reflect the actual meteorological conditions at the location of the photovoltaic system. The meteorological characteristic data of the weather forecast is obtained through a system or API related to the weather forecast, and can be used as supplementary data for some small photovoltaic systems that lack data monitoring equipment, or in scenarios where it is necessary to consider the impact of more macroscopic meteorological conditions on the load, the weather forecast data can supplement the relevant information. Through this flexible method of obtaining meteorological characteristic data, the impact of meteorological factors on the load of the substation can be considered more comprehensively and accurately, providing richer and more reliable data support for subsequent load forecasting.
[0051] In some embodiments of the present application, a method for obtaining meteorological characteristic data of a weather forecast is also provided, such as Figure 3 As shown, the following steps are included:
[0052] Step S301: Obtain a weather forecast map and obtain location information of a distributed photovoltaic system;
[0053] Step S302: steplessly zooming in on the weather forecast map and gridding the zoomed weather forecast map;
[0054] Step S303: Acquire the deployment information of the photovoltaic system, and set the irradiation grid corresponding to the photovoltaic system according to the deployment information;
[0055] Step S304: Acquire meteorological characteristic data of the irradiation grid.
[0056] Specifically, the weather forecast map can be used to query the meteorological characteristic data of the location of the distributed photovoltaic system. The weather forecast map can be zoomed in steplessly, and the zoomed map can be gridded. Stepless zooming can obtain more detailed map information, while gridding divides the map into small grid cells, which facilitates subsequent queries on the location data of the distributed photovoltaic system. The deployment information of the photovoltaic system is obtained, and the corresponding irradiation grid of the photovoltaic system is set based on this information. The deployment information of the photovoltaic system includes factors such as its installation location, orientation, and area, which determine the situation of the photovoltaic system receiving solar radiation. By setting the corresponding irradiation grid, the meteorological area directly related to the photovoltaic system can be more accurately determined, thereby obtaining the meteorological characteristic data of the area.
[0057] For example, based on the GPS coordinates of the photovoltaic system, the target area is located on the weather forecast map at a 0.1°×0.1° grid (about 11km resolution), and a 5-fold stepless magnification is performed to obtain a 1km×1km high-precision meteorological grid.
[0058] Calculate the effective light receiving area according to the inclination and azimuth of the photovoltaic panel:
[0059] G eff =G 水平 cos(θ-α) / cosθ
[0060] Among them, G eff Indicates the effective light receiving area, G 水平 It represents the solar radiation intensity on the horizontal plane, θ is the solar altitude angle, and α is the inclination angle of the photovoltaic panel.
[0061] Furthermore, the location of the specific irradiation grid corresponding to the photovoltaic system in the gridded weather forecast map is determined in combination with the effective light-receiving area and the installation location information of the photovoltaic system. Due to the differences in installation locations, orientations, areas and other factors of different photovoltaic systems, the corresponding irradiation grids will also be different. After determining the irradiation grid, the meteorological characteristic data contained in the irradiation grid is extracted from the gridded weather forecast map, which usually includes but is not limited to information such as temperature, humidity, wind speed, wind direction, and light intensity. Among them, the temperature will affect the performance of photovoltaic cells, and thus affect the photovoltaic power generation. Higher temperatures may cause the efficiency of photovoltaic cells to decrease; humidity has a certain impact on the electrical performance and stability of the photovoltaic system. Excessive humidity may increase the risk of equipment damage due to moisture; wind speed and wind direction will affect the heat dissipation of photovoltaic modules and the stress of mechanical structures; light intensity directly determines the power generation of the photovoltaic system. The stronger the light intensity, the higher the photovoltaic power generation.
[0062] In an embodiment of the present application, by obtaining meteorological characteristic data from weather forecasts, it can be used as supplementary data for some small photovoltaic systems that lack data monitoring equipment, or in scenarios where the impact of more macroscopic meteorological conditions on load needs to be considered, weather forecast data can supplement relevant information and more comprehensively and accurately consider the impact of meteorological factors on substation load, providing richer and more reliable data support for subsequent load forecasts.
[0063] In some embodiments of the present application, a method for determining the credibility of a data set is also provided, such as Figure 4 As shown, the following process is included:
[0064] Step S401: perform basic credibility verification on the initial multidimensional dataset.
[0065] Step S402: performing prediction credibility verification on the initial multidimensional dataset that has passed the basic credibility verification.
[0066] Step S403: Determine the credibility of the multidimensional dataset based on the verification result of the prediction credibility verification.
[0067] Specifically, we conduct a preliminary assessment of dataset quality based on fundamental aspects such as data integrity and consistency, conducting basic credibility verification on multidimensional datasets. For example, we check whether there are large numbers of missing values, whether the data format meets requirements, and whether the logical relationships between data in different dimensions are reasonable. These checks provide a preliminary assessment of the dataset's fundamental quality and allow us to exclude data sets with obvious problems.
[0068] Verify the prediction credibility of multidimensional datasets that have passed basic credibility verification. Using preliminary prediction models or statistical methods, perform simple predictions based on the current multidimensional dataset and compare the prediction results with some known actual conditions. For example, you can use some simple linear regression models to perform preliminary predictions and observe the degree of deviation between the predicted values and the actual values. This helps further evaluate whether the dataset can effectively support accurate load forecasting from the perspective of prediction effectiveness.
[0069] The credibility of the multidimensional dataset is determined based on the prediction credibility verification results. If the verification results indicate that the dataset has a high credibility, meaning that the data quality is good and can provide reliable support for prediction, then the multidimensional dataset will be used in the subsequent feature selection steps, laying the foundation for building an accurate load forecasting model. Conversely, if the credibility of the multidimensional dataset does not provide predictive value, it indicates that the current dataset may be insufficient. In this case, meteorological characteristic data from weather forecasts will be obtained as supplementary data for the multidimensional dataset. By supplementing this new data, the quality and predictive value of the dataset will be improved to meet the needs of subsequent load forecasting.
[0070] In some embodiments of the present application, a method for determining a feature subset is also provided, such as Figure 5 As shown, the following process is included:
[0071] Step S501: Calculate the mutual information value between each feature in the target multidimensional dataset and the substation load.
[0072] Step S502: taking features whose mutual information values are greater than the mutual information threshold as a feature subset.
[0073] Specifically, the target multidimensional dataset is discretized. Many machine learning algorithms perform better when processing discrete data, and the mutual information method is more direct and efficient when calculating discrete data. Continuous features are converted into discrete features using appropriate discretization methods (such as equal-width binning or equal-frequency binning). This divides the data's range into several intervals, each corresponding to a discrete value, facilitating subsequent mutual information calculations.
[0074] Mutual information (MI) is a measure of the degree of dependence between two random variables. In feature selection, the importance of each feature to load forecasting is assessed by calculating the MI between each feature and the target variable (substation load). After calculating the MI values for all features, a MI threshold is set. This threshold is determined based on practical experience and serves as a criterion for feature selection. Features with MI values greater than the threshold are retained, and these features form the feature subset used for load forecasting.
[0075] In the embodiment of the present application, a feature selection method based on the mutual information method is used to select the most representative and predictive features from a multidimensional data set, thereby reducing the interference of irrelevant or redundant features on the model and improving the efficiency and accuracy of the model.
[0076] In some embodiments of the present application, after the substation load prediction result is obtained in step S103, in order to improve the accuracy of the prediction result, a prediction result correction model is further constructed to correct the substation load result and obtain the substation load correction result.
[0077] The following provides a method for constructing a prediction result correction model, such as Figure 6 As shown, the following steps are included:
[0078] Step S601: Acquire a multidimensional data set, a substation load forecast result, and an actual substation load.
[0079] Step S602: Calculate the residual based on the load forecast result and the actual load of the substation;
[0080] Step S603: construct an initial prediction result correction model, and train the prediction result correction model using the multidimensional data set and residuals.
[0081] Specifically, the multidimensional dataset contains information on various factors affecting the substation load, such as historical load data, distributed photovoltaic power generation data, and meteorological characteristics. The substation load forecast is a preliminary prediction obtained using a load forecasting model based on the LGBM algorithm. The actual substation load is the actual load data of the photovoltaic system.
[0082] Residuals are calculated based on the load forecast results and the actual load at the substation. Residuals are the difference between the predicted and actual values, quantifying the degree of deviation between the load forecast model's predictions and actual conditions. Residuals reflect the model's prediction error based on the current data. Analyzing residuals can help identify model deficiencies and provide guidance for subsequent corrections.
[0083] Furthermore, an initial prediction result correction model is constructed, and the initial prediction result correction model can be constructed using an LGBM decision tree. The initial prediction result correction model is trained using a multidimensional data set and its corresponding residuals. During the training process, the correction model learns the relationship between the features and residuals in the multidimensional data set, and uses this information to optimize the load forecast results of the substation. For example, the prediction result correction model compares the prediction results obtained under the combination of meteorological conditions and load historical data with the actual load of the photovoltaic system to see if the prediction value of the LGBM decision tree is too high or too low, and thus makes corresponding adjustments in subsequent predictions, so that the corrected substation load forecast results are more reliable, reducing the error between the prediction results and the actual load, thereby improving the accuracy and practicality of the entire load forecasting method.
[0084] In some embodiments of the present application, a device for predicting loads in a substation based on the LGBM algorithm is provided. The device for predicting loads in a substation based on the LGBM algorithm corresponds one-to-one to the method for predicting loads in a substation based on the LGBM algorithm in the above embodiment. Figure 7 As shown, the substation load prediction device based on the LGBM algorithm includes an acquisition module 101, a feature selection module 102, a prediction module 103 and a correction module 104.
[0085] An acquisition module 101 is configured to acquire a target multidimensional dataset, wherein the target multidimensional dataset includes historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data;
[0086] A feature selection module 102 is configured to perform feature selection on the target multidimensional dataset based on a mutual information method to determine a feature subset associated with the substation load;
[0087] The prediction module 103 is configured to input the feature subset into a load prediction model based on the LGBM algorithm to obtain a load prediction result of the substation predicted by the load prediction model.
[0088] In some embodiments of the present application, in the above-mentioned device, the acquisition module 101 is specifically used to obtain an initial multidimensional data set, wherein the initial multidimensional data set includes historical load data of the substation, distributed photovoltaic power generation data and meteorological characteristic data of the distributed photovoltaic system; perform data cleaning and preprocessing on the initial multidimensional data set; determine the credibility of the initial multidimensional data set after data cleaning and preprocessing; and determine the target multidimensional data set based on the credibility of the initial multidimensional data set.
[0089] In some embodiments of the present application, in the above-mentioned device, the acquisition module 101 is specifically used to use the initial multidimensional dataset as the target multidimensional dataset when the credibility indicates that the initial multidimensional dataset has predictive value; or when the credibility indicates that the initial multidimensional dataset does not have predictive value, obtain meteorological characteristic data of the weather forecast, use the meteorological characteristic data of the weather forecast as supplementary data of the initial multidimensional dataset, and generate the target multidimensional dataset.
[0090] In some embodiments of the present application, in the above-mentioned device, the acquisition module 101 is specifically used to obtain a weather forecast map and obtain the location information of the distributed photovoltaic system; steplessly magnify the weather forecast map and grid the magnified weather forecast map; obtain the deployment information of the photovoltaic system, and set the irradiation grid corresponding to the photovoltaic system according to the deployment information; and obtain the meteorological characteristic data of the irradiation grid.
[0091] In some embodiments of the present application, in the above-mentioned device, the acquisition module 101 is specifically used to perform basic credibility verification on the initial multidimensional dataset; perform predictive credibility verification on the initial multidimensional dataset that passes the basic credibility verification; and determine the credibility of the initial multidimensional dataset based on the verification result of the predictive credibility verification.
[0092] In some embodiments of the present application, in the above-mentioned device, the feature selection module 102 is specifically used to calculate the mutual information value between each feature in the target multidimensional data set and the substation load; and take the features whose mutual information value is greater than the mutual information threshold as the feature subset.
[0093] In some embodiments of the present application, in the above-mentioned device, a correction module is used to construct a prediction result correction model; the target multidimensional data set and the substation load prediction result are input into the prediction result correction model to obtain the substation load correction result.
[0094] It should be noted that any of the above-mentioned substation load prediction devices based on the LGBM algorithm can implement the above-mentioned substation load prediction method based on the LGBM algorithm one-to-one, which will not be repeated here.
[0095] Figure 8 FIG. 1 shows a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 8As shown, at the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), and may also include non-volatile memory, such as at least one disk storage device. Of course, the electronic device may also include other hardware required for its services.
[0096] The processor, network interface, and memory can be interconnected through an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 8 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0097] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0098] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a substation load forecasting device based on the LGBM algorithm at the logical level. The processor executes the program stored in the memory and is specifically used to perform the aforementioned method.
[0099] The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The various methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium mature in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0100] The electronic device can execute the load prediction method based on the LGBM algorithm provided in multiple embodiments of the present application, and realize the load prediction device based on the LGBM algorithm in the area. Figure 7 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.
[0101] An embodiment of the present application also proposes a computer-readable storage medium, which stores one or more programs, and the one or more programs include instructions. When the instructions are executed by an electronic device including multiple application programs, the electronic device can execute the substation load prediction method based on the LGBM algorithm provided by multiple embodiments of the present application.
[0102] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0103] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0104] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0107] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0108] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0110] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A method for load forecasting in a substation based on the LGBM algorithm, characterized in that: The method comprises: Obtaining a target multidimensional dataset, wherein the target multidimensional dataset includes historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data; Performing feature selection on the target multidimensional data set based on a mutual information method to determine a feature subset associated with the substation load; The feature subset is input into a load forecasting model based on the LGBM algorithm to obtain a load forecasting result of the substation predicted by the load forecasting model.
2. The method according to claim 1, characterized in that The obtaining of the target multidimensional dataset includes: Acquire an initial multidimensional data set, the initial multidimensional data set including historical load data of the substation, distributed photovoltaic power generation data, and meteorological characteristic data of the distributed photovoltaic system; performing data cleaning and preprocessing on the initial multidimensional data set; Determine the credibility of the initial multidimensional dataset after data cleaning and preprocessing; The target multidimensional dataset is determined based on the credibility of the initial multidimensional dataset.
3. The method according to claim 2, characterized in that The step of determining the target multidimensional dataset based on the credibility of the initial multidimensional dataset comprises: When the confidence level indicates that the initial multidimensional dataset has predictive value, using the initial multidimensional dataset as the target multidimensional dataset; or When the credibility indicates that the initial multidimensional dataset has no predictive value, meteorological characteristic data of the weather forecast is obtained, and the meteorological characteristic data of the weather forecast is used as supplementary data of the initial multidimensional dataset to generate the target multidimensional dataset.
4. The method according to claim 3, characterized in that The obtaining of meteorological characteristic data of the weather forecast includes: Obtain weather forecast maps and location information of distributed photovoltaic systems; Steplessly zoom in on the weather forecast map and grid the zoomed weather forecast map; Obtain the deployment information of the photovoltaic system and set the irradiation grid corresponding to the photovoltaic system according to the deployment information; Obtain meteorological characteristic data for the irradiance grid.
5. The method according to any one of claims 2 to 4, characterized in that The step of determining the target multidimensional dataset based on the credibility of the initial multidimensional dataset comprises: Performing basic credibility verification on the initial multidimensional dataset; Performing prediction credibility verification on the initial multidimensional dataset that has passed the basic credibility verification; Based on the verification result of the prediction credibility verification, the credibility of the initial multidimensional dataset is determined.
6. The method according to claim 1, characterized in that The performing feature selection on the target multidimensional data set based on the mutual information method to determine a feature subset associated with the substation load includes: Calculating the mutual information value between each feature in the target multidimensional data set and the substation load; The features whose mutual information values are greater than the mutual information threshold are taken as feature subsets.
7. The method according to claim 1, characterized in that The method further comprises: Construct a prediction result correction model; The target multidimensional data set and the substation load forecast result are input into the forecast result correction model to obtain the substation load correction result.
8. A load forecasting device for a substation based on the LGBM algorithm, characterized in that: The device comprises: An acquisition module is used to acquire a target multidimensional data set, wherein the target multidimensional data set includes historical load data of the substation area, distributed photovoltaic power generation data and meteorological characteristic data; A feature selection module, configured to perform feature selection on the target multidimensional data set based on a mutual information method, and determine a feature subset associated with the substation load; The prediction module is used to input the feature subset into the load prediction model based on the LGBM algorithm to obtain the load prediction result of the substation predicted by the load prediction model.
9. An electronic device comprising: processor; as well as A memory arranged to store computer-executable instructions, wherein when the executable instructions are executed, the processor executes the steps of the substation load forecasting method based on the LGBM algorithm according to any one of claims 1 to 7.
10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device comprising a plurality of application programs, enables the electronic device to perform the steps of the substation load forecasting method based on the LGBM algorithm as described in any one of claims 1 to 7.