Electric power system load prediction method based on artificial intelligence technology
By dividing the load data collection areas in the power system, assigning unique identifiers and recording attributes, building data sets, calculating quality indicators, and dynamically switching proofreading strategies, the problem that existing technologies cannot adapt to regional differences and real-time changes is solved, and the accuracy and reliability of load forecasting are improved.
Patent Information
- Application Number
- CN202510972615.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-09-12
AI Technical Summary
Existing power system load forecasting methods are unable to dynamically adjust the calibration strategy based on regional differences and real-time changes, resulting in reduced load forecasting accuracy and reliability.
Based on artificial intelligence technology, the load data collection areas are divided, a unique identifier is assigned to each area, the regional attributes are recorded, the regional data set is constructed, the data quality indicators are calculated, and the proofreading strategy is dynamically switched according to the real-time data quality to perform load forecasting.
The accuracy and reliability of load forecasting are improved, and the accuracy of forecast results is improved by dynamically adjusting strategies to adapt to regional differences and real-time changes.
Smart Images

Figure CN120633943A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power system load forecasting, and in particular to a power system load forecasting method based on artificial intelligence technology. Background Art
[0002] The accuracy of power system load forecasts is crucial for grid dispatch, energy distribution, and market transactions. However, the actual load data collection process faces the following challenges: 1. Significant regional differences in data quality: Data quality (such as missing rate, outlier ratio, and noise level) varies significantly across collection areas (e.g., urban, rural, and industrial areas) due to differences in equipment configuration, communication conditions, and user electricity usage behavior. 2. Dynamic data quality: Data quality fluctuates over time (e.g., equipment failure, communication interruptions, and extreme weather conditions lead to data anomalies or missing data).
[0003] Most existing methods are static rules (such as global interpolation) or global unified processing, but they are unable to dynamically adjust the correction strategy according to regional differences and real-time changes, resulting in reduced load forecasting accuracy and reliability. Summary of the Invention
[0004] The purpose of the present invention is to provide a power system load forecasting method based on artificial intelligence technology, aiming to solve the technical problem in the existing technology that the calibration strategy cannot be dynamically adjusted according to regional differences and real-time changes, resulting in reduced load forecasting accuracy and prediction reliability.
[0005] To achieve the above object, the present invention adopts an artificial intelligence technology-based power system load forecasting method, comprising the following steps:
[0006] Divide the load data collection area for the power system, assign a unique identifier to each collection area, and record the area attributes;
[0007] Collect data from each collection area, build regional data sets, calculate data quality indicators for each area, and output data quality levels;
[0008] In response to real-time changes in data quality, the proofreading strategy is switched dynamically, load forecasting is performed based on the proofread data, and the load forecast results are output.
[0009] Among them, in the steps of dividing the load data collection area for the power system, assigning a unique identifier to each collection area, and recording the area attributes:
[0010] Divide the power system into multiple load data collection areas according to the grid topology, and obtain basic information data for each load data collection area;
[0011] Identify the regional attributes of each load data collection area based on basic information data;
[0012] Assign a unique identifier to each load data collection area, match the area attribute data, and establish a mapping table.
[0013] Wherein, in the step of identifying the regional attributes of each load data collection area based on the basic information data:
[0014] According to the load type information in the basic information data, the load in the load data collection area is classified, and the proportion of each load type in the total load of the area is calculated to determine the load type of the area;
[0015] Based on the load capacity information in the basic information data, the regional electricity consumption scale is judged, and the regional load density is calculated to determine the regional electricity consumption scale;
[0016] Determine the power supply reliability of the region based on the power supply equipment, grid structure, and historical power outage data in the basic information data.
[0017] Wherein, in the step of identifying the regional attributes of each load data collection area based on the basic information data:
[0018] Obtain the load type, power consumption scale, and power supply reliability of the region to determine the regional attributes of the load data collection area.
[0019] Among them, in the steps of collecting data from each collection area, constructing regional data sets, calculating data quality indicators for each area, and outputting data quality levels:
[0020] Collect load data, meteorological data, and equipment status data from each region, align multi-source data by timestamp, and build regional datasets;
[0021] Calculate the missing rate, outlier ratio, and signal-to-noise ratio of each region respectively, and output the calculation results;
[0022] According to the calculation results, the data quality level is output.
[0023] Among them, in the steps of collecting load data, meteorological data, and equipment status data of each region, aligning multi-source data by timestamp, and building regional data sets:
[0024] Set the collection frequency, use the collection time as the timestamp, and match and integrate data from different data sources at the same timestamp.
[0025] Among them, in the step of outputting the data quality level according to the calculation results:
[0026] The missing rate, outlier ratio, and signal-to-noise ratio levels are determined separately, and the data quality level is obtained by integrating the missing rate, outlier ratio, and signal-to-noise ratio levels.
[0027] Among them, in the steps of respectively determining the missing rate, the outlier ratio, and the signal-to-noise ratio level, and integrating the missing rate, the outlier ratio, and the signal-to-noise ratio level to obtain the data quality level:
[0028] If any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is poor, the data quality level is poor;
[0029] If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not poor, then if any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is fair, the data quality level is fair;
[0030] If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not average, then if any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is good, the data quality level is good;
[0031] If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not good, the data quality level is excellent.
[0032] Among them, in the steps of dynamically switching the proofreading strategy in response to real-time changes in data quality, performing load forecasting based on the proofread data, and outputting the load forecast results:
[0033] Dynamically switch the proofreading strategy according to the data quality level into high data quality area, medium data quality area, and low data quality area;
[0034] Load forecast results based on the calibrated regional data.
[0035] Among them, after the step of predicting load results based on the calibrated regional data:
[0036] Dynamically adjust the proofreading strategy based on the comparison feedback between the predicted results and the actual load.
[0037] The present invention provides an artificial intelligence-based load forecasting method for an electric power system. The method divides the electric power system into load data collection areas, assigns a unique identifier to each collection area, and records area attributes. The method collects data from each collection area, constructs a regional data set, calculates data quality indicators for each area, and outputs a data quality level. The method dynamically switches the proofreading strategy based on real-time changes in data quality, performs load forecasting based on the proofread data, and outputs a load forecasting result. The method divides the data collection areas, detects the data quality in each collection area, and dynamically adjusts the proofreading strategy for each data collection area based on real-time data quality, thereby improving load forecasting accuracy and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0039] Figure 1 The present invention is a flowchart of the steps of the power system load forecasting method based on artificial intelligence technology.
[0040] Figure 2 It is a step flow chart of S100 of the present invention.
[0041] Figure 3 It is a step flow chart of S200 of the present invention.
[0042] Figure 4 It is a step flow chart of S300 of the present invention.
[0043] Figure 5 It is a structural principle diagram of the power system load forecasting system based on artificial intelligence technology of the present invention.
[0044] Figure 6 It is a structural principle diagram of the electronic device of the present invention.
[0045] 401-acquisition area division module, 402-data quality index calculation module, 403-dynamic proofreading module. DETAILED DESCRIPTION
[0046] Exemplary embodiments are described in detail herein, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different drawings represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with this application.
[0047] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. As used in this application and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0048] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0049] See also Figures 1 to 4 The present invention provides a method for power system load forecasting based on artificial intelligence technology, comprising the following steps:
[0050] S100: Divide the power system into load data collection areas, assign a unique identifier to each collection area, and record area attributes.
[0051] In this embodiment, the load data collection area is divided into two parts for the power system, each of which is assigned a unique identifier and the area attributes are recorded. The specific process is as follows:
[0052] S101: Divide the power system into multiple load data collection areas according to the power grid topology, and obtain basic information data of each load data collection area;
[0053] S102: Identifying regional attributes of each load data collection area based on basic information data;
[0054] S103: Assign a unique identifier to each load data collection area, match the area attribute data, and establish a mapping table.
[0055] Furthermore, in the step of identifying the regional attributes of each load data collection region based on the basic information data:
[0056] According to the load type information in the basic information data, the load in the load data collection area is classified, and the proportion of each load type in the total load of the area is calculated to determine the load type of the area;
[0057] Based on the load capacity information in the basic information data, the regional electricity consumption scale is judged, and the regional load density is calculated to determine the regional electricity consumption scale;
[0058] Determine the regional power supply reliability based on the basic information data including power supply equipment, grid structure, and historical power outage data;
[0059] Obtain the load type, power consumption scale, and power supply reliability of the region to determine the regional attributes of the load data collection area.
[0060] In the above process, load data collection areas are divided based on the grid topology. Grid topology describes the connections and hierarchical structure between various components in the grid (such as substations, transmission lines, distribution lines, and load nodes). By analyzing the grid topology, areas with similar electrical characteristics, geographical proximity, or close power supply relationships can be divided into load data collection areas.
[0061] After demarcating the load data collection areas, collect basic information for each area. This information may include the number of substations within the area, transmission line length, distribution line type, load type (such as industrial, commercial, and residential loads), load capacity, and power supply range. This data is crucial for identifying regional attributes and establishing mapping tables.
[0062] Analyze basic information data: Conduct in-depth analysis of the basic information data of each load data collection area, and determine the regional attributes of each load data collection area based on the analysis results of the basic information data; the specific process is as follows:
[0063] Preprocess the collected data to ensure data quality and consistency and provide a reliable basis for subsequent analysis; preprocessing includes:
[0064] Missing Value Handling: Check for missing values in basic information data. For key data such as load capacity and number of substations, if missing, various approaches can be used to address this. For example, if load capacity data for a particular area is missing, but that area is similar to adjacent areas in terms of grid structure and load type, an estimate can be made by referencing the load capacity data of adjacent areas. Alternatively, missing values can be filled by querying historical data or contacting relevant personnel to obtain accurate information.
[0065] Outlier handling: Identify and address outliers in the data. For example, if the length of power transmission lines in a certain area far exceeds the normal range, this could be due to data entry errors or special circumstances. Clearly erroneous data should be corrected or deleted. Outliers that may be caused by special circumstances require further investigation, and a decision on whether to retain or adjust the data will be made based on the actual situation.
[0066] Data standardization: Because basic information data may contain data in different units and magnitudes, such as load capacity in kilowatts (kW) or megawatts (MW), and transmission line length in meters (m) or kilometers (km), data standardization is necessary to facilitate subsequent analysis and comparison. For example, all load capacity data can be uniformly converted to megawatts (MW), and transmission line length data can be uniformly converted to kilometers (km).
[0067] After data preprocessing, the following analyses are performed in sequence:
[0068] Load type analysis is an important component of regional attributes. Different load types have different power consumption characteristics and requirements. Based on the load type information in the basic information data, loads within the load data collection area are classified. Common load types include industrial loads, commercial loads, residential loads, and agricultural loads. For example, by reviewing data such as the list of businesses in the area, commercial building information, and residential electricity registrations, the type of each load can be determined.
[0069] Proportion calculation, calculate the proportion of each load type in the total load of the area. Assume that there are n different types of loads in an area, and the power of the i-th load is P i , the total load power is:
[0070]
[0071] Then the proportion ri of the i-th load is:
[0072]
[0073] The dominant load type is determined based on the calculated load type ratio. For example, if the industrial load ratio rindustry is the largest and exceeds a certain proportion (e.g., 50%), the dominant load type in the area is determined to be industrial load, and the area can be preliminarily identified as an industrial zone.
[0074] Electricity consumption scale analysis, the electricity consumption scale reflects the size of regional electricity demand, which is of great significance to the planning and operation of the power system, including load capacity assessment and load density calculation and analysis.
[0075] Load capacity assessment directly checks the load capacity information in the basic information data to assess the scale of regional electricity consumption. Load capacity can be divided into different levels according to certain standards, such as small load area (load capacity less than 10MW), medium load area (load capacity between 10 and 50MW), and large load area (load capacity greater than 50MW).
[0076] Among them, load density calculation and analysis, the load density of the calculation area, the formula is:
[0077]
[0078] Where P is the total load power in the region (unit: MW), S is the power supply area of the region (unit: km 2Load density can more accurately reflect the electricity demand per unit area. For example, if two regions have the same total load power but different power supply areas, the region with a higher load density will have more concentrated electricity demand and may place higher demands on the grid's power supply capacity and reliability. Based on the load density value, electricity consumption can be further categorized into high, medium, and low load density areas.
[0079] Power supply reliability analysis: Power supply reliability is an important indicator to measure the power system's ability to provide users with continuous and stable power supply, including equipment reliability assessment, grid structure analysis, and historical power outage data analysis.
[0080] Equipment reliability assessment involves analyzing basic data related to power supply equipment, such as the operating age, maintenance records, and failure rates of substation equipment. Equipment with a long operating life and a high failure rate may affect power supply reliability. For example, if the average operating age of substation equipment in a region exceeds 20 years and the failure rate is high, the region's power supply reliability may be relatively low.
[0081] Grid structure analysis combines grid topology data to analyze the regional grid structure. Regions with ring networks and multiple power sources typically have higher power supply reliability, while regions with single power sources and radial grid structures have relatively lower reliability. For example, if a region is supplied by two different transmission lines, forming a ring network structure within the region, if one line fails, the other line can continue to supply power, ensuring the region's electricity needs.
[0082] The historical power outage data analysis reviews regional power outage data, including the number of outages, duration, and scope. Regions with high frequency and duration of power outages are associated with poor power supply reliability. For example, if one region experienced five power outages in the past year, each lasting an average of two hours, while another region experienced only one outage lasting 0.5 hours, the former's power supply reliability is significantly lower than the latter's.
[0083] After completing the above analyses, a comprehensive evaluation of the analysis results is required to comprehensively and accurately determine the attributes of the load data collection area.
[0084] Assign weights to each analysis indicator based on its importance to the regional attributes. For example, load type and power consumption scale may have a greater impact on regional attributes and thus be given higher weights; while power supply reliability is also important, it may be given a relatively lower weight in some cases.
[0085] Comprehensive score calculation: Calculate the comprehensive score for each region based on the scores and weights of each analysis indicator. For example, assuming the weights of load type, power consumption scale, and power supply reliability are w1, w2, and w3, respectively, and w1 + w2 + w3 = 1, and the scores of each indicator are s1, s2, and s3, respectively, then the comprehensive score S for the region is: S = w1s1 + w2s2 + w3s3.
[0086] Based on the comprehensive score and pre-set judgment criteria, the attributes of each load data collection area are determined. For example, different scoring intervals can be set to correspond to different regional attributes, such as high-load industrial areas, medium-sized commercial areas, and low-load residential areas.
[0087] To uniquely identify each load data collection region, a unique identifier is assigned to each region. This identifier can be numbers, letters, or a combination of numbers and letters, such as region numbers, such as "Region_001" and "Region_002."
[0088] Matching regional attribute data: Match the unique identifier of each load data collection region with its corresponding regional attribute data. For example, associate "Region_001" with the region's load type attributes (such as industrial area), power consumption scale attributes (such as large load area), power supply reliability attributes (such as high reliability power supply area), etc.
[0089] Based on the matching results, a mapping table is created. This table is in tabular format and contains two or more columns: one column is the unique identifier of the load data collection area, and the other column or columns are the corresponding area attribute data. This mapping table allows for easy querying of the attribute information for each load data collection area.
[0090] S200: Collect data from each collection area, build a regional data set, calculate the data quality index of each area, and output the data quality level.
[0091] In this embodiment, data from each collection area is collected, a regional data set is constructed, the data quality index of each area is calculated, and the data quality level is output. The specific process is as follows:
[0092] S201: Collect load data, meteorological data, and equipment status data from each region, align multi-source data by timestamp, and construct a regional dataset;
[0093] S202: Calculate the missing rate, outlier ratio, and signal-to-noise ratio of each region respectively, and output the calculation results;
[0094] S203: Output the data quality level according to the calculation result.
[0095] Furthermore, in the step of collecting load data, meteorological data, and equipment status data from each region, aligning multi-source data by timestamp, and constructing a regional dataset:
[0096] Set the collection frequency, use the collection time as the timestamp, and match and integrate data from different data sources at the same timestamp.
[0097] Furthermore, in the step of outputting the data quality level according to the calculation result:
[0098] The missing rate, outlier ratio, and signal-to-noise ratio levels are determined separately, and the data quality level is obtained by integrating the missing rate, outlier ratio, and signal-to-noise ratio levels.
[0099] In the above process, data is collected and regional datasets are constructed:
[0100] Load data collection: Smart meters, load monitoring terminals, and other devices deployed in each collection area collect load data at a set frequency (for example, every 15 minutes). This data includes active power (P, in kilowatts, kW), reactive power (Q, in kilovars, kvar), current (I, in amperes, A), voltage (U, in volts, V), and electricity consumption (E, in kilowatt-hours, kWh). For example, smart meters measure and record these parameters in real time and then transmit the data to the data collection system at the collection frequency.
[0101] Meteorological data collection: With the help of weather stations or weather data interfaces, meteorological data is also collected according to the collection frequency. Meteorological data includes temperature (T, unit: Celsius, ℃), humidity (H, unit: percentage, %), wind speed (V w , unit: meter / second, m / s), wind direction (expressed in angle) and rainfall (R, unit: millimeter, mm), etc. The weather station uses various sensors to monitor meteorological parameters in real time and transmit the data to the data collection platform.
[0102] Equipment status data collection: Use sensors and monitoring systems installed on power equipment to collect equipment status data. For example, for transformers, collect oil temperature (Toil, unit: ℃), winding temperature (T winding , unit: ℃) and load rate (L, unit: percentage, %) and other data; for switchgear, collect the opening and closing status (expressed in binary, such as 1 for closing and 0 for opening) and the number of operations (N action , unit: times) and other data.
[0103] Set appropriate collection frequencies for different types of data according to actual requirements and system resources. For example, for load data and meteorological data, since the changes in load data and meteorological data are relatively fast, a higher collection frequency can be set, such as once every 15 minutes; while for equipment status data, such as the number of operations of switchgear, the changes are relatively slow, and the collection frequency can be appropriately reduced, such as once an hour.
[0104] Use the collection time as the timestamp to ensure that each data point has a clear time identifier. The format of the timestamp can adopt the standard date and time format, such as "YYYY-MM-DD HH:MM:SS".
[0105] Since load data, meteorological data, and equipment status data come from different data sources, there may be differences in collection time and frequency. Therefore, it is necessary to align multi-source data according to the timestamp. Taking a unified time base (such as one time point every 15 minutes) as the standard, match and integrate the data of different data sources at the same timestamp.
[0106] For the case where there is no data at the corresponding time point, interpolation methods can be used for estimation. For example, the linear interpolation formula is: if the data values at known time points t1 and t2 (t1 < t2) are y1 and y2 respectively, for the time point t (t1 < t < t2), the interpolation result y is:
[0107]
[0108] Classify and store the aligned and integrated multi-source data according to the collection area to construct an area dataset. The dataset for each area should include the load data, meteorological data, and equipment status data of that area at each time point, and record information such as the collection time and source of the data. For example, a database table can be established, and the fields in the table include area name, timestamp, active power, reactive power, temperature, humidity, oil temperature, load rate, etc.
[0109] The missing rate refers to the proportion of missing data in the dataset and is used to measure the integrity of the data.
[0110] Calculation formula: Assume that there are a total of N data items in a certain area within a period of time, and M of them are missing. Then the missing rate R of this area missing is:
[0111]
[0112] Record and output the calculation results of the missing rate for each area. For example, list the names of each area and the corresponding missing rates in tabular form.
[0113] The proportion of outliers refers to the proportion of the number of outliers in the dataset to the total number of data items and is used to evaluate the accuracy of the data.
[0114] Calculation method: First, determine the criteria for judging outliers based on the data distribution characteristics and business rules. For example, for the active power in the load data, a reasonable range can be set, such as: based on the mean μ of the historical data plus or minus k times the standard deviation σ), that is, [μ-kσ,μ+kσ]. Data outside this range is considered an outlier. Then, count the number of outliers K in each regional data set and divide it by the total number of data items N to obtain the outlier ratio R. abnormal ,Right now:
[0115]
[0116] The outlier ratio of each region is recorded and output, for example, the outlier ratio of each region is listed in a data quality report.
[0117] The signal-to-noise ratio (SNR) is an indicator that measures the relative strength of the effective signal and noise in the data, reflecting the clarity and reliability of the data.
[0118] Calculation method: The signal-to-noise ratio can be calculated using the frequency domain analysis method. First, perform Fourier transform on the data to convert the data from the time domain to the frequency domain. Assume that the signal power is P s , the noise power is P n , then the calculation formula of signal-to-noise ratio SNR is:
[0119]
[0120] In actual calculations, the spectrum ranges of the signal and noise can be determined by analyzing the spectrum diagram, and then the signal power and noise power can be calculated respectively.
[0121] The signal-to-noise ratio calculation results of each region are output, for example, the signal-to-noise ratio of each region is displayed in a graphical form.
[0122] Set thresholds: Based on actual application requirements and the importance of data quality indicators, set different threshold ranges for missing rate, outlier ratio, and signal-to-noise ratio to classify data quality levels. For example, if the data quality level is divided into four levels: excellent, good, fair, and poor, the specific threshold settings are as follows:
[0123] Excellent: Missing rate R missing <5%, outlier ratio R abnormal <3%, signal-to-noise ratio SNR>20dB.
[0124] Good: 5% ≤ R missing <10%, 3%≤R abnormal <6%, 15≤SNR≤20dB.
[0125] Generally: 10%≤R missing <15%, 6%≤R abnormal <10%, 10≤SNR<15dB.
[0126] Poor: R missing ≥15%, R abnormal ≥10%, SNR<10dB.
[0127] The missing rate, outlier ratio, and signal-to-noise ratio calculated for each region are compared with pre-set thresholds to output a missing rate, outlier ratio, and signal-to-noise ratio rating. If any of these indicators is poor, the data quality is rated poor. If none of these indicators is poor, then the data quality is rated fair. If none of these indicators is fair, then the data quality is rated good. If none of these indicators is good, then the data quality is rated excellent. For example, if a region has a missing rate of 4%, an outlier ratio of 2%, and a signal-to-noise ratio of 22dB, then all indicators in this region meet the excellent standard based on the thresholds set above.
[0128] Based on the comparison results, the data quality level for each region is determined and output in the form of reports, charts, or database records. For example, a data quality report can be generated, listing information such as the name of each region, missing rate, outlier ratio, signal-to-noise ratio, and data quality level. This allows power system managers to intuitively understand the quality status of data in each region and provide a reference for subsequent data processing, analysis, and decision-making.
[0129] S300: Dynamically switch the proofreading strategy in response to real-time changes in data quality, perform load forecasting based on the proofread data, and output the load forecast results.
[0130] In this embodiment, the proofreading strategy is switched dynamically in response to real-time changes in data quality, and load forecasting is performed based on the proofread data, and the load forecast results are output. The specific process is as follows:
[0131] S301: Dynamically switch the proofreading strategy to a high data quality area, a medium data quality area, and a low data quality area according to the data quality level;
[0132] S302: forecasting load results based on the verified regional data;
[0133] S303: Dynamically adjust the calibration strategy based on the comparison feedback between the prediction result and the actual load.
[0134] In the above process, each acquisition area is divided into a high data quality area, a medium data quality area, and a low data quality area based on the data quality level output in step S200. For example, an area with an excellent data quality level is divided into a high data quality area, a good data quality level is divided into a medium data quality area, and an average or poor data quality level is divided into a low data quality area.
[0135] Proofreading strategy development for different levels of areas:
[0136] In high-quality data areas, due to the high data quality, there are fewer cases of missing data and outliers, so a more basic proofreading strategy is adopted. For example, a simple data integrity check is performed to ensure that there are no missed data points; the data is smoothed to remove possible minor noise. The moving average method can be used, and its formula is:
[0137]
[0138] in is the smoothed value at time point t, y t-i is the original data value at time point ti, and n is the window size of the moving average.
[0139] In the medium data quality area, the data contains a certain degree of missing data and outliers, and a more complex proofreading strategy is adopted. In addition to data integrity checks and smoothing, missing data also need to be interpolated. This method can be used to fit the data by constructing a piecewise cubic polynomial to ensure the continuity and smoothness of the interpolation function at the interpolation points. For outliers, statistical methods can be used to identify and correct them. For example, data outside the range of the mean plus or minus three standard deviations can be considered outliers and replaced with the mean or median.
[0140] In low-quality data areas, where data quality is poor and there are many missing and outliers, a strict proofreading strategy is adopted. First, comprehensive data cleaning is performed to remove obvious erroneous data and noise. For large amounts of missing data, in addition to interpolation methods, comprehensive estimates can be made by combining historical data and data from adjacent areas. Outliers are analyzed in depth and corrected or eliminated based on business rules and actual conditions.
[0141] Dynamically switch proofreading strategies and establish a real-time monitoring system to continuously track changes in data quality levels across regions. When a region's data quality level changes, the system automatically switches to the corresponding proofreading strategy. For example, if a region that was originally rated as medium data quality improves to high data quality due to a data acquisition device failure repair, the system will automatically switch the proofreading strategy for that region from a medium data quality strategy to a high data quality strategy.
[0142] Select a load forecasting model based on the characteristics of the verified regional data and the load forecasting requirements. Common load forecasting models include time series models (such as ARIMA models), machine learning models (such as support vector machines and random forests), and deep learning models (such as long short-term memory networks (LSTMs). For example, the ARIMA model may be more suitable for load data with significant seasonality and trends. For load data with high dimensionality and complex relationships, deep learning models may provide better forecasting results.
[0143] Model training and parameter tuning: Use calibrated historical regional data to train the selected load forecasting model. During training, cross-validation and other methods are used to evaluate model performance, and model parameters are adjusted to optimize forecasting results. For example, for ARIMA models, the autoregressive order p, differencing order d, and moving average order q need to be determined; for deep learning models, parameters such as the network structure, learning rate, and batch size need to be adjusted.
[0144] Input the corrected real-time regional data into the trained load forecasting model to perform load forecasting. Based on the characteristics of the input data and historical patterns, the model outputs load forecasts for the future. For example, it can predict hourly load values for the next 24 hours.
[0145] Output load forecast results in the form of charts, reports, or data interfaces. Charts can intuitively display load trends, reports can list specific forecast values and time points, and data interfaces can facilitate other systems to access forecast results.
[0146] The real-time monitoring system can be used to obtain the actual load data of each area. The actual load data can be obtained from the power system's SCADA (Supervisory Control and Data Acquisition) or other related systems to ensure the accuracy and timeliness of the data.
[0147] Compare and analyze the load forecast results with the actual load data to calculate the forecast error. Commonly used error indicators include mean absolute error (MAE) and root mean square error (RMSE). The calculation formula for MAE is:
[0148]
[0149] The calculation formula of RMSE is:
[0150]
[0151] where y i is the actual load value, is the predicted load value, and n is the number of samples.
[0152] Analyze the causes of the errors based on their size and distribution. If the forecast error in a particular region is large, it could be due to data quality issues in that region, resulting in deviations in the corrected data, or the load forecast model may not accurately fit the data characteristics of that region.
[0153] Based on the error analysis results, the calibration strategy for that area is dynamically adjusted. For example, if it is found that the large forecast error is due to the presence of incompletely identified outliers in the data, the method for identifying and correcting the outliers can be further optimized. If it is found that the load forecast model is the problem, the model can be replaced or its parameters can be adjusted.
[0154] Establish a continuous optimization mechanism to regularly compare and analyze load forecast results with actual load data, and continuously adjust and optimize calibration strategies and load forecast models to improve the accuracy and reliability of load forecasts. For example, conduct comprehensive error analysis and strategy adjustments weekly or monthly.
[0155] Corresponding to the aforementioned embodiment of the power system load forecasting method based on artificial intelligence technology, the present application also provides an embodiment of the power system load forecasting system based on artificial intelligence technology.
[0156] Figure 5 This is a block diagram of a power system load forecasting system based on artificial intelligence technology according to an exemplary embodiment. Figure 5 The system may include: a collection area division module 401, a data quality index calculation module 402, and a dynamic proofreading module 403; wherein:
[0157] The collection area division module 401 is used to divide the load data collection area for the power system, assign a unique identifier to each collection area, and record area attributes;
[0158] The data quality index calculation module 402 is used to collect data from each collection area, construct a regional data set, calculate the data quality index of each area, and output the data quality level;
[0159] The dynamic proofreading module 403 is used to dynamically switch proofreading strategies according to real-time changes in data quality, perform load forecasting based on the proofread data, and output load forecasting results.
[0160] In this embodiment, the collection area division module 401 divides the load data collection area for the power system, assigns a unique identifier to each collection area, and records the area attributes; the data quality index calculation module 402 collects data from each collection area, constructs a regional data set, calculates the data quality index of each area, and outputs the data quality level; the dynamic proofreading module 403 dynamically switches the proofreading strategy according to the real-time changes in data quality, performs load forecasting based on the proofread data, and outputs the load forecasting results; by dividing the data collection area, detecting the data quality in each collection area, and dynamically adjusting the proofreading strategy of each data collection area according to the real-time data quality, the load forecasting accuracy and forecasting reliability are improved.
[0161] Regarding the system in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0162] For the system embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0163] Accordingly, the present application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned power system load forecasting method based on artificial intelligence technology. Figure 6 As shown in the figure, a hardware structure diagram of a power system load forecasting system based on artificial intelligence technology provided by an embodiment of the present invention is provided, in which any device with data processing capability is provided. Figure 6 In addition to the processor, memory, and network interface shown, any device with data processing capabilities in which the apparatus in the embodiment is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.
[0164] Accordingly, the present application also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the power system load forecasting method based on artificial intelligence technology as described above. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the aforementioned embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), an SD card, a flash card (Flash Card), etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capabilities. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or is to be output.
[0165] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the contents disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of this application and include common knowledge or customary techniques in the art that are not disclosed in this application.
[0166] It will be understood that the present application is not limited to the exact construction that has been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof.
Claims
1. A power system load forecasting method based on artificial intelligence technology, characterized in that: The steps include: Divide the load data collection area for the power system, assign a unique identifier to each collection area, and record the area attributes; Collect data from each collection area, build regional data sets, calculate data quality indicators for each area, and output data quality levels; In response to real-time changes in data quality, the proofreading strategy is switched dynamically, load forecasting is performed based on the proofread data, and the load forecast results are output.
2. The power system load forecasting method based on artificial intelligence technology according to claim 1, characterized in that: In the steps of dividing the load data collection area for the power system, assigning a unique identifier to each collection area, and recording the area attributes: Divide the power system into multiple load data collection areas according to the grid topology, and obtain basic information data for each load data collection area; Identify the regional attributes of each load data collection area based on basic information data; Assign a unique identifier to each load data collection area, match the area attribute data, and establish a mapping table.
3. The power system load forecasting method based on artificial intelligence technology according to claim 2, characterized in that: In the step of identifying the area attributes of each load data collection area for the basic information data: According to the load type information in the basic information data, the load in the load data collection area is classified, and the proportion of each load type in the total load of the area is calculated to determine the load type of the area; Based on the load capacity information in the basic information data, the regional electricity consumption scale is judged, and the regional load density is calculated to determine the regional electricity consumption scale; Determine the power supply reliability of the region based on the power supply equipment, grid structure, and historical power outage data in the basic information data.
4. The power system load forecasting method based on artificial intelligence technology according to claim 3, characterized in that: In the step of identifying the area attributes of each load data collection area for the basic information data: Obtain the load type, power consumption scale, and power supply reliability of the region to determine the regional attributes of the load data collection area.
5. The power system load forecasting method based on artificial intelligence technology according to claim 1, characterized in that: In the steps of collecting data from each collection area, building regional datasets, calculating data quality indicators for each area, and outputting data quality levels: Collect load data, meteorological data, and equipment status data from each region, align multi-source data by timestamp, and build regional datasets; Calculate the missing rate, outlier ratio, and signal-to-noise ratio of each region respectively, and output the calculation results; According to the calculation results, the data quality level is output.
6. The power system load forecasting method based on artificial intelligence technology according to claim 5, characterized in that: In the steps of collecting load data, meteorological data, and equipment status data from each region, aligning multi-source data by timestamp, and building regional datasets: Set the collection frequency, use the collection time as the timestamp, and match and integrate data from different data sources at the same timestamp.
7. The power system load forecasting method based on artificial intelligence technology according to claim 5, characterized in that: In the step of outputting the data quality level based on the calculation results: The missing rate, outlier ratio, and signal-to-noise ratio levels are determined separately, and the data quality level is obtained by integrating the missing rate, outlier ratio, and signal-to-noise ratio levels.
8. The power system load forecasting method based on artificial intelligence technology according to claim 7, characterized in that: In the steps of determining the missing rate, outlier ratio, and signal-to-noise ratio level respectively, and integrating the missing rate, outlier ratio, and signal-to-noise ratio levels to obtain the data quality level: If any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is poor, the data quality level is poor; If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not poor, then if any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is fair, the data quality level is fair; If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not average, then if any one of the missing rate level, outlier ratio level, and signal-to-noise ratio level is good, the data quality level is good; If the missing rate level, outlier ratio level, and signal-to-noise ratio level are not good, the data quality level is excellent.
9. The power system load forecasting method based on artificial intelligence technology according to claim 1, characterized in that: In the steps of dynamically switching the proofreading strategy in response to real-time changes in data quality, performing load forecasting based on the proofread data, and outputting the load forecast results: Dynamically switch the proofreading strategy according to the data quality level into high data quality area, medium data quality area, and low data quality area; Load forecast results based on the calibrated regional data.
10. The power system load forecasting method based on artificial intelligence technology according to claim 9, characterized in that: After the step of forecasting load results for the calibrated area data: Dynamically adjust the proofreading strategy based on the comparison feedback between the predicted results and the actual load.