A power grid data intelligent prediction system and method based on data analysis
By generating and storing typical prediction scenarios of power grid data, the problem of real-time difficulty in grid data prediction is solved, and the efficiency and accuracy of short-term prediction are improved.
Patent Information
- Application Number
- CN202411667842.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-11-21
AI Technical Summary
The prior art is difficult to achieve real-time performance in power grid data prediction, resulting in lag in short-term power grid data prediction results, affecting the stable operation of the power system.
By obtaining historical data of the power grid, predicted scenarios are generated and classified into typical prediction scenarios, stored in caches and databases. When prediction is required, first match with the typical prediction scenario in the cache. If it matches, the prediction value will be directly obtained. Otherwise, match with the prediction scenario in the database or directly calculate.
It improves the short-term prediction efficiency of power grid data, reduces the time consumption of real-time calculations, and ensures the timeliness and accuracy of prediction results.
Smart Images

Figure CN119167113B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power grid data prediction, and in particular to a power grid data intelligent prediction system and method based on data analysis. Background Art
[0002] With the rapid development of technologies such as big data and artificial intelligence, the accuracy and efficiency of power grid data prediction are constantly improving. Various advanced algorithms and models are widely used in power grid data prediction, providing strong support for the stable operation of the power system; power grid data prediction has multi-scale characteristics, requiring not only medium- and long-term predictions, but also short-term predictions to meet the needs of power grid planning and operation at different time scales; short-term power grid data prediction requires real-time or near real-time analysis results, however, due to the complexity of data processing and analysis, the real-time requirements may be difficult to fully meet, resulting in delayed prediction results. Therefore, how to improve the efficiency of short-term power grid data prediction has become an urgent problem to be solved. Summary of the invention
[0003] The purpose of the present invention is to provide a power grid data intelligent prediction system and method based on data analysis to solve the problems raised in the prior art.
[0004] To achieve the above object, the present invention provides the following technical solution: a method for intelligent prediction of power grid data based on data analysis, comprising:
[0005] S11, acquiring historical data of the power grid, and generating a prediction scenario based on the historical data of the power grid;
[0006] S12, obtaining a typical prediction scenario based on the prediction scenario of the power grid data;
[0007] S13, selecting typical prediction scenarios and storing them in the cache, and storing all prediction scenarios in the database;
[0008] S14, when it is necessary to predict the power grid data, firstly match the power grid data with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, the prediction value of the power grid data is obtained according to the typical prediction scenario in the cache; if there is no matching typical prediction scenario in the cache, match the power grid data with the prediction scenarios in the database; if there is a matching prediction scenario in the data, the prediction value of the power grid data is obtained according to the prediction scenario in the database, otherwise the prediction value is directly calculated for the power grid data.
[0009] In step S11, the generation of prediction scenarios based on historical data of the power grid further includes the following steps:
[0010] Extracting grid data features from historical data of the grid, the grid data features include continuous features and discrete features, normalizing the continuous features to obtain input features, and labeling the discrete features;
[0011] A prediction model for power grid data is established and trained and optimized. The input of the prediction model is the input features with the same discrete features, and the output of the prediction model is the actual measured value of the power grid data. Through the prediction model of power grid data, the input features obtained through historical data and the input features set manually are optimized and calculated to obtain the predicted value of the power grid data. The input features are spliced with the predicted value of the power grid data to obtain the prediction scenario.
[0012] The discrete features of the power grid data represent the characteristics of the scene. Optionally, when predicting the power load, the continuous features include temperature, humidity, load growth rate and load at the previous moment, and the discrete features include period, time period, and weather. The period is used to reflect holiday information, the time period is used to reflect peak and valley information of power consumption, and the weather is used to reflect weather condition information. Different values are assigned to different scenes to distinguish them. For different scenes, historical data can be used to train the prediction model when the system load is low, and historical data and existing prediction models can be used for pre-calculation to reduce the impact on real-time system performance. After the pre-calculation is completed, the results are stored in a fast-access cache, which can be a high-speed storage system to ensure rapid retrieval when needed.
[0013] In step S12, obtaining a typical prediction scenario based on the prediction scenario of the power grid data further includes the following steps:
[0014] S31, let i=0, A(i,1), A(i,2),…,A(i,n i ) represents the prediction scenario that is not merged into the typical prediction scenario; n i Indicates the number of forecast scenarios that are not incorporated by the typical forecast scenario;
[0015] S32, taking A(i,1) as a typical prediction scenario, calculating the Euclidean distance between other prediction scenarios and A(i,1), and if the Euclidean distance between other prediction scenarios and A(i,1) is less than a set threshold r, merging other prediction scenarios into the typical prediction scenario A(i,1);
[0016] S33, increase the value of i by one, return to step S32, continue to merge the prediction scenarios, until all prediction scenarios are merged with the typical prediction scenarios, extract all the typical prediction scenarios, and obtain the typical prediction scenarios of the power grid data.
[0017] Each set of historical data can generate a prediction scenario. In order to reduce the number of prediction scenarios, similar prediction scenarios are merged. When the Euclidean distance between two prediction scenarios is less than the set threshold, the two prediction scenarios are merged. At the same time, in order to prevent error transmission, that is, the Euclidean distance between prediction scenarios A (0,1) and A (0,2), A (0,3) and A (0,2) is less than the set threshold, but the Euclidean distance between prediction scenarios A (0,1) and A (0,3) is greater than the set threshold, only one prediction scenario is used as the benchmark each time it is merged. When A (0,1) is used as the benchmark, A (0,3) and A (0,2) will not be merged.
[0018] In step S13, the selecting of typical prediction scenarios and storing them in the cache further includes the following steps:
[0019] S41, let Occ represent the cache occupancy rate, Vum represent the capacity of the cache space, c1 represent the space occupied by a single typical prediction scenario in the cache, calculate the number num of typical prediction scenarios that can be stored in the cache under the occupancy rate Occ, num=Vum×Occ / c1;
[0020] S42, calculate the contribution value tscr of each typical prediction scenario to the response time, tscr=n×[space1+∑(1 / k×spacek)], where n is the number of prediction scenarios merged by the typical prediction scenario, the typical prediction scenario itself is also a prediction scenario, and all n are calculated starting from 1; space1 is the contribution of the typical prediction scenario itself to the response time, spacek is the contribution of the typical prediction scenario and other prediction scenarios to the response time, and k is a coefficient;
[0021] S43, extract input features from typical prediction scenarios, extract continuous features from input features, let B1, B2, ..., Bm represent the extracted continuous features, m is the number of continuous features, take B1, B2, ..., Bm as the sphere center, and the threshold r as the radius, to obtain the expression of m sphere models, each sphere model corresponds to a typical prediction scenario; let Dj represent the jth sphere model, combine Dj with other sphere models in pairs, and obtain other sphere models intersecting with Dj, according to the intersection of Dj and other sphere models, divide the sphere corresponding to Dj into different areas, if the area does not intersect with other sphere models, then calculate the proportion of the area in the sphere and record it as space1; if the area intersects with other k sphere models, then calculate the proportion of the sum of all the areas intersecting with other k sphere models in the sphere and record it as spacek; sort the contribution value of the response time according to the typical prediction scenario, and select the typical prediction scenarios from large to small to store them in the cache.
[0022] When the Euclidean distance between the input features of the power grid data and the input features of the typical prediction scenario or the prediction scenario is less than r, it is judged that the typical prediction scenario or the prediction scenario matches the power grid data; for two typical prediction scenarios, the Euclidean distance between the input features of the two is greater than r, but on the sphere formed with the input features of the two as the center and r as the radius, there may be an addition situation; when the input features of the power grid data are in the intersecting area, any one of the two typical prediction scenarios can be matched with the power grid data. In this case, the two typical prediction scenarios share the contribution value, and the area where multiple typical prediction scenarios intersect is similar; the contribution of the typical prediction scenario itself to the response time is the part that does not intersect with other prediction scenarios; the number of prediction scenarios merged by the typical prediction scenario reflects the frequency of occurrence of the typical prediction scenario. The higher the frequency, the higher the probability that the power grid data matches the typical prediction scenario, and the contribution value of the typical scenario is positively correlated with the number of prediction scenarios merged by the typical prediction scenario; space1 and spacek can be determined by the volume of the area occupying the volume of the spherical model corresponding to the typical prediction scenario.
[0023] In step S13, the selecting of typical prediction scenarios and storing them in the cache further includes the following steps:
[0024] S51, preset the cache occupancy rate Occ, execute steps S41, S42 and S43, determine the typical prediction scenarios stored in the cache; randomly extract historical power grid data and match the scenarios with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, it is judged that the cache hits, otherwise it is judged that the cache misses, and the cache hit rate Acc of the historical power grid data is determined, and the comprehensive efficiency P under the current cache occupancy rate is calculated, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to match the scenarios from the database and determine the predicted value of the power grid data, and V1 represents the time required to match the scenarios from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when the scenarios are matched in the cache;
[0025] S52, set the time step, add the cache occupancy rate Occ to the time step, re-execute steps S41, S42 and S43, and determine the comprehensive efficiency after adding the time step in the same manner as step S51; under all cache occupancies, take the typical prediction scenario storage method when the comprehensive efficiency reaches the highest as the final result.
[0026] The time for scene matching from the database and from the cache, and the cache space occupied by a single set of power grid data when scene matching is performed in the cache can be determined by measuring historical data; after the typical prediction scenario stored in the cache is determined, there is no need to actually store it. By randomly extracting historical data and performing scene matching with the typical prediction scenario determined to be stored in the cache, the cache hit rate can be obtained; the cache occupancy rate does affect the speed of scene matching and the ability to process matches simultaneously. At a high cache occupancy rate, there are many typical prediction scenarios stored in the cache and the cache hit rate is high, but the amount of power grid data that can be matched in parallel is reduced; anyway, at a low cache occupancy rate, there are few typical prediction scenarios stored in the cache and the cache hit rate is low, but the amount of power grid data that can be matched in parallel is increased, and both need to be taken into account to achieve fast response and maximize concurrent scene matching.
[0027] In step S14, the scenario matching of the power grid data with the typical prediction scenarios in the cache further includes the following steps:
[0028] S61, obtaining the power grid data to be predicted and extracting the target output features, calculating the Euclidean distance dis between the target output features and the input features of the typical prediction scenario, if there is a typical prediction scenario with a Euclidean distance dis less than r, proceeding to step S62; if there is no typical prediction scenario with a Euclidean distance dis less than r, matching the target output features with the prediction scenarios in the database, if there is no prediction scenario with a Euclidean distance less than r in the database, directly calculating the target output features using the prediction model to obtain a prediction value, if there is a prediction scenario with a Euclidean distance less than r in the database, taking the prediction value of the prediction scenario with the smallest Euclidean distance as the prediction value of the power grid data to be predicted;
[0029] S62, when the number of typical prediction scenarios with Euclidean distance dis less than r is 1, take the prediction value of the typical prediction scenario with Euclidean distance dis less than r as the prediction value of the power grid data to be predicted; when the number of typical prediction scenarios with Euclidean distance dis less than r is not 1, let dis1, dis2, …, disu represent the Euclidean distances between the target output features less than r and the typical prediction scenarios, u represents the number of typical prediction scenarios with Euclidean distance dis less than r, y1, y2, …, yu represent the prediction values corresponding to the typical prediction scenarios, and obtain the prediction value pre of the power grid data to be predicted by the average value, pre=(dis1+dis2+…+disu) / u.
[0030] For two typical prediction scenarios, the Euclidean distance between their input features is greater than r, but the Euclidean distance between the power grid data and the input features of the two typical prediction scenarios may be less than r, in which case both of them match the power grid data.
[0031] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: a data analysis-based intelligent prediction system for power grid data, comprising a power grid data storage module, a power grid data acquisition module, a power grid data analysis module and a response module; the power grid data storage module is interconnected with the power grid data analysis module, and is used to store prediction scenarios and typical prediction scenarios of power grid data; the output end of the power grid data acquisition module is connected to the input end of the power grid data analysis module, and is used to obtain various types of data of the power grid; the output end of the power grid data analysis module is connected to the input end of the response module, and is used to predict the data of the power grid, calculate the prediction scenario in advance through historical power grid data and store it in the power grid data storage module, and quickly provide prediction results according to the prediction scenarios stored in advance when there is a need for power grid data prediction; the response module allocates and manages power grid resources based on the prediction results of the power grid data.
[0032] The power grid data storage module also includes a database and a cache unit, wherein the database is used to store the prediction scenarios generated by the power grid data analysis module; the cache unit is used to store the typical prediction scenarios generated by the power grid data analysis module. The data analysis module also includes a scenario generation unit, a scenario merging unit, a prediction unit, a scenario matching unit and a scenario selection unit; the scenario generation unit generates prediction scenarios based on power grid data and prediction values, and the scenario merging unit is used to merge the prediction scenarios to obtain typical prediction scenarios; the prediction unit trains a prediction model for power grid data based on historical power grid data and predicts the power grid data; the scenario selection unit is used to determine the typical prediction scenarios that need to be stored in the cache; the scenario matching unit is used to match the typical prediction scenarios in the cache with the power grid data that needs to be predicted, and to match the prediction scenarios in the database with the power grid data that needs to be predicted. The scenario selection unit determines the final cache occupancy rate and the typical prediction scenario stored in the cache by calculating the comprehensive efficiency P under different cache occupancy rates Occ, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to perform scenario matching from the database and determine the predicted value of the power grid data, V1 represents the time required to perform scenario matching from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when performing scenario matching in the cache, and Acc represents the cache hit rate of historical power grid data.
[0033] Compared with the prior art, the beneficial effects of the present invention are: identifying common prediction scenarios in power grid operation, using historical data to train prediction models and perform pre-calculations, storing the results in a fast-access cache, and being able to quickly retrieve them when needed. When the prediction results of a certain scenario are needed, they can be obtained directly from the cache without having to execute the time-consuming calculation process again, thereby improving the prediction efficiency of power grid data; reasonably managing the cache occupancy rate, taking into account the cache hit rate and the ability to process matches simultaneously, while ensuring a fast response, maximizing the number of matching scenarios processed concurrently. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 The present invention is a schematic diagram of the structure of a power grid data intelligent prediction system based on data analysis. DETAILED DESCRIPTION
[0035] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0036] Example: Figure 1 As shown, the present invention provides a technical solution, a power grid data intelligent prediction system based on data analysis, including a power grid data storage module, a power grid data acquisition module, a power grid data analysis module and a response module; the power grid data storage module is interconnected with the power grid data analysis module, and is used to store prediction scenarios and typical prediction scenarios of power grid data; the output end of the power grid data acquisition module is connected to the input end of the power grid data analysis module, and is used to obtain various types of data of the power grid; the output end of the power grid data analysis module is connected to the input end of the response module, and is used to predict the data of the power grid, calculate the prediction scenario in advance through historical power grid data and store it in the power grid data storage module, and quickly provide prediction results according to the prediction scenarios stored in advance when there is a need for power grid data prediction; the response module allocates and manages power grid resources based on the prediction results of the power grid data.
[0037] The power grid data storage module also includes a database and a cache unit, wherein the database is used to store the prediction scenarios generated by the power grid data analysis module; the cache unit is used to store the typical prediction scenarios generated by the power grid data analysis module. The data analysis module also includes a scenario generation unit, a scenario merging unit, a prediction unit, a scenario matching unit and a scenario selection unit; the scenario generation unit generates prediction scenarios based on power grid data and prediction values, and the scenario merging unit is used to merge the prediction scenarios to obtain typical prediction scenarios; the prediction unit trains a prediction model for power grid data based on historical power grid data and predicts the power grid data; the scenario selection unit is used to determine the typical prediction scenarios that need to be stored in the cache; the scenario matching unit is used to match the typical prediction scenarios in the cache with the power grid data that needs to be predicted, and to match the prediction scenarios in the database with the power grid data that needs to be predicted. The scenario selection unit determines the final cache occupancy rate and the typical prediction scenario stored in the cache by calculating the comprehensive efficiency P under different cache occupancy rates Occ, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to perform scenario matching from the database and determine the predicted value of the power grid data, V1 represents the time required to perform scenario matching from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when performing scenario matching in the cache, and Acc represents the cache hit rate of historical power grid data.
[0038] Embodiment: The present invention provides a technical solution, a method for intelligent prediction of power grid data based on data analysis, comprising:
[0039] S11, acquiring historical data of the power grid, and generating a prediction scenario based on the historical data of the power grid;
[0040] Extracting grid data features from historical data of the grid, the grid data features include continuous features and discrete features, normalizing the continuous features to obtain input features, and labeling the discrete features;
[0041] A prediction model for power grid data is established and trained and optimized. The input of the prediction model is the input features with the same discrete features, and the output of the prediction model is the actual measured value of the power grid data. Through the prediction model of power grid data, the input features obtained through historical data and the input features set manually are optimized and calculated to obtain the predicted value of the power grid data. The input features are spliced with the predicted value of the power grid data to obtain the prediction scenario.
[0042] When predicting electricity load, continuous features include temperature, humidity, load growth rate and load at the previous moment, and discrete features include period, time period and weather; period is used to reflect holiday information, time period is used to reflect peak and valley time information of electricity consumption, and weather is used to reflect weather condition information; different values are assigned to different scenes to distinguish them; ordinary time period, peak time period and valley time period are assigned values of 1, 2 and 3 respectively, weekends and weekdays are assigned values of 1 and 2 respectively, and different weather conditions are assigned different values. In this way, the characteristics of the scene are determined according to the combination of discrete features; in the historical data with the same discrete feature combination, the historical data is divided into training set and test set, the prediction model is trained with the historical data of the training set, and optimized with the data of the test set, and the deep learning model can be used as the prediction model; in the short-term electricity symbol prediction at the minute level, the current moment is time t, and the input of the model is the load, load growth rate, temperature and humidity before time t. The order of the input, that is, how many load and load growth rate data are used for prediction, is adjusted during the model training process.
[0043] S12, according to the prediction scenario of the power grid data, execute steps S31, S32 and S33 to obtain a typical prediction scenario:
[0044] S31, let i=0, A(i,1), A(i,2),…,A(i,n i ) represents the prediction scenario that is not merged into the typical prediction scenario; n i Indicates the number of forecast scenarios that are not incorporated by the typical forecast scenario;
[0045] S32, taking A(i,1) as a typical prediction scenario, calculating the Euclidean distance between other prediction scenarios and A(i,1), and if the Euclidean distance between other prediction scenarios and A(i,1) is less than a set threshold r, merging other prediction scenarios into the typical prediction scenario A(i,1);
[0046] S33, increase the value of i by one, return to step S32, continue to merge the prediction scenarios, until all prediction scenarios are merged with the typical prediction scenarios, extract all the typical prediction scenarios, and obtain the typical prediction scenarios of the power grid data.
[0047] S13, selecting typical prediction scenarios and storing them in the cache, and storing all prediction scenarios in the database;
[0048] First, the prediction scenarios that are not merged into the typical prediction scenarios are A(0,1), A(0,2), ..., A(0,n 0 ), taking A(0,1) as the typical prediction scenario, the merging process begins. After the merging, the prediction scenarios that are not merged by the typical prediction scenarios become A(1,1), A(1,2), …, A(1,n 1), A(1,1) is taken as the typical prediction scenario and merged again, and so on, until all prediction scenarios are merged; in particular, A(0,1) and A(1,1) are also considered to be merged by the typical prediction scenario;
[0049] Let Occ represent the cache occupancy, Vum represent the capacity of the cache space, c1 represent the space occupied by a single typical prediction scenario in the cache, calculate the number of typical prediction scenarios num that can be stored in the cache under the occupancy Occ, num=Vum×Occ / c1;
[0050] Calculate the contribution value tscr of each typical prediction scenario to the response time, tscr=n×[space1+∑(1 / k×spacek)], where n is the number of prediction scenarios merged by the typical prediction scenario, and the typical prediction scenario itself is also a prediction scenario, so all n are calculated starting from 1; space1 is the contribution of the typical prediction scenario itself to the response time, spacek is the contribution of the typical prediction scenario and other prediction scenarios to the response time, and k is the coefficient;
[0051] Extract input features from typical prediction scenarios, extract continuous features from input features, let B1, B2, ..., Bm represent the extracted continuous features, m is the number of continuous features, take B1, B2, ..., Bm as the sphere center, and the threshold r as the radius, and get the expression of m sphere models, each sphere model corresponds to a typical prediction scenario; let Dj represent the jth sphere model, and combine Dj with other sphere models to obtain other sphere models that intersect with Dj. According to the intersection of Dj and other sphere models, the sphere corresponding to Dj is divided into different regions. If the region does not intersect with other sphere models, calculate the proportion of the region in the sphere and record it as space1; if the region intersects with other k sphere models, calculate the proportion of the sum of all the intersecting regions with other k sphere models in the sphere and record it as spacek; sort the contribution values of the response time according to the typical prediction scenarios, and select the typical prediction scenarios from large to small and store them in the cache. In addition to directly solving the equations jointly, whether they intersect can also be judged based on the distance between the sphere centers.
[0052] The cache occupancy rate Occ is found to have the optimal value through grid search method:
[0053] Preset the cache occupancy rate Occ, execute steps S41, S42 and S43, determine the typical prediction scenarios stored in the cache; randomly extract historical power grid data and match the scenarios with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, it is judged that the cache hits, otherwise it is judged that the cache misses, and the cache hit rate Acc of the historical power grid data is determined, and the comprehensive efficiency P under the current cache occupancy rate is calculated, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to match the scenarios from the database and determine the predicted value of the power grid data, and V1 represents the time required to match the scenarios from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when the scenarios are matched in the cache;
[0054] Set the time step, add the cache occupancy rate Occ to the time step, re-execute steps S41, S42 and S43, and determine the comprehensive efficiency after adding the time step in the same manner as step S51; under all cache occupancies, take the typical prediction scenario storage method when the comprehensive efficiency reaches the highest as the final result.
[0055] S14, when it is necessary to predict the power grid data, firstly, the power grid data is scene-matched with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, the prediction value of the power grid data is obtained according to the typical prediction scenario in the cache; if there is no matching typical prediction scenario in the cache, the power grid data is scene-matched with the prediction scenario in the database; if there is a matching prediction scenario in the data, the prediction value of the power grid data is obtained according to the prediction scenario in the database, otherwise the prediction value is directly calculated for the power grid data;
[0056] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, the embodiments should be considered exemplary and non-restrictive in all respects, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present invention.
Claims
1. A method for intelligent prediction of power grid data based on data analysis, characterized in that: The following steps are involved: S11, acquiring historical data of the power grid, and generating a prediction scenario based on the historical data of the power grid; S12, obtaining a typical prediction scenario based on the prediction scenario of the power grid data; S13, selecting typical prediction scenarios to store in the cache, and storing all prediction scenarios in the database; comprising the following steps: S41, let Occ represent the cache occupancy rate, Vum represent the capacity of the cache space, c1 represent the space occupied by a single typical prediction scenario in the cache, calculate the number num of typical prediction scenarios that can be stored in the cache under the occupancy rate Occ, num=Vum×Occ / c1; S42, calculate the contribution value tscr of each typical prediction scenario to the response time, tscr=n×[space1+∑(1 / k×spacek)], where n is the number of prediction scenarios merged by the typical prediction scenario, the typical prediction scenario itself is also a prediction scenario, and all n are calculated starting from 1; space1 is the contribution of the typical prediction scenario itself to the response time, spacek is the contribution of the typical prediction scenario and other prediction scenarios to the response time, and k is a coefficient; S43, extract input features from typical prediction scenarios, extract continuous features from input features, let B1, B2, ..., Bm represent the extracted continuous features, m is the number of continuous features, take B1, B2, ..., Bm as the sphere center, and the threshold r as the radius, to obtain the expression of m sphere models, each sphere model corresponds to a typical prediction scenario; let Dj represent the jth sphere model, combine Dj with other sphere models in pairs, and obtain other sphere models intersecting with Dj, according to the intersection of Dj and other sphere models, divide the sphere corresponding to Dj into different regions, if the region does not intersect with other sphere models, then calculate the proportion of the region in the sphere and record it as space1; if the region intersects with other k sphere models, then calculate the proportion of the sum of all the intersecting regions with other k sphere models in the sphere and record it as spacek; sort the contribution value of the response time according to the typical prediction scenario, select the typical prediction scenarios from large to small and store them in the cache; S14, when it is necessary to predict the power grid data, firstly match the power grid data with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, the prediction value of the power grid data is obtained according to the typical prediction scenario in the cache; if there is no matching typical prediction scenario in the cache, match the power grid data with the prediction scenarios in the database; if there is a matching prediction scenario in the data, the prediction value of the power grid data is obtained according to the prediction scenario in the database, otherwise the prediction value is directly calculated for the power grid data.
2. The method for intelligent prediction of power grid data based on data analysis according to claim 1, characterized in that: In step S11, the generation of prediction scenarios based on historical data of the power grid further includes the following steps: Extracting grid data features from historical data of the grid, the grid data features include continuous features and discrete features, normalizing the continuous features to obtain input features, and labeling the discrete features; A prediction model for power grid data is established and trained and optimized. The input of the prediction model is the input features with the same discrete features, and the output of the prediction model is the actual measured value of the power grid data. Through the prediction model of power grid data, the input features obtained through historical data and the input features set manually are optimized and calculated to obtain the predicted value of the power grid data. The input features are spliced with the predicted value of the power grid data to obtain the prediction scenario.
3. The method for intelligent prediction of power grid data based on data analysis according to claim 2 is characterized in that: In step S12, obtaining a typical prediction scenario based on the prediction scenario of the power grid data further includes the following steps: S31, let i=0, A(i,1), A(i,2),…,A(i,n i ) represents the prediction scenario that is not merged into the typical prediction scenario; n i Indicates the number of forecast scenarios that are not incorporated by the typical forecast scenario; S32, taking A(i,1) as a typical prediction scenario, calculating the Euclidean distance between other prediction scenarios and A(i,1), and if the Euclidean distance between other prediction scenarios and A(i,1) is less than a set threshold r, merging other prediction scenarios into the typical prediction scenario A(i,1); S33, increase the value of i by one, return to step S32, continue to merge the prediction scenarios, until all prediction scenarios are merged with the typical prediction scenarios, extract all the typical prediction scenarios, and obtain the typical prediction scenarios of the power grid data.
4. The method for intelligent prediction of power grid data based on data analysis according to claim 3 is characterized in that: In step S13, the selecting of typical prediction scenarios and storing them in the cache further includes the following steps: S51, preset the cache occupancy rate Occ, execute steps S41, S42 and S43, determine the typical prediction scenarios stored in the cache; randomly extract historical power grid data and match the scenarios with the typical prediction scenarios in the cache. If there is a matching typical prediction scenario in the cache, it is judged that the cache hits, otherwise it is judged that the cache misses, and the cache hit rate Acc of the historical power grid data is determined, and the comprehensive efficiency P under the current cache occupancy rate is calculated, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to match the scenarios from the database and determine the predicted value of the power grid data, and V1 represents the time required to match the scenarios from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when the scenarios are matched in the cache; S52, set the time step, add the cache occupancy rate Occ to the time step, re-execute steps S41, S42 and S43, and determine the comprehensive efficiency after adding the time step in the same manner as step S51; under all cache occupancies, take the typical prediction scenario storage method when the comprehensive efficiency reaches the highest as the final result.
5. The method for intelligent prediction of power grid data based on data analysis according to claim 4 is characterized in that: In step S14, the scenario matching of the power grid data with the typical prediction scenarios in the cache further includes the following steps: S61, obtaining the power grid data to be predicted and extracting the target output features, calculating the Euclidean distance dis between the target output features and the input features of the typical prediction scenario, if there is a typical prediction scenario with a Euclidean distance dis less than r, proceeding to step S62; if there is no typical prediction scenario with a Euclidean distance dis less than r, matching the target output features with the prediction scenarios in the database, if there is no prediction scenario with a Euclidean distance less than r in the database, directly calculating the target output features using the prediction model to obtain a prediction value, if there is a prediction scenario with a Euclidean distance less than r in the database, taking the prediction value of the prediction scenario with the smallest Euclidean distance as the prediction value of the power grid data to be predicted; S62, when the number of typical prediction scenarios with Euclidean distance dis less than r is 1, take the prediction value of the typical prediction scenario with Euclidean distance dis less than r as the prediction value of the power grid data to be predicted; when the number of typical prediction scenarios with Euclidean distance dis less than r is not 1, let dis1, dis2, …, disu represent the Euclidean distances between the target output features less than r and the typical prediction scenarios, u represents the number of typical prediction scenarios with Euclidean distance dis less than r, y1, y2, …, yu represent the prediction values corresponding to the typical prediction scenarios, and obtain the prediction value pre of the power grid data to be predicted by the average value, pre=(dis1+dis2+…+disu) / u.
6. A power grid data intelligent prediction system based on data analysis, using a power grid data intelligent prediction method based on data analysis as described in any one of claims 1 to 5, characterized in that: include: Power grid data storage module, power grid data acquisition module, power grid data analysis module and response module; The power grid data storage module is interconnected with the power grid data analysis module, and is used to store prediction scenarios and typical prediction scenarios of power grid data; the output end of the power grid data acquisition module is connected to the input end of the power grid data analysis module, and is used to obtain various types of data of the power grid; the output end of the power grid data analysis module is connected to the input end of the response module, and is used to predict the data of the power grid, calculate the prediction scenario in advance through historical power grid data and store it in the power grid data storage module, and quickly provide prediction results according to the prediction scenarios stored in advance when there is a need for power grid data prediction; the response module allocates and manages power grid resources based on the prediction results of the power grid data.
7. The power grid data intelligent prediction system based on data analysis according to claim 6 is characterized by: The power grid data storage module also includes a database and a cache unit. The database is used to store the prediction scenarios generated by the power grid data analysis module; the cache unit is used to store the typical prediction scenarios generated by the power grid data analysis module.
8. The power grid data intelligent prediction system based on data analysis according to claim 7 is characterized by: The data analysis module further includes a scenario generation unit, a scenario merging unit, a prediction unit, a scenario matching unit and a scenario selection unit; the scenario generation unit generates a prediction scenario based on the power grid data and the prediction value, and the scenario merging unit is used to merge the prediction scenarios to obtain a typical prediction scenario; the prediction unit trains a prediction model for power grid data based on historical power grid data to predict the power grid data; The scene selection unit is used to determine the typical prediction scenes that need to be stored in the cache; The scenario matching unit is used to match the typical prediction scenarios in the cache with the power grid data that needs to be predicted, and to match the prediction scenarios in the database with the power grid data that needs to be predicted.
9. The power grid data intelligent prediction system based on data analysis according to claim 8 is characterized in that: The scenario selection unit determines the final cache occupancy rate and the typical prediction scenario stored in the cache by calculating the comprehensive efficiency P under different cache occupancy rates Occ, P=Acc×(V2-V1)×Vum×(1-Occ) / c2, where V2 represents the time required to match the scenario from the database and determine the predicted value of the power grid data, and V1 represents the time required to match the scenario from the cache and determine the predicted value of the power grid data when there is a typical prediction scenario matching the current power grid data in the cache; c2 is the cache space occupied by a single set of power grid data when scene matching is performed in the cache, and Acc represents the cache hit rate of historical power grid data.
Citation Information
Patent Citations
Search history-based content automatic filling system
CN107748804A
Microgrid two-stage distribution robust optimization scheduling method based on Hausdorff distance
CN116957229A
Digital construction method and system for typical planning scene of county power distribution network
CN118133341A