A method, related device, equipment and storage medium for power grid load forecasting
Through the deep learning model combining time units and periodic information, multiple periodic features are extracted, which solves various periodic problems in the existing model's difficulty in predicting grid load, and improves prediction accuracy.
Patent Information
- Application Number
- CN202011189883.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-10-30
AI Technical Summary
Existing machine learning models are difficult to effectively predict multiple periodic time series in grid load, resulting in insufficient prediction accuracy.
By obtaining historical and target feature matrices and vectors, combining information from time units and periods, deep learning models are used to predict grid loads, especially extracting multiple periodic features through local sequence modules and global context modules.
It improves the accuracy of grid load prediction, can better capture complex periodic characteristics, and enhances the ability to predict grid load changes.
Smart Images

Figure CN112288595B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of deep learning technology, and in particular, to a method for predicting power grid load, related devices, equipment, and storage media. Background Art
[0002] The smart grid aims to create an automated and efficient energy delivery network, which can meet the requirements of network security, energy efficiency, and demand-side management by improving the reliability and quality of power delivery. Predicting the power grid load is the key to ensuring the safe and reliable operation of the power grid and reducing the economic losses of the power grid. Improving the accuracy of load prediction has been the focus of research for many years.
[0003] Currently, there are many prediction methods for power grid load. For example, machine learning models such as feedforward neural networks, recurrent neural networks, sequence-to-sequence neural networks, or temporal convolutional neural networks are used to predict the power grid load. Applying machine learning algorithms to the prediction of power grid load can handle problems such as the volatility and randomness of load data, and improve the accuracy of power grid load prediction to a certain extent, providing a basis for the management and scheduling of the power grid.
[0004] Due to the economic activities of people in the city and the seasonality of the climate, the load sequence often shows strong periodicity and trend. Existing machine learning models only provide solutions for single-period time series, and there is no solution for the power grid load prediction of multi-period time series. Summary of the Invention
[0005] The embodiments of the present application provide a method for predicting power grid load, related devices, equipment, and storage media, which can jointly predict the power grid load based on time units and periods, and can better capture complex periodicity, thereby improving the accuracy of power grid load prediction.
[0006] In view of this, on the one hand, the present application provides a method for predicting power grid load, including:
[0007] Obtain a historical feature matrix, where the historical feature matrix includes the periodic load of the historical period and the periodic description features, the historical period includes D time units, each time unit includes T time sub-units, and both D and T are integers greater than 1;
[0008] Obtain a historical feature vector, where the historical feature vector includes the historical load of the historical time unit and the historical description features, the historical time unit is the next time unit after the historical period appears, and the historical time unit includes T time sub-units;
[0009] Obtain a target feature matrix, where the target feature matrix includes the periodic load of the target period and the periodic description features, and the target period includes D time units;
[0010] Obtain a target feature vector, where the target feature vector includes the historical load of the target time unit and the description features of the first predicted time unit. The target time unit is an adjacent time unit that appears before the first predicted time unit. The first predicted time unit is the next time unit that appears after the target period. The target time unit includes T time sub-units, and the first predicted time unit includes T time sub-units;
[0011] Predict the grid load of the first predicted time unit based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector to obtain the first grid load prediction result.
[0012] On the other hand, the present application provides a grid load prediction device, including:
[0013] An acquisition module, configured to acquire a historical feature matrix, where the historical feature matrix includes the periodic load of the historical period and the periodic description features. The historical period includes D time units, and each time unit includes T time sub-units. Both D and T are integers greater than 1;
[0014] The acquisition module is further configured to acquire a historical feature vector, where the historical feature vector includes the historical load of the historical time unit and the historical description features. The historical time unit is the next time unit that appears after the historical period, and the historical time unit includes T time sub-units;
[0015] The acquisition module is further configured to acquire a target feature matrix, where the target feature matrix includes the periodic load of the target period and the periodic description features, and the target period includes D time units;
[0016] The acquisition module is further configured to acquire a target feature vector, where the target feature vector includes the historical load of the target time unit and the description features of the first predicted time unit. The target time unit is an adjacent time unit that appears before the first predicted time unit. The first predicted time unit is the next time unit that appears after the target period. The target time unit includes T time sub-units, and the first predicted time unit includes T time sub-units;
[0017] A prediction module, configured to predict the grid load of the first predicted time unit based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector to obtain the first grid load prediction result.
[0018] In a possible design, in an implementation manner of the other aspect of the embodiments of the present application,
[0019] An acquisition module, specifically used to acquire the historical load corresponding to each time unit within a historical period;
[0020] Acquire the historical description information corresponding to each time unit within a historical period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0021] Encode the historical load corresponding to each time unit within a historical period and the historical description information corresponding to each time unit within a historical period to obtain a historical feature matrix.
[0022] In a possible design, in another implementation manner on the other hand of the embodiments of the present application,
[0023] An acquisition module, specifically used to acquire the historical load corresponding to a historical time unit;
[0024] Acquire the historical description information corresponding to a historical time unit, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0025] Encode the historical load corresponding to a historical time unit and the historical description information corresponding to a historical time unit to obtain a historical feature vector.
[0026] In a possible design, in another implementation manner on the other hand of the embodiments of the present application,
[0027] An acquisition module, specifically used to acquire the historical load corresponding to each time unit within a target period;
[0028] Acquire the historical description information corresponding to each time unit within a target period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0029] Encode the historical load corresponding to each time unit within a target period and the historical description information corresponding to each time unit within a target period to obtain a target feature matrix.
[0030] In a possible design, in another implementation manner on the other hand of the embodiments of the present application,
[0031] An acquisition module, specifically used to acquire the historical load corresponding to a target time unit;
[0032] Acquire the description information corresponding to a first prediction time unit, where the description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0033] Encode the historical load corresponding to the target time unit and the description information corresponding to the first prediction time unit to obtain a target feature vector.
[0034] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0035] The prediction module is specifically configured to obtain a first feature matrix through a local sequence module based on the target feature matrix and the target feature vector;
[0036] Obtain a second feature matrix through a global context module based on the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector;
[0037] Obtain the power grid load prediction result corresponding to the first prediction time unit through a linear layer based on the first feature matrix and the second feature matrix, where the power grid load prediction result includes the load corresponding to each time subunit in the first prediction time unit.
[0038] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0039] The prediction module is specifically configured to perform attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain an encoding result of the target feature matrix;
[0040] Perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain a first feature matrix.
[0041] In a possible design, in another implementation of another aspect of the embodiments of the present application, the encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least one first attention head and one second attention head;
[0042] The prediction module is specifically configured to, for each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the first attention head to obtain a first encoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0043] For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the second attention head to obtain a second encoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0044] Obtain the encoding result of the target feature matrix according to the first encoding result and the second encoding result.
[0045] In a possible design, in another implementation of another aspect of the embodiments of the present application, the decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head;
[0046] A prediction module, specifically configured to, for each multi-head attention layer in the decoder, perform attention decoding processing on the encoded result of the target feature matrix by the first attention head and the target feature vector to obtain a first decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0047] For each multi-head attention layer in the decoder, perform attention decoding processing on the encoded result of the target feature matrix by the second attention head and the target feature vector to obtain a second decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0048] Obtain a first feature matrix according to the first decoding result and the second decoding result.
[0049] In a possible design, in another implementation of another aspect of the embodiments of the present application,
[0050] A prediction module, specifically configured to perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain an encoded result of the historical feature matrix;
[0051] Perform attention encoding processing on the historical feature vector according to the context encoder included in the global context module to obtain an encoded result of the historical feature vector;
[0052] Based on the context attention mechanism, calculate the encoded result of the historical feature matrix, the encoded result of the historical feature vector, and the encoded result of the target feature matrix to obtain an attention encoded result;
[0053] According to the context decoder included in the global context module, perform attention decoding processing on the attention encoded result and the target feature vector to obtain a second feature matrix.
[0054] In a possible design, in another implementation of another aspect of the embodiments of the present application, the context encoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head;
[0055] A prediction module, specifically configured to perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain an encoded result of the historical feature matrix, including:
[0056] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix based on the first attention head to obtain a third encoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0057] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix based on the second attention head to obtain a fourth encoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0058] Obtain the encoding result of the historical feature matrix according to the third encoding result and the fourth encoding result;
[0059] The prediction module is specifically configured to, for each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vectors based on the first attention head to obtain a fifth encoding result;
[0060] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vectors based on the second attention head to obtain a sixth encoding result;
[0061] Obtain the encoding result of the historical feature vectors according to the fifth encoding result and the sixth encoding result.
[0062] In a possible design, in another implementation manner of another aspect of the embodiments of the present application, the context decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head;
[0063] The prediction module is specifically configured to, for each multi-head attention layer in the context decoder, perform attention decoding processing on the attention encoding result and the target feature vectors based on the first attention head to obtain a third decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0064] For each multi-head attention layer in the decoder, perform attention decoding processing on the attention encoding result and the target feature vectors based on the second attention head to obtain a fourth decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0065] Obtain the second feature matrix according to the third decoding result and the fourth decoding result.
[0066] In a possible design, in another implementation manner of another aspect of the embodiments of the present application,
[0067] The acquisition module is further configured to, after the prediction module predicts the grid load of the first prediction time unit based on the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector to obtain the first grid load prediction result, acquire the updated feature matrix, where the updated feature matrix includes the periodic load and the periodic description features of the updated period, the updated period includes D time units, and the updated period includes the first prediction time unit;
[0068] The acquisition module is further configured to acquire the updated feature vector, where the updated feature vector includes the grid load prediction result corresponding to the first prediction time unit and the description features of the second prediction time unit, the first prediction time unit is an adjacent time unit that appears before the second prediction time unit, the second prediction time unit is the next time unit that appears after the updated period, and the second prediction time unit includes T time sub-units;
[0069] The prediction module is further configured to predict the grid load of the second prediction time unit based on the historical feature matrix, the historical feature vector, the updated feature matrix, and the updated feature vector to obtain the second grid load prediction result.
[0070] Another aspect of the present application provides a computer device, including: a memory and a processor;
[0071] Wherein, the memory is used to store programs;
[0072] The processor is configured to execute the programs in the memory, and the processor is configured to execute the methods in the above aspects according to the instructions in the program code.
[0073] Another aspect of the present application provides a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the methods in the above aspects.
[0074] Another aspect of the present application provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the above aspects.
[0075] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0076] In an embodiment of the present application, a method for power grid load forecasting is provided. A historical feature matrix is obtained, which includes the periodic load and periodic description features of a historical period. A historical feature vector is obtained, which includes the historical load and historical description features of a historical time unit. The historical time unit is the next time unit after the historical period appears. A target feature matrix is obtained, which includes the periodic load and periodic description features of a target period. A target feature vector is obtained, which includes the historical load of a target time unit and the description features of a first predicted time unit. Based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector, the power grid load of the first predicted time unit is predicted to obtain a first power grid load forecasting result. By the above method, in the long term, based on the target feature matrix and a large number of historical feature matrices, the power grid load of a future time unit can be predicted. The historical feature matrix within the historical period explicitly provides global information, which helps to query and locate a period with a similar load trend in a relatively long historical period for the target period. In the short term, the load sequence has the characteristic of multi-periodicity. Therefore, jointly predicting the power grid load based on the time unit and the period can better capture the complex periodicity, thereby improving the accuracy of power grid load forecasting. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 is a schematic diagram of an architecture of a smart grid in an embodiment of the present application;
[0078] Figure 2 is a schematic diagram of load sequence decomposition based on a public dataset in an embodiment of the present application;
[0079] Figure 3 is a schematic flowchart of a power grid load forecasting method in an embodiment of the present application;
[0080] Figure 4 is a schematic diagram of an embodiment for generating a power grid load forecasting result in an embodiment of the present application;
[0081] Figure 5 is a schematic diagram of a hierarchical network structure of a load forecasting model in an embodiment of the present application;
[0082] Figure 6 is a schematic diagram of a network structure of a local sequence module in an embodiment of the present application;
[0083] Figure 7 is a schematic diagram of a network structure of an encoder in a local sequence module in an embodiment of the present application;
[0084] Figure 8 is a schematic diagram of a network structure of a decoder in a local sequence module in an embodiment of the present application;
[0085] Figure 9 It is a schematic diagram of a network structure of the global context module in the embodiment of the present application;
[0086] Figure 10 It is a schematic diagram of a network structure of the context encoder in the global context module in the embodiment of the present application;
[0087] Figure 11 It is a schematic diagram of a network structure of the context decoder in the global context module in the embodiment of the present application;
[0088] Figure 12 It is a schematic diagram of an embodiment of the power grid load prediction device in the embodiment of the present application;
[0089] Figure 13 It is a schematic diagram of a structure of the terminal device in the embodiment of the present application;
[0090] Figure 14 It is a schematic diagram of a structure of the server in the embodiment of the present application. Detailed implementation manners
[0091] The embodiment of the present application provides a method, related device, equipment and storage medium for power grid load prediction, which jointly predicts the power grid load based on time units and cycles, and can better capture complex periodicity, thereby improving the accuracy of power grid load prediction.
[0092] The terms "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "correspond to" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or equipment that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.
[0093] Simply put, the smart grid is the combination of the power grid and Internet technology (IT). It collects real-time data from various digital sensors in the grid, combines the asset data of power equipment, and monitors, analyzes, and statistically processes this data through IT means to achieve intelligent and automated control of the grid. The smart grid aims to create an automated and efficient energy delivery network that meets requirements in aspects such as network security, energy efficiency, and demand-side management by improving the reliability and quality of power transmission. Modern power distribution systems are equipped with advanced monitoring infrastructure capable of generating a large amount of data for fine-grained analysis and improved prediction performance. In the energy field, power load forecasting is a key task as it can provide support for decision-making, assist relevant staff in formulating pricing strategies, and enable seamless integration of renewable energy and reduction of maintenance costs, etc.
[0094] The smart grid can collect data faster and more accurately. Technicians can observe the state of power flow across the entire network in real time, obtain data on high-fault areas of equipment, and make timely adjustments to make the grid more intelligent. For ease of understanding, please refer to Figure 1 , Figure 1 which is a schematic diagram of the architecture of the smart grid in the embodiments of this application. As shown in the figure, the smart grid includes processes such as smart power generation, smart power transmission, smart power transformation, smart power distribution, and smart power consumption.
[0095] In addition to wind power generation and hydropower generation, smart power generation also involves power generation from conventional energy sources, clean energy, and large-capacity energy storage applications. In terms of conventional energy sources, it mainly involves key technologies for the coordination of conventional power grid plants, the development of coordinated control systems and equipment for connecting large-scale energy base unit groups to the grid, and optimization control systems for hydropower, thermal power, and nuclear power units. In terms of clean energy, it mainly conducts research and development and popularization of advanced technologies such as modeling, system simulation, power prediction, and grid-connected operation control of wind farms and photovoltaic power plants, and develops safety and stability control systems for large-scale renewable energy access to the grid, integrated control and reliability assessment systems for renewable energy power generation stations, renewable energy power prediction systems, and complementary power generation and access systems for wind, light, and storage. In terms of energy storage applications, large-capacity energy storage devices need to be developed.
[0096] Smart power transmission requires reducing the risk of large-scale outages, mainly including transmission congestion management, supervisory control and data acquisition (SCADA) systems for power transmission, wide-area measurement systems (WAMS), Geographic Information System (GIS) technology for power transmission, advanced alarm visualization of energy management systems (EMS), and power transmission system simulation and modeling, etc.
[0097] Smart substation automatically completes basic functions such as information collection, measurement, control, protection, metering, and detection based on computer equipment. The power grid load forecasting method provided in this application can predict the power grid load based on artificial intelligence (AI) technology, construct a load forecasting model using machine learning (ML), and realize functions such as real-time automatic control, intelligent regulation, online analysis and decision-making, and collaborative interaction of the power grid through computer equipment.
[0098] AI uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, AI is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making. AI technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. AI basic technologies generally include technologies such as sensors, dedicated AI chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. AI software technologies mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and ML / deep learning.
[0099] Machine learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. ML is the core of AI and the fundamental way to make computers intelligent, and its applications cover all fields of AI. ML and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning.
[0100] Intelligent power distribution is a power management system re-developed according to the needs of users and following the standard specifications of the power distribution system. It has the characteristics of strong professionalism, high automation, easy to use, high performance, and high reliability. Through telemetry and remote control, the load can be reasonably allocated, optimized operation can be achieved, electric energy can be effectively saved, and peak and valley electricity consumption records are available, thus providing the necessary conditions for energy management.
[0101] Intelligent power consumption mainly includes applications such as industrial power consumption, residential power consumption, wind-solar energy storage, and electric vehicles. For example, a power fiber to the home communication network is constructed to cover the power fiber to the home network in the community, realizing the triple play of TV, telephone, and Internet for residential users in the community. A unified network management system is built in the community to manage each node of the communication network, realizing real-time monitoring of the network and equipment and rapid fault location. Another example is the electric vehicle charging facility, which includes charging piles, metering devices, control devices, power quality governance facilities, and a charging control system is deployed.
[0102] It should be noted that the computer device provided in this application can be a server or a terminal device. The server involved in this application can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The terminal device can be a smart phone, a tablet computer, a notebook computer, a palm computer, a personal computer, a smart TV, a smart watch, etc., but is not limited thereto.
[0103] Based on the above intelligent power grid, this application provides a power grid load forecasting method, which can jointly predict the power grid load based on time units and cycles. Among them, the forecasting of the power grid load can be carried out within different time spans, and its forecasting range varies from milliseconds to several years. Long term load forecasting (LTLF) usually refers to the forecasting of peaks and trends for one year or several years, and this kind of forecasting can even be extended to a range of several decades. This application mainly focuses on the power grid load forecasting one day in advance, which is also called short term load forecasting (STLF). Generally speaking, the demand side management and power supply in the next few hours or days have relatively high accuracy requirements for this short term load forecasting.
[0104] Due to the economic activities of people in the city and the seasonality of the climate, the load sequence often exhibits strong periodicity and trend. The seasonal and trend decomposition using loess (STL) method can be used to decompose the load sequence. Please refer to Figure 2 , Figure 2 This is a schematic diagram of the load sequence decomposition based on the public dataset in the embodiments of this application. As shown in the figure, the load sequence shows periodicity every day or every week. Within a one-year cycle, the load sequences in winter and summer show obvious characteristic differences, which is also called seasonality. The part remaining after deducting the trend and seasonality components of the load sequence can be regarded as aperiodic disturbance. Predicting the load sequence is a very challenging task because it has the following characteristics:
[0105] I. It has the characteristics of multi-periodicity and seasonality;
[0106] Since the load sequence often contains multiple periodicities, it is difficult to capture the periodic load sequence. The load sequence of a day can be simply divided into working hours and home hours, showing periodicity on a daily basis, with obvious characteristics of short-term power consumption peaks and valleys. In a longer time interval, the load sequence also has periodic characteristics on a weekly basis. In a year, regions with high electricity heating or cooling demand often show seasonal characteristics in summer or winter.
[0107] II. The characteristics of aperiodic disturbances;
[0108] In addition to the periodic components, the load sequence also includes an aperiodic disturbance part that varies with weather, holidays, and other random factors. In most cases, the load sequence shows relatively stable periodic characteristics, and obvious disturbances contrary to the periodicity only occur at a few moments. Therefore, some work also regards the prediction of this aperiodic disturbance as a data imbalance problem and designs some methods to evaluate the transient and steady states of the load sequence. However, it is still very difficult to capture the multi-periodic characteristics and aperiodic disturbances of the time series simultaneously.
[0109] Combined with the above introduction, the method for predicting the power grid load in this application will be introduced below. Please refer to Figure 3 , an embodiment of the power grid load prediction method in the embodiments of this application includes:
[0110] 101. Obtain a historical feature matrix, where the historical feature matrix includes the periodic load of the historical period and the periodic description features. The historical period includes D time units, and each time unit includes T time sub-units. Both D and T are integers greater than 1;
[0111] In this embodiment, the power grid load prediction device can obtain the historical feature matrix corresponding to the historical period. Each historical period has D time units, and each time unit has T time sub-units. Specifically, the historical feature matrix can be the power grid-related data of a certain historical period in the past year. Taking one time sub-unit as 1 hour, one time unit as 1 day, and one historical period as 7 days (i.e., one week) as an example, then D is 7 and T is 24. Based on this, the periodic load corresponding to one historical period includes the load of 7 * 24 hours, and the periodic description features corresponding to one historical period include 7 * M description features, where M represents the dimension of the description features. For example, the temperature corresponding to each day within 7 days. Optionally, the periodic description features corresponding to one historical period can also include 7 * 24 * M description features. For example, the temperature corresponding to each hour of each day within 7 days.
[0112] It can be understood that during the actual prediction process, the lengths of the time sub-units, the time units, and the periods can also be set according to requirements. This application uses D as 7 and T as 24 as an example for introduction, but this should not be construed as a limitation of this application. Generally, the power grid load prediction device needs to obtain N historical feature matrices within a long time period. For example, obtain the historical feature matrices of 52 historical periods (i.e., 52 weeks) within one year.
[0113] It should be noted that the power grid load prediction device provided in this application can be deployed on a server, or on a terminal device, or on a system composed of a server and a terminal device. This application does not make any limitations.
[0114] 102. Obtain a historical feature vector, where the historical feature vector includes the historical load of the historical time unit and the historical description features. The historical time unit is the next time unit after the historical period appears, and the historical time unit includes T time sub-units;
[0115] In this embodiment, the power grid load prediction device can obtain the historical feature vectors corresponding to a large number of historical time units. Each historical time unit also has T time sub-units. Based on this, for example, if a certain historical period is from July 6, 2020 to July 12, 2020, then the historical time unit corresponding to this historical period is July 13, 2020, that is, the historical time unit is the next time unit after the historical period appears.
[0116] Specifically, the historical load corresponding to a historical time unit includes the load for 1 * 24 hours, and the historical description features corresponding to a historical time unit include 1 * M description features, where M represents the dimension of the description features. For example, the temperature corresponding to 1 day. Optionally, the historical description features corresponding to a historical time unit can also include 1 * 24 * M description features. For example, the temperature corresponding to each hour within 1 day.
[0117] It can be understood that, usually, the power grid load prediction device needs to obtain N historical feature matrices within a long time period. For example, obtain the historical feature matrices for 52 historical cycles (i.e., 52 weeks) within one year. Therefore, the number of historical feature vectors is the same as that of the historical feature matrices, that is, N historical feature vectors are obtained.
[0118] 103. Obtain a target feature matrix, where the target feature matrix includes the periodic load and periodic description features of the target period, and the target period includes D time units;
[0119] In this embodiment, the power grid load prediction device can obtain the target feature matrix corresponding to the target period. The target period has D time units, and each time unit has T time subunits. Specifically, the target feature matrix can be the power grid-related data for the cycle closest to the current time. Taking one time subunit as 1 hour, one time unit as 1 day, and one target period as 7 days (i.e., one week) as an example, then D is 7 and T is 24. Based on this, the periodic load corresponding to one target period includes the load for 7 * 24 hours, and the periodic description features corresponding to one target period include 7 * M description features, or 7 * 24 * M description features.
[0120] 104. Obtain a target feature vector, where the target feature vector includes the historical load of the target time unit and the description features of the first prediction time unit. The target time unit is an adjacent time unit that appears before the first prediction time unit, the first prediction time unit is the next time unit that appears after the target period, the target time unit includes T time subunits, and the first prediction time unit includes T time subunits;
[0121] In this embodiment, the power grid load prediction device can obtain the target feature vector corresponding to the target time unit, and the target time unit also has T time subunits. Based on this, for example, if the target period is from October 12, 2020 to October 18, 2020, then the target time unit corresponding to this target period is October 18, 2020, and the first prediction time unit corresponding to this target period is October 19, 2020, that is, both the target time unit and the first prediction time unit include T time subunits (for example, both include 24 hours).
[0122] Specifically, the historical load corresponding to a target time unit includes the load for 1 * 24 hours, and the description features corresponding to the first prediction time unit include 1 * M description features, where M represents the dimension of the description features. For example, the temperature corresponding to 1 day. Optionally, the description features of the first prediction time unit may also include 1 * 24 * M description features. For example, the temperature corresponding to each hour within 1 day.
[0123] It should be noted that assuming the first prediction time unit is October 19, 2020, which is Monday, then the target time unit is October 18, 2020, which is Sunday. The target period is from October 12, 2020 to October 18, 2020, that is, from Monday to Sunday. Based on this, the historical period can be sourced from data of the past year. To better exhibit the characteristics of cycle similarity, the historical period and the target period can have the same cycle arrangement, that is, each historical period is from Monday to Sunday.
[0124] 105. Predict the grid load of the first prediction time unit based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector to obtain the first grid load prediction result.
[0125] In this embodiment, the grid load prediction device can input the historical feature matrix, historical feature vector, target feature matrix, and target feature vector into the trained load prediction model, and the load prediction model outputs the first grid load prediction result, that is, predicts the grid load of the next time unit.
[0126] Specifically, load prediction can be regarded as a supervised learning problem, and the input historical load sequence can be defined as It consists of load data with a fixed time lag. The lag window is selected empirically. Taking D as 7 and T as 24 as an example, the lag window can be the load sequence of the past week, and there is one load for each hour. The lag window size is n T , then n T = 7 × 24. When given the historical load data of a lag window, based on the constructed load prediction model f, the values of the next n O moments can be predicted, where n T represents the input sequence length of the load prediction model f, n O represents the output sequence length of the load prediction model f, n O can be set to 1 hour, or it can be in half-hour units. This application takes 1 hour as an example for illustration, but it should not be construed as a limitation of this application. Based on this, the sequence is defined as the load sequence output by the load prediction model f.
[0127] Meanwhile, assume that there are M dimensions of descriptive features per hour, and the matrix is used to represent the descriptive features corresponding to the input for n T hours, is the descriptive features corresponding to the output for n O hours. Thus, the problem of power grid load forecasting is represented by a mathematical expression as follows:
[0128]
[0129] where θ represents the model parameters of the load forecasting model f, which can be represented as a 1×24 dimensional feature vector.
[0130] It should be noted that the load forecasting model f can be a deep neural network based on the attention mechanism, or it can also be based on a recurrent neural network, a temporal convolutional neural network, or a multi-layer feedforward neural network, etc., which are network structures suitable for load forecasting. The load forecasting model f can also be a model constructed based on a non-deep neural network. For example, extreme gradient boosting (XGBoost), autoregressive integrated moving average model (ARIMA), or Seasonal ARIMA, etc.
[0131] For ease of understanding, please refer to Figure 4 , Figure 4 which is a schematic diagram of an embodiment for generating the power grid load forecasting result in an embodiment of this application. As shown in the figure, the server can extract a large amount of historical data from the database. For example, historical data for the past year, and generate periodic loads and periodic descriptive features based on these historical data. The user can view today's descriptive information through a terminal device. For example, data such as today's date, weather conditions, temperature, season, and power grid load. If the user wishes to forecast the power grid load for the next day, then they can fill in the power grid load for forecasting the next "1" day and send a forecasting instruction to the server through the terminal device. The server uses the trained load forecasting model to forecast the power grid load for the next 1 day, and finally pushes the forecast result to the terminal device.
[0132] In the embodiments of the present application, a method for predicting power grid load is provided. Through the above method, in the long term, the power grid load of a future time unit can be predicted based on the target feature matrix and a large number of historical feature matrices. The historical feature matrices within the historical period explicitly provide global information, which helps to query and locate the periods with similar load trends in a relatively long historical period during the target period. In the short term, the load sequence has the characteristic of multi-periodicity. Therefore, jointly predicting the power grid load based on the time unit and the period can better capture the complex periodicity, thereby improving the accuracy of power grid load prediction.
[0133] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, obtaining the historical feature matrix specifically includes the following steps:
[0134] Obtain the historical load corresponding to each time unit within the historical period;
[0135] Obtain the historical description information corresponding to each time unit within the historical period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0136] Encode the historical load corresponding to each time unit within the historical period and the historical description information corresponding to each time unit within the historical period to obtain the historical feature matrix.
[0137] In this embodiment, a method for extracting the historical feature matrix is introduced. The load prediction model proposed in the present application can be a multi-step prediction network based on the attention mechanism, and it is introduced by taking D equal to 7 and T equal to 24 as an example. Based on this, the load of 24 hours in a day can be modeled and predicted as a whole for the load sequence.
[0138] Specifically, define a historical feature matrix as X t =[L t-D:t ,Z t-D:t . For the convenience of explanation, please refer to Table 1, which is a schematic diagram of the historical load and historical description information of the historical period.
[0139] Table 1
[0140] Input of the context encoder Dimension of the historical feature matrix Description <![CDATA[L t-D:t > D×24 Historical load for D days going back from the t-th day <![CDATA[T t-D:t > D×1 Temperatures for D days <![CDATA[S t-D:t > D×4 One-hot encoded season information <![CDATA[W t-D:t > D×7 One-hot encoded week information <![CDATA[E t-D:t > D×2 One-hot encoded weekday information <![CDATA[H t-D:t > D×2 One-hot encoded holiday information
[0141] As can be seen from Table 1, obtaining the historical load corresponding to each time unit within the historical period means obtaining the load corresponding to each hour of each day. Among them, taking a day as the time unit, that is, L t-D:t ∈R D×24 , represents the historical load of D days pushed forward from the t-th day. The t-th day is a certain day within the historical period, and Zt-D:t ∈R D×16 represents the result obtained after concatenating [T t-D:t , S t-D:t , W t-D:t , E t-D:t , H t-D:t . M is 16, that is, the dimension of the described feature is 16.
[0142] It should be noted that the context encoder belongs to the global context module, and the global context module is a part of the load prediction model, which will be specifically described in the following embodiments.
[0143] Secondly, in the embodiments of the present application, a method for extracting a historical feature matrix is provided. Through the above method, feature encoding is performed based on historical description information in multiple dimensions, which can improve the feature diversity of the historical feature matrix. In addition, based on temperature information, season information, week information, weekday information, and holiday information, it is possible to better construct multi-periodic features, thereby improving the accuracy of power grid load prediction.
[0144] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, obtaining a historical feature vector specifically includes the following steps:
[0145] Obtain the historical load corresponding to the historical time unit;
[0146] Obtain the historical description information corresponding to the historical time unit, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0147] Encode the historical load corresponding to the historical time unit and the historical description information corresponding to the historical time unit to obtain a historical feature vector.
[0148] In this embodiment, a method for extracting a historical feature vector is introduced. The load for 24 hours in a day can be regarded as a whole for load sequence modeling and prediction.
[0149] Specifically, define a historical feature vector as Y t = [L t:t+1 , Z t:t+1 . For ease of explanation, please refer to Table 2, which is a schematic representation of the historical load and historical description information of the historical time unit.
[0150] Table 2
[0151] Input of the context encoder Dimension of the historical feature vector Description <![CDATA[L t:t+1 > 1×24 Historical load on the t-th day <![CDATA[T t:t+1 > 1×1 Temperature on the t-th day <![CDATA[S t:t+1 > 1×4 One-hot encoded season information <![CDATA[W t:t+1 > 1×7 One-hot encoded week information <![CDATA[E t:t+1 > 1×2 One-hot encoded weekday information <![CDATA[H t:t+1 > 1×2 One-hot encoded holiday information
[0152] As can be seen from Table 2, the historical load corresponding to the historical time unit is obtained, that is, the load corresponding to each hour within the t-th day is obtained. In addition, the feature matrix Z with the day as the time unit t:t+1 ∈R 1×16 , is obtained by connecting [T t:t+1 , S t:t+1 , W t:t+1 , E t:t+1 , H t:t:t+1 in the feature dimension. M is 16, that is, the dimension of the described feature is 16.
[0153] Exemplarily, an example will be used to illustrate below:
[0154]
[0155] That is, the 24 elements in L t:t+1 are the historical loads for each hour within the t-th day. For example, 280 represents 2.8 million kilowatts, and 275 represents 2.75 million kilowatts. Optionally, the 24 elements in L t:t+1 can also be normalized, which is not limited here.
[0156] Exemplarily, an example will be used to illustrate below:
[0157] T t:t+1 = [25.5];
[0158] That is, 25.5 in T t:t+1 represents that the temperature on the t-th day is 25.5 degrees Celsius.
[0159] Exemplarily, an example will be used to illustrate below:
[0160] S t:t+1 = [0, 0, 1, 0];
[0161] Assume that when the position corresponding to the first element is "1", it represents "spring", when the position corresponding to the second element is "1", it represents "summer", when the position corresponding to the third element is "1", it represents "autumn", and when the position corresponding to the fourth element is "1", it represents "winter". That is, S t:t+1 = [0, 0, 1, 0] represents that the t-th day belongs to autumn.
[0162] Exemplarily, an example will be used to illustrate below:
[0163] W t:t+1 = [0, 0, 0, 0, 1, 0, 0];
[0164] Assume that when the position corresponding to the first element is "1", it represents "Monday"; when the position corresponding to the second element is "1", it represents "Tuesday"; when the position corresponding to the third element is "1", it represents "Wednesday"; when the position corresponding to the fourth element is "1", it represents "Thursday"; when the position corresponding to the fifth element is "1", it represents "Friday"; when the position corresponding to the sixth element is "1", it represents "Saturday"; when the position corresponding to the seventh element is "1", it represents "Sunday", that is, W t:t+1 =[0, 0, 0, 0, 1, 0, 0] indicates that the t-th day is Friday.
[0165] Exemplarily, an example will be combined for illustration below:
[0166] E t:t+1 =[1, 0];
[0167] Assume that when the position corresponding to the first element is "1", it represents "working day"; when the position corresponding to the second element is "1", it represents "rest day", that is, E t:t+1 =[1, 0] indicates that the t-th day is a working day.
[0168] Exemplarily, an example will be combined for illustration below:
[0169] H t:t+1 =[0, 1];
[0170] Assume that when the position corresponding to the first element is "1", it represents "holiday"; when the position corresponding to the second element is "1", it represents "non-holiday", that is, H t:t+1 =[0, 1] indicates that the t-th day is a non-holiday.
[0171] Secondly, in the embodiments of the present application, a method for extracting historical feature vectors is provided. Through the above method, feature encoding is performed based on historical description information in multiple dimensions, which can improve the feature diversity of historical feature vectors. In addition, based on temperature information, season information, week information, working day information, and holiday information, multi-periodic features can be better constructed, thereby improving the accuracy of power grid load forecasting.
[0172] Optionally, on the basis of the above Figure 3 corresponding various embodiments, in another optional embodiment provided by the embodiments of the present application, obtaining a target feature matrix specifically includes the following steps:
[0173] Obtain the historical load corresponding to each time unit within the target period;
[0174] Obtain the historical description information corresponding to each time unit within the target period, where the historical description information includes at least one of temperature information, season information, day-of-week information, weekday information, and holiday information;
[0175] Encode the historical load corresponding to each time unit within the target period and the historical description information corresponding to each time unit within the target period to obtain a target feature matrix.
[0176] In this embodiment, a method for extracting a target feature matrix is introduced. Taking D equal to 7 and T equal to 24 as an example, based on this, the load for 24 hours in a day can be used as a whole for modeling and predicting the load sequence.
[0177] Specifically, define a target feature matrix as X d =[L d-D:d ,Z d-D:d , for the convenience of explanation, please refer to Table 3. Table 3 is a schematic of the historical load and historical description information of the target period.
[0178] Table 3
[0179]
[0180]
[0181] As can be seen from Table 3, obtaining the historical load corresponding to each time unit within the target period means obtaining the load corresponding to each hour of each day. Among them, taking a day as the time unit, that is, L d-D:d ∈R D×24 , represents the historical load for a total of D days backward from the d-th day. The d-th day is the target time unit, and Z d-D:d ∈R D×16 represents the result obtained by concatenating [T d-D:d ,S d-D:d ,W d-D:d ,E d-D:d ,H d-D:d in the feature dimension. M is 16, that is, the dimension of the description feature is 16.
[0182] It should be noted that the encoder belongs to the local sequence module, and the local sequence module is another part of the load prediction model, which will be specifically described in subsequent embodiments.
[0183] Secondly, in the embodiments of the present application, a method for extracting a target feature matrix is provided. Through the above method, feature encoding is performed based on historical description information in multiple dimensions, which can improve the feature diversity of the target feature matrix. In addition, based on temperature information, season information, day-of-week information, weekday information, and holiday information, it is possible to better construct multi-periodic features, thereby improving the accuracy of power grid load forecasting.
[0184] Optionally, based on the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, obtaining a target feature vector specifically includes the following steps:
[0185] Obtain the historical load corresponding to the target time unit;
[0186] Obtain the description information corresponding to the first prediction time unit, where the description information includes at least one of temperature information, season information, day-of-week information, weekday information, and holiday information;
[0187] Encode the historical load corresponding to the target time unit and the description information corresponding to the first prediction time unit to obtain a target feature vector.
[0188] In this embodiment, a method for extracting a target feature vector is introduced. The load for 24 hours in a day can be used as a whole for modeling and predicting the load sequence.
[0189] Specifically, define a target feature vector as Y d =[L d-1:d ,Z d:d+1 , for ease of explanation, please refer to Table 4, which is a schematic diagram of the historical load of the target time unit and the description information of the first prediction time unit.
[0190] Table 4
[0191]
[0192]
[0193] As can be seen from Table 4, obtaining the historical load corresponding to the target time unit means obtaining the load corresponding to each hour within the (d - 1)-th day. In addition, the feature matrix Z d:d+1 ∈R 1×16 in the time unit of day is obtained by concatenating [T d:d+1 ,S d:d+1 ,W d:d+1 ,E d:d+1 ,H d:d+1 on the feature dimension, and M is 16, that is, the dimension of the description feature is 16.
[0194] It should be noted that the decoder belongs to the local sequence module, which is another part of the load prediction model and will be specifically described in the following embodiments.
[0195] Secondly, in the embodiments of the present application, a method for extracting target feature vectors is provided. Through the above method, feature encoding is performed based on description information in multiple dimensions, which can enhance the feature diversity of the target feature vectors. In addition, based on temperature information, season information, week information, weekday information, and holiday information, it is possible to better construct multi-periodic features, thereby improving the accuracy of power grid load prediction.
[0196] Optionally, based on the corresponding embodiments above, in another optional embodiment provided by the embodiments of the present application, the power grid load of the first prediction time unit is predicted according to the historical feature matrix, historical feature vector, target feature matrix, and target feature vector to obtain the first power grid load prediction result, which specifically includes the following steps: Figure 3 Based on the target feature matrix and the target feature vector, obtain the first feature matrix through the local sequence module;
[0197] Based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector, obtain the second feature matrix through the global context module;
[0198] Based on the first feature matrix and the second feature matrix, obtain the power grid load prediction result corresponding to the first prediction time unit through the linear layer, where the power grid load prediction result includes the load corresponding to each time sub-unit in the first prediction time unit.
[0199] In this embodiment, a method for power grid load prediction based on a hierarchical network structure is introduced. In the long term, for the power grid load prediction of the first prediction time unit (for example, the d-th day), it can be predicted based on the input target feature matrix (X
[0200] In this embodiment, a method for power grid load prediction based on a hierarchical network structure is introduced. In the long term, for the power grid load prediction of the first prediction time unit (for example, the d-th day), it can be predicted based on the input target feature matrix (X d ) and a large number of context pairs (that is, the historical feature matrix and the historical feature vector), where the context pair can be expressed as (X C , Y C ) := (X i , Y i ), X i∈C represents the i-th historical feature matrix, Y i represents the i-th historical feature vector, := means defined as. C represents the total number of context pairs. It can be understood that a large number of context pairs can be context pairs for one year or multiple years, and these context pairs explicitly provide global information, which helps the target feature matrix (X i ), X d)Query and locate days with similar load trends over a relatively long period.
[0201] For ease of understanding, please refer to Figure 5 , Figure 5 which is a schematic diagram of a hierarchical network structure of the load prediction model in the embodiment of the present application. As shown in the figure, the load prediction model includes a local sequence module and a global context module. Among them, the local sequence module includes an encoder and a decoder, and the global context module includes a context encoder, a context attention layer, and a context decoder. Input the target feature matrix into the encoder included in the local sequence module, and then input the target feature vector and the encoded target feature matrix into the decoder included in the local sequence module. Through this local sequence module, a first feature matrix is output. The first feature matrix can be expressed as 1×Hd f .
[0202] Similarly, input the historical feature matrix and the historical feature vector into the global context module, and then perform attention calculation on the encoded target feature matrix through the context attention layer included in the global context module. Finally, input the attention calculation result and the target feature vector into the context decoder included in the global context module, and thus output a second feature matrix. The second feature matrix can be expressed as 1×Hd f .
[0203] Finally, input the first feature matrix output by the local sequence module and the second feature matrix input by the global context module into a linear layer. Among them, the linear layer can be a fully connected layer. After passing through the linear layer, the grid load prediction result corresponding to the first prediction time unit is output Grid load prediction result can be expressed as a 1×24 feature vector, and each element in the feature vector corresponds to the load corresponding to one hour.
[0204] Secondly, in the embodiments of the present application, a power grid load forecasting method based on a hierarchical network structure is provided. Through the above method, considering that the load sequence often contains complex periodic characteristics and it is difficult to capture such complex periodicity in terms of time points. If the similarity between each moment in the load sequence is simply calculated in the sequence model, although the load forecasting problem can be modeled from the perspective of time series, the multi-periodic characteristics of the load sequence still cannot be fully captured. Therefore, the present application adopts a hierarchical periodicity extraction mechanism, and the local sequence module and the global context module respectively extract the periodic characteristics of the load sequence in the long term and the short term, thereby improving the accuracy of power grid load forecasting. The local sequence module and the global context module can hierarchically extract multiple periodicities in the load sequence.
[0205] Optionally, based on the above Figure 3 corresponding respective embodiments, in another optional embodiment provided by the embodiments of the present application, based on the target feature matrix and the target feature vector, obtaining the first feature matrix through the local sequence module specifically includes the following steps:
[0206] Performing attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain the encoding result of the target feature matrix;
[0207] Performing attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain the first feature matrix.
[0208] In this embodiment, a method for outputting the first feature matrix based on the local sequence module is introduced. Combining the content described in the foregoing embodiments, the local sequence module includes an encoder and a decoder, and the following will be combined with Figure 6 for introduction.
[0209] Specifically, please refer to Figure 6 , Figure 6 which is a schematic diagram of a network structure of the local sequence module in the embodiments of the present application. As shown in the figure, input the target feature matrix X d into the input embedding layer, where the input embedding layer can be implemented by a fully connected layer, and it maps the input to a feature dimension of Hd through the parameter matrix f , thereby realizing a linear mapping, where H represents the number of attention heads in the multi-head attention mechanism, and d f represents the feature dimension of each attention head. Since the target feature matrix (X d)The description features and loads of each day need to be processed in parallel. Therefore, position encoding needs to be added to encode the time sequence of daily data appearance.
[0210] Input the linearly mapped target feature matrix into the encoder. After attention auto-encoding processing, the encoded result of the target feature matrix is obtained. Among them, during the attention auto-encoding processing, the query (Q), key (K), and value (V) all come from the target feature matrix. Similarly, input the target feature vector Y d into the input embedding layer to achieve linear mapping. Then, jointly input the linearly mapped target feature vector and the encoded result of the target feature matrix into the decoder to obtain the first feature matrix.
[0211] Again, in the embodiments of the present application, a method for outputting the first feature matrix based on the local sequence module is provided. Through the above method, an encoder and a decoder based on the attention mechanism are set in the local sequence module. The encoder can better focus on the attention relationship inside the target feature matrix, and the decoder can predict the attention relationship between the target feature vector and the target feature matrix, thereby improving the accuracy of power grid load prediction.
[0212] Optionally, based on the corresponding various embodiments above, in another optional embodiment provided by the embodiments of the present application, the encoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head; Figure 3 Based on the encoder included in the local sequence module, perform attention encoding processing on the target feature matrix to obtain the encoded result of the target feature matrix, which specifically includes the following steps:
[0213] For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the first attention head to obtain the first encoded result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0214] For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the second attention head to obtain the second encoded result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0215] Based on the first encoded result and the second encoded result, obtain the encoded result of the target feature matrix.
[0216] According to the first encoded result and the second encoded result, obtain the encoded result of the target feature matrix.
[0217] In this embodiment, a coding method based on the multi-head attention mechanism is introduced. The encoder included in the local sequence module is stacked by multiple layers of multi-head attention mechanism layers. For the convenience of introduction, a single-layer multi-head attention layer is taken as an example for introduction. In practical applications, it may also include two or more multi-head attention layers, which is not limited here. And for the convenience of introduction, each multi-head attention layer includes at least two attention heads, namely the first attention head and the second attention head. The first attention head adopts the standard multi-head attention mechanism, while the second attention head adopts the non-periodic perturbation enhanced multi-head attention mechanism. In this application, four attention heads are taken as an example, among which there are two first attention heads and two second attention heads.
[0218] For the convenience of understanding, please refer to Figure 7 , Figure 7 which is a schematic diagram of a network structure of the encoder in the local sequence module in the embodiment of the present application. As shown in the figure, for the encoder, in each layer of the multi-head attention layer, the first attention head is used to perform attention coding processing on the target feature matrix, and thus the first coding result is obtained. If two first attention heads are used to perform attention coding processing on the target feature matrix, two first coding results are obtained. Similarly, the second attention head is used to perform attention coding processing on the target feature matrix, and thus the second coding result is obtained. If two second attention heads are used to perform attention coding processing on the target feature matrix, two second coding results are obtained.
[0219] The attention coding processing method of the first attention head is as follows:
[0220]
[0221] The attention coding processing method of the second attention head is as follows:
[0222]
[0223] Among them, O represents the coding result, d f represents the feature dimension of each attention head, Q represents the query matrix, K represents the keyword matrix, V represents the value matrix, h ∈ {1, 2,..., H} represents the index of each attention head, and H represents the number of attention heads in the multi-head attention mechanism.
[0224] Based on this, in each layer of the multi-head attention layer, the intermediate coding result is obtained by combining the coding results corresponding to the four attention heads, that is, the four attention heads are concatenated together as the intermediate coding result. Finally, the intermediate coding result is continued as the input of the next multi-head attention layer until the coding result of the target feature matrix is obtained after being processed by all the multi-head attention layers.
[0225] It should be noted that in the standard attention mechanism, the first attention head is used to calculate the positive correlation between two feature vectors. For example, if it is necessary to predict the 24 loads within a day on Tuesday, and the input is the target feature matrix of the past week, then there will be a strong correlation between the Tuesday to be predicted and the Tuesday in the previous week. However, if the Tuesday to be predicted happens to be the Spring Festival holiday, then the electricity consumption will be very different compared to the Tuesday in the previous week. Therefore, by adding a second attention head, the network can learn the negative correlation between two feature vectors, thereby paying attention to the impact of aperiodic perturbations.
[0226] Since the load sequence is also affected by many factors, for example, sudden changes may occur during holidays or extreme weather. And such changes often show significant differences from the previously observed load trends. The standard multi-head attention mechanism can pay attention to information from different representation subspaces through multiple heads, but it cannot guarantee that adding more heads can capture more useful features. The standard multi-head attention mechanism calculates a weight distribution based on each head. The query matrix and key matrix of each head calculate the similarity between days in the load sequence through dot product, and obtain the weight distribution through the softmax function. Combining with the aperiodic perturbation enhanced multi-head attention mechanism provided in this application, the negative correlation between load sequences can be paid attention to.
[0227] Furthermore, in the embodiments of the present application, a coding method based on the multi-head attention mechanism is provided. Through the above method, the encoder in the local sequence module is stacked by multiple layers of multi-head attention mechanism layers. In addition, for the encoder, a part uses the standard multi-head attention mechanism, which can use multiple heads to fully extract the patterns of the sequence, and another part uses the aperiodic perturbation enhanced multi-head attention mechanism, which can accurately predict the turning points generated by the load sequence when affected by aperiodic perturbations. The combined use of the two has a better effect, thereby improving the accuracy of power grid load prediction.
[0228] Optionally, based on the above Figure 3 In another optional embodiment provided by the embodiments of the present application corresponding to the respective embodiments, the decoder includes at least one layer of multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head;
[0229] According to the decoder included in the local sequence module, perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector to obtain a first feature matrix, which specifically includes the following steps:
[0230] For each multi - head attention layer in the decoder, based on the encoding result of the target feature matrix by the first attention head and the target feature vector, perform attention decoding processing to obtain a first decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0231] For each multi - head attention layer in the decoder, based on the encoding result of the target feature matrix by the second attention head and the target feature vector, perform attention decoding processing to obtain a second decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0232] According to the first decoding result and the second decoding result, obtain a first feature matrix.
[0233] In this embodiment, a decoding method based on the multi - head attention mechanism is introduced. The decoder included in the local sequence module is stacked by multiple layers of multi - head attention mechanism layers. For the sake of convenience in introduction, one layer of multi - head attention layer is taken as an example for introduction. In practical applications, it may also include two or more multi - head attention layers, which is not limited here. And for the sake of convenience in introduction, each multi - head attention layer includes at least two attention heads, namely the first attention head and the second attention head. The first attention head adopts the standard multi - head attention mechanism, while the second attention head adopts the non - periodic perturbation enhanced multi - head attention mechanism. In this application, taking four attention heads as an example, there are two first attention heads and two second attention heads.
[0234] For the sake of easy understanding, please refer to Figure 8 , Figure 8 is a schematic diagram of a network structure of the decoder in the local sequence module in the embodiment of the present application. As shown in the figure, for the decoder, in each layer of multi - head attention layer, use the first attention head to perform attention decoding processing on the decoding result of the target feature matrix and the target feature vector, thereby obtaining a first decoding result. Similarly, use the second attention head to perform attention decoding processing on the decoding result of the target feature matrix and the target feature vector, thereby obtaining a second decoding result.
[0235] It should be noted that the attention encoding processing methods of the first attention head and the second attention head are as described in the foregoing embodiments, so they will not be elaborated here.
[0236] Based on this, in each layer of multi - head attention layer, combine the decoding results corresponding to the four attention heads to obtain an intermediate decoding result, that is, splice the four attention heads together as the intermediate result, and finally use the intermediate decoding result as the input of the next multi - head attention layer until the first feature matrix is obtained after the processing of all multi - head attention layers.
[0237] Furthermore, in the embodiments of the present application, a decoding method based on the multi-head attention mechanism is provided. Through the above method, the decoder in the local sequence module is stacked by multiple layers of multi-head attention mechanism layers. In addition, for the decoder, a part adopts the standard multi-head attention mechanism, which can fully extract the patterns of the sequence by using multiple heads, and another part adopts the aperiodic perturbation enhanced multi-head attention mechanism, which can accurately predict the turning points generated when the load sequence is affected by aperiodic perturbations. The combined use of the two has a better effect, thus improving the accuracy of power grid load forecasting.
[0238] Optionally, on the basis of the above Figure 3 corresponding embodiments, in another optional embodiment provided by the embodiments of the present application, based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector, a second feature matrix is obtained through the global context module, which specifically includes the following steps:
[0239] Perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain the encoded result of the historical feature matrix;
[0240] Perform attention encoding processing on the historical feature vector according to the context encoder included in the global context module to obtain the encoded result of the historical feature vector;
[0241] Based on the context attention mechanism, calculate the encoded results of the historical feature matrix, historical feature vector, and target feature matrix to obtain the attention encoded result;
[0242] Perform attention decoding processing on the attention encoded result and the target feature vector according to the context decoder included in the global context module to obtain the second feature matrix.
[0243] In this embodiment, a method for outputting the second feature matrix based on the global context module is introduced. Combining the content described in the foregoing embodiments, the global context module includes a context encoder and a context decoder, which will be combined with Figure 9 for introduction.
[0244] Specifically, please refer to Figure 9 , Figure 9 which is a schematic diagram of a network structure of the global context module in the embodiments of the present application. As shown in the figure, the historical feature matrix (X 1 to X N ) is input to the input embedding layer, and the historical feature vector (Y 1 to Y N ) is also input to the input embedding layer. Among them, the input embedding layer can be implemented by a fully connected layer, and it passes through the parameter matrix Map the input to a feature dimension of Hd f , thus achieving a linear mapping, where H represents the number of attention heads in the multi-head attention mechanism, and d f represents the feature dimension of each attention head. Since the historical feature matrices (X 1 to X N ) need to process the description features and loads of each day in parallel, therefore, position encoding needs to be added to achieve the effect of encoding the time sequence of daily data appearance.
[0245] Input the linearly mapped historical feature matrix and historical feature vector into the context encoder, and after attention auto-encoding processing, obtain the encoded results (k 1 to k N ) of the historical feature matrix, as well as the encoded results of the historical feature vector. Among them, during the attention auto-encoding process of the historical feature matrix, Q, K, and V all come from the historical feature matrix. During the attention auto-encoding process of the historical feature vector, Q, K, and V all come from the historical feature vector.
[0246] After obtaining the encoded results (k 1 to k N ) of the historical feature matrix, input it and the encoded results (q d ) of the target feature matrix into the context attention layer together, perform similarity calculation through the context attention layer, weight the context pairs in the set, and obtain the attention encoded result (v d ) according to the encoded results of the historical feature vector. When the target feature matrix (X d ) is relatively high compared to a certain context point (X i ) in the historical feature matrix, its corresponding target feature vector (Y d ) should also have a relatively high similarity with the context point (Y i ). In the short term, the load sequence has the characteristics of daily and weekly cycles. Therefore, use one week as the input lag window. In the local sequence module, use the attention mechanism to model the sequence relationship between daily load and weekly load, rather than exploring the sequence relationship with each time point as the unit.
[0247] Finally, use the context decoder to perform attention decoding processing on the attention encoded result and the target feature vector to obtain the second feature matrix.
[0248] Again, in the embodiments of the present application, a method for outputting a second feature matrix based on a global context module is provided. Through the above method, a context encoder and a context decoder based on an attention mechanism are provided in the global context module. The context encoder can better focus on the attention relationships within the historical feature matrix and the historical feature vector, and the context decoder can predict the attention relationship between the attention encoding result and the target feature vector, thereby improving the accuracy of power grid load forecasting.
[0249] Optionally, on the basis of the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, the context encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least one first attention head and one second attention head; Figure 3 Based on the context encoder included in the global context module, performing attention encoding processing on the historical feature matrix to obtain an encoding result of the historical feature matrix, which specifically includes the following steps:
[0250] For each multi-head attention layer in the context encoder, performing attention encoding processing on the historical feature matrix based on the first attention head to obtain a third encoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0251] For each multi-head attention layer in the context encoder, performing attention encoding processing on the historical feature matrix based on the second attention head to obtain a fourth encoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0252] Based on the third encoding result and the fourth encoding result, obtaining the encoding result of the historical feature matrix;
[0253] Based on the context encoder included in the global context module, performing attention encoding processing on the historical feature vector to obtain an encoding result of the historical feature vector, which specifically includes the following steps:
[0254] For each multi-head attention layer in the context encoder, performing attention encoding processing on the historical feature vector based on the first attention head to obtain a fifth encoding result;
[0255] For each multi-head attention layer in the context encoder, performing attention encoding processing on the historical feature vector based on the second attention head to obtain a sixth encoding result;
[0256] Based on the fifth encoding result and the sixth encoding result, obtaining the encoding result of the historical feature vector.
[0257] Based on the fifth encoding result and the sixth encoding result, obtaining the encoding result of the historical feature vector.
[0258] In this embodiment, a coding method based on the multi-head attention mechanism is provided. The context encoder included in the global context module is stacked by multiple layers of multi-head attention mechanism layers. For the convenience of introduction, one layer of multi-head attention layer is used for introduction. In practical applications, it may also include two or more multi-head attention layers, which is not limited here. And for the convenience of introduction, each multi-head attention layer includes at least two attention heads, namely the first attention head and the second attention head. The first attention head adopts the standard multi-head attention mechanism, while the second attention head adopts the aperiodic perturbation enhanced multi-head attention mechanism. In this application, an example of using four attention heads is taken, among which, there are two first attention heads and two second attention heads.
[0259] For the convenience of understanding, please refer to Figure 10 , Figure 10 FIG. is a schematic diagram of a network structure of the context encoder in the global context module in the embodiment of the present application. As shown in the figure, for the context encoder, in each layer of multi-head attention layer, the first attention head is used to perform attention coding processing on the historical feature matrix, thereby obtaining a third coding result. In each layer of multi-head attention layer, the second attention head is used to perform attention coding processing on the historical feature matrix, thereby obtaining a fourth coding result.
[0260] Based on this, in each layer of multi-head attention layer, the intermediate coding result is obtained by combining the coding results corresponding to the four attention heads, that is, the four attention heads are concatenated together as the intermediate coding result. Finally, the intermediate coding result is continued as the input of the next multi-head attention layer until the coding result of the historical feature matrix is obtained after the processing of all multi-head attention layers.
[0261] Similarly, in each layer of multi-head attention layer, the first attention head is used to perform attention coding processing on the historical feature vector, thereby obtaining a fifth coding result. In each layer of multi-head attention layer, the second attention head is used to perform attention coding processing on the historical feature vector, thereby obtaining a sixth coding result.
[0262] Based on this, in each layer of multi-head attention layer, the intermediate coding result is obtained by combining the coding results corresponding to the four attention heads, that is, the four attention heads are concatenated together and converted into an intermediate coding result similar to that of a single head through a linear layer. Finally, the intermediate coding result is continued as the input of the next multi-head attention layer until the coding result of the historical feature vector is obtained after the processing of all multi-head attention layers.
[0263] It should be noted that the attention coding processing methods of the first attention head and the second attention head are as described in the foregoing embodiments, so they will not be elaborated here.
[0264] Again, in the embodiments of the present application, a coding method based on the multi-head attention mechanism is provided. Through the above method, the context encoder in the global context module is stacked by multiple layers of multi-head attention mechanism layers. In addition, for the context encoder, a part adopts the standard multi-head attention mechanism, which can fully extract the patterns of the sequence by using multiple heads, and another part adopts the non-periodic perturbation enhanced multi-head attention mechanism, which can accurately predict the turning points generated by the load sequence when affected by non-periodic perturbations. The combined use of the two has a better effect, thereby improving the accuracy of power grid load forecasting.
[0265] Optionally, on the basis of the corresponding respective embodiments above, in another optional embodiment provided by the embodiments of the present application, the context decoder includes at least one layer of multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head; Figure 3 According to the context decoder included in the global context module, perform attention decoding processing on the attention coding result and the target feature vector to obtain a second feature matrix, which specifically includes the following steps:
[0266] For each multi-head attention layer in the context decoder, perform attention decoding processing on the attention coding result and the target feature vector based on the first attention head to obtain a third decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0267] For each multi-head attention layer in the decoder, perform attention decoding processing on the attention coding result and the target feature vector based on the second attention head to obtain a fourth decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0268] According to the third decoding result and the fourth decoding result, obtain the second feature matrix.
[0269] In this embodiment, a decoding method based on the multi-head attention mechanism is introduced. The context decoder included in the global context module is stacked by multiple layers of multi-head attention mechanism layers. For the convenience of introduction, one layer of multi-head attention layer is used for introduction. In practical applications, it may also include three or more multi-head attention layers, which are not limited here. And for the convenience of introduction, each multi-head attention layer includes at least two attention heads, namely the first attention head and the second attention head. The first attention head adopts the standard multi-head attention mechanism, while the second attention head adopts the non-periodic perturbation enhanced multi-head attention mechanism. In this application, four attention heads are used as an example, where there are two first attention heads and two second attention heads.
[0270] For the convenience of understanding, please refer to
[0271] For the convenience of understanding, please refer toFigure 11 , Figure 11 This is a schematic diagram of a network structure of a context decoder in the global context module in an embodiment of the present application. As shown in the figure, for the context decoder, in each multi-head attention layer, the first attention head is used to perform attention decoding processing on the attention encoding result and the target feature vector, thereby obtaining a third decoding result. Similarly, the second attention head is used to perform attention decoding processing on the attention encoding result and the target feature vector, thereby obtaining a fourth decoding result.
[0272] It should be noted that the attention encoding processing methods of the first attention head and the second attention head are as described in the foregoing embodiments, so details are not described herein.
[0273] Based on this, in each multi-head attention layer, the intermediate decoding results are obtained by combining the decoding results corresponding to the four attention heads, that is, the four attention heads are concatenated together and converted into an intermediate decoding result similar to that of a single head through a linear layer. Finally, the intermediate decoding result is continued as the input of the next multi-head attention layer until the second feature matrix is obtained after the processing of all multi-head attention layers.
[0274] Again, in an embodiment of the present application, a decoding method based on a multi-head attention mechanism is provided. Through the above method, the context decoder in the global context module is stacked by multiple layers of multi-head attention mechanism layers. In addition, for the context decoder, a part uses a standard multi-head attention mechanism, which can use multiple heads to fully extract the patterns of the sequence, and another part uses a non-periodic perturbation enhanced multi-head attention mechanism, which can accurately predict the turning points generated when the load sequence is affected by non-periodic perturbations. The combined use of the two has a better effect, thereby improving the accuracy of power grid load prediction.
[0275] Optionally, on the basis of the corresponding embodiments above, Figure 3 In another optional embodiment provided by the embodiment of the present application, after predicting the power grid load of the first prediction time unit according to the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector to obtain the first power grid load prediction result, the following steps may further be included:
[0276] Obtain an updated feature matrix, where the updated feature matrix includes the periodic load and the periodic description features of the updated period, the updated period includes D time units, and the updated period includes the first prediction time unit;
[0277] Obtain the updated feature vector, where the updated feature vector includes the power grid load prediction result corresponding to the first prediction time unit and the description features of the second prediction time unit. The first prediction time unit is an adjacent time unit that appears before the second prediction time unit, and the second prediction time unit is the next time unit after the updated period appears. The second prediction time unit includes T time sub-units;
[0278] Predict the power grid load of the second prediction time unit based on the historical feature matrix, historical feature vector, updated feature matrix, and updated feature vector to obtain the second power grid load prediction result.
[0279] In this embodiment, a method for predicting the load at the second prediction time is introduced. After predicting the first power grid load prediction result, the updated feature matrix corresponding to the updated period can be obtained. The difference between the updated period and the target period is that the updated period includes the first prediction time unit, while the target period does not include the first prediction time unit. The same is that the updated period also has D time units, and each time unit has T time sub-units. Specifically, the updated feature matrix can be the power grid-related data of the cycle closest to the current time. Taking one time sub-unit as 1 hour, one time unit as 1 day, and one updated feature matrix as 7 days as an example, then D is 7 and T is 24. Based on this, the cycle load corresponding to one updated feature matrix includes 7 * 24 hours of load, and the cycle description features corresponding to one updated feature matrix include 7 * M description features, or 7 * 24 * M description features.
[0280] After predicting the first power grid load prediction result, the updated feature vector can also be obtained. The updated feature vector includes the power grid load prediction result corresponding to the first prediction time unit and the description features of the second prediction time unit. Based on this, for example, if the updated period is from October 13, 2020 to October 19, 2020, then the first prediction time unit corresponding to this updated period is October 19, 2020, and the second prediction time unit corresponding to this updated period is October 20, 2020, that is, both the first prediction time unit and the second prediction time unit include T time sub-units (for example, both include 24 hours).
[0281] Specifically, the power grid load prediction result corresponding to the first prediction time unit includes 1 * 24 hours of load, and the description features corresponding to the second prediction time unit include 1 * M description features, where M represents the dimension of the description features.
[0282] Finally, input the historical feature matrix, historical feature vector, updated feature matrix, and updated feature vector into the trained load prediction model, and the load prediction model outputs the second power grid load prediction result, that is, the power grid load for the next time unit is predicted.
[0283] Secondly, in the embodiments of the present application, a method for predicting the load at the second prediction time is provided. Through the above method, a hierarchical multi-period extraction mechanism and an aperiodic perturbation enhanced multi-head attention mechanism are constructed, so that in the load prediction problem, the short-term load values for the next one to two days can be accurately predicted, and it is applicable to institutions with large energy management requirements such as power grids and can be integrated into their existing systems to achieve the advantages of small computational cost and low cost during prediction.
[0284] The power grid load prediction device in the present application will be described in detail below. Please refer to Figure 12 , Figure 12 FIG. is a schematic diagram of an embodiment of the power grid load prediction device in the embodiments of the present application. The power grid load prediction device 20 includes:
[0285] An acquisition module 201, configured to acquire a historical feature matrix, where the historical feature matrix includes the periodic load and periodic description features of a historical period. The historical period includes D time units, and each time unit includes T time sub-units. Both D and T are integers greater than 1;
[0286] The acquisition module 201 is further configured to acquire a historical feature vector, where the historical feature vector includes the historical load and historical description features of a historical time unit. The historical time unit is the next time unit after the historical period appears, and the historical time unit includes T time sub-units;
[0287] The acquisition module 201 is further configured to acquire a target feature matrix, where the target feature matrix includes the periodic load and periodic description features of a target period. The target period includes D time units;
[0288] The acquisition module 201 is further configured to acquire a target feature vector, where the target feature vector includes the historical load of a target time unit and the description features of a first prediction time unit. The target time unit is an adjacent time unit before the first prediction time unit appears. The first prediction time unit is the next time unit after the target period appears. The target time unit includes T time sub-units, and the first prediction time unit includes T time sub-units;
[0289] A prediction module 202, configured to predict the power grid load of the first prediction time unit according to the historical feature matrix, historical feature vector, target feature matrix, and target feature vector, and obtain a first power grid load prediction result.
[0290] Optionally, based on the corresponding embodiments above, Figure 12 in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0291] an acquisition module 201 is specifically configured to acquire historical loads corresponding to each time unit within a historical period;
[0292] acquire historical description information corresponding to each time unit within a historical period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0293] encode the historical loads corresponding to each time unit within a historical period and the historical description information corresponding to each time unit within a historical period to obtain a historical feature matrix.
[0294] Optionally, based on the corresponding embodiments above, Figure 12 in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0295] an acquisition module 201 is specifically configured to acquire historical loads corresponding to a historical time unit;
[0296] acquire historical description information corresponding to a historical time unit, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0297] encode the historical loads corresponding to a historical time unit and the historical description information corresponding to a historical time unit to obtain a historical feature vector.
[0298] Optionally, based on the corresponding embodiments above, Figure 12 in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0299] an acquisition module 201 is specifically configured to acquire historical loads corresponding to each time unit within a target period;
[0300] acquire historical description information corresponding to each time unit within a target period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0301] encode the historical loads corresponding to each time unit within a target period and the historical description information corresponding to each time unit within a target period to obtain a target feature matrix.
[0302] Optionally, based on the corresponding embodiments above, Figure 12Based on the corresponding embodiment, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0303] An acquisition module 201, specifically configured to acquire the historical load corresponding to the target time unit;
[0304] Acquire the description information corresponding to the first prediction time unit, where the description information includes at least one of temperature information, season information, week information, weekday information, and holiday information;
[0305] Encode the historical load corresponding to the target time unit and the description information corresponding to the first prediction time unit to obtain a target feature vector.
[0306] Optionally, based on the above Figure 12 Based on the corresponding embodiment, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0307] A prediction module 202, specifically configured to obtain a first feature matrix through a local sequence module based on the target feature matrix and the target feature vector;
[0308] Obtain a second feature matrix through a global context module based on the historical feature matrix, historical feature vector, target feature matrix, and target feature vector;
[0309] Based on the first feature matrix and the second feature matrix, obtain the power grid load prediction result corresponding to the first prediction time unit through a linear layer, where the power grid load prediction result includes the load corresponding to each time sub-unit in the first prediction time unit.
[0310] Optionally, based on the above Figure 12 Based on the corresponding embodiment, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application,
[0311] The prediction module 202 is specifically configured to perform attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain an encoding result of the target feature matrix;
[0312] Perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain the first feature matrix.
[0313] Optionally, based on the above Figure 12 Based on the corresponding embodiment, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application, the encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least two attention heads and a second attention head;
[0314] The prediction module 202 is specifically configured to, for each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the first attention head to obtain a first encoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0315] For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the second attention head to obtain a second encoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0316] According to the first encoding result and the second encoding result, obtain the encoding result of the target feature matrix.
[0317] Optionally, on the basis of the corresponding embodiment above, Figure 12 In another embodiment of the power grid load prediction device 20 provided by the embodiment of the present application, the decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head;
[0318] The prediction module 202 is specifically configured to, for each multi-head attention layer in the decoder, perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector based on the first attention head to obtain a first decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0319] For each multi-head attention layer in the decoder, perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector based on the second attention head to obtain a second decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0320] According to the first decoding result and the second decoding result, obtain the first feature matrix.
[0321] Optionally, on the basis of the corresponding embodiment above, Figure 12 In another embodiment of the power grid load prediction device 20 provided by the embodiment of the present application,
[0322] The prediction module 202 is specifically configured to perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain the encoding result of the historical feature matrix;
[0323] Perform attention encoding processing on the historical feature vector according to the context encoder included in the global context module to obtain the encoding result of the historical feature vector;
[0324] Based on the context attention mechanism, calculate the encoding results of the historical feature matrix, the encoding results of the historical feature vector, and the encoding results of the target feature matrix to obtain the attention encoding results;
[0325] According to the context decoder included in the global context module, perform attention decoding processing on the attention encoding results and the target feature vector to obtain the second feature matrix.
[0326] Optionally, based on the above Figure 12 In another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application, on the basis of the corresponding embodiments
[0327] The context encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least a first attention head and a second attention head;
[0328] The prediction module 202 is specifically configured to perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain the encoding results of the historical feature matrix, including:
[0329] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix based on the first attention head to obtain a third encoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0330] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix based on the second attention head to obtain a fourth encoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0331] According to the third encoding result and the fourth encoding result, obtain the encoding results of the historical feature matrix;
[0332] The prediction module 202 is specifically configured to, for each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector based on the first attention head to obtain a fifth encoding result;
[0333] For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector based on the second attention head to obtain a sixth encoding result;
[0334] According to the fifth encoding result and the sixth encoding result, obtain the encoding results of the historical feature vector.
[0335] Optionally, in the above Figure 12Based on the corresponding embodiment, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application, the context decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head;
[0336] The prediction module 202 is specifically configured to, for each multi-head attention layer in the context decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the first attention head to obtain a third decoding result, where the first attention head is used to calculate the positive correlation between feature vectors;
[0337] For each multi-head attention layer in the decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the second attention head to obtain a fourth decoding result, where the second attention head is used to calculate the negative correlation between feature vectors;
[0338] Obtain a second feature matrix according to the third decoding result and the fourth decoding result.
[0339] Optionally, based on the corresponding embodiment above, in another embodiment of the power grid load prediction device 20 provided by the embodiments of the present application, Figure 12 After the prediction module 202 predicts the power grid load of the first prediction time unit according to the historical feature matrix, historical feature vector, target feature matrix, and target feature vector to obtain a first power grid load prediction result, the acquisition module 201 is further configured to obtain an updated feature matrix, where the updated feature matrix includes the periodic load and periodic description features of the updated period, the updated period includes D time units, and the updated period includes the first prediction time unit;
[0340] The acquisition module 201 is further configured to obtain an updated feature vector, where the updated feature vector includes the power grid load prediction result corresponding to the first prediction time unit and the description features of the second prediction time unit. The first prediction time unit is an adjacent time unit that appears before the second prediction time unit, and the second prediction time unit is the next time unit that appears after the updated period. The second prediction time unit includes T time sub-units;
[0341] The prediction module 202 is further configured to predict the power grid load of the second prediction time unit according to the historical feature matrix, historical feature vector, updated feature matrix, and updated feature vector to obtain a second power grid load prediction result.
[0342] The embodiments of the present application further provide another power grid load prediction device, which can be deployed on a terminal device, such as
[0343] Another power grid load prediction device provided by the embodiments of the present application can be deployed on a terminal device, such as Figure 13As shown, for the sake of convenience in explanation, only the parts related to the embodiments of the present application are shown. For the specific technical details not disclosed, please refer to the method part of the embodiments of the present application. The terminal device can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS) device, an in-vehicle computer, etc. Taking the terminal device as a mobile phone as an example:
[0344] Figure 13 The block diagram of a part of the structure of the mobile phone related to the terminal device provided by the embodiments of the present application is shown. Refer to Figure 13 , the mobile phone includes: a radio frequency (RF) circuit 310, a memory 320, an input unit 330, a display unit 340, a sensor 350, an audio circuit 360, a wireless fidelity (WiFi) module 370, a processor 380, and a power supply 390 and other components. Those skilled in the art can understand that Figure 13 the structure of the mobile phone shown in
[0345] does not limit the mobile phone, and it may include more or fewer components than shown in the figure, or combine some components, or have different component arrangements. Figure 13 The following specifically introduces each component of the mobile phone:
[0346] The RF circuit 310 can be used for receiving and sending signals during information reception or call processes. Specifically, after receiving the downlink information from the base station, it is given to the processor 380 for processing; in addition, the uplink data designed is sent to the base station. Usually, the RF circuit 310 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 310 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the Global System of Mobile communication (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, Short Messaging Service (SMS), etc.
[0347] The memory 320 can be used to store software programs and modules. The processor 380 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 320. The memory 320 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 320 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0348] The input unit 330 can be used to receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the mobile phone. Specifically, the input unit 330 may include a touch panel 331 and other input devices 332. The touch panel 331, also known as a touch screen, can collect touch operations of the user on or near it (such as operations of the user using a finger, a stylus, or any suitable object or accessory on or near the touch panel 331), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 331 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch orientation of the user, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 380, and can receive and execute the commands sent by the processor 380. In addition, various types such as resistive, capacitive, infrared, and surface acoustic wave can be used to implement the touch panel 331. In addition to the touch panel 331, the input unit 330 may further include other input devices 332. Specifically, the other input devices 332 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power on / off keys, etc.), a trackball, a mouse, a joystick, etc.
[0349] The display unit 340 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 340 may include a display panel 341. Optionally, the display panel 341 can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 331 can cover the display panel 341. When the touch panel 331 detects a touch operation on or near it, it is transmitted to the processor 380 to determine the type of touch event. Subsequently, the processor 380 provides a corresponding visual output on the display panel 341 according to the type of touch event. Although in Figure 13 the touch panel 331 and the display panel 341 are implemented as two independent components to realize the input and input functions of the mobile phone, in some embodiments, the touch panel 331 and the display panel 341 can be integrated to realize the input and output functions of the mobile phone.
[0350] The mobile phone may further include at least one sensor 350, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. Among them, the ambient light sensor can adjust the brightness of the display panel 341 according to the brightness of the ambient light. The proximity sensor can turn off the display panel 341 and / or the backlight when the mobile phone is moved to the ear. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity, and can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. As for other sensors such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors that the mobile phone can also be configured with, they will not be elaborated here.
[0351] The audio circuit 360, the speaker 361, and the microphone 362 can provide an audio interface between the user and the mobile phone. The audio circuit 360 can transmit the electrical signal converted from the received audio data to the speaker 361, and the speaker 361 converts it into a sound signal for output. On the other hand, the microphone 362 converts the collected sound signal into an electrical signal, which is received by the audio circuit 360 and converted into audio data. After the audio data is output to the processor 380 for processing, it is sent through the RF circuit 310 to, for example, another mobile phone, or the audio data is output to the memory 320 for further processing.
[0352] WiFi belongs to short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse the web, and access streaming media through the WiFi module 370. It provides users with wireless broadband Internet access. Although Figure 13The WiFi module 370 is shown, but it can be understood that it does not belong to the essential components of the mobile phone and can be completely omitted within the scope of not changing the essence of the invention as needed.
[0353] The processor 380 is the control center of the mobile phone, connecting various parts of the entire mobile phone through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 320, and by calling the data stored in the memory 320, it executes various functions of the mobile phone and processes data. Optionally, the processor 380 may include one or more processing units; optionally, the processor 380 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 380 either.
[0354] The mobile phone also includes a power supply 390 (such as a battery) for supplying power to each component. Optionally, the power supply can be logically connected to the processor 380 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system.
[0355] Although not shown, the mobile phone may also include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0356] In the above embodiments, the steps executed by the terminal device can be based on the Figure 13 shown terminal device structure.
[0357] The embodiment of the present application also provides another power grid load prediction device. This power grid load prediction device can be deployed on a server. Figure 14 It is a schematic diagram of a server structure provided by the embodiment of the present application. The server 400 may vary greatly due to configuration or performance differences and may include one or more central processing units (CPUs) 422 (for example, one or more processors) and a memory 432, and one or more storage media 430 (for example, one or more mass storage devices) for storing application programs 442 or data 444. Among them, the memory 432 and the storage media 430 can be transient storage or persistent storage. The programs stored in the storage media 430 may include one or more modules (not marked in the figure), and each module may include a series of instruction operations on the server. Further, the central processor 422 can be set to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the server 400.
[0358] The server 400 may also include one or more power supplies 426, one or more wired or wireless network interfaces 450, one or more input / output interfaces 458, and / or one or more operating systems 441, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM and so on.
[0359] The steps performed by the server in the above embodiments may be based on the Figure 14 server structure shown.
[0360] An embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When it runs on a computer, it causes the computer to execute the methods described in the foregoing embodiments.
[0361] An embodiment of the present application also provides a computer program product including a program. When it runs on a computer, it causes the computer to execute the methods described in the foregoing embodiments.
[0362] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0363] In the several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0364] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0365] In addition, in each embodiment of the present application, each functional unit may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0366] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0367] As described above, the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for power grid load forecasting, characterized in that, it includes: Obtain a historical feature matrix, where the historical feature matrix includes the periodic load and periodic description features of a historical period. The historical period includes D time units, and each time unit includes T time sub-units. Both D and T are integers greater than 1; Obtain a historical feature vector, where the historical feature vector includes the historical load and historical description features of a historical time unit. The historical time unit is the next time unit after the occurrence of the historical period, and the historical time unit includes T time sub-units; Obtain a target feature matrix, where the target feature matrix includes the periodic load and periodic description features of a target period. The target period includes D time units; Obtain a target feature vector, where the target feature vector includes the historical load of a target time unit and the description features of a first prediction time unit. The target time unit is an adjacent time unit before the first prediction time unit. The first prediction time unit is the next time unit after the occurrence of the target period. The target time unit includes T time sub-units, and the first prediction time unit includes T time sub-units; Based on the target feature matrix and the target feature vector, obtain a first feature matrix through a local sequence module; Based on the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector, obtain a second feature matrix through a global context module; Based on the first feature matrix and the second feature matrix, obtain the power grid load forecasting result corresponding to the first prediction time unit through a linear layer, where the power grid load forecasting result includes the load corresponding to each time sub-unit in the first prediction time unit.
2. The method according to claim 1, characterized in that, the obtaining of the historical feature matrix includes: Obtain the historical load corresponding to each time unit within the historical period; Obtain the historical description information corresponding to each time unit within the historical period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to each time unit within the historical period and the historical description information corresponding to each time unit within the historical period to obtain the historical feature matrix.
3. The method according to claim 1, characterized in that, the obtaining of the historical feature vector includes: Obtain the historical load corresponding to the historical time unit; Obtain the historical description information corresponding to the historical time unit, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to the historical time unit and the historical description information corresponding to the historical time unit to obtain the historical feature vector.
4. The method according to claim 1, characterized in that, the obtaining of the target feature matrix includes: Obtain the historical load corresponding to each time unit within the target period; Obtain the historical description information corresponding to each time unit within the target period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to each time unit within the target period and the historical description information corresponding to each time unit within the target period to obtain the target feature matrix.
5. The method according to claim 1, characterized in that, the obtaining of the target feature vector includes: Obtain the historical load corresponding to the target time unit; Obtain the description information corresponding to the first predicted time unit, where the description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to the target time unit and the description information corresponding to the first predicted time unit to obtain the target feature vector.
6. The method according to claim 1, characterized in that, the obtaining of the first feature matrix based on the target feature matrix and the target feature vector through the local sequence module includes: Perform attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain the encoding result of the target feature matrix; Perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain the first feature matrix.
7. The method according to claim 6, characterized in that, the encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least one first attention head and one second attention head; the performing of attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain the encoding result of the target feature matrix includes: For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the first attention head to obtain a first encoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the second attention head to obtain a second encoding result, where the second attention head is used to calculate the negative correlation between feature vectors; Obtain the encoding result of the target feature matrix according to the first encoding result and the second encoding result.
8. The method according to claim 6, characterized in that, the decoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least one first attention head and one second attention head; the performing of attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain the first feature matrix includes: For each multi-head attention layer in the decoder, based on the encoding result of the target feature matrix by the first attention head and the target feature vector, perform attention decoding processing to obtain a first decoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the decoder, based on the encoding result of the target feature matrix by the second attention head and the target feature vector, perform attention decoding processing to obtain a second decoding result, where the second attention head is used to calculate the negative correlation between feature vectors; According to the first decoding result and the second decoding result, obtain the first feature matrix.
9. The method according to any one of claims 1 or 6 to 8, wherein, The obtaining the second feature matrix through the global context module based on the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector includes: According to the context encoder included in the global context module, perform attention encoding processing on the historical feature matrix to obtain an encoding result of the historical feature matrix; According to the context encoder included in the global context module, perform attention encoding processing on the historical feature vector to obtain an encoding result of the historical feature vector; Based on the context attention mechanism, calculate the encoding result of the historical feature matrix, the encoding result of the historical feature vector, and the encoding result of the target feature matrix to obtain an attention encoding result; According to the context decoder included in the global context module, perform attention decoding processing on the attention encoding result and the target feature vector to obtain the second feature matrix.
10. The method according to claim 9, wherein, The context encoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head; The according to the context encoder included in the global context module, perform attention encoding processing on the historical feature matrix to obtain an encoding result of the historical feature matrix, includes: For each multi-head attention layer in the context encoder, based on the first attention head, perform attention encoding processing on the historical feature matrix to obtain a third encoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the context encoder, based on the second attention head, perform attention encoding processing on the historical feature matrix to obtain a fourth encoding result, where the second attention head is used to calculate the negative correlation between feature vectors; According to the third encoding result and the fourth encoding result, obtain the encoding result of the historical feature matrix; The according to the context encoder included in the global context module, perform attention encoding processing on the historical feature vector to obtain an encoding result of the historical feature vector, includes: For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector based on the first attention head to obtain a fifth encoding result; For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector based on the second attention head to obtain a sixth encoding result; According to the fifth encoding result and the sixth encoding result, obtain the encoding result of the historical feature vector.
11. The method according to claim 9, wherein, the context decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least one first attention head and one second attention head; The step of performing attention decoding processing on the attention encoding result and the target feature vector according to the context decoder included in the global context module to obtain the second feature matrix includes: For each multi-head attention layer in the context decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the first attention head to obtain a third decoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the second attention head to obtain a fourth decoding result, where the second attention head is used to calculate the negative correlation between feature vectors; According to the third decoding result and the fourth decoding result, obtain the second feature matrix.
12. The method according to claim 1, wherein, after predicting the grid load of the first prediction time unit according to the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector to obtain a first grid load prediction result, the method further includes: Obtain an updated feature matrix, where the updated feature matrix includes the periodic load and periodic description features of the updated period, the updated period includes D time units, and the updated period includes the first prediction time unit; Obtain an updated feature vector, where the updated feature vector includes the grid load prediction result corresponding to the first prediction time unit and the description features of the second prediction time unit, the first prediction time unit is an adjacent time unit before the second prediction time unit, the second prediction time unit is the next time unit after the updated period appears, and the second prediction time unit includes T time sub-units; Predict the grid load of the second prediction time unit according to the historical feature matrix, the historical feature vector, the updated feature matrix, and the updated feature vector to obtain a second grid load prediction result.
13. A grid load prediction device, wherein, comprising: An acquisition module, configured to acquire a historical feature matrix, where the historical feature matrix includes the periodic load and periodic description features of a historical period, the historical period includes D time units, each time unit includes T time subunits, and both D and T are integers greater than 1; The acquisition module is further configured to acquire a historical feature vector, where the historical feature vector includes the historical load and historical description features of a historical time unit, the historical time unit is the next time unit after the historical period appears, and the historical time unit includes T time subunits; The acquisition module is further configured to acquire a target feature matrix, where the target feature matrix includes the periodic load and periodic description features of a target period, and the target period includes D time units; The acquisition module is further configured to acquire a target feature vector, where the target feature vector includes the historical load of a target time unit and the description features of a first prediction time unit, the target time unit is an adjacent time unit before the first prediction time unit, the first prediction time unit is the next time unit after the target period appears, the target time unit includes T time subunits, and the first prediction time unit includes T time subunits; A prediction module, configured to predict the grid load of the first prediction time unit according to the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector, to obtain a first grid load prediction result; The prediction module is specifically configured to obtain a first feature matrix through a local sequence module based on the target feature matrix and the target feature vector; Obtain a second feature matrix through a global context module based on the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector; Based on the first feature matrix and the second feature matrix, obtain the grid load prediction result corresponding to the first prediction time unit through a linear layer, where the grid load prediction result includes the load corresponding to each time subunit in the first prediction time unit.
14. The apparatus according to claim 13, wherein, The acquisition module is specifically configured to acquire the historical load corresponding to each time unit in the historical period; Acquire the historical description information corresponding to each time unit in the historical period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to each time unit in the historical period and the historical description information corresponding to each time unit in the historical period to obtain the historical feature matrix.
15. The apparatus according to claim 13, wherein, The acquisition module is specifically configured to acquire the historical load corresponding to the historical time unit; Obtain the historical description information corresponding to the historical time unit, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to the historical time unit and the historical description information corresponding to the historical time unit to obtain the historical feature vector.
16. The apparatus according to claim 13, wherein, the obtaining module is specifically configured to obtain the historical load corresponding to each time unit within the target period; obtain the historical description information corresponding to each time unit within the target period, where the historical description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to each time unit within the target period and the historical description information corresponding to each time unit within the target period to obtain the target feature matrix.
17. The apparatus according to claim 13, wherein, the obtaining module is specifically configured to obtain the historical load corresponding to the target time unit; obtain the description information corresponding to the first prediction time unit, where the description information includes at least one of temperature information, season information, week information, weekday information, and holiday information; Encode the historical load corresponding to the target time unit and the description information corresponding to the first prediction time unit to obtain the target feature vector.
18. The apparatus according to claim 13, wherein, the prediction module is specifically configured to perform attention encoding processing on the target feature matrix according to the encoder included in the local sequence module to obtain an encoding result of the target feature matrix; Perform attention decoding processing on the encoding result of the target feature matrix and the target feature vector according to the decoder included in the local sequence module to obtain the first feature matrix.
19. The apparatus according to claim 18, wherein, the encoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least a first attention head and a second attention head; the prediction module is specifically configured to, for each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the first attention head to obtain a first encoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the encoder, perform attention encoding processing on the target feature matrix based on the second attention head to obtain a second encoding result, where the second attention head is used to calculate the negative correlation between feature vectors; Obtain the encoding result of the target feature matrix according to the first encoding result and the second encoding result.
20. The apparatus according to claim 18, wherein, the decoder includes at least one layer of multi-head attention layers, and each multi-head attention layer includes at least a first attention head and a second attention head; The prediction module is specifically configured to, for each multi-head attention layer in the decoder, perform attention decoding processing on the encoded result of the target feature matrix by the first attention head and the target feature vector to obtain a first decoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the decoder, perform attention decoding processing on the encoded result of the target feature matrix by the second attention head and the target feature vector to obtain a second decoding result, where the second attention head is used to calculate the negative correlation between feature vectors; Obtain the first feature matrix according to the first decoding result and the second decoding result.
21. The apparatus according to any one of claims 13 or 18 to 20, wherein, The prediction module is specifically configured to perform attention encoding processing on the historical feature matrix according to the context encoder included in the global context module to obtain an encoded result of the historical feature matrix; Perform attention encoding processing on the historical feature vector according to the context encoder included in the global context module to obtain an encoded result of the historical feature vector; Based on the context attention mechanism, calculate the encoded result of the historical feature matrix, the encoded result of the historical feature vector, and the encoded result of the target feature matrix to obtain an attention encoding result; Perform attention decoding processing on the attention encoding result and the target feature vector according to the context decoder included in the global context module to obtain the second feature matrix.
22. The apparatus according to claim 21, wherein, The context encoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head; The prediction module is specifically configured to, for each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix by the first attention head to obtain a third encoding result, where the first attention head is used to calculate the positive correlation between feature vectors; For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature matrix by the second attention head to obtain a fourth encoding result, where the second attention head is used to calculate the negative correlation between feature vectors; Obtain the encoded result of the historical feature matrix according to the third encoding result and the fourth encoding result; The prediction module is specifically configured to, for each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector by the first attention head to obtain a fifth encoding result; For each multi-head attention layer in the context encoder, perform attention encoding processing on the historical feature vector by the second attention head to obtain a sixth encoding result; Obtain the encoded result of the historical feature vector according to the fifth encoding result and the sixth encoding result.
23. The apparatus according to claim 21, wherein, the context decoder includes at least one multi-head attention layer, and each multi-head attention layer includes at least a first attention head and a second attention head; the prediction module is specifically configured to, for each multi-head attention layer in the context decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the first attention head to obtain a third decoding result, wherein the first attention head is used to calculate the positive correlation between feature vectors; for each multi-head attention layer in the decoder, perform attention decoding processing on the attention encoding result and the target feature vector based on the second attention head to obtain a fourth decoding result, wherein the second attention head is used to calculate the negative correlation between feature vectors; obtain the second feature matrix according to the third decoding result and the fourth decoding result.
24. The apparatus according to claim 13, wherein, the obtaining module is further configured to, after the prediction module predicts the grid load of the first prediction time unit according to the historical feature matrix, the historical feature vector, the target feature matrix, and the target feature vector to obtain a first grid load prediction result, obtain an updated feature matrix, where the updated feature matrix includes the periodic load and the periodic description feature of the updated period, the updated period includes D time units, and the updated period includes the first prediction time unit; the obtaining module is further configured to obtain an updated feature vector, where the updated feature vector includes the grid load prediction result corresponding to the first prediction time unit and the description feature of the second prediction time unit, the first prediction time unit is an adjacent time unit that appears before the second prediction time unit, the second prediction time unit is the next time unit that appears after the updated period, and the second prediction time unit includes T time sub-units; the prediction module is further configured to predict the grid load of the second prediction time unit according to the historical feature matrix, the historical feature vector, the updated feature matrix, and the updated feature vector to obtain a second grid load prediction result.
25. A computer device, wherein, it includes: a memory and a processor; wherein, the memory is used to store a program; the processor is used to execute the program in the memory, and the processor is used to execute the method according to any one of claims 1 to 12 according to the instructions in the program code.
26. A computer-readable storage medium, including instructions, which when running on a computer, cause the computer to execute the method according to any one of claims 1 to 12.
27. A computer program product, wherein, the computer program product includes instructions, which when running on a computer device, cause the computer device to execute the method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Power load prediction method and device, computer equipment and storage medium
CN111160625A