Knowledge graph-guided generating capacity prediction method
By constructing a regional power timing knowledge graph and extracting the timing and structural characteristics of the power entity, the problem of low accuracy in power generation prediction of traditional machine learning models is solved, and more accurate power generation prediction and more reasonable power supply solutions are achieved.
Patent Information
- Application Number
- CN202510004734.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional machine learning models are difficult to accurately reflect the changing laws of power generation in new scenarios, resulting in low accuracy of power generation prediction results.
Using the knowledge graph guidance method, by acquiring and preprocessing regional power data, a regional power timing knowledge graph is constructed, the timing characteristics and structural characteristics of power entities are extracted, and a globally embedded prediction model is constructed to predict power generation.
By comprehensively considering the structural characteristics and timing characteristics of the power entity, the model architecture of the timing convolution network is optimized, the accuracy of power generation prediction is improved, and a more reasonable regional power supply solution is provided.
Smart Images

Figure CN119940960A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to the field of data analysis technology. More specifically, the present application relates to a knowledge graph-guided power generation prediction method. Background Art
[0002] In order to meet the electricity demand in the power supply area, power supply companies need to formulate reasonable power supply plans based on the power generation capacity of power generation companies. Therefore, accurate power generation forecasting is the prerequisite for ensuring the balance of power supply and demand. The power generation forecasting task aims to accurately predict the power generation capacity of generators at different time granularities such as seasons and months, which is an important support for the reasonable dispatch of regional power.
[0003] Generally speaking, the power generation of a generator set is affected by many factors such as natural conditions and equipment status, and there is a high degree of uncertainty. At the same time, since power data is scattered in different power systems such as power stations and weather stations, it is difficult to mine and express the relationship between these data.
[0004] The commonly used method at present is to predict power generation based on traditional machine learning. However, traditional machine learning models are based on specific data distribution assumptions and are difficult to accurately reflect the changing patterns of power generation under new scenarios, resulting in low accuracy of power generation prediction results. Summary of the invention
[0005] In order to at least solve the technical problems mentioned above, this application proposes a knowledge graph-guided power generation prediction method in multiple aspects, including:
[0006] Acquire preprocessed regional power data, and construct a regional power time series knowledge graph based on the regional power data;
[0007] Extracting time series features and structural features of power entities from the regional power time series knowledge graph;
[0008] Constructing a global embedded prediction model of the power entity according to the time series and structural characteristics of the power entity;
[0009] Based on the prediction model and regional power data, power generation is predicted.
[0010] In some embodiments, before obtaining the pre-processed regional power data and constructing the regional power time series knowledge graph based on the regional power data, the method further includes:
[0011] Acquiring comprehensive power data, wherein the comprehensive power data includes text data and numerical data;
[0012] Performing semantic division on the text data and defining entity types;
[0013] Based on the data points in the numerical data, the abnormal data therein are corrected to complete the preprocessing of the comprehensive power data and obtain the regional power data.
[0014] In some embodiments, constructing a regional power time series knowledge graph based on the regional power data includes:
[0015] identifying power entities in the regional power data;
[0016] extracting power relations from the regional power data;
[0017] Based on the power entities and the power relationships, the regional power time series knowledge graph is constructed.
[0018] In some embodiments, extracting power relations from the regional power data comprises:
[0019] Marking the entity type of the regional power data;
[0020] Specially extract the marked entity types to obtain corresponding word embedding matrices;
[0021] Determining, according to the word embedding matrix, a probability that the power relationship exists between power entities;
[0022] The time type entity of the regional power data is used as the timestamp of the power relationship to obtain a power relationship set.
[0023] In some embodiments, extracting the time series features and structural features of power entities from the regional power time series knowledge graph includes:
[0024] The power entities of the regional power time series knowledge graph and the power relations between the power entities are used as initial embeddings and embedded into a low-dimensional vector space based on a preset model;
[0025] Based on the preset model, learning the time series dependency characteristics of the power relationship corresponding to the power entity to obtain a time series embedding set containing the time series characteristics of the power entity and a power relationship time series embedding set;
[0026] The triples representing the regional power time series knowledge graph are concatenated with the initial linear vector in a dimensionally expanded manner;
[0027] Performing power feature mapping on the initial linear vector to obtain a triplet embedding vector;
[0028] Based on the preset model and the triple embedding vector, an entity structure embedding set containing the structural features of the power entity is output.
[0029] In some embodiments, constructing a prediction model for global embedding of the power entity according to the timing and structural characteristics of the power entity includes:
[0030] Based on the temporal embedding set corresponding to the temporal features and the entity structure embedding set of the structural features, calculating the attention weight vector corresponding to the prediction model;
[0031] According to the attention weight vector, weighted summing is performed on the structural embedding of the power entity at each moment to obtain a global structural embedding vector of the power entity;
[0032] The time series set embedding vector and the global structure embedding vector are concatenated, and feature fusion is performed using a perceptron with a preset number of layers to obtain a global embedding set of the power entity;
[0033] Based on the global embedding set, the prediction model is constructed.
[0034] In some embodiments, predicting power generation based on the prediction model and regional power data includes:
[0035] Based on the prediction model, extracting features of the regional power data to obtain a local feature matrix;
[0036] Normalizing the local feature matrix to obtain a global feature matrix;
[0037] The global characteristic matrix is transformed bilinearly to calculate the predicted value of each power generation influencing factor, and the predicted results of the time series entities are summed to obtain the predicted power generation.
[0038] In some embodiments, the normalizing the local feature matrix to obtain a global feature matrix includes:
[0039] Converting the distribution of each dimension of the local feature matrix into a standard normal distribution to obtain a standardized local feature matrix;
[0040] The local feature matrix is adjusted to a target distribution based on a scaling factor and an offset factor to obtain the global feature matrix.
[0041] In a second aspect, the present application also provides a device, comprising: a memory, a processor, and a knowledge graph-guided power generation prediction program stored on the memory and executable on the processor, the knowledge graph-guided power generation prediction program being configured to implement the steps of the knowledge graph-guided power generation prediction method as described in any one of claims 1 to 8.
[0042] In a third aspect, the present application also provides a computer-readable storage medium, on which a knowledge graph-guided power generation prediction program is stored, and when the knowledge graph-guided power generation prediction program is executed by a processor, the steps of the knowledge graph-guided power generation prediction method as described in any of the above items are implemented.
[0043] The embodiment of the present application provides a method for predicting power generation guided by a knowledge graph. Aiming at the close connection between factors affecting power generation, a knowledge graph describing power timing information is constructed to establish associations between power entities. Then, the triples in the knowledge graph are used to capture the structural associations between entities, thereby improving the timing characteristics and structural characteristics of power entities. Based on these two characteristics, the model architecture of the timing convolutional network is optimized to construct a prediction model. Finally, the prediction model is used to predict power generation. Through the above method, the impact of the structural characteristics and timing characteristics of power entities on power generation is comprehensively considered, making the prediction results more accurate and providing technical support for power supply companies to formulate reasonable regional power supply plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present application will become easy to understand. In the accompanying drawings, several embodiments of the present application are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0045] Figure 1 This is a flow chart of a first embodiment of a knowledge graph-guided power generation prediction method of the present application;
[0046] Figure 2 This is a flow chart of a second embodiment of a knowledge graph-guided power generation prediction method of the present application;
[0047] Figure 3 This is a table diagram of regional power comprehensive data in the first embodiment of the present application;
[0048] Figure 4 A schematic diagram of the process of constructing a regional power time series knowledge graph in the example of this application;
[0049] Figure 5 This is an example diagram of the regional power time series knowledge graph in the example of this application;
[0050] Figure 6 This is a schematic diagram of a numerical data table related to power generation in the example of this application;
[0051] Figure 7 This is a schematic diagram of a timing embedding set table in an example of this application;
[0052] Figure 8This is a schematic diagram of the feature matrix table output by the TCN model in the example of this application;
[0053] Fig. 9 It is a schematic diagram of the terminal structure of the hardware operating environment involved in the embodiment of the present application. DETAILED DESCRIPTION
[0054] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0055] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present application indicate the presence of described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or collections thereof.
[0056] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this application specification and claims, unless the context clearly indicates otherwise, the singular forms of "a", "an" and "the" are intended to include plural forms. It should also be further understood that the term "and / or" used in this application specification and claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes these combinations.
[0057] As used in this specification and claims, the term "if" may be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]," depending on the context.
[0058] The specific implementation of the present application is described in detail below with reference to the accompanying drawings.
[0059] The present application embodiment provides a knowledge graph-guided power generation prediction method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of a knowledge graph-guided power generation prediction method of the present application.
[0060] Step S101: Obtain preprocessed regional power data, and construct a regional power time series knowledge graph based on the regional power data.
[0061] In this embodiment, the regional power data can be obtained by collecting the power system operation documents of a specific power supply area. Generally, the power system operation documents contain regional power comprehensive data. Figure 3 ,like Figure 3 The regional power comprehensive data table shown contains comprehensive data such as basic information of power stations, operating parameters of generator sets, meteorological and hydrological data, and electricity consumption information of enterprises. Regional power comprehensive data includes text data and numerical data. Regional power data can be obtained by preprocessing the power comprehensive data.
[0062] In order to explain more clearly how to construct a regional power time series knowledge graph based on regional power data, the following first introduces how to preprocess the regional power comprehensive data. It can be understood that the preprocessing step is performed before step S101.
[0063] In the regional power comprehensive data, text data describes the detailed information of the power system operation, and numerical data reflects the physical quantity and performance indicators of the power elements. Acquiring these two types of data at the same time can fully understand the operating status of the power system in the power supply area from both qualitative and quantitative perspectives. For text data, regular expressions are used to match common date and time formats, and dates in different formats are parsed into standard date objects and uniformly output in ISO-8601 time format. After that, the word segmentation tool is used to semantically divide the text data and define entity types, such as "power generation facilities", "power supply units", "power users" and "natural conditions". Optionally, the BIO annotation system can also be used to annotate the text after word segmentation, and define the use of "B-entity category" to mark the beginning of the entity, "I-entity category" to mark the rest of the entity, and "O" to mark the non-entity part, so that the entity type definition is clearer.
[0064] For numerical data, considering that regional power data involves different dimensions, it is easy to be missing and wrong. In a feasible implementation, Lagrange interpolation and Min-Max Scaling can be used to pre-process the numerical data. The data points in the numerical data are represented as in, It represents the measured value of numerical data at time t (1≤t≤T), T (T≥1) is the length of the time series, and the corresponding Lagrangian basis function L(t) is calculated according to the data point, and then the interpolation polynomial P(t) is constructed using the data point and the Lagrangian basis function. Finally, the time when there is an outlier is Substitute into P(t) and calculate its corresponding value This is used to correct abnormal data. Then, according to the maximum value of each numerical data and minimum value Use the following formula to convert the measured value of numeric data into Normalized to the interval [0,1]:
[0065]
[0066] After preprocessing all text data and numerical data, we can get the text data set and numerical datasets That is, the pre-processed regional power data. is the λth (1≤λ≤C) text data, w i for The i-th sentence in For the Numeric data for The measured value at time q (1≤q≤T).
[0067] Next, based on the regional power data, power entities can be identified and power relationships between power entities can be extracted. Power relationships reflect the connections between power entities, and through the power relationships between power entities, a regional power time series map can be constructed. In one embodiment, a specific model can be used to extract power entities, which can improve the accuracy of power entity identification.
[0068] In the above embodiment, by constructing a regional power time series knowledge graph, the power supply and demand relationship in the power supply area and the potential connection between the factors affecting the power volume can be revealed, which can be used in subsequent power generation forecasting work to improve the accuracy of the forecast.
[0069] Step S102: extracting the timing characteristics and structural characteristics of power entities from the regional power timing knowledge graph.
[0070] In this embodiment, the power entities of the regional power time series knowledge graph include time series power entities. A time series power entity refers to an entity whose state or attribute changes over time, such as the real-time power and real-time voltage of a motor in a power system. The measured value of the time series power entity reflects the state of the entity at different times. By embedding and updating the time series power entity using the measured value, the state change law of the power entity can be mined, thereby improving the accuracy of power generation prediction. In addition, in the regional power time series knowledge graph, the structured features of the power triples can intuitively reflect the connections and relationships between power entities, so they can be used to extract structural features.
[0071] Furthermore, when extracting time series features, the regional power knowledge graph (hereinafter referred to as The power entity and the power relationship between the power entities are used as the initial embedding, and are embedded into the low-dimensional vector space through a preset model. The prediction model can be a DistMult model, which has the characteristics of high computational efficiency and strong expression ability. Based on the DistMult model, the power entity is embedded into the low-dimensional vector space. First, The power entities and the power relations between power entities in the data are used to generate random d (d>0) dimensional vectors as the initial embedding. Then, the semantic coherence of entities and relations is measured by a bilinear scoring function, and the entity vector is updated based on the marginal loss to optimize the embedding representation of the entity and obtain the initial embedding matrix of the entity. and the relation initial embedding matrix Among them, e u is the uth (1≤u≤l)th entity e u The initial embedding of For the Relationship The initial embedding of .
[0072] Using electric power u The measured value at time t For the initial embedding e u To maintain the integrity of the data in the time series, when there are missing measurements at certain moments, the measurements corresponding to these moments are set to their historical mean. The feedforward neural network of the DistMult model is used to measure the measurements. For the initial embedding e u The impact of this update u In different dimensions, e u The embedding can be dynamically adjusted based on the measurement information. u Each dimension of e u The temporal embedding at time t
[0073]
[0074] in, Represents the vector e u The a-th (1≤a≤d)-th dimension element at time t, d is the dimension of entity embedding, [a] is used to indicate the a-th dimension element in the vector, is the Sigmoid activation function, w u and b u is a weight vector, which is initialized according to the statistical characteristics of the key influencing factor data, reflecting the contribution weight of the measurement values at different times to the entity embedding.
[0075] In order to more accurately reflect the dynamic connection between power entities, the measurement values of the head and tail entities are also used to update the relation embedding to learn the temporal dependency characteristics of the entity correspondence relationship, thereby improving the accuracy of power generation prediction. The triples in Among them, e u and e ν is a pair of entities in the power entity set E, is the first relationship, first through e u With e ν The measured value and Update power relation embedding element by element Get The representation at time t and Then, and Splicing is performed and obtained through a feedforward neural network The relation embedding at time t
[0076]
[0077] Among them, w ν and b ν is the weight vector, W r is the weight matrix.
[0078] The above method is used to calculate the time series embedding of all power entities and relationships in the regional power time series knowledge graph, and obtain the entity time series embedding set containing time series features. and relational temporal embedding sets in, and Entity e u and relationship The set of time series embeddings at T time moments.
[0079] Furthermore, when extracting structural features, we use The embedding of triples in the tuple represents the structural information of the entity. u The set of triples is in, for e u The γth (1≤γ≤|S u |) triples. For the power triple Assume that the embeddings of the entities and relations at time t are and The three embeddings are concatenated into a new vector in element order by dimensional expansion, and the concatenated result is mapped to power features through a linear structure to obtain a triplet embedding
[0080]
[0081] Among them, W t is the weight matrix.
[0082] Next, calculate S u The representation of China Electric Power triples, and the triple embedding set Adjust e through the attention mechanism u For S u The degree of attention of different triples embedded in the , thus more accurately reflecting the dependency relationship between power entities. u The temporal embedding at time t First calculate Embedding with different triples The attention weight Then After normalization, we get the attention weight vector Then according to α u For S u The triple embeddings in are weighted summed to obtain the power entity e u The structural embedding at time t
[0083]
[0084] Calculate according to the above method The structural features of all time-series power entities are obtained to obtain the entity structure embedding set containing the structural features. in, For entity e u The set of structural embeddings at each time instant.
[0085] In the above embodiment, the time series characteristics of the power entity can reflect the dynamic change law of the power system. And the structural characteristics of the power entity can reflect the dependency information between the power entities. By extracting the above two characteristics and applying them to the structural adjustment of the prediction model, the prediction model can capture the deep correlation between the power entities, thereby improving the reliability of the prediction results.
[0086] Step S103: construct a prediction model for global embedding of power entities according to the time series and structural characteristics of the power entities.
[0087] In this embodiment, in order to obtain a more comprehensive feature representation of the power entity and improve the accuracy of power generation prediction, the attention mechanism can be used to focus on the important structural features of the power entity, and a multi-layer perceptron can be used to capture the complex interaction between the timing and structural features. u The temporal embedding at time t First calculate Embedded with entity structures at different times The attention weight Get the attention weight vector According to For u The structural embeddings at each moment are weighted summed to obtain the global structural embedding of the entity
[0088]
[0089] Then, the timing of the power entity is embedded in and global structural embedding Concatenate into a new vector in element order, that is, concatenate in a dimensionally expanded manner and In order to make nonlinear changes to power features and prevent overfitting, a perceptron with a preset number of layers (such as two layers) is used to perform feature fusion to obtain the global embedding of the power entity.
[0090]
[0091] in, is the weight matrix of the first layer of the multilayer perceptron, is the weight matrix of the second layer.
[0092] According to the above process, calculate The global embedding set of all power entities in in, For entity e u The global embedding at T time moments. This set can be used to build and adjust the model structure of the prediction model.
[0093] Step S104: predicting power generation based on the prediction model and regional power data.
[0094] In this embodiment, in order to achieve accurate prediction of power generation, considering that the temporal convolutional network (TCN) model has good time series feature extraction capabilities, the feature extraction process of TCN is optimized, and power generation prediction is performed based on the extended TCN model.
[0095] First, for the power entities in the region u Global embedding of First, extract the dilation factor convolution Local features of multiple time steps, set the convolution kernel The size in the horizontal and vertical directions is The expansion factor is η (η≥1), according to and η to calculate the expansion factor Use ζ(ζ≥0) to separate the matrices Expand the rows and columns of to get the expanded matrix from Starting from the upper left corner, the convolution kernel and The local area elements of are multiplied and summed. Scanning row by row completes the expansion matrix Get e u The local feature matrix F u :
[0096]
[0097] Then, batch normalization is used to improve the expressiveness of the local feature matrix. First, F u The distribution of each dimension of data in is transformed into a standard normal distribution to obtain the standardized local feature matrix Introducing the learnable scaling factor δ and offset factor κ will Adjust to a suitable distribution to more effectively reflect the time series characteristics within the data and obtain the batch normalized local feature matrix
[0098]
[0099] Here, “×” means element-by-element multiplication of the vectors, that is, each element in the vector is scaled independently.
[0100] Considering that the TCN model gives the same weight to each moment during feature extraction, it is difficult to adjust the importance of different moments according to the changes in power data in the region. This application introduces the attention mechanism. The global information dynamically adjusts the representation of each time step to obtain the global feature matrix
[0101]
[0102] Among them, W Q , W K and W V They are the Query, Key and Value weight matrices respectively.
[0103] Will Calculate the predicted value of each power generation influencing factor through bilinear transformation And sum up the prediction results of l time series entities to get the final power generation
[0104]
[0105] Among them, w k and w l is the weight vector.
[0106] In this embodiment, the timing and structural characteristics of power entities are integrated through the attention mechanism, the model architecture of the temporal convolutional network is optimized, the long-distance dependency information of the entity is better captured, and the accurate prediction of power generation is achieved.
[0107] Reference Figure 2 , Figure 2 This is a flow chart of the second embodiment of a knowledge graph-guided power generation prediction method of this application. It should be noted that the process steps of this embodiment can be understood as Figure 1 Therefore, the above description can also be applied to this embodiment.
[0108] Step S201: Identify power entities in regional power data.
[0109] In an embodiment, a specific BERT pre-trained model can be used to extract power entities. The BERT pre-trained model has a strong ability to capture contextual information, so as to accurately understand the specific semantics of power terms in various contexts, and effectively improve the accuracy of power entity recognition. The bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, referred to as BiLSTM) network can simultaneously capture the information in the front and back directions of the power text, and better understand the overall semantics of the text. The conditional random field (Conditional Random Field, referred to as CRF) considers the dependency between power tags, thereby optimizing the results of sequence labeling. This application adopts the BERT-BiLSTM-CRF model to realize entity recognition and improve the accuracy of power entity recognition.
[0110] For sentences BERT is For each word in the text, a corresponding word vector, position vector, and segment vector are generated. Then the three vector representations are summed and input into the encoding layer of BERT for feature extraction. The encoding layer consists of multiple Transformer structures. Transformer captures the contextual relationship between words through the self-attention mechanism, which can better reflect the accurate semantics of the word in the entire power text context, and finally outputs the context-related word vector representation for each word:
[0111] h1,…,h i …,h n =BERT(w1,…,w i ,…w n )
[0112] The BiLSTM network consists of a forward LSTM and a backward LSTM to further extract the semantic information of the sentence. For a given vector sequence h1,…,h i …,h n , the forward LSTM extracts features in the order of the sequence, while the reverse LSTM processes the data in reverse order from the end of the sequence. The bidirectional features can provide a more comprehensive and accurate semantic feature representation, which helps to improve the accuracy and reliability of text information processing related to the power system. Finally, the output vectors of the two directions are concatenated as the output of the BiLSTM network:
[0113] h'1,…,h' i ,…,h' n =BiLSTM(h1,…,h i …,h n )
[0114] The CRF model is based on the output vector h'1,…,h' of BiLSTM i ,…,h' n , calculate the probability of each possible label sequence under the given power text, the CRF model selects the sequence with the highest probability from all possible label sequences, and obtains the final power entity labeling result:
[0115] x1,…,x i ,…,x n =CRF(h'1,…,h' i ,…,h' n )
[0116] Scan the power entity tag sequence x1,…,x from the beginning of the sentence i ,…,x n , determine the boundary of the power entity based on the label information, and then obtain the power entity set in the sentence by concatenating the text within the boundary Among them, e u for The u-th (1≤u≤c) entity in .
[0117] Step S202: extract power relations from regional power data.
[0118] There are many types of relationships in the power system and they are highly professional. This application inserts tags with entity type information into the power text to highlight the head and tail entities, thereby effectively injecting prior knowledge of power entities. The power relationship types are defined as "attribute", "connection", "influence", and "supply". and Any pair of entities e in u and e ν (1≤u≤ν≤c), according to the previous step, we get e u and e ν The entity types are type u and type ν , thereby generating head and tail entity type tags <S:type u >、 < / S:type u >、 <O:type ν > and < / O:type ν >. Insert the above markup into Zhonge u and e ν The front and back positions of the corresponding entity, and get the sentence with entity type tag Will Input BERT for feature extraction and get the sentence The word embedding matrix of :
[0119]
[0120] The entity e u and e ν Embedding at the start position and Concatenate into a new vector in element order, that is, concatenate in a dimensionally expanded manner and Get entity pair embedding h u,ν Input the feedforward network to get the relationship prediction vector r u,ν :
[0121] r u,ν =W r h u,ν +b r
[0122] Among them, W r is the weight matrix, b r is the bias vector.
[0123] r u,ν The value of each dimension element in represents the entity e u and e ν The probability that there is a certain power relationship between them is u,ν If the values of all elements in are less than the given threshold ε (0≤ε≤1), then e u and e ν There is no relationship between them, otherwise output r u,ν The relationship corresponding to the largest term in u,ν At the same time, the time type entity identified in step 1.2 is used as the timestamp of the relationship to form entity e u and e ν The temporal relationship prediction results of Entities and relationships in the grid, get the power entity set Power Relationship Set and power triplet set
[0124] Step S203: construct a regional power time series knowledge graph based on the power entities and the power relationships.
[0125] In this embodiment, based on the above steps, according to S, the power entity is used as the graph node, the measurement value of the entity is used as the attribute value of the node, the power relationship is used as the edge of the node in the graph, the edge type corresponds to the power relationship in R, and the timestamp corresponding to the entity is used as the index of the entity measurement value, and the timestamp corresponding to the relationship is used as the time mark of the power event, thereby reflecting the state of the power system at different time points, and obtaining the regional power time series knowledge graph
[0126] To facilitate understanding of the technical solution of the present application, the technical solution of the present application is explained below with reference to an example of power generation forecast for a certain power supply area in Yunnan Province in the fourth quarter of 2022.
[0127] Reference Figure 4 , the construction process of regional power time series knowledge graph is as follows Figure 4 As shown in the figure, for the 92 natural days in the fourth quarter of 2022, the power information in the power system operation report during this period is obtained, including text data and numerical data. The text data involves a total of 6124 sentences, and the sampling period of the numerical data is 8:00-18:00 every day, and the sampling frequency is once an hour, resulting in a total of 1104 sets of time series data.
[0128] For text data, standardize dates in different formats into ISO-8601 time format. For example, standardize the date "October 7, 2022" into "2022-10-07". Use the open source word segmentation tool Jieba to perform semantic segmentation on the text. For example, for the sentence "On November 10, 2022, the average wind speed in the area was 3.5 m / s" in the weather report, the result after word segmentation is "On November 10, 2022, the average wind speed in the area was 3.5 m / s". Define entity types as "power generation facilities", "transmission units", "power units", and "natural conditions", and use the BIO annotation system for annotation. For example, for the entity "Hydroelectric Power Station A", annotate "A" as "B-Power Station", "Water" as "I-Power Station", etc.
[0129] For numerical data, the Lagrange interpolation method is first used to interpolate missing values and outliers. Taking the evaporation data as an example, for the data point (3,?) with missing values, the corresponding Lagrange basis function L(t) is calculated based on the effective value of the evaporation, and then the Lagrange basis function is used to construct the interpolation polynomial P(t). Then, the time t=3 with missing values is substituted into P(t), and the evaporation at that time is 3.1mm. The interpolated power data is normalized. Taking the data point (12,630) of runoff at time 12 as an example, according to the maximum value of the runoff and minimum value The data range of the calculated runoff data is [0, 991], and then the difference between the runoff value of 630 at time 12 and is 345. Divide the difference of 345 by the upper limit of the data range, 991, so as to scale the data proportionally to the interval [0, 1]. After preprocessing the text data and the numerical dataset, the text dataset and the numerical dataset Among them, the sentences in describe the occurrence of power events, shows the time-series changes of power generation factors. Some numerical data are as shown in Figure 6 shown.
[0130] Next, use the BERT-BiLSTM-CRF model for entity recognition. Take the power text "The No. 1 generator set of Jia Hydroelectric Power Station failed on November 15, 2022, resulting in a decrease in power generation" as an example. Use BERT to capture the context semantic information of words, and use BiLSTM to further obtain the bidirectional semantic dependency relationship of the power text. Use CRF to consider the adjacent label connection, and find the tag sequence with the highest probability through the Viterbi algorithm. Finally, determine entities such as "Jia Hydroelectric Power Station" and "No. 1 Generator Set". For all perform entity recognition to obtain the power entity set
[0131] Define the power relationship types as "attribute", "connection", "influence", and "supply". For the power text "The power generation of Jia Hydroelectric Power Station is affected by the runoff", according to the entities e 10 "Jia Hydroelectric Power Station" and e 13 "runoff" entity types generate tags <S: power station>, < / S: power station>, <O: natural condition>, < / O: natural condition>. Insert the above tags into the original sentence to get the sentence with entity tag information "<S: power station>Jia Hydroelectric Power Station< / S: power station>'s power generation is affected by <O: natural condition>runoff< / O: natural condition>". Input the sentence with entity type tags into the BERT model for encoding to obtain the context representation with tag information. Concatenate the start and end positions of the head and tail entities to output the representation to get the entity pair embedding h 10,13 , set the threshold ε = 0.7, and input h 10,13 into the feedforward neural network for prediction to obtain that the relationship type between e 10 and e 13 is r2 "influence". Perform relation extraction on 6124 sentences in the text data to obtain the power triple set Finally, generate the regional power time-series knowledge graph according to S Among them, is a collection of power entities, It is a set of power relations. The entity types include "power generation facilities", "transmission units", "power consumption units" and "natural conditions". The relationship types include "attribute", "connection", "influence" and "supply". Regional power time series knowledge graph like Figure 5 shown.
[0132] Furthermore, the DistMult model is used to The entities and relations in are embedded into a low-dimensional vector space to obtain the initial embedding matrix of the entity and the relation initial embedding matrix The dimension of entity and relationship embedding is 512. Then, the initial embedding vector is updated using the measured value of the entity. Taking entity e7 "flow rate" as an example, the influence of its measured value 0.725 at time t=25 on the embedding is calculated, and the values of different dimensions of e7 are updated to obtain the representation of e7. Thus, the influence of the change of flow velocity in time series on the hydropower generation is fully considered. For the triplet < flow velocity e7, influence r2, motor speed e 61 >, using the measured values of flow velocity and motor speed 0.725 and 0.870 respectively to update the relation embedding, we get the representation of the relation “influence” r2 at time t=25 and Will and Feature mapping is performed through a feedforward neural network to obtain the complete embedding of r2 Calculated by the above method The temporal embedding of all entities and relations in , and the entity temporal embedding set is obtained and relational temporal embedding sets in, and Entity e u and relationship The time series embedding set at 1104 moments is used to reflect the dynamic characteristics of power entities and relations. u The temporal embedding set like Figure 7 shown.
[0133] Furthermore, the elements in the triples are concatenated in a dimensionally expanded manner, and the concatenated results are subjected to power feature mapping through a linear structure to obtain triple embedding. The attention mechanism is used to calculate the degree of attention of e7 to different triples, so as to focus more accurately on the relevant triple information. For the time series embedding of the entity "flow rate" at time t = 71 calculate Embedding with different triples The attention weight Will After normalization, we get the attention weight vector β 7 = Through Beta 7 We perform weighted summation of the triple embeddings in S7 to obtain the structural embedding of e7 According to the above method, the structural characteristics of all time-series power entities are calculated to obtain the entity structure embedding set
[0134] Take the embedding of "Power Plant Entity A" at time t=64 Take 2 as an example, the corresponding structure vector at each moment is Calculate the attention weights of the structure vector and the current time series embedding use The weighted sum of the structure vectors is used to obtain the global structure embedding This focuses on the key structural information of power entities and Perform the above operations to obtain the global embedding set of power entities A multi-layer perceptron is used to capture the complex interaction between temporal and structural features. The entity e is obtained by concatenating the temporal embedding and global structural embedding of the entity and inputting them into two layers of perceptrons for feature fusion. u Global embedding at time t=64 According to the above method, the global embedding of all entities is calculated to obtain the global embedding set in, For entity e u Global embedding at 1104 moments. It aims to reflect the overall structure of the regional power system and the dynamic relationship between various entities.
[0135] The local features of entity embedding are extracted through the TCN model. 33 Taking "precipitation" as an example, capturing The long-term dependence information of the precipitation can be accurately reflected in the temporal variation trend of the precipitation, and the local feature matrix F of the precipitation can be obtained. 33 Then, for F 33 Perform batch normalization to adjust the characteristic distribution of precipitation to better suit the power generation prediction task, and obtain the output matrix of the TCN model Some outputs are as follows Figure 8 Next, calculate The attention weights are updated to reflect the contribution of precipitation to power generation at different times, and the global feature matrix is obtained. Finally, the power generation forecast for precipitation is calculated The prediction results of the remaining 871 entities are summed up to get the final power generation
[0136] Reference Fig. 9 , the figure is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present application.
[0137] like Fig. 9 As shown, the device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (WIreless-FIdelity, WI-FI) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM) memory, or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk memory. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.
[0138] Those skilled in the art will understand that Fig. 9 The structure shown in the figure does not constitute a limitation of the device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0139] like Fig. 9 As shown, the memory 1005 as a storage medium may include an operating system, a data storage module, a network communication module, a user interface module, and a knowledge graph-guided power generation prediction program, and execute the knowledge graph-guided power generation prediction method provided in the embodiment of the present application.
[0140] Although multiple embodiments of the present application have been shown and described herein, it is obvious to those skilled in the art that such embodiments are provided only by way of example. Those skilled in the art can think of many changes, modifications and alternatives without departing from the thought and spirit of the present application. It should be understood that in the process of practicing the present application, various alternatives to the embodiments of the present application described herein can be adopted. The attached claims are intended to limit the scope of protection of the present application, and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A knowledge graph-guided power generation prediction method, characterized in that: include: Acquire preprocessed regional power data, and construct a regional power time series knowledge graph based on the regional power data; Extracting time series features and structural features of power entities from the regional power time series knowledge graph; Constructing a global embedded prediction model of the power entity according to the time series and structural characteristics of the power entity; Based on the prediction model and regional power data, power generation is predicted.
2. The method according to claim 1, characterized in that Before obtaining the pre-processed regional power data and constructing the regional power time series knowledge graph based on the regional power data, the method further includes: Acquiring comprehensive power data, wherein the comprehensive power data includes text data and numerical data; Performing semantic division on the text data and defining entity types; Based on the data points in the numerical data, the abnormal data therein are corrected to complete the preprocessing of the comprehensive power data and obtain the regional power data.
3. The method according to claim 1, characterized in that The constructing of a regional power time series knowledge graph based on the regional power data includes: identifying power entities in the regional power data; extracting power relations from the regional power data; Based on the power entities and the power relationships, the regional power time series knowledge graph is constructed.
4. The method according to claim 3, characterized in that: The extracting power relationship from the regional power data comprises: Marking the entity type of the regional power data; Specially extract the marked entity types to obtain corresponding word embedding matrices; Determining, according to the word embedding matrix, a probability that the power relationship exists between power entities; The time type entity of the regional power data is used as the timestamp of the power relationship to obtain a power relationship set.
5. The method according to claim 1, characterized in that The extracting of time series features and structural features of power entities from the regional power time series knowledge graph includes: The power entities of the regional power time series knowledge graph and the power relations between the power entities are used as initial embeddings and embedded into a low-dimensional vector space based on a preset model; Based on the preset model, learning the time series dependency characteristics of the power relationship corresponding to the power entity to obtain a time series embedding set containing the time series characteristics of the power entity and a power relationship time series embedding set; The triples representing the regional power time series knowledge graph are concatenated with the initial linear vector in a dimensionally expanded manner; Performing power feature mapping on the initial linear vector to obtain a triplet embedding vector; Based on the preset model and the triple embedding vector, an entity structure embedding set containing the structural features of the power entity is output.
6. The method according to claim 1, characterized in that The step of constructing a prediction model for global embedding of the power entity according to the time series and structural characteristics of the power entity includes: Based on the temporal embedding set corresponding to the temporal features and the entity structure embedding set of the structural features, calculating the attention weight vector corresponding to the prediction model; According to the attention weight vector, weighted summing is performed on the structural embedding of the power entity at each moment to obtain a global structural embedding vector of the power entity; The time series set embedding vector and the global structure embedding vector are concatenated, and feature fusion is performed using a perceptron with a preset number of layers to obtain a global embedding set of the power entity; Based on the global embedding set, the prediction model is constructed.
7. The method according to claim 1, characterized in that The predicting of power generation based on the prediction model and regional power data includes: Based on the prediction model, extracting features of the regional power data to obtain a local feature matrix; Normalizing the local feature matrix to obtain a global feature matrix; The global characteristic matrix is transformed bilinearly to calculate the predicted value of each power generation influencing factor, and the predicted results of the time series entities are summed to obtain the predicted power generation.
8. The method according to claim 7, characterized in that The normalizing the local feature matrix to obtain a global feature matrix includes: Converting the distribution of each dimension of the local feature matrix into a standard normal distribution to obtain a standardized local feature matrix; The local feature matrix is adjusted to a target distribution based on a scaling factor and an offset factor to obtain the global feature matrix.
9. A device, characterized in that: The device includes: a memory, a processor, and a knowledge graph-guided power generation prediction program stored in the memory and executable on the processor, wherein the knowledge graph-guided power generation prediction program is configured to implement the steps of the knowledge graph-guided power generation prediction method as described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a knowledge graph-guided power generation prediction program, which, when executed by a processor, implements the steps of the knowledge graph-guided power generation prediction method as described in any one of claims 1 to 8.