A real-time query method, system and storage medium for massive power grid data
By adopting a top-down global index and local index structure in the power grid system, combined with the index update mechanism of the sliding time window, the problem of real-time query and rapid update of power grid data is solved, and efficient grid data processing and query is achieved.
Patent Information
- Application Number
- CN202210705194.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-06-21
AI Technical Summary
It is difficult for the prior art to realize real-time query and rapid update of power data in power grid systems, especially when processing massive real-time data streams, there are shortcomings in query efficiency and update efficiency.
A top-down global index structure is adopted, and a local index is established at the leaf nodes of the global index. Combined with the index update mechanism of the sliding time window, real-time update and fast query of the power grid data trajectory flow are realized.
Through real-time updates and fast query methods, the processing speed and query efficiency of power grid data are significantly improved, resource overhead is reduced, and rapid and real-time query of massive power grid data is achieved.
Smart Images

Figure CN115098543B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method, a system and a storage medium for querying power grid data. Background Art
[0002] In the power grid system, substation centralized control plays a vital role, because it not only drives the transmission and distribution work, but also is responsible for the assembly of electrical equipment, the adjustment of voltage and the connection of important tasks. A large amount of power data is generated in real time and needs to be processed continuously. For example, users can make quick decisions or predictions on power centralized control through real-time analysis of historical and current information. In the real-time processing system, each input primitive plays the role of retrieval and update. Specifically, for such a system, all existing tree indexes (such as R-tree) can only meet the requirements of query efficiency, but perform poorly in terms of update efficiency.
[0003] Big data platforms (such as Flink and Spark) provide a set of relatively convenient data management and development capabilities. Although these platforms can quickly customize corresponding services according to target needs, this simple and crude method is often high-cost and low-performance, which is a loss-making behavior. The essence of power data monitoring in the power grid system is to process a large amount of graph data in real time, and the similarity in graph data often implies important meanings. In order to solve this graph similarity problem, there are many methods for graph similarity calculation, mainly centralized and distributed methods, which are briefly described in the following two paragraphs.
[0004] Centralized graph similarity methods often define different distance measurement methods and prove the adaptability of different distance metrics from a mathematical perspective. For example, the distance between people is often replaced by "Euclidean distance" in daily life, but in actual industrial applications, there may also be "Manhattan distance", "Edit distance", etc.
[0005] Distributed graph similarity is often built on a big data platform based on Hadoop. The entire graph similarity processing system is handled by a master-slave distributed cluster. The master node mainly completes the coordination work, allowing different slave nodes to perform local computing tasks. However, this simple MapReduce architecture may still have certain limitations: (1) It cannot adaptively respond to the instant updates of real-time power grid data streams; (2) It cannot perceive load skew in time and space; (3) The query performance is not suitable for real-time distributed stream processing environments (limited by the cost of frequent updates).
[0006] "Distributed trajectory similarity search" (D. Xie, F. Li, and J. Phillips. Distributed trajectory similarity search. PVLDB, 10: 1478–1489, 2017.) discloses a distributed spatial data computing framework DFT, which is oriented to distributed large-scale computing of trajectory data, but the screening process is implemented using distributed spatial indexes. This requires that each node of the global index and the local index must use a roaring bitmap to record the set of trajectory IDs passing through the node, which results in a very large distributed spatial index in DFT, as DFT needs to know all the trajectory IDs in the trajectory set.
[0007] The patent named "A spatial indexing method based on social perception in a distributed environment" and publication number CN109190052A discloses a spatial indexing method based on R-tree, which supports queries in a distributed environment. However, the index is established for static data, and the index update and storage overhead are large, and it is impossible to realize real-time query on power grid data streams, that is, the index update overhead is much greater than the query process overhead. Summary of the invention
[0008] Purpose of the invention: The first purpose of the present invention is to provide a method for real-time updating and rapid query of the index structure of power grid data trajectory stream. The second purpose of the present invention is to provide a system for real-time and rapid query of power grid data.
[0009] Technical solution: The real-time query method for massive power grid data of the present invention comprises the following steps:
[0010] (1) Obtain the continuous query requirements and query graph Q issued by the user;
[0011] (2) Load the global index and local index. The global index is used to locate the candidate partitions similar to the query graph Q, and the local index is used to search for candidate graphs similar to the query graph Q on the node.
[0012] (3) Filter the global index and local index to obtain candidate graphs, calculate the similarity of the candidate graphs, randomly read the graphs in the overlapping part of the current time window and the historical time window, calculate the distance between the newly entered graph and the currently valid graph, and use it as the latest state increment update. At the same time, delete the graphs and their states that are not in the time window;
[0013] (4) Aggregate the graph similarity calculation results of each local node to obtain a result graph that meets the query requirements.
[0014] Furthermore, in step (3), the Hausdorff distance is used to calculate the graph similarity, and the graph set T = <t1,t2,…,t m > and query graph Q = <q1,q2,…,q n >The distance between them is:
[0015] HAU(T,Q)=max{max dis(t i ,Q),max dis(q i ,T)}
[0016] The distance between the two points required for the HAU(T,Q) calculation process is retained, so that the overall state information table maintenance cost will be minimized and reading will be faster.
[0017] Furthermore, the query requirement in step (1) is a k-nearest neighbor query. In step (4), after calculating the similarity of the candidate graphs, a mini-heap is constructed and the top k results are returned. The global node receives the results returned by each local node and performs merge sort to obtain k result graphs that meet the query requirements.
[0018] Furthermore, the global index in step (2) partitions the power grid data using a top-down structure. The partitioning method is: mapping the power grid data stream data into a two-dimensional grid space, sorting by the horizontal coordinates of the center points of the rectangles, cutting the two-dimensional grid space with vertical slices, and then sorting by the vertical coordinates of the center points of the rectangles. Several rectangles form a leaf node. This makes it easier to distribute similar graphs in a partition, and the number of graphs in each partition is roughly the same.
[0019] Furthermore, the local index is established in each leaf node of the global index, and the local index is an R-tree index. However, for its sorting algorithm, the present invention uses the first pivot point determination of the quick sorting algorithm as the termination condition.
[0020] Furthermore, the global index and local index in step (2) are updated based on a sliding time window during the inflow of real-time power grid data, and the index is rebuilt for the power grid data that cannot be reused.
[0021] The real-time query system for massive power grid data described in the present invention includes a coordinator node and several job nodes. The sampler in the coordinator node is used to map the current power grid data trajectory distribution into a two-dimensional grid space and monitor the changes in the power grid data distribution in real time. The sampler sends the mapped data to the distributor to establish a global index, partition the data, and distribute the data to the corresponding job nodes based on the partitions; each of the job nodes includes a local index and a job executor, and the job executor is used to perform data filtering and refinement processes during the query process.
[0022] Furthermore, the refinement process in the job executor includes graph similarity calculation, and the method of graph similarity calculation is: randomly read the calculated state of the overlapping part of the current time window and the historical time window, calculate the distance between the newly entered graph and the currently valid graph, and update it as the latest state increment, while deleting the graph and its state that are not in the time window.
[0023] Beneficial effects: Compared with the prior art, the significant advantages of the present invention are: (1) a global index with a top-down structure is established, and a local index is established at the leaf nodes of the global index, which reduces the volume of the spatial index, improves the processing speed of the trajectory flow, and reduces the resource overhead level of the monitoring equipment; (2) the query of the present invention is based on a secondary index that is updated in real time, which realizes real-time query of data and greatly improves the query speed; (3) when querying, a reused graph similarity calculation is used to avoid the repeated calculation of the state of the overlapping part of the window when the power grid data flow changes in real time, thereby improving the query speed and realizing fast and real-time query of massive power grid data, so that the data subject can bear dynamic workload capacity, making the centralized control and monitoring of the power grid more mature and sensitive. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 is a flow chart of the query method of the present invention;
[0025] Figure 2 A flow chart of the global index construction method of the present invention;
[0026] Figure 3 A flow chart of the local index construction method of the present invention;
[0027] Figure 4 It is a schematic diagram of the secondary index structure of the present invention;
[0028] Figure 5 This is a schematic diagram of the global index partitioning result of the present invention;
[0029] Figure 6 This is a schematic diagram of updating a sliding time window according to an embodiment of the present invention;
[0030] Figure 7 This is the state reuse calculation process of an embodiment of the present invention;
[0031] Figure 8 This is a query system architecture diagram of the present invention;
[0032] Fig. 9 This is a time delay comparison diagram of three frameworks on the Beijing data set when k is the independent variable in an embodiment of the present invention;
[0033] Fig.10This is a time delay comparison diagram of three frameworks on the Chengdu data set when k is the independent variable in an embodiment of the present invention;
[0034] Fig.11 This is a throughput comparison chart of the three frameworks on the Beijing dataset when k is used as an independent variable in an embodiment of the present invention;
[0035] Fig.12 This is a throughput comparison chart of the three frameworks on the Chengdu dataset when k is used as an independent variable in an embodiment of the present invention;
[0036] Fig.13 The figure shows the throughput comparison of the three frameworks on the Beijing dataset when L is used as the independent variable in an embodiment of the present invention;
[0037] Fig.14 The figure shows a throughput comparison chart of the three frameworks on the Beijing dataset when L is used as the independent variable in an embodiment of the present invention. DETAILED DESCRIPTION
[0038] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.
[0039] like Figure 1 As shown, the real-time query method for power grid data includes the following steps:
[0040] (1) The query demand issued by the user, including the operation based on the visual graphical interface (such as inputting the target location) or SQL statement, is parsed and converted into an executable physical plan. The query in this embodiment is to query the grid points with low current similarity, that is, the grid points where abnormalities may occur, and perform continuous query. The input of the continuous query is a query graph Q, and the graph set T = {T1, T2, ..., T m}, a distance measure D and an integer k.
[0041] (2) Loading an index, which is a master-slave secondary index, including a global index and a local index. The global index is used to quickly locate candidate partitions similar to the query graph Q, and the local index is used to search for candidate graphs similar to the query graph Q on the node.
[0042] like Figure 2 and Figure 3 The following are the construction processes of global index and local index respectively. Figure 4 The above is a schematic diagram of the secondary index structure. The construction and update of the index specifically include the following contents:
[0043] (21) Multi-threading is used to pull multi-source power grid data and integrate power grid data streams in real time. The present invention targets the generic representation of power grid data - graph structure data.
[0044] (22) All graph data is divided into N non-overlapping buckets, and the graph in each bucket is further divided into N sub-buckets. Finally, each sub-bucket is regarded as a partition, so there are N*N partitions. The STR (Sort Tile Recursive) partitioning method is used to partition the graph, and the bottom-up structure of the traditional STR is changed to top-down to reduce overhead and improve data flow processing performance. The basic idea is: the current trajectory distribution is mapped to a two-dimensional grid density. There are N rectangles in the two-dimensional grid. M is the maximum capacity of the R-tree node. The rectangles are sorted by the x-coordinate of the center point of the rectangle. The data space is cut into vertical slices, and then in each vertical slice, the rectangles are sorted by the y coordinates of the rectangle center points. Every M rectangles form a leaf node. In this way, similar graphs are more likely to be distributed in a partition, and the number of graphs in each partition is roughly the same. The partitions are as follows: Figure 5 shown.
[0045] (23) For the data in each partition, we can obtain two MBRs (minimum bounding rectangles) - the MBR formed by the first point in each time window and the MBR formed by the last point. The MBRs are stored in the leaf nodes.
[0046] The function R(s,t) represents an MBR in space, where s = {s1,s2,…,s d} is the minimum point in MBR, t={t1,t2,……,t d} is the maximum point in the MBR. In space, s is the lower left corner of the MBR, and t is the upper right corner of the MBR. For any two given MBRs in space: R1(s,t) and R2(p,q), MinMinDist(R1(s,t),R2(p,q)) is defined as:
[0047]
[0048] For any two given MBRs in space: R1(s,t) and R2(p,q), MaxMaxDist(R1(s,t),R2(p,q)) is defined as:
[0049]
[0050] MinDist(q,MBR) is the minimum distance from point q to MBR (four corners and four edges). For a query graph Q of length n, find the MBR in the R-tree for each point i in Q that satisfies the condition MinDist(q,MBR)≤τ, and obtain the corresponding partition set P. Using this method, we can obtain two partition sets P1 and P2 in the above two MBRs. For each partition p∈P1∩P2, if MinDist(q1,MBR1)+MinDist(q2,MBR2)≤τ, then p is a relevant partition.
[0051] (24) A local index based on an R-tree is established in each leaf node of the global index. When querying based on the local index, the priority of the first point and the last point is increased to further improve the query speed, which is suitable for massive, real-time updated power grid data.
[0052] (25) As real-time data continues to flow in, the corresponding index structure will undergo certain changes. The changes in the index are mainly to re-index the non-reused parts of the index according to steps (22) to (24). Because the original index has always been top-down, it has performance advantages.
[0053] The present invention updates the index based on a sliding window, where a sliding window means that a certain amount of data overlaps between two time windows. Figure 6 The following is an example of a sliding window update. The current time window is from 9:00 to 9:30, and the window size L is 30 minutes. The data at 9:02 will be calculated within this time window. If the time window slides for 5 minutes, the data at 9:02 will no longer be included in the calculation range. However, the data at 9:16 will still be included in the calculation.
[0054] (26) Store indexes for loading in a query. For persistent storage of historical data.
[0055] (3) Complete the filtering and refinement process based on the index. First, determine the data partition according to the global index. Then, perform further pruning operations according to the local indexes in the partitions specified by the global index to obtain the candidate graph. The refinement operation is to perform similarity calculation and other operations on the candidate graph.
[0056] The similarity operation is completed by parallel calculation of the physical nodes of the above data partitions. For different query requirements, the execution logic of the physical nodes is roughly similar. The similarity between graphs is the distance between graphs in the current time window, which will be stored as a state in Flink's state backend (generally implemented with RocksDB). During the continuous query processing execution, the overlapping state of the current time window and the historical time window will be reused in the current calculation. For example, the calculation result of 8.30-8.50 and the current calculation task of 8.40-9.00 that needs to be calculated have a 10-minute overlap. The present invention does not repeat the calculation of the overlapping 10-minute spatiotemporal data, but directly reads the information randomly from RocksDB.
[0057] Specifically, the state reuse calculation process is divided into three parts:
[0058] 1) The distance of the graph is obtained by compound calculation of the distances between multiple points. In this embodiment, the distance of the space-time graph is calculated by using the Hausdorff distance as an example. Figure 7 As shown, this metric takes the maximum-minimum distance to get the final value, Figure T = <t1,t2,…,t m > and Figure Q = <q1,q2,…,q n >The distance between them is defined as: HAU(T,Q)=max{max dis(t i ,Q),max dis(q i ,T)}. During the calculation process, not every distance between two points can be used to calculate the distance between graphs. The distance between two points that needs to be used is called the contributing distance. The state storage requires space. In order to reduce the space overhead and the scanning cost, the present invention only saves the distance that contributes to the final similarity result. In this way, the overall state information table maintenance cost will be minimized and the reading will be faster.
[0059] 2) Calculate the distance between the newly entered space-time graph and the currently valid space-time graph as the latest state increment update. That is, the current state table is continuously expanded incrementally during this process. In this process, just do not perform useless distance calculations with expired space-time data marked as invalid.
[0060] 3) The space-time graph that is not within the time window, including its associated states, is removed to save space overhead.
[0061] (4) Query requirements are generally divided into general queries and k-nearest neighbor queries, and the processes are as follows:
[0062] (41) General query process
[0063] This embodiment takes a special case of a space-time graph, a space-time sequence, as an example for introduction:
[0064] (411) Given the query spatiotemporal sequence τ q Generally, a unique identifier is input as a means, or a graphical interface can be used to indicate a time-space sequence. The time-space diagram can be deduced by analogy, and what finally reaches the background is the unique identifier ID.
[0065] (412) Read the constructed global index G and local index I, first according to τ q Location information, find the partition to which the sub-area belongs from G. For example, τ q If it belongs to Beijing, G first locates the partition range of Beijing, and then reads the local index I from the partition pointed to by the leaf node of G.
[0066] (413) Eliminate non-candidate data according to I to obtain a candidate data set.
[0067] (414) In the candidate data set, a refinement process is performed, including graph similarity calculation based on state reuse calculation.
[0068] (415) The spatiotemporal sequence ID that finally meets the query requirements is returned and passed to the front-end system, which can be rendered into any form (graphics, target users, etc.).
[0069] (42) k nearest neighbor process
[0070] The general query does not pass in any additional parameters, while the k-nearest neighbor query requires returning the k nearest graphs for a given query graph.
[0071] (421) The global node stores the global index G, the local node stores the local index I and performs pruning operations.
[0072] (422) Perform refinement operations in the local node, build a small top heap, and return the top k results.
[0073] (423) The global node receives the lists returned by N local nodes, performs merge sort, and obtains the final k results.
[0074] (5) Collect the results in each local node, aggregate them and return the result graph to the user. In this embodiment, if you need to query which grid points have abnormal phenomena, the unique identifiers of these grid points are returned, and the front-end page finally displays the specific grid point information.
[0075] like Figure 8As shown, the real-time query system for power grid data described in the present invention includes a coordinator node and several job nodes. When the power grid data enters the system, the sampler will first map the current trajectory distribution into a two-dimensional grid space, and monitor the changes in the distribution of power grid data in real time. At the coordinator node, the distributor partitions the trajectory stream data according to the structure of the global index, and distributes the data to the corresponding job nodes based on the partitions. At each job node, a local index and a job executor are provided, and the local index is responsible for efficiently executing candidate power grid data queries (using the Filter-and-Refine idea). The job executor consists of a clipping mechanism and a state reuse mechanism to fully improve resource utilization and reduce computational complexity.
[0076] The query method of the present invention is verified by experiments below.
[0077] Experimental environment: All methods in the experiment are run on Apache Flink 1.9.1. All comparison methods have the same configuration by default and are implemented in Java. The experiment is deployed on a cluster of HP blade servers with 128 processing nodes (Intel(R) Xeon(R) CPU E7-8860 v3@2.20GHz), 2TB@1600MHz memory, and all server nodes run the CentOS 7.4 operating system.
[0078] Experimental data source: The experiment uses two real datasets, Chengdu and Beijing, of which the sampling points in Beijing are more unevenly distributed. Chengdu is provided by Didi Chuxing's Gaia data plan, and the trajectories are all from a certain area in Chengdu, China. The dataset size is about 900GB; Beijing is part of the taxi trajectory data in Beijing, China in 2008. In general, the data structure of these two datasets is the same, both consisting of vehicle ID, timestamp, and longitude and latitude. In the experiment, these static datasets will be published through Apache Kafka to simulate the generation of real-time data streams.
[0079] Experimental setup: The experiment selects the k-nearest neighbor query as the scenario, which is more likely to produce a real-time processing performance bottleneck. This query, as the name implies, returns the first k objects that meet the user's query requirements (for example, users' queries for Weibo are often Top-10 hot searches). Since there is a need for sorting in a large amount of data in this scenario, it is likely that the ideal performance cannot be achieved in the real-time scenario. Therefore, compared with general queries, such as range queries and point queries, this query is easier to evaluate the performance difference of the present invention. In addition, we selected DFT (Distributed Trajectory Similarity Search, VLDB 17) technology as the benchmark method for comparison of this method. DFT can obtain competitive results in offline scenarios, but it is obviously not suitable in real-time scenarios, so the present invention adds a state reuse mechanism to DFT, and the method after introducing this mechanism is called CDFT.
[0080] like Figures 9 to 12 The performance comparison of the three frameworks when k is used as an independent variable is shown. As k increases, the computational complexity increases, causing the throughput of the left and right frameworks to begin to decrease and the latency to increase. However, in both cases, the performance ranking of the three frameworks is: the present invention > CDFT > DFT.
[0081] Fig.13 and Fig.14 The performance comparison of the three frameworks is shown when the window size L of the sliding window is used as an independent variable. As L increases, the number of sample points of the graph in the window will also increase. As L increases, the performance of CDFT and the present invention decreases slower than DFT. This shows that the method of the present invention is more suitable for incremental similarity calculation between long power grid data.
Claims
1. A real-time query method for massive power grid data, characterized in that: The steps include: (1) Obtain the continuous query requirements and query graph Q issued by the user; (2) Loading a global index and a local index, wherein the global index is used to locate candidate partitions similar to the query graph Q, and the local index is used to search for candidate graphs similar to the query graph Q on nodes; the global index partitions the power grid data using a top-down structure, and the partitioning method is as follows: mapping the power grid data stream data into a two-dimensional grid space, sorting them using the horizontal coordinates of the center points of the rectangles, then cutting the two-dimensional grid space using vertical slices, and then sorting them using the vertical coordinates of the center points of the rectangles, and a number of rectangles form a leaf node; the local index is established in each leaf node of the global index, and the local index is an R-tree index; (3) Filter the global index and local index to obtain candidate graphs, calculate the similarity of the candidate graphs, randomly read the graphs in the overlapping part of the current time window and the historical time window, calculate the distance between the newly entered graph and the currently valid graph, and use it as the latest state increment update. At the same time, delete the graphs and their states that are not in the time window; Use Hausdorff distance to calculate graph similarity, atlas and query graph The distance between them is: ,reserve The distance between the two points that need to be used in the calculation process; (4) Aggregate the graph similarity calculation results of each local node to obtain a result set that meets the query requirements.
2. The real-time query method for massive power grid data according to claim 1 is characterized in that: The query requirement in step (1) is k Nearest neighbor query, in step (4), after calculating the similarity of the candidate graphs, a small top heap is constructed and the previous k The global node receives the results returned by each local node and performs merge sorting to obtain k A result graph that meets the query requirements.
3. The real-time query method for massive power grid data according to claim 1 is characterized in that: The global index and local index in step (2) are updated based on a sliding time window during the inflow of real-time power grid data, and the index is rebuilt for the power grid data that cannot be reused. The sliding time window contains overlapping data within two time windows.
4. A real-time query system for massive power grid data based on the method according to any one of claims 1 to 3, characterized in that: It includes a coordinator node and several job nodes. The sampler in the coordinator node is used to map the current power grid data trajectory distribution into a two-dimensional grid space and monitor the changes in the power grid data distribution in real time. The sampler sends the mapped data to the distributor to establish a global index, partition the data, and distribute the data to the corresponding job nodes based on the partitions; each of the job nodes contains a local index and a job executor, and the job executor is used to perform data filtering and refinement processes during the query process.
5. The real-time query system for massive power grid data according to claim 4 is characterized in that: The refinement process in the job executor includes graph similarity calculation, and the method of graph similarity calculation is: randomly read the calculated state of the overlapping part of the current time window and the historical time window from RocksDB, calculate the distance between the newly entered graph and the currently valid graph, and update it as the latest state increment, while deleting the graph and its state that are not in the time window.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the query method according to any one of claims 1 to 3 is implemented.
Citation Information
Patent Citations
A spatial indexing method based on social perception in a distributed environment
CN109190052A
Massive structured data storage and query methods and systems supporting high-speed loading
CN102521405A
Method for storing and searching mass sensor data
CN102651020A