A time-series parameter data storage system for power equipment based on a statistical tree structure

By adopting a clustered index statistical tree based on the statistical tree structure in the power equipment timing parameter data storage system, pre-aggregation of data and building leaf node link lists are realized, the problem of time delay in the existing technology is solved, data query efficiency is improved, and real-time monitoring needs of power equipment are met.

CN119669299BActive Publication Date: 2025-05-27JIANGSU NJUSNGCHENG HI TECH INDAL +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510194425.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-05-27
Estimated Expiration
2045-02-21

AI Technical Summary

Technical Problem

The prior art has a large delay when querying the timing parameter data of power equipment, especially when the data scale is large, resulting in the inability to detect equipment abnormalities in time, affecting the real-time monitoring and maintenance of power equipment.

Method used

The power equipment timing parameter data storage system based on the statistical tree structure is adopted. The data is pre-aggregated when data is written through the clustered index statistical tree, and a leaf node link list is constructed at the leaf node layer to improve the query efficiency of data points and their aggregation indicators.

Benefits of technology

On the premise of satisfying the real-time writing of data generated by the power equipment network, the query efficiency of timing parameter data points and their aggregation indicators is significantly improved, and the real-time performance of power equipment monitoring and maintenance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119669299B_ABST
    Figure CN119669299B_ABST
Patent Text Reader

Abstract

The present invention discloses a time-series parameter data storage system for power equipment based on a statistical tree structure, belonging to the technical field of database storage. The storage system is used to solve the problem of monitoring data collection delay caused by low query efficiency of power equipment data in a large-scale power equipment network scenario. The present invention includes a storage structure based on a clustered index statistical tree and a computing device including a storage management module. The clustered index statistical tree preprocesses the aggregated data of time-series parameter data points and adopts a hanging leaf node and non-uniform segmentation technology, thus reducing the data point writing delay. The storage management module utilizes the leaf node linked list of the clustered index statistical tree and the clustered index statistical technology to improve the data point query and aggregated query efficiency. This storage technology improves the data point query and aggregated query efficiency and reduces the delay of collecting data during equipment monitoring on the premise of meeting the data writing rate requirements of the power equipment network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a time-series parameter data storage system for power equipment based on a statistical tree structure, belonging to the technical field of database storage. Background Art

[0002] The time-series parameter data of power equipment is a set of data points formed by monitoring and recording the operating status and key performance parameters of the equipment during operation and arranging them in chronological order. These time-series parameter data play a key role in the monitoring and maintenance of power equipment, so the storage of data is an important topic in the field. At present, with the popular application of power equipment across the country, the number of power equipment continues to increase, and the scale of the time-series parameter data generated by it also expands accordingly.

[0003] Generally speaking, the time-series parameter data of power equipment is stored in a time-series database. The time-series database uses an LSM tree file organization structure, which can receive and store the time-series parameter data points generated by power equipment at high speed. However, when querying the time-series parameter data points within a certain period of time, this structure needs to access multiple files in the LSM file tree and integrate them, which introduces a large delay for querying the data points within a certain period of time and their aggregation metrics, and the larger the data scale, the greater the delay. In the scenario of monitoring the health status of power equipment, this delay is likely to cause the failure to detect equipment abnormalities in time. Therefore, on the premise of ensuring that the data writing rate meets the data rate generated by the power equipment network, how to improve the data query efficiency has become a key topic in the field. Summary of the Invention

[0004] Object of the Invention: Aiming at the problems and deficiencies in the prior art, the present invention provides a time-series parameter data storage structure and a storage management module for power equipment based on a statistical tree to improve the query efficiency of time-series parameter data points and their aggregation metrics and meet the real-time query requirements in the power equipment detection scenario.

[0005] Technical Solution: A time-series parameter data storage system for power equipment based on a statistical tree structure includes:

[0006] A storage structure for storing the time-series parameter data generated by power equipment; the time-series parameter data is stored in a clustered index statistical tree, and one statistical tree stores a single time series of a single device;

[0007] A computing device including a storage management module, and the storage management module is operable to:

[0008] Insert data points, and insert the newly generated time-series parameter data of the device into a statistical tree to which the device belongs;

[0009] Data point query, to obtain all time series parameter data within a period of time in a statistical tree to which the device belongs;

[0010] Aggregate data query, to obtain the aggregate metrics of time series parameter data within a period of time in a statistical tree to which the device belongs, including the count, latest value, oldest value of the parameter data, and the sum, maximum value, minimum value, and average value of the numerical type parameter data. Power equipment time series parameter data storage structure and storage management module.

[0011] The storage structure includes a clustered index statistical tree composed of a root node, non-root non-leaf nodes, and leaf nodes. The root node serves as the only interface to access the clustered index statistical tree, and at the same time stores the starting point of the time domain maintenance interval of the hanging leaf nodes, as well as the metadata of the child nodes. The metadata includes the starting point of the time domain maintenance interval and the aggregate data of the data points within the time domain maintenance interval. The non-root non-leaf nodes store the metadata of the child nodes. The leaf nodes store the time series parameter data points themselves and at the same time store pointers to the successor leaf nodes.

[0012] The clustered index statistical tree includes:

[0013] Statistical clustered index, the index is a clustered statistical tree. The leaf nodes of the statistical tree store the time series parameter data points, and the non-leaf nodes store the metadata of their child nodes, where each child node stores a copy of the metadata;

[0014] Hanging leaf nodes are the leaf nodes that store the latest time series parameter data. The parent node of the hanging leaf nodes is fixed as the root node;

[0015] Leaf node linked list, the leaf nodes of the clustered statistical tree are arranged in ascending order of timestamp, and form a leaf node linked list through the pointers from the leaf nodes in the adjacent time intervals from the smaller time interval to the larger time interval;

[0016] Uneven division, when the amount of data stored in a node exceeds a pre-set threshold, the node will be divided into two nodes, and the free space generated during the division process is allocated to the two nodes formed by the division according to a set ratio.

[0017] Each node of the clustered index statistical tree has its time domain maintenance interval. The meaning of the time domain maintenance interval is: among the time series parameter data produced by the power equipment corresponding to the statistical tree, the data points whose timestamps fall within the time domain maintenance interval are stored and only stored in the node to which the time domain maintenance interval belongs and its direct or indirect child nodes;

[0018] The time series parameter data stored in the leaf nodes includes the time series parameter data points produced by the power equipment associated with the tree-like index within its time domain maintenance interval. The data points are composed of timestamps and the measured values of the measuring points at the corresponding times of the timestamps; the time series parameter data points stored in the leaf nodes are arranged in ascending order of timestamp;

[0019] The non-leaf nodes store the metadata of the child nodes. Each child node stores a copy of its own metadata in its direct parent node. The metadata stored in the non-leaf nodes is sorted in ascending order according to the start point of the time domain maintenance interval; the metadata of the nodes includes:

[0020] If the node is a leaf node, its metadata includes the address of the node, the start point of the time domain maintenance interval, and the time stamp of the latest data point, the data value of the latest data point, the time stamp of the oldest data point, the data value of the oldest data point, and the total count of data points in the data points stored in the leaf node; if the data point type stored in the leaf node is numeric, the metadata also includes the sum of the data point values, the maximum value of the data point values, and the minimum value of the data point values;

[0021] If the node is a non-leaf node, its metadata includes the address of the node and an aggregation of the metadata of all its child nodes; the aggregation method of the child node metadata is as follows: for the time stamp of the latest data point and the maximum value of the data point values, the aggregated value takes the maximum value of all the metadata; for the start point of the time domain maintenance interval, the time stamp of the oldest data point, and the minimum value of the data point values, the aggregated value takes the minimum value of all the metadata; for the total count of data points and the sum of the data point values, the aggregated value takes the sum of all the metadata; for the data value of the latest data point, the aggregated value takes the data value of the latest data point of the metadata with the largest time stamp of the latest data point; for the data value of the oldest data point, the aggregated value takes the data value of the oldest data point of the metadata with the smallest time stamp of the oldest data point.

[0022] In the data writing process of the power equipment time series parameter data storage system based on the statistical tree according to the present invention, data is pre-aggregated at different granularities, and a leaf node linked list is constructed at the leaf node layer, so as to improve the query efficiency of data points and their aggregated metrics on the premise of satisfying the real-time writing of data generated by the power equipment network. Description of the Drawings

[0023] Figure 1 is a schematic diagram of the system structure of an embodiment of the present invention;

[0024] Figure 2 is a schematic diagram of the initial storage structure of the system structure of an embodiment of the present invention;

[0025] Figure 3 is a flowchart of checking the number of non-leaf node metadata entries and splitting of the system structure of an embodiment of the present invention;

[0026] Figure 4 is a flowchart of checking the occupied space of leaf node data points and splitting of the system structure of an embodiment of the present invention;

[0027] Figure 5It is the flowchart of inserting a downgraded leaf node into the cluster index statistical tree of the system structure according to an embodiment of the present invention;

[0028] Figure 6 It is the flowchart of the data point insertion module of the storage management module according to an embodiment of the present invention;

[0029] Figure 7 It is the flowchart of the data point query module of the storage management module according to an embodiment of the present invention;

[0030] Figure 8 It is the flowchart of the aggregated data query module of the storage management module according to an embodiment of the present invention;

[0031] Figure 9 It is the flowchart of aggregating and querying the statistical tree subtree excluding hanging nodes of the storage management module according to an embodiment of the present invention;

[0032] Figure 10 It is the schematic diagram of the storage technology according to an embodiment of the present invention.

[0033] Reference numerals: 101: root node; 102: non-root and non-leaf node; 103: leaf node; 104: hanging leaf node; 105: list of child node metadata; 106: metadata; 107: pointer; 108: start point of the time domain maintenance interval of the hanging leaf node; 109: pointer to the hanging leaf node; 110: list of time series parameter data points; 111: successor pointer pointing to the successor leaf node; A01: computer system; A02: processor; A03: storage management module; A04: data point writing module; A05: data point query module; A06: aggregated data query module; A07: storage medium; A08: cluster index statistical tree. Detailed implementation manners

[0034] The present invention will be further clarified below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. After reading the present invention, various equivalent forms of modification of the present invention by those skilled in the art all fall within the scope defined by the appended claims of this application.

[0035] Such as Figure 1As shown in the figure, a clustered index statistical tree containing a certain amount of time series parameter data points includes a root node 101, non-root non-leaf nodes 102, leaf nodes 103, and hanging leaf nodes 104. In the figure, 105 is a list of child node metadata, which stores metadata 106. The metadata contains a pointer 107 to the child node; the hanging leaf node is abbreviated as "hanging node". The start point 108 of the time domain maintenance interval of the hanging leaf node and the pointer 109 to the hanging leaf node are stored in the root node; in the leaf node, there is a list 110 of time series parameter data points. In addition to the hanging leaf node, the leaf node additionally stores a successor pointer 111 to the successor leaf node. Among various types of nodes, in the root node and non-root non-leaf nodes, the metadata of the child nodes except the hanging leaf node are stored in ascending order of the start point of the time domain maintenance interval; the root node additionally stores the start point of the time domain maintenance interval of the hanging leaf node and the hanging leaf node pointer; the leaf node and the hanging leaf node store the time series parameter data points in ascending order of the timestamp. Each node has its time domain maintenance interval, which is a left-closed right-open interval, indicating the interval where the timestamps of the time series parameter data points maintained by this node are located; the time series parameter data points whose timestamps fall within the time domain maintenance interval of this node are stored and only stored in this node or its child nodes. The time domain maintenance intervals of the direct child nodes (including the hanging leaf node) of the parent node do not intersect with each other, and their union is equal to the time domain maintenance interval of the parent node. The metadata includes a pointer to the child node, the start point of the time domain maintenance interval of the child node, and the aggregation metrics of the data points within the time domain maintenance interval of the child node. The aggregation metrics include the latest timestamp, the oldest timestamp, the latest value, the oldest value, and the number of data points; for time series parameter data points with numerical values, the aggregation metrics additionally include the maximum value, the minimum value, and the sum of the data values.

[0036] As Figure 1 shown, this section will describe the method for inferring the time domain maintenance interval of each node. The time domain maintenance interval of the root node is fixed as [0, +∞). The time domain maintenance interval of the hanging leaf node is inferred through the start point T h of the time domain maintenance interval of the hanging leaf node stored in the root node, and its time domain maintenance interval is [T h , +∞). The time domain maintenance intervals of the remaining nodes are inferred through the metadata stored in their parent nodes: Let the time domain maintenance interval of the parent node p be [L p , R p ), and it has n direct child nodes, denoted as p 1 , p 2 ,..., p n . The start points of the time domain maintenance intervals in their metadata are , where ; then the time domain maintenance interval of the node p 1 is , and the time domain maintenance interval of the node p 2 is , and so on; the last child node p n has a time-domain maintenance interval of . Specifically, when the parent node p is the root node, due to the existence of hanging leaf nodes, the end point (non-inclusive) of the time-domain maintenance interval of the last child node is T h instead of +∞.

[0037] As Figure 1 shown, this section will describe the archived nodes and the newly born nodes. If the end point of the time-domain maintenance interval of a non-root and non-leaf node is equal to the start point T of the time-domain maintenance interval of the hanging leaf node h , or this node is a hanging leaf node, it indicates that this node is the node with the largest start point of the time-domain maintenance interval at its level. Such nodes are denoted as newly born nodes; newly born nodes are more likely to receive the time-series parameter data points newly generated by power equipment, so as much space as possible needs to be allocated. Other leaf nodes and non-root and non-leaf nodes except newly born nodes are denoted as archived nodes. These nodes only receive the time-series parameter data points that arrive out of order, and the time-domain distribution of the out-of-order data points is relatively uniform, so the space requirements of each archived node are basically the same.

[0038] As Figure 2 shown, an empty clustered index statistical tree contains two nodes, the root node 101 and the hanging leaf node 104. The time-domain maintenance intervals of the root node and the hanging leaf node are both [0, +∞). Since there are no other child nodes of the root node except the hanging leaf node at this time, no metadata entries are stored in its metadata list, and the value T h stored in the start point entry of its hanging leaf node is 0. No data points are stored in the hanging leaf node.

[0039] As Figure 3 shown, it describes the process of checking the number of metadata entries of non-leaf nodes and splitting. As shown in the parameter input of step S301, this process accepts a parameter X as the non-leaf node to be checked. In the figure, |X| represents the number of entries in the child node metadata list of the non-leaf node X, L m represents the threshold of the number of metadata entries of the non-leaf node, T h represents the start point of the time-domain maintenance interval of the hanging leaf node, R old represents the splitting weight of the archived node, R new represents the splitting weight of the newly born node. There is a threshold L for the number of metadata entries in the metadata lists of the root node and non-root and non-leaf nodes m, its default value is 16. In the splitting condition determination of step S302, when the number of metadata entries in a node exceeds the upper limit, the non-leaf node splitting process will be triggered. First, in the root node determination of step S303, if the node to be split is the root node, then step S304 root node migration is executed. A non-root and non-leaf node of the root node is newly created, the child node metadata in the root node is moved to the newly created non-root and non-leaf node, and this non-root and non-leaf node is used as the only child node of the root node, and its metadata is stored in the child node metadata list of the root node. Subsequently, step S305 target node change is executed, and a splitting operation is performed on this non-root and non-leaf node. The splitting operation of the non-root and non-leaf node starts from variable initialization in step S306 and ends with establishing node connections in S310. Specifically, a time point T is determined, a non-root and non-leaf node is newly created, and the child node metadata in the node to be split whose start time of the time domain maintenance interval is greater than or equal to T is moved to the new non-root and non-leaf node, and the metadata of the new non-root and non-leaf node is added to the child node metadata list of the parent node of the node to be split. The time point T is determined in such a way that after splitting according to this time point, the absolute value of the difference between the ratio of the remaining available metadata entry numbers in the split node and the new node and the splitting weight R of this node is the smallest. The splitting weight R is determined according to the following rules: According to the new node determination in step S307, if the node to be split is a new node, after splitting, an archived node and a new node are generated, where the newly created non-root and non-leaf node is the new node. The new node requires more space compared to the archived node. Therefore, according to the new node splitting in step S308, the splitting weight of the new node is R new , its default value is 1:9; if the node to be split is an archived node, after splitting, two archived nodes are generated. Among them, for the newly created non-root and non-leaf node, the start time of its time domain maintenance interval is relatively large, and the probability of receiving out-of-order data points is relatively high. Therefore, according to the archived node splitting in step S309, the splitting weight of the archived node is R old , its default value is 4:6. After the node splitting step is completed, in step 311 node check, node connections are established according to process 300, and the number of metadata entries in the parent node is checked. If necessary, splitting continues. Using the above non-uniform splitting process, this storage technology can reduce the waste of pre-allocated space, reduce the number of nodes generated by the clustered index statistical tree in the same data scenario, and improve the writing efficiency of data points.

[0040] As Figure 4 shown, the process of checking the occupied space of leaf node data points and splitting is described. As shown in parameter input of step S401, this process receives a parameter X as the leaf node to be checked. In the figure, |X| represents the occupied space of the data points of leaf node X, L d represents the threshold of the occupied space of leaf node data points, T h represents the start time of the time domain maintenance interval of the hanging leaf node, R oldRepresents the splitting weight of the archived node, R new Represents the splitting weight of the newly generated node. See the sub - process "Insert Degraded Leaf Node" Figure 5 There is a data point occupancy space threshold L for the data point list of the leaf node and the hanging leaf node d Its default value is 4KB. In the splitting condition determination of step S402, when the space occupied by the data points in the leaf node reaches the threshold, the leaf node splitting process will be triggered. In the hanging node determination of step S403, the splitting operation is different according to whether the leaf node is a hanging leaf node. The splitting operation of the hanging leaf node is as follows: in step S404 to determine the splitting time point, a time point T is determined; in step S405 to split the node, a new hanging leaf node is created, and the data points in the node to be split with a timestamp greater than or equal to T are moved to the new hanging leaf node; where the time point T satisfies that after splitting according to this time point, the absolute value of the difference between the ratio of the remaining available space in the split node and the new hanging leaf node and the splitting weight R new is the smallest. After splitting, the original hanging leaf node degrades to a leaf node, and this node takes the newly created hanging leaf node as its successor node; at the same time, update the start point of the time domain maintenance interval of the hanging leaf node in the root node to the start point of the time domain maintenance interval of the newly created leaf node; then, in step S406 to insert the degraded leaf node, according to the Figure 5 insert - degraded - leaf - node process, insert the degraded leaf node into the clustered index statistical tree. The splitting operation of the non - hanging leaf node is as follows: in step S407 to determine the splitting time point, a time point T is determined; in steps S408 variable initialization and S409 to split the node, a new leaf node is created, and the data points in the node to be split with a timestamp greater than or equal to T are moved to the new leaf node; where the time point T satisfies that after splitting according to this time point, the absolute value of the difference between the ratio of the remaining available space in the split node and the new leaf node and the splitting weight R old is the smallest; after splitting, add the metadata of the newly created leaf node to the sub - node metadata list of the parent node of the split node; in step S410 node check, according to the Figure 3 process of checking the number of non - leaf node metadata entries and splitting, check the number of metadata entries of the parent node of the split node. Using the above non - equal splitting process, this storage technology can reserve more space for the newly generated time - series parameter data points; at the same time, by inserting the degraded leaf node as a whole into the clustered index statistical tree, this storage technology can batch - write the time - series parameter data points, thereby improving the writing efficiency of the time - series parameter data points.

[0041] Such as Figure 5As shown, the process of inserting a hanging leaf node to be demoted to a leaf node into the clustered index statistics tree is described. As shown in the parameter input of step S501, this process receives a parameter X as the demoted leaf node to be inserted. In step S502, when determining an empty root node, if the child node metadata list of the root node is empty, then in step S503, when inserting an empty node, the metadata of this demoted leaf node is added to the child node metadata list of the root node. If the child node metadata list of the root node is not empty, the demoted leaf node is the leaf node with the largest starting point of the time domain maintenance interval except for the hanging leaf nodes, and its time domain maintenance interval is included in each newly generated non-root and non-leaf node. Therefore, the demoted leaf node is maintained by the root node and each newly generated non-root and non-leaf node. In this case, in step S504, when initializing variables, starting from the root node, the demoted leaf node indexes downward along the newly generated non-root and non-leaf nodes through steps S505 and S506. During the indexing process, in step S508, when aggregating metadata, the metadata of the demoted leaf node is aggregated into the metadata stored in its parent node by the newly generated non-root and non-leaf node; S505: Index query; S506: Leaf node determination; after indexing to a leaf node, in step S507, when writing metadata, the metadata of the demoted leaf node is added to the parent node of this leaf node. Subsequently, in step S509, when checking nodes, in accordance with the process Figure 3 of checking the number of non-leaf node metadata entries and splitting, the number of metadata entries of this parent node is checked. Using the above process, during the process of inserting the demoted leaf node into the clustered index statistics tree, the metadata of the newly generated non-root and non-leaf nodes is updated along the way. This storage technology improves the data writing efficiency while accelerating the aggregated query of data.

[0042] Next, each module of the storage management module corresponding to this storage technology will be described. As Figure 6 shown, the working process of the data point insertion module is described. As shown in the parameter input of step S601, this process receives two parameters. One of the parameters is T, representing the statistics tree of the time series parameter data point to be inserted, and one parameter is D = ⟨t, v⟩ representing the time series parameter data point to be inserted, where t represents the data point timestamp and v represents the value of the data point. In the figure, T h represents the starting point of the time domain maintenance interval of the hanging leaf node. The specific process of inserting a data point is that in step S602, when determining a hanging node, if the data point timestamp falls within the interval [T h, within [0, +∞), this data point is maintained by the hanging leaf node. In the initialization of the target node in step S606, this data point is inserted into the hanging leaf node; otherwise, in the initialization of the target node in step S603, starting from the root node, in steps S604 and S605, the leaf node that maintains this data point is indexed downward along the clustering index, and in step S607, the data point is inserted into this leaf node during data point writing. S604: Leaf node determination; S605: Index query. After the insertion is completed, in step S608, the node check, check the data point occupancy space of the node where the data point is inserted. If it exceeds the threshold, the node is split according to the Figure 4 shown process. As Figure 7 shown, it describes the working process of the data point query module. As shown in step S701 parameter input, this process receives two parameters. One parameter is T, representing the statistical tree to be queried, and one parameter is I = [s, e], representing the time stamp interval to be queried, where s represents the start point of the time stamp interval and e represents the end point of the time stamp interval. In the figure, T h represents the start point of the time domain maintenance interval of the hanging leaf node. The specific process of data point query is as follows: In step S702 result set initialization, initialize the query result set. In step S703 hanging node determination, if the start point s of the time stamp interval to be queried is greater than or equal to the start point T h of the time domain maintenance interval of the hanging leaf node, then the time stamp interval of this query is included in the time domain maintenance interval of the hanging leaf node. In step 707 start node initialization, set the start node of the query to the hanging leaf node; otherwise, in step S704 start node initialization, starting from the root node, in steps S705 leaf node determination and S706 index query, index along the statistical tree to the leaf node whose time domain maintenance interval contains the start point s of the time stamp interval to be queried, and use this leaf node as the start node of the query. Subsequently, starting from the query start node, in steps S708 data point query, S709 query end determination, and S710 subsequent query, along the linked list formed by the successor pointers of the leaf nodes, query to the leaf node whose time domain maintenance end is strictly greater than the end point e of the time stamp interval to be queried. During the query process, the time series parameter data points whose time stamps in the leaf nodes are included in [s, e] are added to the query result set. Finally, in step S711 output result set, output the query result set. The above data point query process makes full use of the leaf node linked list formed by the leaf nodes and their successor pointers. On the premise of ensuring the ordered property of the output data, it improves the data point query efficiency of this storage management module through a linear collection method, and improves the situation of insufficient data point query performance in the power equipment scenario.

[0043] As Figure 8As shown, the working process of the aggregated data query module is described. As shown in parameter input of step S801, this process receives two parameters. One parameter is T, representing the statistical tree to be queried, and one parameter is I = [s, e], representing the timestamp interval to be queried, where s represents the start point of the timestamp interval and e represents the end point of the timestamp interval. In the figure, T h represents the start point of the time domain maintenance interval of the hanging leaf node. The specific process of the aggregated data query is as follows. In the hanging node determination of step S802, if the end point e of the timestamp interval to be queried is greater than or equal to the start point T h of the time domain maintenance interval of the hanging leaf node, then there are data points involved in this query within the hanging node. Therefore, in the hanging node query and result set initialization of step S803, the metadata of the data points with timestamps within [T h , e] in the hanging node is used as the initial query result; otherwise, in the result set initialization of step S804, the query result is initialized as an empty set. Subsequently, in the statistical tree subtree query of step S805, according to Figure 9 the aggregated query process, the aggregated data of the data points with timestamps within I among the data points in T except those in the hanging node is collected, and in step S806, it is aggregated into the aggregated result. Finally, in step S807, the aggregated result is output.

[0044] As Figure 9 shown, the process of aggregated query with a subtree of the statistical tree as the query object after removing the hanging node is described. As shown in parameter input of step S901, this process receives two parameters. One parameter is T, representing the root node of the statistical tree subtree to be queried, and one parameter is I = [s, e], representing the timestamp interval to be queried, where s represents the start point of the timestamp interval and e represents the end point of the timestamp interval. The specific process of the statistical tree subtree aggregated query is as follows. In the leaf node determination of step S902, if T is a leaf node, then in the result set initialization of step S903, the metadata of the data points with timestamps within I in T is directly returned; in the child node determination of step S904, if T is a non-leaf node and has no child nodes, then in the result set initialization of step S905, an empty aggregated data is directly returned. If T is a non-leaf node and has child nodes, then in steps S906 to S913, its child nodes are traversed. During the traversal process, in the time domain intersection determination of step S907, if the intersection of the time domain maintenance interval of the child node and I is an empty set, then this node is skipped; in the time domain inclusion determination of step S908, if the time domain maintenance interval of the child node is included in I, then in the metadata aggregation of step S909, the metadata of this child node is aggregated into the aggregated result; if the intersection of the time domain maintenance interval of the child node and I is not an empty set and is not included in I, then in the statistical tree subtree query of step S910, according to Figure 9In the process, the aggregated data of the subtree with this child node as the root node is queried, and the result is aggregated into the total aggregation result in the result aggregation of step S911. Finally, the aggregation result is output in step S914. The above aggregation query process makes full use of the child node metadata stored in the non-leaf nodes of the statistical tree, and improves the aggregation data query efficiency of this storage management module by using the pre-aggregated metadata, thus improving the insufficient aggregation data query performance in the power equipment scenario. Figure 9 Among them, S906: Variable initialization, S912: Iteration end determination; S913: Iterative query.

[0045] As Figure 10 shown, the overall architecture of this time series parameter data storage system is described. Among them, A01 is the computer system on which this storage system is deployed. The storage management module A03 of this storage system runs in the processor A02 of the computer system. The storage management module includes a data point writing module A04, a data point query module A05, and an aggregation data query module A06. The storage management module and its sub-modules read and write the storage medium A07 of the computer system to complete data writing and querying. A number of cluster index statistical trees A08 are stored in the storage medium, and each cluster index statistical tree corresponds to a time series. When using this time series parameter data storage system, the user sends data point insertion, data point query, and data point aggregation query requests to the storage management module. The storage management module reads and writes the cluster index statistical trees in the computer storage medium, collects relevant data, and returns it to the user.

Claims

1. A power equipment timing parameter data storage system based on a statistical tree structure, characterized in that: include: A storage structure for storing timing parameter data generated by the power device; The time series parameter data is stored in a clustered index statistical tree, and one statistical tree stores a single time series of a single device; A computing device comprising a storage management module, the storage management module being operable to: Data point insertion: insert the newly generated time series parameter data of the device into a statistical tree to which the device belongs; Data point query, obtain all time series parameter data within a period of time in a statistical tree to which the device belongs; Aggregate data query, obtain the aggregate indicators of time series parameter data within a period of time in a statistical tree to which the device belongs, including the count, latest value, oldest value of parameter data, and the sum, maximum value, minimum value, and average value of numerical type parameter data; The clustered index statistics tree includes: Statistical cluster index, the index is a clustered statistical tree, the leaf nodes of the statistical tree store time series parameter data points, and the non-leaf nodes store the metadata of their child nodes, where each child node stores a copy of the metadata; The hanging leaf node is a leaf node that stores the latest time series parameter data. The parent node of the hanging leaf node is fixed as the root node; Leaf node linked list,The leaf nodes of the cluster statistics tree are arranged in ascending order of timestamps, and the leaf node linked list is formed through the pointers of the time intervals from small to large between the leaf nodes of adjacent time intervals; Unequal partitioning: when the amount of data stored in a node exceeds a preset threshold, the node will be partitioned into two nodes, and the free space generated during the partitioning process will be allocated to the two nodes formed by the partitioning according to a set ratio.

2. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: Each node of the clustered index statistical tree has its time domain maintenance interval. The time domain maintenance interval means that among the time series parameter data produced by the power equipment corresponding to the statistical tree, the data points whose timestamps fall within the time domain maintenance interval are stored and only stored in the node to which the time domain maintenance interval belongs and its direct or indirect child nodes; The time series parameter data stored in the leaf node includes the time series parameter data points generated by the power equipment associated with the tree index in its time domain maintenance interval. The data points are composed of a timestamp and a measurement point measurement value at the time corresponding to the timestamp; the time series parameter data points stored in the leaf node are arranged in ascending order according to the timestamp; The metadata of the child nodes is stored in non-leaf nodes. Each child node stores a copy of its own metadata in its direct parent node. The metadata stored in non-leaf nodes is arranged in ascending order according to the starting point of their time domain maintenance interval. The metadata of the node includes: If the node is a leaf node, its metadata includes the address of the node, the starting point of the time domain maintenance interval, and the timestamp of the latest data point, the data value of the latest data point, the timestamp of the oldest data point, the data value of the oldest time point, and the total count of data points among the data points stored in the leaf node; if the data point type stored in the leaf node is numeric, the metadata also includes the sum of the data point values, the maximum value of the data point values, and the minimum value of the data point values; If the node is a non-leaf node, its metadata includes the address of the node and an aggregate of the metadata of all its child nodes; the aggregation method of the child node metadata is as follows: for the maximum value of the latest data point timestamp and data point value, the aggregate value is the maximum value of the corresponding values ​​of all child node metadata; for the starting point of the time domain maintenance interval, the oldest data point timestamp, and the minimum value of the data point value, the aggregate value is the minimum value of the corresponding values ​​of all child node metadata; for the total count of data points and the sum of data point values, the aggregate value is the sum of the corresponding values ​​of all child node metadata; For the latest data point data value, the aggregate value is the latest data point data value of the child node metadata with the largest latest data point timestamp; for the oldest data point data value, the aggregate value is the oldest data point data value of the child node metadata with the smallest oldest data point timestamp.

3. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The characteristics of the hanging leaf node are: The leaf node that stores the latest time series parameter data point has a fixed parent node as the root node of the statistical tree. The root node, as the parent node of the suspended leaf node, does not store the metadata of the suspended leaf node, but stores the starting point of the time domain maintenance interval of the suspended leaf node.

4. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The characteristics of the leaf node linked list are: The suspended leaf nodes store the latest time series parameter data. Each leaf node stores a successor pointer that points to another leaf node. Except for the leaf node whose time domain maintenance interval starts at 0, each leaf node is pointed to by only one successor pointer. All leaf nodes form a linked list with the leaf node whose time domain maintenance interval starts at 0 as the first node and the suspended leaf node as the last node. The leaf nodes in the linked list are arranged in ascending order according to the starting point of the time domain maintenance interval.

5. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The characteristics of the unequal segmentation are: When the amount of data stored in a node exceeds a preset threshold, the node will be split into two nodes. The threshold indicators include: For leaf nodes, the threshold indicator is the size of the storage space occupied by the time series parameter data points stored in the leaf node; For non-leaf nodes, the threshold indicator is the number of sub-node metadata entries stored in the non-leaf node, where the starting point of the time domain maintenance interval stored in the root node of the suspended leaf node is not counted in the number of sub-node metadata entries; When performing node segmentation, according to the category of the segmented node, the specific features of the node segmentation are: For non-hanging leaf nodes, when a node is split, a new non-hanging leaf node is created, and a time point is selected to move the time series parameter data points in the split leaf node whose timestamp is greater than or equal to the time point to the new leaf node; the selected time point should satisfy: after the split, the absolute value of the difference between the ratio of the remaining space in the split leaf node to the remaining space in the newly created leaf node and the archive node split weight is the smallest; after the split, the metadata entry of the newly created leaf node is inserted into the parent node of the split node, and the successor pointers of the newly created node and the split leaf node are updated; For hanging leaf nodes, when nodes are split, a new hanging leaf node is created, and the original hanging leaf node is downgraded to a leaf node; at the same time, a time point is selected, and the time series parameter data points with timestamps less than the time point in the original hanging leaf node are moved to the new hanging leaf node; the selected time point should satisfy: after the split, the absolute value of the difference between the ratio of the remaining space in the downgraded leaf node to the remaining space in the newly created hanging leaf node and the split weight of the new node is the smallest; after the split, the downgraded leaf node is inserted into the cluster statistical tree; For non-leaf nodes that are not root nodes, when a node is split, a new non-leaf node is created, and a time point is selected at the same time, and the metadata entries of the child nodes whose time domain maintenance interval starting point is greater than or equal to the time point in the split non-leaf node are moved to the new non-leaf node. After the move, the parent node of the corresponding child node becomes the newly created non-leaf node; the selected time point should satisfy the following: after the split, if the split node is the node with the largest time domain maintenance interval starting point among the non-leaf nodes of the same depth before the split, then the absolute value of the difference between the ratio of the number of remaining storable child node metadata entries in the split non-leaf node to the number of remaining storable child node metadata entries in the newly created non-leaf node and the split weight of the new node is the smallest; otherwise, the absolute value of the difference between the ratio of the number of remaining storable child node metadata entries in the split non-leaf node to the number of remaining storable child node metadata entries in the newly created non-leaf node and the split weight of the archive node is the smallest; after the split, the metadata entries of the newly created non-leaf node are inserted into the parent node of the split node; For the root node, when the node is split, a new non-leaf node is created, and all the metadata entries of the child nodes in the root node except the starting point of the time domain maintenance interval of the hanging leaf node are moved to the new non-leaf node. At the same time, the metadata entries of the new non-leaf node are stored in the root node; then, the new non-leaf node is split according to the standard of the node with the largest starting point of the time domain maintenance interval among the non-leaf nodes of the same depth.

6. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The storage management module includes the following steps to implement the data point insertion operation: If the timestamp of the newly generated time series parameter data point is greater than or equal to the starting point of the time domain maintenance interval of the hanging leaf node stored in the root node of the statistical tree, the time series parameter data point is inserted into the hanging leaf node. If the amount of data in the hanging leaf node exceeds the threshold after the insertion, the node splitting is triggered; If the timestamp of the newly generated time series parameter data point is less than the starting point of the time domain maintenance interval of the hanging leaf node, the root node is used as the current node, and the insertion position of the time series parameter data point is searched along the statistical tree. The insertion position is the leaf node whose time domain maintenance interval contains the timestamp of the data point. The operation steps include: If the current node is a non-leaf node, search from the sub-node metadata entries for a sub-node metadata entry whose time domain maintenance interval starting point is less than or equal to the timestamp of the data point to be inserted and whose time domain maintenance interval starting point is the largest, and set the current node as the sub-node; If the current node is a leaf node, search for the data point with the largest timestamp whose timestamp is less than or equal to the timestamp of the data point to be inserted from the time series parameter data points stored in the leaf node; Case 1: If the timestamps of all data points are greater than the timestamp of the data point to be inserted, insert the data point to be inserted into the head of the list; Case 2: If there is a data point with the same timestamp as the node to be inserted, replace the existing data point with the data point to be inserted, or ignore the data point to be inserted according to the configuration; If neither of the above two cases exists, insert the data point to be inserted after the found data point; After inserting the data point into the leaf node, the child node metadata entry stored in the parent node of the leaf node is updated; if the storage space occupied by the leaf node after inserting the data point exceeds the threshold, the node splitting of the leaf node is triggered.

7. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The storage management module includes the following steps to implement the data point query operation: If the starting point of the queried time interval is greater than or equal to the starting point of the time domain maintenance interval of the suspended leaf node stored in the root node, then the data points in the suspended leaf node whose timestamps fall within the time interval are retrieved and returned; If the starting point of the queried time interval is less than the starting point of the time domain maintenance interval of the suspended leaf node stored in the root node, then the root node is taken as the current node, and the leaf node whose time domain maintenance interval includes the starting point of the queried time interval is searched along the statistical tree; Then, along the leaf node linked list, data points whose timestamps fall within the queried interval are collected until a data point whose timestamp is greater than the end point of the queried interval is encountered, or the last data point of the hanging leaf node is iterated to, and the data points collected in the process are returned.

8. The power equipment timing parameter data storage system based on the statistical tree structure according to claim 1 is characterized in that: The specific implementation process of the aggregate data query operation contained in the storage management module is as follows: Take the root node as the current node and collect the aggregated data of the data points within the specified time interval along the statistical tree. The specific steps of the collection operation include: If the current node is a non-leaf node, traverse the sub-node metadata entries stored in the current non-leaf node; if the oldest data point timestamp and the latest data point timestamp in the sub-node metadata entry are both included in the specified time interval, aggregate the sub-node metadata into the collected aggregate data; if the closed interval with the oldest data point timestamp in the sub-node metadata entry as the left boundary and the latest data point timestamp as the right boundary is not included in the specified time interval, but the intersection of the two is not an empty set, recursively collect aggregate data from the sub-tree with the sub-node corresponding to the sub-node metadata as the root, and aggregate it into the collected aggregate data; If the current node is a leaf node, then the data points with timestamps within the specified time interval are taken out from the time series parameter data points stored in the leaf node, and these data points are aggregated into the collected aggregate data; After collecting aggregate data from the statistical tree, if the starting point of the time domain maintenance interval of the hanging leaf node is included in the specified time interval, the data points with timestamps in the specified time interval are obtained from the hanging leaf node, and their aggregate data are calculated and added to the collected aggregate data; finally, the collected aggregate data is returned.

Citation Information

Patent Citations

  • Method for optimizing static power consumption by utilizing related relation graph

    CN110471522A

  • Multi-sampling-stream-oriented time series data management method and system

    CN110825733A