Compression method of intelligent power distribution and utilization terminal data in time sequence database
By separating static and dynamic data, merging timelines, and adopting differentiated data compression methods, the storage redundancy problem of time series databases is solved, and query performance and compression rate are improved.
Patent Information
- Application Number
- CN202510452558.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
When the existing timing database stores data collected by smart power distribution terminals, there are memory and disk storage redundancy problems, and the data query performance is insufficient.
By separating static and dynamic data, merging timelines, defining measurement data structures, registering structures, and using differentiated data compression methods, compressing data by columns.
Reduces data storage redundancy in timing databases, and improves query performance and compression rate.
Smart Images

Figure CN120386798A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a compression method for intelligent power distribution and utilization terminal data in a time series database. Background Art
[0002] The data collected by intelligent power distribution and utilization terminals is characterized by being massive, having an extremely fast data update speed, and being widely distributed geographically. With the advancement of the construction of smart grids, intelligent power distribution and utilization advanced application technologies, as well as cloud computing, Internet technologies, etc., have been increasingly applied in the power system. The application of these emerging technologies involves a large amount of input data with different structures, different sources, and different dimensions, and the basic measurement data structure storage provided by existing time series databases cannot meet their actual operation requirements.
[0003] Existing time series databases generally adopt two methods to store complex power distribution and utilization terminal collection data. One is to split the composite structure data into multiple pieces of data for storage, in which case the timestamps of multiple measurement values of one entity are the same; the other is to simply convert the composite structure into a json format string or a binary string and then store it. Splitting the composite structure data into multiple pieces of data for storage, where the timestamps of multiple measurement values of one entity are the same, will cause a large amount of redundant disk and memory space occupation in the database; converting the composite structure into a json format string or a binary string and then storing it, although this method can reduce the number of measurement points, it requires additional column information, and the data compression ratio is very low, which will also cause space waste, and character data parsing and filtering are required during data analysis, and the timeliness of data query cannot be guaranteed. Summary of the Invention
[0004] Object of the Invention: The object of the present invention is to provide a compression method for intelligent power distribution and utilization terminal data in a time series database, to solve the problem of a large amount of redundancy in memory and disk storage of intelligent power distribution and utilization terminal collection data in the time series database, and at the same time improve the database query performance.
[0005] Technical Solution: A compression method for intelligent power distribution and utilization terminal data in a time series database according to the present invention includes the following steps:
[0006] (1) The original terminal collects data and performs preprocessing to generate a matrix-style original data group;
[0007] (2) Analyze the static attribute values and dynamic attribute values in the data, and separate the static and dynamic data;
[0008] (3) Organize the dynamic data, merge multiple time lines collected by the same power distribution and utilization terminal device into one time line; define the measurement data structure, and register the structure definition with the time series database;
[0009] (4) Organize the static data, define the logical measurement points, write the logical measurement points corresponding to each intelligent power distribution and consumption acquisition terminal into the time series database, and the time series database loads the key information into the memory and creates an index.
[0010] (5) Regularize the dynamic data according to the defined data structure and send it to the time series database. The time series database stores the measurement data in the composite structure in the cache according to the time window.
[0011] (6) When the time window of a certain timeline is completed, the time series data in the cache is packaged, compressed and archived in the way of column aggregation and compression.
[0012] Further, in step (1), the preprocessing includes eliminating the data without timestamps, sorting the data by timestamps, deleting the duplicate time data, retaining the latest data, and converting the character-type data in the formats of numeric characters, date and time data, encoded data, and numerical information and categorical data in text data into numeric data.
[0013] Further, in step (2), the static attribute value is the attribute value that does not update, and the dynamic attribute value is the measurement attribute value that is continuously collected, changed, and updated over time.
[0014] Further, in step (3), the timeline is a sequence in which the measurement attribute values of a certain power distribution and consumption terminal device are sorted by time.
[0015] Further, in step (3), the structure definition includes the attribute names, attribute types, and attribute lengths of each measurement component; among them, the structure attributes include at least the timestamp and one or more measurement attribute components.
[0016] Further, in step (4), one logical measurement point corresponds to one merged timeline in step (3), and the definition of the logical measurement point includes the static data of the power distribution and consumption acquisition terminal, the storage duration of the timeline data, the lossless compression accuracy, and the structure definition in step (3) associated with the timeline data.
[0017] Further, in step (4), the key information loaded by the time series database is the key column information commonly used as query filtering conditions.
[0018] Further, in step (5), the format after regularizing the dynamic data is [logical measurement point name + timestamp + binary data], where the binary data is formed by converting the values of each measurement component into binary according to the size and order of the structure definition and then splicing them.
[0019] Further, in step (5), the time series database applies for a buffer with a dynamic size for each logical measurement point to store the structured data of a time window; the time window is a time unit configured by the database.
[0020] Further, in step (5), the column - wise aggregation and compression of the time - series data in the cache includes the following steps:
[0021] (61) Obtain the structure definition information corresponding to the timeline data according to the logical measurement point name;
[0022] (62) Parse the data in the cache according to the structure definition to form multiple column - value arrays, including arrays of each component value and a timestamp array, where the number of array elements is exactly the same;
[0023] (63) Select different compression processing methods and compression algorithms according to the types of different column data respectively;
[0024] (64) After each component is compressed, add length and algorithm information to the header and splice them in order. The spliced compressed block also needs to add a packet header, including the compressed block length, logical measurement point identifier, data type, number of data items, data time range, and relevant statistical information.
[0025] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: The present invention analyzes almost unchanging static archive data (such as device location, etc.) and dynamic measurement data that changes over time in the data collected by the analysis terminal; The static data is organized and stored in the time - series database, and the key information resides in memory; Define a data structure for the dynamic measurement data and register the structure in the time - series database; The time - series database caches the measurement data of the composite structure and processes the component values according to the associated structure definition; The time - series database adopts a differential data compression method according to the component data type; The present invention reduces the data storage redundancy in the time - series database and improves the compression ratio and read - write performance. Description of the Drawings
[0026] Figure 1 It is a schematic diagram of the present invention;
[0027] Figure 2 It is a schematic diagram of the data pre - processing of the present invention;
[0028] Figure 3 It is a data - processing flowchart of the present invention;
[0029] Figure 4 It is a compression - processing method of the present invention. Detailed Embodiments
[0030] The technical solution of the present invention will be further described below with reference to the drawings.
[0031] As Figure 1 shown, an embodiment of the present invention provides a compression method for intelligent power distribution and consumption terminal data in a time - series database, including the following steps:
[0032] Step 1: Clean and organize the data collected by the original terminal to generate a matrix-style raw data group; data cleaning includes deleting duplicates and handling outliers; data organization includes sorting the data, format and type conversion.
[0033] Deleting duplicates means identifying and eliminating duplicate or redundant entries in the dataset to ensure that the analysis is performed on unique and accurate data. For a time-series database, for the same logical measurement point, there can only be one data value at each moment. If there are multiple data values at the same moment for a certain logical measurement point, only the last one is retained.
[0034] Handling outliers mainly targets null data or data with invalid value ranges. To ensure data validity, this type of data needs to be cleaned before being stored in the database. To ensure the integrity of the database data, extreme value data is not filtered.
[0035] Data sorting means sorting the data of the same logical measurement point according to the time stamp corresponding to the measured value, ensuring that the data of the same logical measurement point is in time order when stored in the database, which can reduce the impact of time disorder on the write performance of the time-series database.
[0036] Format and type conversion refers to the process of converting one data format to another format or data type. For example, converting a string to a number, or formatting a number into a specific string form. In principle, the data stored in the time-series database should use non-character data as much as possible to ensure data read performance. For example, the time stamp input is generally various standardized strings, and it needs to be converted to a 64-bit long integer data before being stored in the database. The time stamp uses UTC time.
[0037] The matrix-style raw data group can be regarded as a two-dimensional temporary table and has two representation methods. One is that the row is the sorted data set of a certain logical measurement point, and the column represents different logical measurement points. In this way, when writing to the time-series database, batch writing is used, and one logical measurement point is written at a time; the other is that the row is the data set of all logical measurement points with values at a certain moment, and the column represents different time stamps. In this way, when writing to the time-series database, section writing is used.
[0038] Step 2: Separate the static attribute values and dynamic attribute values in the data. Some attribute values in the device model are basically not updated (such as device location), which can be called static archive attributes; the measurement attributes in the device model are continuously collected over time, change rapidly, and are updated frequently, which are called dynamic measurement attributes. The static archive attributes and dynamic measurement attributes of the device model are distinguished and processed. The archive data is separately stored in the archive table, corresponding to the measurement point table in the time-series database, and a memory index is established. The measurement data is subsequently stored in the time-series data management module. This method can reduce the duplicate storage of static data values. In addition, because the amount of static archive data is small, caching the static archive attributes in memory can improve data read and write performance.
[0039] Step 3: As Figure 1 shown, establish a tree model based on the matrix - type raw data group, corresponding data sources, attribute levels, attribute coverage, and the subordinate relationship between data. The important data source for the intelligent power distribution and consumption terminal to collect data is device collection. The bottom layer of the tree model is a single measurement attribute value sequence (in the time - series database, this sequence is called a timeline). The nodes above the bottom layer correspond to a certain data source, that is, one data source contains multiple measurement attribute value sequences. For the measurement attribute value sequences under the same data source, if their timestamp sequences are exactly the same, they can be merged into one timeline, that is, one data value contains multiple measurement attribute component values. This method can reduce the number of timelines, greatly save the memory and disk space of the time - series database, and thus improve the database performance.
[0040] Step 4: As Figure 2 shown, define a structure according to the structured time and data sequence merged in Step 3. The structure definition includes the attribute names, attribute types, and attribute lengths of each member variable. The structure attributes at least include timestamps and one or more measurement attribute component values. In Step 1, the timestamp has been converted into a 64 - bit long - integer data, and the measurement attribute component values support floating - point type, double - precision type, integer type, boolean type, and character type. After determining the structure definition, register the structure definition with the time - series database. Because the measurement attributes included in different data sources can be different, there can be multiple structure definitions in the time - series database, but not many. The time - series database can cache all structure definitions in memory for fast retrieval.
[0041] Step 5: According to the static archive data separated from the original data group in Step 2, combined with the timelines merged in Step 3, define logical measurement points. Each logical measurement point corresponds to a timeline in the time - series database. The definition of the logical measurement point includes the static archive data of the data source and relevant additional information. The additional information includes the storage duration of the timeline data, the lossless compression accuracy, and the structure definition in Step 4 associated with the timeline data. As Figure 2 shown, after determining the definition of the logical measurement point, write the logical measurement point information to the time - series database. The time - series database archives the logical measurement point information into a file for long - term storage. At the same time, the key columns that are often used as query filtering conditions are loaded into memory as logical measurement point tags, and label indexes are established respectively to accelerate the retrieval performance.
[0042] Step 6: Regularize the dynamically measured data combined in Step 3 according to the structural definition of the dynamic data of this data source in Step 4. The regularized format is [logical measurement point name + timestamp + binary data], and submit it to the time series database. After receiving the data, the time series database determines the corresponding time line according to the logical measurement point name, caches the dynamic measurement data of different time lines, and dynamically adjusts the cache size according to the real-time data volume of the configured time window. If the time series database receives new measurement data that falls within this time window, the data needs to be inserted into this cache and sorted according to the data time. Here, the data time refers to the value of the timestamp attribute in the measurement data structure, not the data storage time.
[0043] Step 7: When the time window of a certain time line in Step 6 is completed, it is necessary to package, compress and archive the time series data in the cache. As Figure 3 shown, for data with a composite structure, the compression ratio can be improved by using the columnar aggregation compression method. The specific compression processing method is as follows:
[0044] a Obtain the structure definition information corresponding to the data of this time line according to the logical measurement point name (defined in Step 4).
[0045] b Analyze each piece of data in the cache according to the type and length of each measurement component in the structure definition, and finally form multiple column value arrays, including arrays of each component value and a timestamp array, with the same number of array elements.
[0046] c Select different compression algorithms for compression processing according to the types of different column data. In particular, for the timestamp column, since the time is ordered and most data sources collect data periodically, the timestamps in this scenario are incremented at equal intervals. Before compression, perform a preprocessing of calculating the difference between the previous and the next values, which has the highest compression ratio. For the compression of floating-point and double-precision type values, the set lossless compression accuracy needs to be obtained from the logical measurement point information in Step 5.
[0047] d After each component is compressed, add the length and compression algorithm information before the compressed string, form a binary string by sequentially combining all compressed strings, and add header information, including the compressed block length, logical measurement point identifier, data type, number of data items, data time range, and relevant statistical information.
[0048] Step 8: After the data block is compressed, store it in a disk file and update the time series database index. The index node contains the offset of the data block in the file. Data retrieval reads according to the data block offset in this index node and decompresses according to the compression algorithm and compressed block length in the compressed block header information, and restores it according to the original binary data organization method of the compressed block.
Claims
1. A compression method for intelligent power distribution and utilization terminal data in a time series database, characterized in that, It includes the following steps: (1) The original terminal collects data and performs preprocessing to generate a matrix - type original data group; (2) Analyze the static attribute values and dynamic attribute values in the data to separate static and dynamic data; (3) Organize the dynamic data, merge multiple time - lines collected by the same power distribution and utilization terminal device into one time - line; define the measurement data structure and register the structure definition with the time - series database; (4) Organize the static data, define logical measurement points, write the logical measurement points corresponding to each intelligent power distribution and utilization collection terminal into the time - series database, and the time - series database loads the key information into the memory and establishes an index; (5) Regularize the dynamic data according to the defined data structure and send it to the time - series database. The time - series database stores the measurement data of the composite structure in the cache according to the time window; (6) When a time - window of a certain time - line is completed, the time - series data in the cache is packaged, compressed and archived in the way of column - wise aggregation and compression.
2. The compression method of the intelligent power distribution and utilization terminal data in the time series database according to claim 1, wherein In step (1), the preprocessing includes eliminating data without timestamps, sorting the data by timestamps, deleting duplicate time data, retaining the latest data, and converting character - type data in the formats of numerical characters, date and time data, encoded data, numerical information in text data, and categorical data into numerical data.
3. The compression method of the intelligent power distribution and utilization terminal data in the time series database according to claim 1, characterized in that, In step (2), the static attribute value is an attribute value that does not update, and the dynamic attribute value is a measured attribute value that is continuously collected, changed and updated over time.
4. A compression method for intelligent power distribution and utilization terminal data in a time series database according to claim 1, characterized in that, In step (3), the time - line is a sequence of measured attribute values of a certain power distribution and utilization terminal device sorted by time.
5. A compression method for intelligent power distribution and utilization terminal data in a time series database according to claim 1, characterized in that, In step (3), the structure definition includes the attribute names, attribute types and attribute lengths of each measurement component; among them, the structure attributes at least include timestamps and one or more measurement attribute components.
6. The compression method of the intelligent power distribution and utilization terminal data in the time series database according to claim 1, wherein In step (4), one logical measurement point corresponds to one merged time - line in step (3). The definition of the logical measurement point includes the static data of the power distribution and utilization collection terminal, the storage duration of the time - line data, the lossless compression accuracy, and the structure definition in step (3) associated with the time - line data.
7. A compression method for intelligent power distribution and utilization terminal data in a time series database according to claim 1, characterized in that, In step (4), the key information loaded by the time - series database is the key column information commonly used as query filtering conditions.
8. A compression method for intelligent power distribution and utilization terminal data in a time series database according to claim 1, characterized in that, In step (5), the format of the regularized dynamic data is [logical measurement point name + timestamp + binary data], where the binary data is formed by splicing the measured component values converted into binary according to the size and order of the structure definition.
9. A compression method for intelligent power distribution and utilization terminal data in a time series database according to claim 1, characterized in that, In step (5), the time - series database applies for a dynamically - sized buffer for each logical measurement point to store the structured data of a time window; the time window is a time unit configured by the database.
10. A method for processing and compressing the data collected by an intelligent power distribution and utilization terminal in a time series database, characterized in that, In step (5), the column - wise aggregation and compression of the time - series data in the cache includes the following steps: (61) Obtain the structure definition information corresponding to the time - line data according to the logical measurement point name; (62) Analyze the data in the cache according to the structure definition to form multiple column - value arrays, including an array of each component value and a timestamp array, and the number of array elements is exactly the same; (63) Select different compression processing methods and compression algorithms according to the types of different column data respectively; (64) After the compression of each component is completed, length and algorithm information are added to the header and concatenated in sequence. The concatenated compressed block also needs to add a packet header, including the compressed block length, logical measurement point identifier, data type, number of data items, data time range, and relevant statistical information.