Online data reorganization method, system and equipment for time sequence database and medium

By using ordered iterators and BlockSpan queues in the timing database, the problem of large storage space occupied by the timing database and low query efficiency after deletion of data marks is solved, and efficient storage and query of data files is achieved.

CN120011366APending Publication Date: 2025-05-16上海沄熹科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139252.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing time-series database is not physically deleted after the data mark is deleted, resulting in large storage space occupies and low query efficiency. The data column type is modified and converted during the reading process, which reduces the reading speed.

Method used

Through an ordered iterator, data is sorted by tags, invalid data is filtered, data is read and sorted by time, written to the data file and replaced the original file. Use the BlockSpan queue to record data block address and line number information, and supports deletion of data filtering, adding column filling, deletion of column filtering, field type change and data sorting.

Benefits of technology

It reduces the storage space of data files, improves query efficiency, solves the problems of wasted storage space and low query speed, and at the same time realizes online data reorganization without affecting the normal use of the database system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011366A_ABST
    Figure CN120011366A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence database online data reorganization method, system and device and a medium, belongs to the technical field of time sequence databases, and aims to solve the technical problem of how to reduce data file storage space occupation, improve query efficiency and do not affect online normal use of a database system. According to the technical scheme, classification of data according to labels, filtering of invalid data, reading of the data and sorting of the data according to time are completed through an ordered iterator; writing the data read by the ordered iterator into the data file and replacing the original file; wherein the ordered iterator queries effective data of a group of entities represented by Tags in a specified time period, and the queried data is returned by using a BlockSpan queue for recording a data block address, a starting line number and an accumulated line number.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time series databases, and in particular to a method, system, device and medium for online data reorganization of a time series database. Background Art

[0002] Time series data refers to a collection of data recorded in chronological order, usually consisting of a timestamp and a corresponding observation. The timestamp indicates the time point of the data point, and the observation is the value, indicator or event measured or recorded at a given point in time. Time series data is widely used in many fields, such as financial market analysis, weather forecasting, traffic flow monitoring, and production process monitoring.

[0003] Time series database, the full name of which is time series database, is a database management system specifically used to store, analyze and process time series data.

[0004] Time series data usually has a large amount of data and a high frequency of collection. Therefore, in order to achieve real-time and efficient large-scale data writing, most time series databases not only improve the writing speed in a targeted manner during the design process, but also generally adopt a mark-delete method for deletion operations, that is, adding a deletion mark to the deleted data row / column instead of complex and time-consuming real physical deletion. Time series data query and analysis requirements are high, especially common query and analysis of data within a period of time. In order to meet fast query and analysis, time series databases generally partition data based on timestamps. In addition, time series databases mostly use columnar files for storage. Since columnar files are used for data storage, when the data column type is modified, in order to reduce the impact on write performance, the data is generally not modified directly, but the column metadata is modified, and data conversion is only performed during the reading process.

[0005] Some basic concepts of time series databases (different time series databases have different expressions, and here only one expression method involved in this patent is listed):

[0006] Metric: A metric is similar to a table in a relational database and represents a collection of time series data of the same type.

[0007] Timestamp: Timestamp.

[0008] Tags: Tags are information describing the characteristics of a data source and generally do not change over time.

[0009] Fields: measurement values, measurement results of data sources, generally changing over time.

[0010] Different combinations of labels actually identify a separate entity, distinguishing it from other entities.

[0011] The existing time series database library has the following problems:

[0012] ① The problem of large storage space usage and storage space waste caused by marking the data in the time series database for deletion instead of actually physically deleting it.

[0013] ②Out-of-order data seriously affects the query speed of time series data.

[0014] ③ After the data column type is modified, it is converted during the data reading process, which reduces the reading speed and affects the query efficiency.

[0015] Therefore, how to reduce the storage space occupied by data files and improve query efficiency without affecting the normal online use of the database system is a technical problem that needs to be solved urgently. Summary of the invention

[0016] The technical task of the present invention is to provide a method, system, device and medium for online data reorganization of a time series database to solve the problem of how to reduce the storage space occupied by data files and improve query efficiency without affecting the normal online use of the database system.

[0017] The technical task of the present invention is achieved in the following way: a method for online data reorganization of a time series database, the method is specifically as follows:

[0018] Use ordered iterators to classify data by label, filter invalid data, read data, and sort data by time;

[0019] Write the data read by the ordered iterator into the data file and replace the original file;

[0020] Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

[0021] Preferably, the structure of the BlockSpan queue is as follows:

[0022] struct BlockSpan{

[0023] BlockItem*block_item;

[0024] uint32_t start_row;

[0025] uint32_t row_count;

[0026] };

[0027] BlockItem is the metadata of a data block, including labels, number of rows, maximum / minimum timestamps, and information about whether there is out-of-order data.

[0028] start_row is the starting row number of the data to be returned by the iterator;

[0029] row_count is the cumulative number of rows returned starting from start_row.

[0030] Preferably, the ordered iterator has the following functionality:

[0031] ①Support deletion of data filtering;

[0032] ② After adding columns, historical data can be filled (usually filled with NULL);

[0033] ③Support historical data filtering after deleting columns;

[0034] ④Support the reading and conversion of historical data after the field type is changed;

[0035] ⑤Support data sorting.

[0036] As a preferred method, the ordered iterator data reading process is as follows:

[0037] S1. Get all corresponding BlockItems according to Tags;

[0038] S2. Check whether BlockItems are out of order:

[0039] S201. If BlockItems are not marked as out of order, traverse each BlockItem, obtain the maximum and minimum timestamps, and check whether the timestamps are within the specified time period:

[0040] ①If it is not within the time range, skip it and continue traversing;

[0041] ② If it is within the time range, traverse each row in the corresponding BlockItem;

[0042] S202, if the BlockItems are marked out of order, execute step S203;

[0043] S203, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. There may be multiple BlockItems with the same start and end time range, so multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems should be traversed and recorded in interval_block_map;

[0044] S204, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64>>In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems;

[0045] S205, traverse each group of timestamps of intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals;

[0046] S206. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing the ordered BlockItems, and the process goes to step S201 to continue execution.

[0047] Preferably, each row in the corresponding BlockItem is traversed as follows:

[0048] Traverse to the first row of data that is within the time range and not marked for deletion, record the BlockItem in the BlockSpan queue, and use the corresponding row number as the starting row number start_row, continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, the consecutive rows end, update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, repeat until all data traversal is completed, and return the BlockSpans after filtering and sorting.

[0049] Preferably, the data read by the ordered iterator is written to the data file and replaces the original file as follows:

[0050] (1) Create a temporary data file partition, and create a table and column file in the temporary data file partition according to the latest metadata;

[0051] (2) Traverse the metadata of the temporary partition column and read the corresponding column of the data returned by the ordered iterator:

[0052] If there is no corresponding column in the read data, it means that the corresponding column is a newly added column, and the corresponding column of the temporary partition is filled with NULL;

[0053] If the column type in the read data is inconsistent with the metadata of the temporary partition column, it means that the corresponding column type is modified, the data in the ordered iterator is read out according to the original type, and the original type is converted to the latest column type in the temporary partition and written into the temporary partition column file;

[0054] (3) After all column files are traversed and written, try to lock the corresponding partition. After successfully locking, use the temporary partition to replace the old partition, clear the cache, and reload the partition and metadata;

[0055] (4) After the replacement is completed, clean up the remaining files.

[0056] A time series database online data reorganization system, the system comprising:

[0057] The reading module is used to classify data by label, filter invalid data, read data, and sort data by time through an ordered iterator;

[0058] The data writing and replacement module is used to write the data read by the ordered iterator into the data file and replace the original file. Specifically, the column type is converted according to the metadata of the original data and the latest version of the metadata and then written. NULL is used to fill the columns that are not newly added in the original data. At the same time, the column descriptions that do not exist in the latest version of the metadata have been deleted by comparing the original data and do not need to be written into the reorganized data file again.

[0059] Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

[0060] Preferably, the working process of the reading module is as follows:

[0061] S1. Get all corresponding BlockItems according to Tags;

[0062] S2. Check whether BlockItems are out of order:

[0063] S2A. If BlockItems is not marked as out of order, execute steps S2A1 to S2A3:

[0064] S2A1. Traverse each BlockItem, obtain the maximum and minimum timestamps of the BlockItem, and check whether the timestamps are within the specified time period:

[0065] If it is not within the time range, skip it and continue traversing;

[0066] If it is within the time range, then traverse each row in the corresponding BlockItem, i.e., execute step S2A2;

[0067] S2A2, first traverse to the first row of data that is within the time range and not marked for deletion, record BlockItem in BlockS pan, and use the corresponding row number as the starting row number start_row, continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, the consecutive rows end, update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, and repeat until all data traversal is completed;

[0068] S2A3, return the BlockSpans after filtering and sorting;

[0069] S2B. If the BlockItems are marked out of order, execute steps S2B1 to S2B:

[0070] S2B1, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. There may be multiple BlockItems with the same start and end time range, so multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems should be traversed and recorded in interval_block_map;

[0071] S2B2, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64>>In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems;

[0072] S2B3, traverse each group of timestamps in intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals;

[0073] S2B4. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing ordered BlockItems, and jumps to step S2A to continue execution.

[0074] An electronic device comprising: a memory and at least one processor;

[0075] Wherein, the memory stores a computer program;

[0076] The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the online data reorganization method for a time series database as described above.

[0077] A computer-readable storage medium stores a computer program, which can be executed by a processor to implement the online data reorganization method for a time series database as described above.

[0078] The time series database online data reorganization method, system, device and medium of the present invention have the following advantages:

[0079] (1) The present invention realizes database data organization, reduces storage space occupation, and uses ordered iterators to classify and sort data during data reorganization, and performs possible type conversion, thereby improving the efficiency of time series data query. At the same time, data reorganization is performed online without downtime, and does not affect the normal use of the database system;

[0080] (ii) The present invention can reduce the storage space occupied by data files and improve query efficiency without affecting the normal online use of the database system. It solves the problems of large storage space occupation and storage space waste caused by data mark deletion in the time series database without actual physical deletion, the serious impact of disordered data on the query speed of time series data, and the conversion of data column types during data reading after modification, which reduces the reading speed and affects the query efficiency.

[0081] (III) The present invention uses an ordered iterator to organize and read the original time series data, completes the marked deleted data filtering during the reading process, and completes the sorting according to the timestamp. When writing and replacing data, the latest metadata is compared with the metadata of the data read by the ordered iterator to implement the marked deleted column filtering, fill the new column, convert the column type, and complete the data reorganization. After the data reorganization is completed, the storage space of the marked deleted data is released, the disk space occupancy is optimized, and the data classification and sorting are completed at the same time, improving the query efficiency;

[0082] (iv) The ordered iterator of the present invention is based on a conventional iterator, and uses a BlockSpan structure queue that records the data block address, the starting row number, and the cumulative number of rows to store the column data that has been read;

[0083] (V) In the process of reading and writing data, the present invention converts the column type according to the metadata of the original data and the metadata of the latest version and then writes it, and fills the newly added columns that are not in the original data with NULL;

[0084] (VI) The present invention compares the original data during the process of reading and writing data, and the column descriptions that do not exist in the latest version of metadata have been deleted and do not need to be written into the reorganized data file again. BRIEF DESCRIPTION OF THE DRAWINGS

[0085] The present invention is further described below in conjunction with the accompanying drawings.

[0086] Attached Figure 1 A schematic diagram of sorting, filtering and sorting out-of-order data for ordered iterators;

[0087] Attached Figure 2 This is a schematic diagram of the effect after data reorganization. DETAILED DESCRIPTION

[0088] The time series database online data reorganization method, system, device and medium of the present invention are described in detail below with reference to the accompanying drawings and specific embodiments of the specification.

[0089] Embodiment 1:

[0090] As attached Figure 1 and 2 As shown, this embodiment provides a method for online data reorganization of a time series database, and the method is specifically as follows:

[0091] (1) Use ordered iterators to classify data by label, filter invalid data, read data, and sort data by time;

[0092] (2) writing the data read by the ordered iterator into the data file and replacing the original file;

[0093] Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

[0094] The structure of the BlockSpan queue in this embodiment is as follows:

[0095]

[0096] BlockItem is the metadata of a data block, including labels, number of rows, maximum / minimum timestamps, and information about whether there is out-of-order data.

[0097] start_row is the starting row number of the data to be returned by the iterator;

[0098] row_count is the cumulative number of rows returned starting from start_row.

[0099] The ordered iterator in this embodiment has the following functions:

[0100] ①Support deletion of data filtering;

[0101] ② After adding columns, historical data can be filled (usually filled with NULL);

[0102] ③Support historical data filtering after deleting columns;

[0103] ④Support the reading and conversion of historical data after the field type is changed;

[0104] ⑤Support data sorting.

[0105] The ordered iterator data reading process in step (1) of this embodiment is specifically as follows:

[0106] S1. Get all corresponding BlockItems according to Tags;

[0107] S2. Check whether BlockItems are out of order:

[0108] S201. If BlockItems are not marked as out of order, traverse each BlockItem, obtain the maximum and minimum timestamps, and check whether the timestamps are within the specified time period:

[0109] ①If it is not within the time range, skip it and continue traversing;

[0110] ② If it is within the time range, traverse each row in the corresponding BlockItem;

[0111] S202, if the BlockItems are marked out of order, execute step S203;

[0112] S203, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. There may be multiple BlockItems with the same start and end time range, so multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems should be traversed and recorded in interval_block_map;

[0113] S204, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64>>In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems;

[0114] S205, traverse each group of timestamps of intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals;

[0115] S206. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing the ordered BlockItems, and the process goes to step S201 to continue execution.

[0116] The specific steps of traversing each row in the corresponding BlockItem in step S201 of this embodiment are as follows:

[0117] Traverse to the first row of data that is within the time range and not marked for deletion, record the BlockItem in the BlockSpan queue, and use the corresponding row number as the starting row number start_row, continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, the consecutive rows end, update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, repeat until all data traversal is completed, and return the BlockSpans after filtering and sorting.

[0118] The specific steps of writing the data read by the ordered iterator into the data file and replacing the original file in step (ii) of this embodiment are as follows:

[0119] (1) Create a temporary data file partition, and create a table and column file in the temporary data file partition according to the latest metadata;

[0120] (2) Traverse the metadata of the temporary partition column and read the corresponding column of the data returned by the ordered iterator:

[0121] If there is no corresponding column in the read data, it means that the corresponding column is a newly added column, and the corresponding column of the temporary partition is filled with NULL;

[0122] If the column type in the read data is inconsistent with the metadata of the temporary partition column, it means that the corresponding column type is modified, the data in the ordered iterator is read out according to the original type, and the original type is converted to the latest column type in the temporary partition and written into the temporary partition column file;

[0123] (3) After all column files are traversed and written, try to lock the corresponding partition. After successfully locking, use the temporary partition to replace the old partition, clear the cache, and reload the partition and metadata;

[0124] (4) After the replacement is completed, clean up the remaining files.

[0125] Embodiment 2:

[0126] This embodiment provides a time series database online data reorganization system, the system comprising:

[0127] The reading module is used to classify data by label, filter invalid data, read data, and sort data by time through an ordered iterator;

[0128] The data writing and replacement module is used to write the data read by the ordered iterator into the data file and replace the original file. Specifically, the column type is converted according to the metadata of the original data and the latest version of the metadata and then written. NULL is used to fill the columns that are not newly added in the original data. At the same time, the column descriptions that do not exist in the latest version of the metadata have been deleted by comparing the original data and do not need to be written into the reorganized data file again.

[0129] Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

[0130] The working process of the reading module in this embodiment is as follows:

[0131] S1. Get all corresponding BlockItems according to Tags;

[0132] S2. Check whether BlockItems are out of order:

[0133] S2A. If BlockItems is not marked as out of order, execute steps S2A1 to S2A3:

[0134] S2A1. Traverse each BlockItem, obtain the maximum and minimum timestamps of the BlockItem, and check whether the timestamps are within the specified time period:

[0135] If it is not within the time range, skip it and continue traversing;

[0136] If it is within the time range, then traverse each row in the corresponding BlockItem, i.e., execute step S2A2;

[0137] S2A2, first traverse to the first row of data that is within the time range and not marked for deletion, record BlockItem in BlockS pan, and use the corresponding row number as the starting row number start_row, continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, the consecutive rows end, update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, and repeat until all data traversal is completed;

[0138] S2A3, return the BlockSpans after filtering and sorting;

[0139] S2B. If the BlockItems are marked out of order, execute steps S2B1 to S2B:

[0140] S2B1, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. There may be multiple BlockItems with the same start and end time range, so multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems should be traversed and recorded in interval_block_map;

[0141] S2B2, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64>>In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems;

[0142] S2B3, traverse each group of timestamps in intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals;

[0143] S2B4. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing ordered BlockItems, and jumps to step S2A to continue execution.

[0144] Embodiment 3:

[0145] This embodiment also provides an electronic device, including: a memory and a processor;

[0146] Wherein, the memory stores computer-executable instructions;

[0147] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the online data reorganization method for a time series database in any embodiment of the present invention.

[0148] The processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor may be a microprocessor or any conventional processor, etc.

[0149] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory can also include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, at least one disk storage period, a flash memory device, or other volatile solid-state storage devices.

[0150] Embodiment 4:

[0151] This embodiment also provides a computer-readable storage medium, in which a plurality of instructions are stored, and the instructions are loaded by a processor, so that the processor executes the online data reorganization method of the time series database in any embodiment of the present invention. Specifically, a system or device equipped with a storage medium can be provided, on which a software program code that implements the functions of any of the above embodiments is stored, and a computer (or CPU or MPU) of the system or device reads and executes the program code stored in the storage medium.

[0152] In this case, the program code itself read from the storage medium can realize the function of any one of the above-mentioned embodiments, and thus the program code and the storage medium storing the program code constitute a part of the present invention.

[0153] The storage medium embodiments for providing the program code include a floppy disk, a hard disk, a magneto-optical disk, an optical disk (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RYM, DVD-RW, DVD+RW), a magnetic tape, a non-volatile memory card, and a ROM. Alternatively, the program code can be downloaded from a server computer via a communication network.

[0154] In addition, it should be clear that the functions of any of the above embodiments can be implemented not only by executing the program code read by the computer, but also by enabling an operating system operating on the computer to complete part or all of the actual operations based on instructions from the program code.

[0155] In addition, it can be understood that the program code read from the storage medium is written to a memory provided in an expansion board inserted into the computer or written to a memory provided in an expansion unit connected to the computer, and then based on the instructions of the program code, a CPU installed on the expansion board or the expansion unit is enabled to perform part or all of the actual operations, thereby realizing the functions of any of the above-mentioned embodiments.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for online data reorganization of a time series database, characterized in that: The method is as follows: Use ordered iterators to classify data by label, filter invalid data, read data, and sort data by time; Write the data read by the ordered iterator into the data file and replace the original file; Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

2. The method for online data reorganization of a time series database according to claim 1, characterized in that: The structure of the BlockSpan queue is as follows: struct BlockSpan{ BlockItem*block_item; uint32_t start_row; uint32_t row_count; }; BlockItem is the metadata of a data block, including labels, number of rows, maximum / minimum timestamps, and information about whether there is out-of-order data. start_row is the starting row number of the data to be returned by the iterator; row_count is the cumulative number of rows returned starting from start_row.

3. The method for online data reorganization of a time series database according to claim 1, characterized in that: An ordered iterator has the following features: ①Support deletion of data filtering; ② After adding columns, historical data can be completed; ③Support historical data filtering after deleting columns; ④Support the reading and conversion of historical data after the field type is changed; ⑤Support data sorting.

4. The method for online data reorganization of a time series database according to claim 1, characterized in that: The ordered iterator data reading process is as follows: S1. Get all corresponding BlockItems according to Tags; S2. Check whether BlockItems are out of order: S201. If BlockItems are not marked as out of order, traverse each BlockItem, obtain the maximum and minimum timestamps, and check whether the timestamps are within the specified time period: ①If it is not within the time range, skip it and continue traversing; ② If it is within the time range, traverse each row in the corresponding BlockItem; S202, if the BlockItems are marked out of order, execute step S203; S203, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. Multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems are traversed and recorded in interval_block_map; S204, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64> >In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems; S205, traverse each group of timestamps of intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals; S206. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing the ordered BlockItems, and the process goes to step S201 to continue execution.

5. The method for online data reorganization of a time series database according to claim 4, characterized in that: Traversing each row in the corresponding BlockItem is as follows: Traverse to the first row of data that is within the time range and not marked for deletion, record the BlockItem in the BlockSpan queue, and use the corresponding row number as the starting row number start_row, continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, the consecutive rows end, update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, repeat until all data traversal is completed, and return the BlockSpans after filtering and sorting.

6. The method for online data reorganization of a time series database according to claim 1, characterized in that: Write the data read by the ordered iterator into the data file and replace the original file as follows: (1) Create a temporary data file partition, and create a table and column file in the temporary data file partition according to the latest metadata; (2) Traverse the metadata of the temporary partition column and read the corresponding column of the data returned by the ordered iterator: If there is no corresponding column in the read data, it means that the corresponding column is a newly added column, and the corresponding column of the temporary partition is filled with NULL; If the column type in the read data is inconsistent with the metadata of the temporary partition column, it means that the corresponding column type is modified, the data in the ordered iterator is read out according to the original type, and the original type is converted to the latest column type in the temporary partition and written into the temporary partition column file; (3) After all column files are traversed and written, try to lock the corresponding partition. After successfully locking, use the temporary partition to replace the old partition, clear the cache, and reload the partition and metadata; (4) After the replacement is completed, clean up the remaining files.

7. An online data reorganization system for a time series database, characterized in that: The system includes: The reading module is used to classify data by label, filter invalid data, read data, and sort data by time through an ordered iterator; The data writing and replacement module is used to write the data read by the ordered iterator into the data file and replace the original file. Specifically, the column type is converted according to the metadata of the original data and the latest version of the metadata and then written. NULL is used to fill the columns that are not newly added in the original data. At the same time, the column descriptions that do not exist in the latest version of the metadata have been deleted by comparing the original data and do not need to be written into the reorganized data file again. Among them, the ordered iterator queries the valid data of the entities represented by a set of Tags within a specified time period, and uses a BlockSpan queue that records the data block address, starting row number and cumulative number of rows to return the queried data.

8. The time series database online data reorganization system according to claim 7, characterized in that: The working process of the reading module is as follows: S1. Get all corresponding BlockItems according to Tags; S2. Check whether BlockItems are out of order: S2A. If BlockItems is not marked as out of order, execute steps S2A1 to S2A3: S2A1. Traverse each BlockItem, obtain the maximum and minimum timestamps of the BlockItem, and check whether the timestamps are within the specified time period: If it is not within the time range, skip it and continue traversing; If it is within the time range, then traverse each row in the corresponding BlockItem, i.e., execute step S2A2; S2A2. First, traverse to the first row of data that is within the time range and not marked for deletion, record BlockItem in BlockSpan, and use the corresponding row number as the starting row number start_row. Continue traversing and recording the cumulative number of consecutive rows until the current BlockItem traversal is completed or any row of data is not within the time range or any row of data is marked for deletion, and the consecutive rows end. Update the cumulative number of consecutive rows row_count in BlockSpan, add BlockSpan to the BlockSpans queue, and repeat until all data traversal is completed. S2A3, return the BlockSpans after filtering and sorting; S2B. If the BlockItems are marked out of order, execute steps S2B1 to S2B: S2B1, initialize a map <pair<timestamp64,timestamp64> ,vector<Block Item*> >interval_block_map, used to record all BlockItems and their start and end timestamps. Multiple BlockItems corresponding to the same set of timestamps should be appended to the vector, and all BlockItems are traversed and recorded in interval_block_map; S2B2, pair of interval_block_map<timestamp64,timestamp64> Sort and store the sort results in vector <vector <timestamp64> >In intervals, intervals ultimately store an ordered queue of timestamps of all BlockItems; S2B3, traverse each group of timestamps in intervals, and find the vector with the corresponding group timestamp as the key in interval_block_map<BlockItem*> , append the key and the corresponding value to the vector <pair<pair<timestamp64,timestamp64> ,vector<BlockItem*> >>bloc k_items, continue to traverse the next set of timestamps of intervals; S2B4. After the intervals traversal is completed, the complete and ordered BlockItems are recorded in block_items. At this time, the problem converges to processing ordered BlockItems, and jumps to step S2A to continue execution.

9. An electronic device, characterized in that: include: memory and at least one processor; Wherein, the memory stores a computer program; The at least one processor executes the computer program stored in the memory, so that the at least one processor executes the online data reorganization method for a time series database according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which can be executed by a processor to implement the online data reorganization method for a time series database according to any one of claims 1 to 6.