Disorder processing method and system for time series data storage engine
By configuring multiple ZdataPoint objects and ZBlocks in the timing data storage engine, dividing timing memory blocks and out-of-order memory blocks, and reserving storage space based on the out-of-order data probability, the problem of performance degradation in the timing data storage engine when processing out-of-order data is solved, and efficient out-of-order data storage and processing is achieved.
Patent Information
- Application Number
- CN202310096745.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-02-07
AI Technical Summary
The performance of the timing data storage engine degrades when processing out-of-order data, and discarding out-of-order data will affect the accuracy of the data.
A time-sequential processing method of a time-sequential data storage engine is adopted. By configuring multiple ZdataPoint objects and ZBlocks in the memory area, dividing timing memory blocks and out-of-sequential memory blocks, and reserving storage space based on the out-of-sequential data probability, the effective storage and processing of sequential and out-of-sequential data is realized.
This method reduces memory management consumption, improves memory utilization, can store and use out-of-order data while ensuring the performance of the storage engine, and expands the functions of the storage engine.
Smart Images

Figure CN116166715B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of time series databases, and in particular to a method and system for processing out-of-order data in a time series data storage engine. Background Art
[0002] Time series data refers to time series data, which is a data column recorded in chronological order. Each data in the same data column must be of the same caliber and must be comparable. Time series data can be a period number or a point in time. Time series data usually has the characteristics of fast writing speed (writing demand is greater than reading demand), large writing capacity, and high attention to recent data.
[0003] In the actual scenario of time series data generation, the collected data will first be transmitted to the decoding processing program and then written into the database. In this process, due to network delays and other reasons, the order in which the data is written into the database may be inconsistent with the order in which the data is generated. In addition, time series data may be retransmitted due to data errors, which will also cause time series data disorder.
[0004] The current pain points are:
[0005] (1) The occurrence of disordered data is uncertain;
[0006] (2) The processing method of reordering time series data through background tasks will affect performance;
[0007] (3) Discarding out-of-order data will affect the accuracy of the data.
[0008] In a time series data storage engine, the processing of out-of-order data will reduce the performance of the storage engine. How to ensure the performance of the storage engine while storing and using out-of-order data is a technical problem that needs to be solved. Summary of the invention
[0009] The technical task of the present invention is to address the above shortcomings and provide a disorder processing method and system for a time series data storage engine to solve the technical problem of how to ensure the performance of the storage engine while storing and using disordered data.
[0010] In a first aspect, the present invention provides a method for out-of-order processing of a time series data storage engine, wherein the storage engine includes a memory area and a disk area, and the time series data includes sequential data and out-of-order data. The method includes the following steps:
[0011] There are multiple different ZdataPoint objects configured in the memory area. Each ZdataPoint object contains multiple ZBlocks. Each ZBlock is used to control the storage of time series data within a period of time. The corresponding time periods between ZBlocks do not overlap.
[0012] Allocate a continuous memory area of a fixed size in the memory area as a cache, divide the cache into a sequential memory block for storing sequential data and a random memory block for storing random data, the memory space of the sequential memory block is larger than the memory space of the random memory block, and reserve a sequential space for storing sequential data and a random space for storing random data in the sequential memory block based on the probability of random data;
[0013] For the time series data collected by the collection point, divide the time series data into corresponding ZdataPoint objects based on the source of the collection point, apply for a ZBlock, and write the time series data into the time series memory block corresponding to the ZBlock through the ZBlock;
[0014] For a ZBlock that is in a full state, the time series data in the time series memory block corresponding to the ZBlock is written to the disk area through the disk writing thread. After the disk writing is successful, the information record of the time series data is added to the file index part;
[0015] Among them, applying for a ZBlock and writing the timing data into the timing memory block through ZBlock follows the following principles:
[0016] Write sequential data into sequential memory blocks and write out-of-order data into out-of-order memory blocks;
[0017] If the current random memory block is full, and the newly generated random data is within the time period corresponding to the current ZBlock, apply for a random memory block through ZBlock and write the newly generated random data into the random memory block;
[0018] If the current sequential memory block is full, apply for a new ZBlock, and write the sequential data into a new sequential memory block corresponding to the new ZBlock through the new ZBlock;
[0019] If the current cache is full, a predetermined number of ZBlocks are recycled based on a preconfigured cache recycling mechanism, and by recycling the predetermined number of ZBlocks, the ZBlocks are deleted and the corresponding sequential memory blocks are released.
[0020] Preferably, in the memory area, the ZBlocks are sorted in order of time periods to form a ZBlock array.
[0021] Preferably, when applying for a ZBlock and writing the timing data into the timing memory block corresponding to the ZBlock through the ZBlock, the ZBlock has eight corresponding states, namely:
[0022] Initial state: the state of ZBlock when it is just created;
[0023] Writing state: when the sequential memory space corresponding to the ZBlock is not used up;
[0024] Unwritable state: when the sequential memory space corresponding to ZBlock is exhausted;
[0025] Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state;
[0026] Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state.
[0027] Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status;
[0028] Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk;
[0029] End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted; if it is in the state of waiting to be written to the disk again, the corresponding Block data block in the disk area is updated and then the ZBlock is deleted;
[0030] The cache recycling mechanism is as follows: when the remaining space in the cache is insufficient to allocate new time series data blocks, cache recycling is triggered, and a predetermined number of old ZBlocks are recycled based on a time period. Before recycling a ZBlock, the ZBlock status is judged. For a ZBlock that is in an unwritable state or is to be rewritten to disk, the time series data in the time series memory block corresponding to the ZBlock is written to disk, and then the ZBlock is recycled.
[0031] Preferably, the out-of-order data includes:
[0032] Memory disordered data, where the ZBlock corresponding to the timestamp of the disordered data still exists in the memory area;
[0033] Disk out-of-order data, where the ZBlock corresponding to the timestamp of the out-of-order data is no longer in the memory area, but exists in a Block in the corresponding disk area;
[0034] For out-of-order data in memory, execute as follows:
[0035] Analyze the ZBlock status corresponding to its timestamp;
[0036] When the ZBlock status is writing, the out-of-order data is directly written into the reserved out-of-order space;
[0037] When the ZBlock status is not writable, first set the ZBlock status to temporary write, then write the out-of-order data to the reserved out-of-order space, and finally set the ZBlock status to not writable;
[0038] When the ZBlock status is flushed, the out-of-order data is written to the reserved out-of-order space, and the ZBlock status is set to pending re-flush;
[0039] For the disk disordered data, the disk disordered data is directly discarded.
[0040] Preferably, the sequential memory block is 64K, the out-of-order memory block is 1K, and each ZBlock is allocated with a corresponding sequential memory block;
[0041] The number of time series records that can be stored in a time series memory block is blockRows:
[0042]
[0043] n random spaces are reserved in the sequential memory block, where
[0044]
[0045] The sequential space in the sequential memory block is used to store sequential data. Sequential data refers to the time stamp of the sequential data when it is inserted into the ZBlock that is greater than the time stamp of all the sequential data stored in the ZBlock. The reserved space for sequential data is:
[0046] blockRows-n
[0047] The probability of out-of-order data is calculated based on the adaptive method, including the following steps:
[0048] When the cache recycles a ZBlock, extract the proportion of out-of-order data contained in the current ZBlock. If the number of out-of-order data in the ZBlock is 0, count the proportion together with the next ZBlock. Set the order of the extracted proportions to R1, R2, R3...Rn, and recalculate the value of the out-of-order data probability based on the collected value of the out-of-order data.
[0049] The value of disorderRatio calculated for the i-th time is Di:
[0050]
[0051] The proportion of out-of-order data = the number of out-of-order data / the total amount of data stored in ZBlock.
[0052] In a second aspect, the present invention provides a disorder processing system for a time series data storage engine, wherein the storage engine includes a memory area and a disk area, the time series data includes sequential data and disordered data, and the system is used to execute the disorder processing method for the time series data storage engine as described in any one of the first aspects, and the system includes a configuration module and a writing module;
[0053] The configuration module is used to perform the following:
[0054] There are multiple different ZdataPoint objects configured in the memory area. Each ZdataPoint object contains multiple ZBlocks. Each ZBlock is used to control the storage of time series data within a period of time. The corresponding time periods between ZBlocks do not overlap.
[0055] Allocate a continuous memory area of a fixed size in the memory area as a cache, divide the cache into a sequential memory block for storing sequential data and a random memory block for storing random data, the memory space of the sequential memory block is larger than the memory space of the random memory block, and reserve a sequential space for storing sequential data and a random space for storing random data in the sequential memory block based on the probability of random data;
[0056] The write module is used to perform the following:
[0057] For the time series data collected by the collection point, divide the time series data into corresponding ZdataPoint objects based on the source of the collection point, apply for a ZBlock, and write the time series data into the time series memory block corresponding to the ZBlock through the ZBlock;
[0058] For a ZBlock that is in a full state, the time series data in the time series memory block corresponding to the ZBlock is written to the disk area through the disk writing thread. After the disk writing is successful, the information record of the time series data is added to the file index part;
[0059] The writing module follows the following principles:
[0060] Write sequential data into sequential memory blocks and write out-of-order data into out-of-order memory blocks;
[0061] If the current random memory block is full, and the newly generated random data is within the time period corresponding to the current ZBlock, apply for a random memory block through ZBlock and write the newly generated random data into the random memory block;
[0062] If the current sequential memory block is full, apply for a new ZBlock, and write the sequential data into a new sequential memory block corresponding to the new ZBlock through the new ZBlock;
[0063] If the current cache is full, a predetermined number of ZBlocks are recycled based on a preconfigured cache recycling mechanism, and by recycling the predetermined number of ZBlocks, the ZBlocks are deleted and the corresponding sequential memory blocks are released.
[0064] Preferably, in the memory area, the ZBlocks are sorted in order of time periods to form a ZBlock array.
[0065] Preferably, the ZBlock corresponds to eight states, namely:
[0066] Initial state: the state of ZBlock when it is just created;
[0067] Writing state: when the sequential memory space corresponding to the ZBlock is not used up;
[0068] Unwritable state: when the sequential memory space corresponding to ZBlock is exhausted;
[0069] Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state;
[0070] Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state.
[0071] Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status;
[0072] Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk;
[0073] End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted; if it is in the state of waiting to be written to the disk again, the corresponding Block data block in the disk area is updated and then the ZBlock is deleted;
[0074] The cache recycling mechanism is as follows: when the remaining space in the cache is insufficient to allocate new time series data blocks, cache recycling is triggered, and a predetermined number of old ZBlocks are recycled based on a time period. Before recycling a ZBlock, the ZBlock status is judged. For a ZBlock that is in an unwritable state or is to be rewritten to disk, the time series data in the time series memory block corresponding to the ZBlock is written to disk, and then the ZBlock is recycled.
[0075] Preferably, the out-of-order data includes:
[0076] Memory disordered data, where the ZBlock corresponding to the timestamp of the disordered data still exists in the memory area;
[0077] Disk out-of-order data, where the ZBlock corresponding to the timestamp of the out-of-order data is no longer in the memory area, but exists in a Block in the corresponding disk area;
[0078] For out-of-order data in memory, the write module is used to perform the following:
[0079] Analyze the ZBlock status corresponding to its timestamp;
[0080] When the ZBlock status is writing, the out-of-order data is directly written into the reserved out-of-order space;
[0081] When the ZBlock status is not writable, first set the ZBlock status to temporary write, then write the out-of-order data to the reserved out-of-order space, and finally set the ZBlock status to not writable;
[0082] When the ZBlock status is flushed, the out-of-order data is written to the reserved out-of-order space, and the ZBlock status is set to pending re-flush;
[0083] For the disk disordered data, the writing module is used to perform the following: directly discarding the disk disordered data.
[0084] Preferably, the sequential memory block is 64K, the out-of-order memory block is 1K, and each ZBlock is allocated with a corresponding sequential memory block;
[0085] The number of time series records that can be stored in a time series memory block is blockRows:
[0086]
[0087] n random spaces are reserved in the sequential memory block, where
[0088]
[0089] The sequential space in the sequential memory block is used to store sequential data. Sequential data refers to the time stamp of the sequential data when it is inserted into the ZBlock that is greater than the time stamp of all the sequential data stored in the ZBlock. The reserved space for sequential data is:
[0090] blockRows-n
[0091] The configuration module is used to calculate the probability of out-of-order data based on an adaptive method, and includes the following steps:
[0092] When the cache recycles a ZBlock, extract the proportion of out-of-order data contained in the current ZBlock. If the number of out-of-order data in the ZBlock is 0, count the proportion together with the next ZBlock. Set the order of the extracted proportions to R1, R2, R3...Rn, and recalculate the value of the out-of-order data probability based on the collected value of the out-of-order data.
[0093] The value of disorderRatio calculated for the i-th time is Di:
[0094]
[0095] The proportion of out-of-order data = the number of out-of-order data / the total amount of data stored in ZBlock.
[0096] The out-of-order processing method and system of the time series data storage engine of the present invention have the following advantages:
[0097] 1. By using a cache mechanism of fixed-size memory blocks and an adaptive out-of-order data probability algorithm, the consumption of memory management is reduced and memory utilization is improved;
[0098] 2. It enables out-of-order data to be stored in the time series data storage engine, which expands the functionality of the storage engine and has little impact on the performance of the storage engine. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0100] The present invention is further described below in conjunction with the accompanying drawings.
[0101] Figure 1 This is a flowchart of the out-of-order processing method of the time series data storage engine in Example 1;
[0102] Figure 2 This is an overall structural diagram of the time series data storage engine in the disorder processing method of the time series data storage engine in Example 1;
[0103] Figure 3 This is a structural diagram of ZBlock in the disorder processing method of the time series data storage engine in Example 1;
[0104] Figure 4 This is a state diagram of ZBlock in the out-of-order processing method of the time series data storage engine in Example 1. DETAILED DESCRIPTION
[0105] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it. However, the embodiments are not intended to limit the present invention. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments may be combined with each other.
[0106] Embodiments of the present invention provide a method and system for out-of-order processing of a time series data storage engine, which are used to solve the technical problem of how to ensure the performance of the storage engine while storing and using out-of-order data.
[0107] Embodiment 1:
[0108] The present invention discloses a method for processing disorder of a time series data storage engine, wherein the storage engine includes a memory area and a disk area, and the time series data includes sequential data and disordered data. Figure 1 As shown, the method comprises the following steps:
[0109] S100, multiple different ZdataPoint objects are configured in the memory area, each ZdataPoint object contains multiple ZBlocks, each ZBlock is used to regulate the storage of time series data within a period of time, and the corresponding time periods between ZBlocks do not overlap;
[0110] Allocate a continuous memory area of a fixed size in the memory area as a cache, divide the cache into a sequential memory block for storing sequential data and a random memory block for storing random data, the memory space of the sequential memory block is larger than the memory space of the random memory block, and reserve a sequential space for storing sequential data and a random space for storing random data in the sequential memory block based on the probability of random data;
[0111] S200, for the time series data collected by the collection point, divide the time series data into corresponding ZdataPoint objects based on the source of the collection point, apply for a ZBlock, and write the time series data into a time series memory block corresponding to the ZBlock through the ZBlock;
[0112] For a ZBlock that is in a full state, the time series data in the time series memory block corresponding to the ZBlock is written to the disk area through the disk writing thread. After the disk writing is successful, the information record of the time series data is added to the file index part.
[0113] The structure of the storage engine in this embodiment is as follows Figure 2 As shown in the figure, the time series data storage engine stores time series data collected from many data points. The collected time series data is divided into different ZdataPoint objects according to different sources of data points. The ZdataPoint object contains many ZBlocks. The structure of ZBlock is as follows: Figure 3 As shown. Each ZBlock stores time series data within a period of time, and there is no overlap between the time periods of ZBlocks. Each ZBlock has a fixed-size memory space of 64K. When it is full, a new ZBlock object needs to be applied to store new time series data. ZBlocks are sorted in the order of time periods to form a ZBlock array.
[0114] The background disk-push thread will flush the time series data in the full ZBlock to the disk. After the flush is successful, the information record of this time series data block will be appended to the file index.
[0115] The cache is a fixed-size continuous memory area allocated in the memory, which is only used to store time series data records. In this embodiment, the cache is divided into 64K time series memory blocks and 1K fixed-size random memory blocks. Each ZBlock object is allocated a 64K fixed-size time series memory block.
[0116] Assume that the number of bytes occupied by a time series data record is rowLength, and the probability of disordered data generation is disorderRatio, which is initially set to 0.001.
[0117] The number of time series records that can be stored in a time series memory block is blockRows:
[0118]
[0119] In the 64K memory block, n spaces are reserved to store out-of-order data.
[0120]
[0121] The orderedtsdata space in the time series memory block is used to store sequential data. Sequential data refers to the time series data whose timestamp when inserted into ZBlock is greater than the timestamp of all time series data stored in ZBlock. The reserved space for this kind of time series data is:
[0122] blockRows-n
[0123] When the reserved n disordered data spaces are used up, but there is new disordered data to be stored, ZBlock applies for 1K disordered space DisOrdered to store these new disordered data.
[0124] The status of ZBlock is as follows Figure 4 As shown, there are eight states:
[0125] Initial state: the state of ZBlock when it is just created;
[0126] Writing state: when the sequential memory space corresponding to the ZBlock is not used up;
[0127] Unwritable state: when the sequential memory space corresponding to ZBlock is exhausted;
[0128] Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state;
[0129] Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state.
[0130] Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status;
[0131] Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk;
[0132] End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted; if it is in the state of waiting to be written to the disk again, the corresponding Block data block in the disk area is updated and then the ZBlock is deleted.
[0133] In this embodiment, the disordered data probability disorderRatio is calculated based on an adaptive method. When the cache reclaims a ZBlock, the proportion of disordered data contained in the current ZBlock is extracted (the number of disordered data divided by the total amount of data stored in the ZBlock). If the number of disordered data in the ZBlock is 0, the proportion is counted together with the next ZBlock. Assume that the order of the extracted proportions is R1, R2, R3...Rn. This embodiment will recalculate the value of disorderRatio based on the collected value of the disordered data.
[0134] The ratio array only stores the ratio values of the most recent 100 random data. Only the most recent 100 ratio values participate in the calculation.
[0135] Assume that the value of disorderRatio calculated for the i-th time is Di:
[0136]
[0137] In step S200 of this embodiment, a ZBlock is applied, and the timing data is written into the timing memory block through the ZBlock, following the following principles:
[0138] (1) Write sequential data into sequential memory blocks, and write out-of-order data into out-of-order memory blocks;
[0139] (2) If the current out-of-order memory block is full and the newly generated out-of-order data is within the time period corresponding to the current ZBlock, apply for a new out-of-order memory block through ZBlock and write the newly generated out-of-order data into the said out-of-order memory block;
[0140] (3) If the current sequential memory block is full, apply for a new ZBlock, and write the sequential data into a new sequential memory block corresponding to the new ZBlock through the new ZBlock;
[0141] (4) If the current cache is full, a predetermined number of ZBlocks are recycled based on a preconfigured cache recycling mechanism. By recycling the predetermined number of ZBlocks, the ZBlocks are deleted and the corresponding sequential memory blocks are released.
[0142] In this embodiment, the cache size is a fixed value determined when the program starts. When the cache cannot allocate new 64k data blocks, some older ZBlock objects need to be released.
[0143] The triggering condition for cache recycling is: when a new data block is requested from the cache and insufficient cache space is reported. The cache recycling algorithm is: recycle the ten oldest ZBlocks, and before recycling, check the state of the ZBlock. If it is in a non-writable state or a state to be rewritten to disk, the time series data in this ZBblock object needs to be written to disk before recycling.
[0144] In this embodiment, out-of-order data is divided into two categories: memory out-of-order data and disk out-of-order data. Memory out-of-order data refers to the ZBlock object corresponding to the timestamp of the out-of-order data still existing in the memory; disk out-of-order data refers to the ZBlock object corresponding to the timestamp of the out-of-order data is no longer in the memory, but corresponds to a certain Block block on the disk.
[0145] For the disordered time series data in memory, this embodiment performs the following: Analyze the ZBlock status corresponding to its timestamp. When the ZBlock status is writing, directly write the disordered data into the space reserved for disordered data; when the ZBlock status is not writable, first set the ZBlock status to temporary writing, then write the disordered data into the disordered space of the ZBlock, and finally set the ZBlock status to not writable; when the ZBlock status is written to disk, after writing the disordered data into the disordered space of the ZBlock, set the ZBlock status to waiting to be written to disk again.
[0146] In this embodiment, out-of-order time series data on the disk is directly discarded.
[0147] This embodiment analyzes the real disorder scenario of time series data and proposes a method for discarding disordered data with a long time to reduce processing consumption. This enables the storage of disordered data in the time series data storage engine, expands the function of the storage engine, and has little impact on the performance of the storage engine. By using a cache mechanism of fixed-size memory blocks and an adaptive disordered data probability algorithm, the consumption of memory management is reduced and memory utilization is improved.
[0148] Embodiment 2:
[0149] The present invention discloses an out-of-order processing system for a time series data storage engine. The storage engine includes a memory area and a disk area. The time series data includes sequential data and out-of-order data. The system includes a configuration module and a write module. The system is used to execute the method disclosed in Example 1 to realize processing of out-of-order data in the storage engine.
[0150] The configuration module is used to perform the following:
[0151] (1) Multiple different ZdataPoint objects are configured in the memory area. Each ZdataPoint object contains multiple ZBlocks. Each ZBlock is used to control the storage of time series data within a period of time. The corresponding time periods between ZBlocks do not overlap;
[0152] (2) Allocate a continuous memory area of a fixed size in the memory area as a cache, and divide the cache into a sequential memory block for storing sequential data and a random memory block for storing random data. The memory space of the sequential memory block is larger than the memory space of the random memory block. Based on the probability of random data, a sequential space for storing sequential data and a random space for storing random data are reserved in the sequential memory block.
[0153] In this embodiment, the structure of the storage engine is as follows: Figure 2 As shown in the figure, the time series data storage engine stores time series data collected from many data points. The collected time series data is divided into different ZdataPoint objects according to different sources of data points. The ZdataPoint object contains many ZBlocks. The structure of ZBlock is as follows: Figure 3 As shown. Each ZBlock stores time series data within a period of time, and there is no overlap between the time periods of ZBlocks. Each ZBlock has a fixed-size memory space of 64K. When it is full, a new ZBlock object needs to be applied to store new time series data. ZBlocks are sorted in the order of time periods to form a ZBlock array.
[0154] The background disk-push thread will flush the time series data in the full ZBlock to the disk. After the flush is successful, the information record of this time series data block will be appended to the file index.
[0155] The cache is a fixed-size continuous memory area allocated in the memory, which is only used to store time series data records. In this embodiment, the cache is divided into 64K time series memory blocks and 1K fixed-size random memory blocks. Each ZBlock object is allocated a 64K fixed-size time series memory block.
[0156] Assume that the number of bytes occupied by a time series data record is rowLength, and the probability of disordered data generation is disorderRatio, which is initially set to 0.001.
[0157] The number of time series records that can be stored in a time series memory block is blockRows:
[0158]
[0159] In the 64K memory block, n spaces are reserved to store out-of-order data.
[0160]
[0161] The orderedtsdata space in the time series memory block is used to store sequential data. Sequential data refers to the time series data whose timestamp when inserted into ZBlock is greater than the timestamp of all time series data stored in ZBlock. The reserved space for this kind of time series data is:
[0162] blockRows-n
[0163] When the reserved n disordered data spaces are used up, but there is new disordered data to be stored, ZBlock applies for 1K disordered space DisOrdered to store these new disordered data.
[0164] The status of ZBlock is as follows Figure 4 As shown, there are eight states:
[0165] Initial state: the state of ZBlock when it is just created;
[0166] Writing state: when the sequential memory space corresponding to the ZBlock is not used up;
[0167] Unwritable state: when the sequential memory space corresponding to ZBlock is exhausted;
[0168] Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state;
[0169] Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state.
[0170] Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status;
[0171] Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk;
[0172] End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted; if it is in the state of waiting to be written to the disk again, the corresponding Block data block in the disk area is updated and then the ZBlock is deleted.
[0173] In this embodiment, the disordered data probability disorderRatio is calculated based on an adaptive method. When the cache reclaims a ZBlock, the proportion of disordered data contained in the current ZBlock is extracted (the number of disordered data divided by the total amount of data stored in the ZBlock). If the number of disordered data in the ZBlock is 0, the proportion is counted together with the next ZBlock. Assume that the order of the extracted proportions is R1, R2, R3,,,,,Rn. This embodiment recalculates the value of disorderRatio based on the collected value of the disordered data.
[0174] The ratio array only stores the ratio values of the most recent 100 random data. Only the most recent 100 ratio values participate in the calculation.
[0175] Assume that the value of disorderRatio calculated for the i-th time is Di:
[0176]
[0177] The write module is used to perform the following:
[0178] (1) For the time series data collected by the collection point, divide the time series data into corresponding ZdataPoint objects based on the source of the collection point, apply for a ZBlock, and write the time series data into the time series memory block corresponding to the ZBlock through the ZBlock;
[0179] (2) For a ZBlock that is in a full state, the time series data in the time series memory block corresponding to the ZBlock is written to the disk area through the disk writing thread. After the disk writing is successful, the information record of the time series data is added to the file index part.
[0180] In this embodiment, the writing module follows the following principles:
[0181] (1) Write sequential data into sequential memory blocks, and write out-of-order data into out-of-order memory blocks;
[0182] (2) If the current out-of-order memory block is full and the newly generated out-of-order data is within the time period corresponding to the current ZBlock, apply for a new out-of-order memory block through ZBlock and write the newly generated out-of-order data into the said out-of-order memory block;
[0183] (3) If the current sequential memory block is full, apply for a new ZBlock, and write the sequential data into a new sequential memory block corresponding to the new ZBlock through the new ZBlock;
[0184] (4) If the current cache is full, a predetermined number of ZBlocks are recycled based on a preconfigured cache recycling mechanism. By recycling the predetermined number of ZBlocks, the ZBlocks are deleted and the corresponding sequential memory blocks are released.
[0185] In this embodiment, the cache size is a fixed value determined when the program is started. When the cache cannot allocate new 64k data blocks, some older ZBlock objects need to be released.
[0186] The triggering condition for cache recycling is: when a new data block is requested from the cache and insufficient cache space is reported. The cache recycling algorithm is: recycle the ten oldest ZBlocks, and before recycling, check the state of the ZBlock. If it is in a non-writable state or a state to be rewritten to disk, the time series data in this ZBblock object needs to be written to disk before recycling.
[0187] In this embodiment, out-of-order data is divided into two categories, namely memory out-of-order data and disk out-of-order data. Memory out-of-order data refers to the ZBlock object corresponding to the timestamp of the out-of-order data still existing in the memory; disk out-of-order data refers to the ZBlock object corresponding to the timestamp of the out-of-order data is no longer in the memory, but corresponds to a certain Block block on the disk.
[0188] For the disordered time series data in memory, the writing module of this embodiment is used to perform the following: Analyze the ZBlock status corresponding to its timestamp. When the ZBlock status is writing, directly write the disordered data into the space reserved for disordered data; when the ZBlock status is not writable, first set the ZBlock status to temporary writing, then write the disordered data into the disordered space of ZBlock, and finally set the ZBlock status to not writable; when the ZBlock status is written to disk, after writing the disordered data into the disordered space of ZBlock, set the ZBlock status to waiting to be written to disk again.
[0189] For out-of-order time series data on the disk, the write module in this embodiment is used to directly discard the out-of-order data on the disk.
[0190] The present invention is shown and described in detail above through the accompanying drawings and preferred embodiments. However, the present invention is not limited to these disclosed embodiments. Based on the above multiple embodiments, those skilled in the art can know that the code review methods in the above different embodiments can be combined to obtain more embodiments of the present invention, and these embodiments are also within the protection scope of the present invention.
Claims
1. A method for out-of-order processing of a time-series data storage engine, characterized in that, the storage engine includes a memory area and a disk area, the time-series data includes sequential data and out-of-order data, and the method includes the following steps: Multiple different ZdataPoint objects are configured in the memory area, each ZdataPoint object contains multiple ZBlocks, each ZBlock is used to control the storage of time-series data for a period of time, and the corresponding time periods between ZBlocks do not intersect; A continuously sized fixed memory area is allocated in the memory area as a cache. In the cache, a time-series memory block for storing time-series data and an out-of-order memory block for storing out-of-order data are divided. The memory space of the time-series memory block is larger than that of the out-of-order memory block, and based on the out-of-order data probability, a sequential space for storing sequential data and an out-of-order space for storing out-of-order data are reserved in the time-series memory block; For the time-series data collected by the collection point, the time-series data is divided into the corresponding ZdataPoint object based on the collection point source, a ZBlock is applied, and the time-series data is written into the time-series memory block corresponding to the ZBlock through the ZBlock; For a ZBlock in a full-written state, the time-series data in the time-series memory block corresponding to the ZBlock is written to the disk area through a disk-writing thread. After the disk writing is successful, an information record of the time-series data is appended to the file index section; Among them, applying for a ZBlock and writing the time-series data into the time-series memory block through the ZBlock follows the following principles: Write sequential data into the sequential space and out-of-order data into the out-of-order space; If the current out-of-order space is full and the newly generated out-of-order data is within the time period corresponding to the current ZBlock, apply for an out-of-order memory block through the ZBlock and write the newly generated out-of-order data into the out-of-order memory block; If the current sequential space is full, apply for a new ZBlock and write the time-series data into the new time-series memory block corresponding to the new ZBlock through the new ZBlock; If the current cache is full, based on a pre-configured cache recycling mechanism, recycle a predetermined number of ZBlocks. By recycling a predetermined number of ZBlocks, delete the ZBlocks and release the corresponding time-series memory blocks.
2. The method for out-of-order processing of a time-series data storage engine according to claim 1, characterized in that, In the memory area, the ZBlocks are sorted in sequence according to the time period to form a ZBlock array.
3. The method for out-of-order processing of a time-series data storage engine according to claim 1 or 2, characterized in that, When applying for a ZBlock and writing the time-series data into the time-series memory block corresponding to the ZBlock, the ZBlock corresponds to eight states, which are respectively: Initial state: The state when the ZBlock is just created; Writing state: When the sequential space corresponding to the ZBlock has not been used up; Unwritable state: When the sequential space corresponding to the ZBlock is exhausted; Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state; Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state. Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status; Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk; End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted. If it is in the state of waiting to be re-written to disk, the corresponding Block data block in the disk area is updated, and then the ZBlock is deleted; The cache recycling mechanism is as follows: when the remaining space in the cache is insufficient to allocate new time series data blocks, cache recycling is triggered, and a predetermined number of old ZBlocks are recycled based on a time period. Before recycling a ZBlock, the ZBlock status is judged. For a ZBlock that is in an unwritable state or is to be rewritten to disk, the time series data in the time series memory block corresponding to the ZBlock is written to disk, and then the ZBlock is recycled.
4. The out-of-order processing method of the time series data storage engine according to claim 3, It is characterized in that The out-of-order data includes: Memory disordered data, where the ZBlock corresponding to the timestamp of the disordered data still exists in the memory area; Disk out-of-order data, where the ZBlock corresponding to the timestamp of the out-of-order data is no longer in the memory area, but exists in a Block in the corresponding disk area; For out-of-order data in memory, execute as follows: Analyze the ZBlock status corresponding to its timestamp; When the ZBlock status is writing, the out-of-order data is directly written into the reserved out-of-order space; When the ZBlock status is not writable, first set the ZBlock status to temporary write, then write the out-of-order data to the reserved out-of-order space, and finally set the ZBlock status to not writable; When the ZBlock status is flushed, the out-of-order data is written to the reserved out-of-order space, and the ZBlock status is set to pending re-flush; For the disk disordered data, the disk disordered data is directly discarded.
5. The out-of-order processing method of the time series data storage engine according to claim 3, It is characterized in that The sequential memory block is 64K, the out-of-order memory block is 1K, and each ZBlock is allocated a corresponding sequential memory block; The number of time series records that can be stored in a time series memory block is blockRows: rowLength indicates the number of bytes occupied by a time series data record; n random spaces are reserved in the sequential memory block, where The sequential space in the sequential memory block is used to store sequential data. Sequential data refers to the time stamp of the sequential data when it is inserted into the ZBlock that is greater than the time stamp of all the sequential data stored in the ZBlock. The reserved space for sequential data is: blockRows-n The probability of out-of-order data is calculated based on the adaptive method, including the following steps: When the cache recycles a ZBlock, extract the proportion of out-of-order data contained in the current ZBlock. If the number of out-of-order data in the ZBlock is 0, count the proportion together with the next ZBlock. Set the order of the extracted proportions to R1, R2, R3...Rn, and recalculate the value of the out-of-order data probability based on the collected value of the out-of-order data. The value of disorderRatio calculated for the i-th time is Di: The proportion of out-of-order data = the number of out-of-order data / the total amount of data stored in ZBlock.
6. A disorder processing system for a time series data storage engine, It is characterized in that The storage engine includes a memory area and a disk area, the time series data includes sequential data and disordered data, the system is used to execute the disorder processing method of the time series data storage engine according to any one of claims 1 to 5, and the system includes a configuration module and a writing module; The configuration module is used to perform the following: There are multiple different ZdataPoint objects configured in the memory area. Each ZdataPoint object contains multiple ZBlocks. Each ZBlock is used to control the storage of time series data within a period of time. The corresponding time periods between ZBlocks do not overlap. Allocate a continuous memory area of a fixed size in the memory area as a cache, divide the cache into a sequential memory block for storing sequential data and a random memory block for storing random data, the memory space of the sequential memory block is larger than the memory space of the random memory block, and reserve a sequential space for storing sequential data and a random space for storing random data in the sequential memory block based on the probability of random data; The write module is used to perform the following: For the time series data collected by the collection point, divide the time series data into corresponding ZdataPoint objects based on the source of the collection point, apply for a ZBlock, and write the time series data into the time series memory block corresponding to the ZBlock through the ZBlock; For a ZBlock that is in a full state, the time series data in the time series memory block corresponding to the ZBlock is written to the disk area through the disk writing thread. After the disk writing is successful, the information record of the time series data is added to the file index part; The writing module follows the following principles: Write sequential data into sequential space and write out-of-order data into out-of-order space; If the current out-of-order space is full, and the newly generated out-of-order data is within the time period corresponding to the current ZBlock, apply for an out-of-order memory block through ZBlock and write the newly generated out-of-order data into the out-of-order memory block; If the current sequential space is full, apply for a new ZBlock, and write the sequential data into a new sequential memory block corresponding to the new ZBlock through the new ZBlock; If the current cache is full, a predetermined number of ZBlocks are recycled based on a preconfigured cache recycling mechanism, and by recycling the predetermined number of ZBlocks, the ZBlocks are deleted and the corresponding sequential memory blocks are released.
7. The out-of-order processing system for the time series data storage engine according to claim 6, It is characterized in that In the memory area, ZBlocks are sorted in the order of time periods to form a ZBlock array.
8. The out-of-order processing system for a time series data storage engine according to claim 6, It is characterized in that The ZBlock corresponds to eight states, namely: Initial state: the state of ZBlock when it is just created; Writing state: when the sequential space corresponding to the ZBlock has not been used up; Unwritable state: when the sequential space corresponding to the ZBlock is exhausted; Temporary write state: when the ZBlock state is in a non-writable state, but there is out-of-order data that needs to be written into the sequential memory block corresponding to the ZBlock, the ZBlock state is first set to a temporary write state. After the out-of-order data is written, the ZBlock state is set back to a non-writable state; Filing to disk: When the fetching thread starts processing a ZBlock, it first sets the ZBlock to the fetching to disk state. After the fetching thread completes the processing, it sets the ZBlock to the fetched to disk state. Flush status: After the flush thread writes the time series data in ZBlock to disk, it updates the ZBlock status to flush status; Waiting for re-writing to disk: When the ZBlock status is already written to disk, but there is out-of-order data written to the memory block corresponding to the ZBlock, after the out-of-order data is written to the ZBlock, the ZBlock status is set to waiting for re-writing to disk; End state: When the cache recycling mechanism recycles a ZBlock, the state of the ZBlock is determined. If it has been written to the disk, the ZBlock is directly deleted. If it is in the state of waiting to be re-written to disk, the corresponding Block data block in the disk area is updated, and then the ZBlock is deleted; The cache recycling mechanism is as follows: when the remaining space in the cache is insufficient to allocate new time series data blocks, cache recycling is triggered, and a predetermined number of old ZBlocks are recycled based on a time period. Before recycling a ZBlock, the ZBlock status is judged. For a ZBlock that is in an unwritable state or is to be rewritten to disk, the time series data in the time series memory block corresponding to the ZBlock is written to disk, and then the ZBlock is recycled.
9. The out-of-order processing system for the time series data storage engine according to claim 8, It is characterized in that The out-of-order data includes: Memory disordered data, where the ZBlock corresponding to the timestamp of the disordered data still exists in the memory area; Disk out-of-order data, where the ZBlock corresponding to the timestamp of the out-of-order data is no longer in the memory area, but exists in a Block in the corresponding disk area; For out-of-order data in memory, the write module is used to perform the following: Analyze the ZBlock status corresponding to its timestamp; When the ZBlock status is writing, the out-of-order data is directly written into the reserved out-of-order space; When the ZBlock status is not writable, first set the ZBlock status to temporary write, then write the out-of-order data to the reserved out-of-order space, and finally set the ZBlock status to not writable; When the ZBlock status is flushed, the out-of-order data is written to the reserved out-of-order space, and the ZBlock status is set to pending re-flush; For the disk disordered data, the writing module is used to perform the following: directly discarding the disk disordered data.
10. The out-of-order processing system for a time series data storage engine according to claim 8, It is characterized in that The sequential memory block is 64K, the out-of-order memory block is 1K, and each ZBlock is allocated a corresponding sequential memory block; The number of time series records that can be stored in a time series memory block is blockRows: rowLength indicates the number of bytes occupied by a time series data record; n random spaces are reserved in the sequential memory block, where The sequential space in the sequential memory block is used to store sequential data. Sequential data refers to the time stamp of the sequential data when it is inserted into the ZBlock that is greater than the time stamp of all the sequential data stored in the ZBlock. The reserved space for sequential data is: blockRows-n The configuration module is used to calculate the probability of out-of-order data based on an adaptive method, and includes the following steps: When the cache recycles a ZBlock, extract the proportion of out-of-order data contained in the current ZBlock. If the number of out-of-order data in the ZBlock is 0, count the proportion together with the next ZBlock. Set the order of the extracted proportions to R1, R2, R3...Rn, and recalculate the value of the out-of-order data probability based on the collected value of the out-of-order data. The value of disorderRatio calculated for the i-th time is Di: The proportion of out-of-order data = the number of out-of-order data / the total amount of data stored in ZBlock.
Citation Information
Patent Citations
Transport accelerator implementing extended transmission control functionality
CN106105141A
Universal real-time data storage management system and implementation method thereof
CN111367880A