Flash timing data storage method and system based on time window segmentation and hot and cold layering

CN122816533APending Publication Date: 2026-09-25QINGPING TECH BEIJING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610965383.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

对于海量数据而言,这种遍历查询的时间复杂度为O (N),其中N为存储的总记录数,当存储的数据量达到百万级甚至千万级时,一次范围查询的耗时可达数百毫秒甚至数秒,完全无法满足用户对快速查询的需求

Benefits of technology

本实施例中,基于时间窗的双层分段索引构建方法,以及冷热分层的自适应归档机制,解决了现有技术中查询效率低、启动恢复时间长、空间利用率低、Flash 寿命短的问题,是本发明实施例的创新点之一。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816533A_ABST
    Figure CN122816533A_ABST
Patent Text Reader

Abstract

The application discloses a Flash timing data storage method and system based on time window segmentation and hot and cold layering, relates to the field of embedded timing data storage technology, and comprises the following steps: determining a target time window and a target data segment; sequentially writing current timing data into the target data segment and adding a secondary index item; when the target data segment is full, adding the target data segment into an archiving queue and ending the current time window; traversing each data segment in the archiving queue at a predetermined period, and performing compression processing on the data segment meeting the hot and cold migration condition; migrating the compressed data segment to a cold data area, and updating a primary index item corresponding to the data segment; writing system state information and the primary index item corresponding to the latest data segment in the cold data area into an anchor point area, and solidifying data in the anchor point area to a Flash memory. The application can simultaneously meet the requirements of efficient writing, quick query, quick recovery, long-term archiving and long service life of timing data in an embedded scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of embedded storage technology, and more specifically, to a Flash time-series data storage method and system based on time window segmentation and hot / cold stratification. Background Technology

[0002] With the rapid development of industrial internet, Internet of Things and smart terminal technologies, embedded devices in fields such as power monitoring, industrial control, vehicle electronics and environmental monitoring are continuously generating massive amounts of time-series data.

[0003] Time-series data is characterized by strict temporal order, continuous writing, and temporal locality of reference. This means data is generated sequentially over time, write operations are primarily sequential appends, and query operations typically target a specific time range. For example, a power monitoring terminal needs to collect parameters such as voltage, current, and power every second, generating over 80,000 records per day; industrial equipment operation logs collect status data every 100ms, generating over 800,000 records per day; and vehicle black boxes can sample data at a frequency of up to 50Hz, generating over 4 million records per day. Faced with such massive amounts of time-series data, the Flash storage media of embedded devices must simultaneously meet the requirements of efficient writing, fast querying, long-term archiving, and long lifespan.

[0004] However, traditional embedded time-series data storage solutions typically employ a circular log storage architecture. This architecture divides the Flash storage space into a fixed-size circular buffer, with new data continuously appended to the end of the buffer. When the buffer is full, the oldest data written is overwritten. While this architecture achieves simple sequential writing, it suffers from several intractable technical drawbacks: First, the query efficiency is extremely low. Because circular logs lack a time-based index structure, when a user needs to query data within a specific time range, the system must traverse all records in the entire buffer, comparing timestamps one by one to find the target data. For massive datasets, the time complexity of this traversal query is O(N), where N is the total number of stored records. When the amount of stored data reaches millions or even tens of millions, a single range query can take hundreds of milliseconds or even several seconds, completely failing to meet users' needs for fast queries.

[0005] Second, the system startup and recovery time is too long. Traditional ring log architectures typically do not persistently store the current write pointer position. When the device loses power and restarts, the system cannot directly obtain the latest write position. It must traverse the entire Flash storage space, scan all storage blocks, and verify the validity of the data one by one before finding the location of the latest valid data. For 16GB Flash storage media, this full scan recovery time can reach more than 5 seconds, severely affecting the device's startup speed and failing to meet the fast startup requirements of embedded devices.

[0006] Third, space utilization is low. Traditional circular log architectures use the same storage method for all data, without distinguishing between hot and cold data, resulting in severe intermingling of written hot and cold data. Real-time hot data requires fast read and write performance, while historical cold data is rarely accessed, yet both occupy the same storage space. Furthermore, traditional solutions do not perform compression processing for the characteristics of time-series data. For slowly changing time-series data such as temperature, voltage, and current, there is a large amount of redundant information, leading to wasted storage space and hindering long-term archiving.

[0007] Fourth, Flash memory has a relatively short lifespan. Flash storage media has a limited number of write / erase cycles. Typically, SLC Flash (Single-Level Cell Flash) has approximately 100,000 write / erase cycles, MLC Flash (Multi-Level Cell Flash) approximately 3,000 cycles, and TLC Flash (Triple-Level Cell Flash) approximately 1,000 cycles. Traditional circular log architectures, lacking optimization for the sequential nature of time-series data, tend to generate a large number of random write operations, resulting in an excessively high write amplification factor, typically exceeding 5.0. This means that writing one byte of valid data requires five bytes of write / erase operations on the Flash memory, significantly accelerating Flash wear and shortening the device's lifespan.

[0008] To address the issues of space reclamation and wear leveling in Flash storage, several Flash management methods have been proposed in existing technologies. For example, Chinese patent application CN121209789A discloses a page state management method for Flash storage. This method divides Flash pages into different states to achieve dirty page migration and space reclamation, thereby improving Flash storage space utilization and lifespan. However, this method is designed for general random write scenarios and does not optimize for the temporal locality and sequentiality of time-series data. It cannot provide a time-based fast index structure for time-series data, nor can it implement a cold / hot tiered archiving strategy for time-series data. Therefore, this method still cannot solve the problems of low query efficiency, long recovery time, and inability to archive long-term data in traditional circular log architectures, and cannot meet the storage management needs of time-series data in embedded scenarios.

[0009] Furthermore, while existing server-side time-series databases, such as InfluxDB and TimescaleDB, offer efficient indexing, compression, and hierarchical storage, their complex architectures require significant memory and CPU resources, making them unsuitable for resource-constrained embedded devices. Embedded devices typically have only a few hundred MB of memory and a few hundred MHz of CPU, insufficient to support the operation of these large time-series databases. Therefore, server-side time-series database solutions cannot be directly ported to embedded environments.

[0010] In summary, existing technologies lack a storage method for embedded Flash storage media that addresses the characteristics of time-series data, and cannot simultaneously meet the requirements of efficient writing, fast querying, fast recovery, long-term archiving, and long lifespan. Therefore, there is an urgent need to propose a new time-series data storage method to solve the aforementioned deficiencies of existing technologies. Summary of the Invention

[0011] This invention provides a Flash time-series data storage method and system based on time window segmentation and hot / cold stratification, in order to overcome at least one technical problem existing in the prior art.

[0012] In a first aspect, embodiments of the present invention provide a Flash time-series data storage method based on time window segmentation and hot / cold stratification, comprising: The Flash storage space is divided into a hot data area, a cold data area, an index area, and an anchor point area; Real-time acquisition of timing data from embedded devices; Determine the target time window based on the current time and the preset time window length; Determine whether a first-level index entry for the target time window exists in the index area. If it does not exist, create a new target data segment in the hot data area and add a corresponding first-level index entry in the index area. If it exists, use the data segment corresponding to the first-level index entry as the target data segment. The current time series data is sequentially written to the target data segment. For every k time series data entries written, a secondary index entry is added to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, the target data segment is added to the archive queue, the current time window ends, and the process returns to the step of determining the target time window. The system iterates through each data segment in the archive queue at a predetermined period and compresses the data segments that meet the cold / hot migration conditions. The compressed data segment is migrated to the cold data area, the first-level index entry corresponding to the data segment is updated, and the data segment is marked as compressed. The system status information and the first-level index item corresponding to the latest data segment in the cold data area are written into the anchor area, and the anchor area data is solidified into the Flash memory.

[0013] Optionally, the hot data area is used to store the latest real-time data; the cold data area is used to store historical data; the index area is used to store first-level index and second-level index data; and the anchor point area is used to store the system's anchor point data.

[0014] Optionally, a one-to-one correspondence is established between time windows and data segments, and the capacity of the data segment corresponding to the target time window is expressed as follows: Where C is in bytes; f is the sampling frequency of the time series data in Hz; T is the length of the time window in seconds; and L is the length of a single time series data entry in bytes.

[0015] Alternatively, the cold and heat migration conditions can be expressed as: ,in, For the current time, This is the end time of the data segment. This represents the time threshold for cold and hot migration.

[0016] Optionally, compression processing specifically includes: Interpolation encoding is performed on the time-series data within the data segment, and the difference between adjacent time-series data is calculated. ,in, , These are the sampled values ​​of the i-th and i-1th time series data points, respectively. This is the i-th difference; The LZ4 compression algorithm is used to perform batch compression on the difference data.

[0017] Optionally, the compression process may be performed on data segments that meet the cold / hot migration conditions, including: Calculate the retention time of the data segment; If the data segment has exceeded its retention period, the data segment is deleted directly; otherwise, it is compressed.

[0018] Optionally, the retention time of the data segment is calculated, specifically including: Calculate the data retention value of the data segment based on its access frequency F, anomaly flag A, and event importance E. Where V takes values ​​in the range [0,1]; α, β, and γ are weighting coefficients, satisfying α+β+γ=1; Based on the retention value The retention time of the data segment is calculated as follows: ,in, The actual retention time of the data segment, in days; The default base retention time, in days.

[0019] Optionally, the first-level index entry includes the start time, end time, physical address, and status of the data segment; the second-level index entry records the timestamp of the corresponding time-series data and its offset address within the data segment.

[0020] Optionally, updating the first-level index entry corresponding to the data segment specifically includes: Update the physical address in the first-level index entry corresponding to the data segment to the address of the cold data area.

[0021] Secondly, embodiments of the present invention also provide a Flash time-series data storage system based on time window segmentation and hot / cold stratification, comprising: The preprocessing module is used to divide the Flash storage space into hot data area, cold data area, index area and anchor point area; The acquisition module is used to acquire time-series data from embedded devices in real time. The determination module is used to determine the target time window based on the current time and the preset time window length. The judgment module is used to determine whether there is a first-level index entry for the target time window in the index area. If it does not exist, a new target data segment is created in the hot data area, and a corresponding first-level index entry is added to the index area. If it exists, the data segment corresponding to the first-level index entry is used as the target data segment. The writing module is used to sequentially write the current time series data into the target data segment. For every k time series data entries written, a secondary index entry is added to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, the target data segment is added to the archive queue, the current time window ends, and the process returns to the target time window determination step. The compression module is used to traverse each data segment in the archive queue at a predetermined period and compress the data segments that meet the cold and hot migration conditions. The migration module is used to migrate the compressed data segment to the cold data area, update the first-level index entry corresponding to the data segment, and mark the data segment as compressed. The solidification module is used to write the system status information and the first-level index item corresponding to the latest data segment in the cold data area into the anchor point area, and solidify the anchor point area data into the Flash memory.

[0022] The innovative aspects of this invention include, but are not limited to, the following: In this embodiment, the time-window-based two-layer segmented index construction method and the cold-hot layered adaptive archiving mechanism solve the problems of low query efficiency, long startup and recovery time, low space utilization and short Flash life in the prior art, which is one of the innovations of this invention. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 A flowchart of a data storage method provided in an embodiment of the present invention; Figure 2 Another flowchart of the data storage method provided in an embodiment of the present invention; Figure 3 A schematic diagram of a data storage system provided in an embodiment of the present invention; Figure 4 This is another schematic diagram of the data storage system provided in an embodiment of the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] It should be noted that the terms "comprising" and "having," and any variations thereof, in the embodiments and drawings of this invention are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the steps or units listed, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0027] To address the limitations of existing Flash storage methods in simultaneously meeting the requirements of high-efficiency writing, fast querying, fast recovery, long-term archiving, and long lifespan, this invention discloses a Flash time-series data storage method and system based on time window segmentation and hot / cold stratification. These will be described in detail below.

[0028] Figure 1 A flowchart of a data storage method provided in an embodiment of the present invention is provided below. Figure 1 The Flash time-series data storage method based on time window segmentation and hot / cold stratification provided in this embodiment of the invention includes: Step 1: Divide the Flash storage space into a hot data area, a cold data area, an index area, and an anchor point area.

[0029] This invention addresses the access locality of time-series data by dividing the Flash storage space into four independent regions based on the data access frequency: a hot data region, a cold data region, an index region, and an anchor region. This partitioning achieves separate storage of hot and cold data.

[0030] The hot data area is used to store the latest real-time data, i.e., hot data. Because hot data is accessed frequently and requires fast read and write performance, the data segments in the hot data area are stored in an uncompressed raw format, supporting fast random reads and sequential writes.

[0031] The cold data area is used to store historical archived data, i.e., cold data. Because cold data is accessed infrequently, the data segments in the cold data area are stored in a compressed format. Compression significantly improves space utilization and enables long-term archiving.

[0032] The index area is used to store global first-level and second-level index data. Since the index data is accessed frequently, it is stored in a separate index area to ensure fast access to the index.

[0033] The anchor area is used to store the system's anchor data, including the address of the latest data segment, system status information, etc. The anchor data is persistently stored and is immediately flushed to Flash after each update to ensure that it is not lost when power is lost.

[0034] Step 2: Real-time acquisition of timing data from the embedded device.

[0035] The real-time collected time-series data includes information such as timestamps, sampled values, and status bits. Each time-series data has a uniform format and a fixed length, such as L. Thus, when writing to the hot data area, there is no need to attach a length identifier to each record, which helps save storage space. At the same time, when using the secondary index for positioning, the offset address within the data segment can be directly calculated based on the record sequence number, providing a basis for the fast positioning of the secondary skip index.

[0036] Step 3: Determine the target time window based on the current time and the preset time window length.

[0037] This invention divides continuous time-series data into independent data segments according to a preset time window length. Each data segment corresponds to a continuous time interval, and all data segments are arranged sequentially in chronological order. Based on this, a two-layer index structure is constructed to achieve rapid location of data from time.

[0038] Therefore, after acquiring time-series data, the target time window is first determined based on the timestamp of the current time-series data and the preset time window length. For example, when the time window length is T seconds, the target time window is the T-second interval where the current timestamp is located.

[0039] Step 4: Determine whether a first-level index entry for the target time window exists in the index area. If it does not exist, create a new target data segment in the hot data area and add the corresponding first-level index entry in the index area. If it exists, use the data segment corresponding to the first-level index entry as the target data segment.

[0040] To address the low query efficiency issue in existing technologies, this invention constructs a time-based first-level index, which maintains a one-to-one mapping between global time windows and data segments. Therefore, after determining a target time window, the system first checks whether a first-level index entry for that target time window exists in the index area.

[0041] If the target data segment exists, the data segment corresponding to that first-level index entry is used as the target data segment. If it does not exist, it means that data has not yet been written to the target time window. Therefore, a free physical storage segment needs to be allocated in the hot data area as the target data segment, and a corresponding first-level index entry needs to be added to the index area for the target time window. The first-level index entry includes the start time, end time, physical address of the data segment in the Flash storage space, and its status (e.g., cold, hot, pending archiving, compressed, etc.). In this way, during subsequent queries, a binary search is performed in the first-level index based on the query time range to quickly locate one or more data segments that intersect with the query time range, without needing to traverse all data segments, thus effectively improving query efficiency and meeting the user's need for fast queries.

[0042] Step 5: Write the current time series data sequentially to the target data segment. For every k time series data entries written, add a secondary index entry to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, add the target data segment to the archive queue, end the current time window, and return to the step of determining the target time window.

[0043] Once the target data segment is obtained, data can be written sequentially to it. To further improve query efficiency, this invention introduces a secondary index. For every k time-series data entries written, a secondary index entry is added to the target data segment as a skip index in the index area. This secondary index entry records the timestamp of the time-series data entry and its offset address within the target data segment. Thus, during data queries, the primary index is used to locate the data segment, and then the secondary index is used to quickly locate the target time-series data within the segment, without needing to traverse all the time-series data within the segment.

[0044] Where k represents the jump step size, and its value can be set according to the specific application scenario, for example, it can be set to 128. This invention does not impose a specific limitation on this.

[0045] When the target data segment is full, it is added to the archive queue to prepare for subsequent data migration. It should be noted that "full" here means the target data segment has reached the preset capacity corresponding to the target time window. That is, the writing of the time-series data to be collected within the target time window has been completed, the current time window ends, and the process returns to step 3 to redetermine the target time window. Steps 3-5 are then executed repeatedly to complete the writing of the real-time collected time-series data.

[0046] In one embodiment, if the preset time window length is T, then the capacity of the data segment corresponding to the target time window is expressed as: Where C is the theoretical capacity of the data segment corresponding to the target time window, in bytes; f is the sampling frequency of the time-series data, in Hz; T is the length of the time window, in seconds; and L is the length of a single time-series data record, in bytes. The time window is considered to end when the data segment reaches its theoretical capacity C. Using this formula, the system can dynamically adjust the length of the time window according to the sampling frequency and record length of different application scenarios, ensuring that the size of each data segment matches the block size of the Flash memory. This enables sequential writing, reduces write amplification, thereby reducing Flash wear and increasing its lifespan.

[0047] It should be noted that when a data segment is full, the actual end time of the data segment is based on the timestamp of the last time-series data and is used as the actual end time of the corresponding time window. At this time, the end time in the corresponding first-level index item is updated, and the status is updated to pending archiving.

[0048] By employing the aforementioned two-level index structure, this invention reduces the time complexity of range queries for time-series data from the traditional O(N) to O(logN), where N is the total number of records, significantly improving query efficiency. Simultaneously, this index structure has extremely low memory consumption; the secondary index for each data segment occupies only O(N / K) space. For example, when K is 128, the index space usage is only 0.78% of the original data, which can be stored within the limited memory of embedded devices.

[0049] Step 6: Iterate through each data segment in the queue to be archived at a predetermined period and compress the data segments that meet the cold and hot migration conditions.

[0050] This invention achieves partitioned storage of hot and cold data by partitioning Flash memory. During the data writing process, the system checks all data segments in the archive queue at predetermined intervals to determine whether each data segment meets the cold / hot migration conditions, and thus adopts different processing strategies.

[0051] In one embodiment, the cold and hot migration conditions are set as follows: ,in, For the current time, This is the end time of the data segment. The time threshold for thermal migration can be set according to the actual scenario, and this invention does not impose any specific limitations on it.

[0052] When the difference between the current time and the end time of the data segment is greater than the time threshold for cold / hot migration If the condition is met, the data segment will be compressed to facilitate subsequent migration. Otherwise, if the condition is not met, the data segment will remain in the archive queue awaiting the next check. At this time, the data segments in the archive queue are still considered hot data and can be accessed normally.

[0053] In one embodiment, when compressing the data segment, an algorithm combining interpolation encoding and batch compression is used, taking into account the characteristics of time-series data. First, interpolation encoding is performed on the time-series data within the data segment, and the difference between adjacent time-series data is calculated. ,in, , These are the sampled values ​​of the i-th and i-1th time series data points, respectively. This is the i-th difference.

[0054] For slowly changing time-series data, such as temperature, voltage, and current, the fluctuation range of the difference is much smaller than that of the original value. Therefore, fewer bytes can be used to store the difference, thus achieving initial compression. Then, the LZ4 compression algorithm is used to batch compress the difference data to further improve the compression ratio.

[0055] It should be noted that the use of the LZ4 compression algorithm is only one implementation method in this embodiment and is not intended to limit the invention. Furthermore, the compression process does not modify the data in the data segment, and a rollback is possible if compression fails.

[0056] Figure 2 For another flowchart of the data storage method provided in this embodiment of the invention, please refer to... Figure 2 In this embodiment, when the data segment meets the cold / hot migration condition, the method further includes: Step 06: Calculate the retention time of the data segment; if the data segment has exceeded the retention time, delete the data segment directly; otherwise, perform compression processing.

[0057] This invention adopts a value-driven adaptive archiving strategy. When a data segment meets the cold-hot migration conditions, the retention time of the data segment is first calculated. If the data segment has already expired, it is directly deleted; if it has not yet expired, it is compressed and then migrated to the cold data area.

[0058] In one embodiment, calculating the retention time of a data segment based on its retention value specifically includes: Calculate the data retention value of a data segment based on its access frequency (F), anomaly flag (A), and event importance (E). Where V ranges from [0,1]; α, β, and γ are the weight coefficients of the corresponding indicators, satisfying α+β+γ=1. The weight coefficients can adjust the weight of different factors on the value. For example, the default settings are access frequency weight α=0.4, anomaly marking weight β=0.3, and event importance weight γ=0.3.

[0059] Based on retention value The retention time for the calculated data segment is... ,in, This is the actual retention time for the data segment, in days. This is the default base retention period, in days. Base Retention Period Specific configurations can be made according to the application scenario, such as setting it in a power scenario. =30 days, vehicle-mounted scenario settings =90 days.

[0060] By performing multi-dimensional value scoring on data segments and calculating retention time based on the value scores, abnormal or accidental data with higher retention value (V value) will have their retention time automatically extended, while normal data will be automatically cleaned up. This dynamic adjustment of data retention time requires no manual intervention, ensuring the long-term preservation of important data while avoiding useless data occupying space, thus significantly improving the utilization rate of storage space.

[0061] Step 7: Move the compressed data segment to the cold data area, update the first-level index entry corresponding to the data segment, and mark the data segment as compressed.

[0062] After compression, the system sequentially migrates the compressed data segments to the cold data area, eliminating random wear and further reducing Flash wear.

[0063] After the data segment is migrated, its corresponding physical storage address, status, etc., all change. Therefore, it is necessary to update the first-level index corresponding to the data segment at the same time, update the physical address of the data segment in the Flash storage space to the address of the cold data area, and mark the data segment as compressed.

[0064] This invention achieves the separation of hot and cold data through a hot-cold tiered archiving mechanism. Hot data maintains its fast read / write characteristics, while cold data improves space utilization through compression. Simultaneously, this mechanism transforms all write operations into sequential writes: writing to hot sectors is a sequential append operation, while migrating to cold sectors is a batch sequential write operation. This significantly reduces the write amplification factor and extends the lifespan of the Flash memory.

[0065] Step 8: Write the system status information and the first-level index item corresponding to the latest data segment in the cold data area into the anchor area, and then save the anchor area data to the Flash memory.

[0066] To ensure data integrity during power outages, an update operation is performed on the anchor point area after the hot and cold migrations are completed. This includes writing the current system status information (such as the address and offset of the data segment being written in the hot data area, and the current time window identifier) ​​and the first-level index entry corresponding to the latest data segment in the cold data area into the anchor point area, and then storing the anchor point area data in the Flash memory. This records the current system operating status, ensuring that after the device restarts, the current system status information and the first-level index entry corresponding to the latest data segment in the cold data area can be directly obtained by reading the anchor point data in the anchor point area. This eliminates the need to traverse the entire Flash storage space, reducing the startup recovery time complexity from O(N) to O(1), significantly improving startup recovery speed and enabling rapid device startup.

[0067] The present invention provides a Flash time-series data storage method based on time window segmented indexing and cold / hot hierarchical archiving, a time window-based two-layer segmented index construction method, and a cold / hot hierarchical adaptive archiving mechanism. These methods solve the problems of low query efficiency, long startup and recovery time, low space utilization, and short Flash lifespan in the prior art. The invention is applicable to various embedded scenarios such as power monitoring, industrial control, and automotive electronics, and meets different application requirements.

[0068] Based on the same inventive concept, this invention also provides a Flash time-series data storage system based on time window segmentation and hot / cold stratification. Figure 3 Please refer to the schematic diagram of a data storage system provided in an embodiment of the present invention. Figure 3 The present invention provides a Flash time-series data storage system 100 based on time window segmentation and hot / cold stratification, comprising: Preprocessing module 101 is used to divide the Flash storage space into hot data area, cold data area, index area and anchor point area; Acquisition module 102 is used to acquire timing data from embedded devices in real time; The determination module 103 is used to determine the target time window based on the current time and the preset time window length; The judgment module 104 is used to determine whether there is a first-level index entry for the target time window in the index area. If it does not exist, a new target data segment is created in the hot data area and the corresponding first-level index entry is added in the index area. If it exists, the data segment corresponding to the first-level index entry is used as the target data segment. The writing module 105 is used to sequentially write the current time series data into the target data segment. Every time k time series data are written, a secondary index entry is added to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, the target data segment is added to the archive queue, the current time window ends, and the step of determining the target time window is returned. Compression module 106 is used to traverse each data segment in the archive queue at a predetermined period and compress the data segments that meet the cold and hot migration conditions. Migration module 107 is used to migrate the compressed data segment to the cold data area, update the first-level index entry corresponding to the data segment, and mark the data segment as compressed; The solidification module 108 is used to write the system status information and the first-level index item corresponding to the latest data segment in the cold data area into the anchor area, and solidify the anchor area data into the Flash memory.

[0069] The above system embodiments correspond to the method embodiments and have the same technical effects as the method embodiments. For details, please refer to the method embodiments, which will not be repeated here.

[0070] Optionally, compression module 106 is specifically configured as follows: Interpolation encoding is performed on the time-series data within the data segment, and the difference between adjacent time-series data is calculated. ,in, , These are the sampled values ​​of the i-th and i-1th time series data points, respectively. This is the i-th difference; The LZ4 compression algorithm is used to perform batch compression on the difference data.

[0071] Optionally, Figure 4 For another structural diagram of the data storage system provided in this embodiment of the invention, please refer to... Figure 4 In this embodiment, the data storage system 100 further includes a computing module 109, which is specifically configured as follows: Calculate the retention time of the data segment; If the data segment has exceeded its retention period, delete the data segment directly; otherwise, compress it.

[0072] Optionally, when calculating the retention time of the data segment, the calculation module 109 is specifically configured as follows: Calculate the data retention value of a data segment based on its access frequency (F), anomaly flag (A), and event importance (E). Where V takes values ​​in the range [0,1]; α, β, and γ are weighting coefficients, satisfying α+β+γ=1; Based on retention value The retention time for the calculated data segment is... ,in, This is the actual retention time for the data segment, in days. The default base retention time, in days.

[0073] Optionally, migration module 107 is specifically configured as follows: Update the physical address in the first-level index entry corresponding to the data segment to the address of the cold data area.

[0074] To further illustrate the present invention, embodiments are provided under the following three different application scenarios (power monitoring terminal scenario, industrial equipment log scenario, and vehicle-mounted black box scenario): Example 1: Power Monitoring Terminal Scenario In this embodiment, the method of the present invention is applied to a power monitoring terminal of a 10kV distribution network. This terminal is used to monitor parameters such as voltage, current, power, and power factor of the line, providing data support for the operation monitoring and fault analysis of the distribution network.

[0075] The power monitoring terminal is configured with the following parameters: sampling frequency f = 1Hz, meaning data is collected once per second; single record length L = 32 bytes, including 4 bytes for timestamp, 4 bytes for voltage, 4 bytes for current, 4 bytes for active power, 4 bytes for reactive power, 4 bytes for power factor, 4 bytes for status bit, and 2 bytes reserved; Flash storage capacity is 16GB, including 4GB for hot data area, 11GB for cold data area, 0.8GB for index area, and 0.2GB for anchor point area; time window length T = 86400 seconds, or 1 day, therefore each data segment corresponds to one day of monitoring data; cold / hot migration threshold. =7 days, meaning that hot data segments written more than 7 days ago will be migrated to the cold data area; basic retention time =30 days, meaning the default retention time for normal data is 30 days; weighting coefficients α=0.4, β=0.3, γ=0.3.

[0076] According to the time window capacity formula Substituting the parameters, we can calculate the result. This means that the theoretical capacity of each data segment is 2.63MB. Since the block size of Flash memory is 4MB, each data segment can be stored within a single Flash block, resulting in a segment utilization rate of [missing information]. This utilization rate is acceptable in the hot data area because the hot data area needs to ensure fast write speeds, while the cold data area can improve utilization by compressing and merging multiple data segments.

[0077] When each data segment is created, the system adds an index entry to the first-level index. For example, the index entry for the data segment 2024-05-01 is: Start Time: 1714521600 (2024-05-01 00:00:00); End Time: 1714608000 (2024-05-02 00:00:00); Physical Address: 0x10000000; Status: Hot Then, within each data segment, a secondary index entry is created for every 128 records. For example, the index entry for the 128th record is: timestamp: 1714521728, offset address: 128*32=4096 bytes; the index entry for the 256th record is: timestamp: 1714521856, offset address: 256*32=8192 bytes. And so on. Each data segment's secondary index has a total of 86400 / 128=675 index entries. Each index entry occupies 8 bytes, therefore each data segment's secondary index only occupies 675*8=5400 bytes, approximately 5.3KB, with extremely low space usage.

[0078] After a data segment is written and 7 days have passed, it meets the migration criteria, and the system begins archiving it. First, the retention value of the data segment is calculated: the data segment represents normal operational data with no abnormal events, therefore A=0, E=0; the access frequency of the data segment is 0.1, meaning it is accessed on average once every 10 days, so after normalization, F=0.1. Substituting these values ​​into the value formula... Then calculate the retention time. This means that the actual retention time for this data segment is 31.2 days, which is slightly longer than the default 30 days.

[0079] Next, segment-level compression is performed on this data segment: First, difference encoding is performed. Parameters such as voltage, current, and power change slowly, and the difference between adjacent records is very small. For example, the original voltage values ​​are 220.1, 220.2, 220.1, 220.0, 220.1, ...; the difference values ​​are 0, 0.1, -0.1, -0.1, 0.1, .... Each original voltage value is stored as a 4-byte floating-point number, while the difference ranges from -1 to 1, so it can be stored as a 1-byte integer. In this way, the storage size of each parameter is reduced from 4 bytes to 1 byte.

[0080] After encoding, the data is batch compressed using LZ4, and the final compressed data segment size is [size missing]. The compression ratio R = (2.63 - 0.79) / 2.63 ≈ 70%, which meets expectations.

[0081] The compressed data segments are written to the cold data area. The flash block size of the cold data area is 4MB, so one block of the cold data area can store 5 compressed data segments. The segment utilization rate of the cold data area is... This significantly improves space utilization.

[0082] Query performance test: When a user needs to query monitoring data for three days, from May 1st to May 3rd, 2024, the traditional circular log solution requires traversing 3 * 86400 = 259200 records and comparing timestamps one by one, taking approximately 250ms. However, using the method of this invention, the query process is as follows: Based on the start and end times of the query, a binary search is performed in the primary index to find the corresponding three data segments: 2024-05-01, 2024-05-02, and 2024-05-03, taking approximately 1µs. For each data segment, a binary search is performed in the secondary index to find the offset address corresponding to the query start time and the offset address corresponding to the query end time, with each segment search taking approximately 1µs. Finally, based on the offset addresses, the target data in these three data segments can be directly read, taking approximately 1ms. The total query time is approximately 2ms, which is 125 times faster than the traditional method, significantly improving query efficiency.

[0083] Startup recovery performance test: When the device powers off and restarts, traditional circular log solutions require traversing the entire 16GB Flash storage space, scanning all blocks, and finding the latest write location, taking approximately 5 seconds. However, using the method of this invention, after system startup, the anchor data at a fixed address in the anchor area can be directly read to obtain the address of the latest data segment, the address of the index, and other information. The entire recovery process takes approximately 0.1ms, which is 50,000 times faster than traditional solutions, achieving rapid startup.

[0084] Example 2: Industrial Equipment Log Scenario In this embodiment, the method of the present invention is applied to an equipment condition monitoring terminal in an industrial production line. This terminal is used to collect parameters such as the operating status, temperature, vibration, and pressure of the equipment, and at the same time record the equipment's operating log, providing data support for equipment fault diagnosis and predictive maintenance.

[0085] The industrial monitoring terminal is configured with the following parameters: sampling frequency f = 10Hz, meaning data is collected every 100ms; single record length L = 64 bytes, including 8 bytes for timestamp, 4 bytes for temperature, 4 bytes for vibration, 4 bytes for pressure, 4 bytes for rotational speed, and 40 bytes for log information; Flash storage capacity is 32GB, including 8GB for hot data, 23GB for cold data, 0.8GB for index, and 0.2GB for anchor points; time window length T = 3600 seconds, or 1 hour, therefore each data segment corresponds to one hour of operating data; cold / hot migration threshold. =24 hours, meaning hot data segments written for more than one day will be migrated to the cold data area; basic retention time =30 days, meaning the default retention time for normal data is 30 days; weighting coefficients α=0.4, β=0.3, γ=0.3.

[0086] According to the time window capacity formula Substituting the parameters, we can calculate the result. That is, the theoretical capacity of each data segment is 2.2MB, which is also compatible with a 4MB Flash block. The segment utilization rate of the hot data area is 2.2 / 4=55%.

[0087] When a device malfunctions, the data segment from the hour preceding the malfunction contains precursory data. This data segment has an anomaly flag A=1, event importance E=1, and access frequency F=0.8 because it is frequently accessed after the malfunction occurs for fault analysis. Substituting these values ​​into the value formula yields... Then calculate the retention time. This means that the retention time for the faulty data segment is 57.6 days, which is nearly twice as long as the normal data, ensuring the long-term preservation of the faulty data and facilitating subsequent fault tracing and analysis.

[0088] For industrial log data, there is a lot of repetitive content in the log information, such as "normal operation" and "normal temperature". Therefore, differential encoding and batch compression are more effective, with a compression rate of up to 60%. That is, the original 2.2MB data segment is compressed to only 0.88MB. A 4MB block in the cold data area can store 4 compressed data segments, with a utilization rate of 88%.

[0089] Fault tracing query performance: When equipment malfunctions, maintenance personnel need to query the operational data for the hour preceding the malfunction. Traditional methods require traversing the entire storage area to find the corresponding time, taking approximately 100ms. However, the method of this invention first uses a primary index to find the data segment corresponding to the hour of the malfunction, then uses a secondary index to find the location of the malfunction time and directly reads the data from that segment. The entire query takes approximately 1ms, a 100-fold improvement, allowing maintenance personnel to quickly obtain the operational data prior to the malfunction for fault diagnosis.

[0090] Example 3: Vehicle-mounted black box scenario In this embodiment, the method of the present invention is applied to an on-board black box (Event Data Recorder, EDR), which is used to collect parameters such as vehicle speed, acceleration, angular velocity, braking status, throttle status, and steering angle for accident liability determination and driving behavior analysis.

[0091] The parameters of this vehicle-mounted black box are as follows: sampling frequency f = 50Hz, i.e., data is collected every 20ms; single record length L = 128 bytes, including 8 bytes for timestamp, 4 bytes for vehicle speed, 12 bytes for acceleration, 12 bytes for angular velocity, 1 byte for braking status, 1 byte for throttle status, 4 bytes for steering angle, 32 bytes for GPS information, and 54 bytes for other status information; Flash storage capacity is 64GB, including 16GB for hot data area, 47GB for cold data area, 0.8GB for index area, and 0.2GB for anchor point area; time window length T = 600 seconds, i.e., 10 minutes, so each data segment corresponds to 10 minutes of running data; cold / hot migration threshold. =24 hours, meaning hot data segments written for more than one day will be migrated to the cold data area; basic retention time =90 days, meaning the default retention time for normal data is 90 days; weighting coefficients α=0.4, β=0.3, γ=0.3.

[0092] Substituting the parameters into the time window capacity formula, we get... That is, the theoretical capacity of each data segment is 3.66MB, which almost fills the 4MB Flash block. The segment utilization rate of the hot data area is 3.66 / 4≈91.5%, which is very high. This is because the sampling frequency of vehicle data is high, and the size of the data segment is just right to match the Flash block, reducing space waste.

[0093] When a vehicle collision occurs, the data segment from the 10 minutes prior to the collision contains vehicle movement data and is crucial evidence for determining liability. This data segment is marked with an anomaly flag A=1, event importance E=1, and access frequency F=1 because it is frequently accessed after the accident for accident analysis. Substituting these values ​​into the value formula yields... Then calculate the retention time. This means that the collision data segment is retained for 180 days, which is nearly twice the normal data retention period, ensuring the long-term preservation of accident data and meeting the traffic management department's requirements for accident data preservation.

[0094] When a vehicle is driving normally, parameters such as speed, acceleration, and angular velocity change slowly, and the difference between adjacent records is very small. The compression rate can reach 65%. The original 3.66MB data segment is compressed to only 1.28MB. A 4MB block in the cold data area can store 3 compressed data segments, with a utilization rate of 96%.

[0095] Accident tracing and query performance: After an accident, traffic management departments need to query vehicle operation data for the 30 seconds prior to the collision. Traditional methods require traversing the entire black box storage to find the corresponding time, taking approximately 200ms. However, the method of this invention first uses a primary index to find the corresponding 10-minute data segment, then uses a secondary index to locate the collision time, directly reading the data from the first 30 seconds of that segment. The entire query takes approximately 1ms, a 200-fold improvement, allowing traffic management departments to quickly obtain accident data and determine liability.

[0096] To verify the technical effects of this invention, tests were conducted on an actual embedded device. The test device was an ARM Cortex-A7-based embedded platform with a main frequency of 1GHz, 512MB of memory, and 16GB of MLC NAND Flash. The test data consisted of actual operating data from the power monitoring terminal, totaling 10 million records. The test results are as follows: Query efficiency: The average time for range queries in traditional ring log solutions is 248ms, while the average time for the solution of this invention is 1.9ms, which improves query efficiency by 130 times.

[0097] Startup recovery time: The traditional solution has a startup recovery time of 4.8 seconds, while the solution of this invention has a startup recovery time of 0.09ms, which is 53,333 times faster.

[0098] Write amplification factor: The traditional solution has a write amplification factor of 4.9, while the solution of this invention has a write amplification factor of 1.2, which reduces write amplification by 75.5% and increases the lifespan of the Flash by 4.08 times.

[0099] Space utilization: The traditional solution has a storage space utilization rate of 32%, while the solution of this invention has a utilization rate of 75%, which is 134% higher. With the same Flash space, this invention can store 2.3 times more data.

[0100] The test results above show that the method of the present invention completely solves the defects of the prior art and can simultaneously meet the needs of efficient writing, fast querying, fast recovery, long-term archiving and long lifespan of time-series data in embedded scenarios.

[0101] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of one embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention.

[0102] Those skilled in the art will understand that the modules in the apparatus of the embodiments can be distributed in the apparatus of the embodiments as described in the embodiments, or they can be located in one or more devices different from this embodiment with corresponding changes. The modules of the above embodiments can be combined into one module, or they can be further divided into multiple sub-modules.

[0103] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A Flash time-series data storage method based on time window segmentation and hot / cold stratification, characterized in that, include: The Flash storage space is divided into a hot data area, a cold data area, an index area, and an anchor point area; Real-time acquisition of timing data from embedded devices; Determine the target time window based on the current time and the preset time window length; Determine whether a first-level index entry for the target time window exists in the index area. If it does not exist, create a new target data segment in the hot data area and add a corresponding first-level index entry in the index area. If it exists, use the data segment corresponding to the first-level index entry as the target data segment. The current time series data is sequentially written to the target data segment. For every k time series data entries written, a secondary index entry is added to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, the target data segment is added to the archive queue, the current time window ends, and the process returns to the step of determining the target time window. The system iterates through each data segment in the archive queue at a predetermined period and compresses the data segments that meet the cold / hot migration conditions. The compressed data segment is migrated to the cold data area, the first-level index entry corresponding to the data segment is updated, and the data segment is marked as compressed. The system status information and the first-level index item corresponding to the latest data segment in the cold data area are written into the anchor area, and the anchor area data is solidified into the Flash memory.

2. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, The hot data area is used to store the latest real-time data; the cold data area is used to store historical data; the index area is used to store first-level index and second-level index data; and the anchor point area is used to store the system's anchor point data.

3. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, A time window corresponds one-to-one with a data segment, and the capacity of the data segment corresponding to the target time window is expressed as follows: Where C is in bytes; f is the sampling frequency of the time series data in Hz; T is the length of the time window in seconds; and L is the length of a single time series data entry in bytes.

4. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, The conditions for thermal migration are expressed as follows: ,in, For the current time, This is the end time of the data segment. This represents the time threshold for cold and hot migration.

5. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, Compression processing specifically includes: Interpolation encoding is performed on the time-series data within the data segment, and the difference between adjacent time-series data is calculated. ,in, , These are the sampled values ​​of the i-th and i-1-th time series data points, respectively. This is the i-th difference; The LZ4 compression algorithm is used to perform batch compression on the difference data.

6. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, Before compressing data segments that meet the cold / hot migration conditions, the following steps are also included: Calculate the retention time of the data segment; If the data segment has exceeded its retention period, the data segment is deleted directly; otherwise, it is compressed.

7. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 6, characterized in that, Calculating the retention time of the data segment specifically includes: Calculate the data retention value of the data segment based on its access frequency F, anomaly flag A, and event importance E. Where V takes values ​​in the range [0,1]; α, β, and γ are weighting coefficients, satisfying α+β+γ=1; Based on the retention value The retention time of the data segment is calculated as follows: ,in, The actual retention time of the data segment, in days; The default base retention time, in days.

8. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 1, characterized in that, The first-level index entry includes the start time, end time, physical address, and status of the data segment; the second-level index entry records the timestamp of the corresponding time-series data and its offset address within the data segment.

9. The Flash time-series data storage method based on time window segmentation and hot / cold stratification according to claim 8, characterized in that, Updating the first-level index entry corresponding to the data segment specifically includes: Update the physical address in the first-level index entry corresponding to the data segment to the address of the cold data area.

10. A Flash time-series data storage system based on time window segmentation and hot / cold stratification, characterized in that, include: The preprocessing module is used to divide the Flash storage space into hot data area, cold data area, index area and anchor point area; The acquisition module is used to acquire time-series data from embedded devices in real time. The determination module is used to determine the target time window based on the current time and the preset time window length. The judgment module is used to determine whether there is a first-level index entry for the target time window in the index area. If it does not exist, a new target data segment is created in the hot data area, and a corresponding first-level index entry is added to the index area. If it exists, the data segment corresponding to the first-level index entry is used as the target data segment. The writing module is used to sequentially write the current time series data into the target data segment. For every k time series data entries written, a secondary index entry is added to the target data segment in the index area, where k represents the jump step size. When the target data segment is full, the target data segment is added to the archive queue, the current time window ends, and the process returns to the target time window determination step. The compression module is used to traverse each data segment in the archive queue at a predetermined period and compress the data segments that meet the cold and hot migration conditions. The migration module is used to migrate the compressed data segment to the cold data area, update the first-level index entry corresponding to the data segment, and mark the data segment as compressed. The solidification module is used to write the system status information and the first-level index item corresponding to the latest data segment in the cold data area into the anchor point area, and solidify the anchor point area data into the Flash memory.

Citation Information

Patent Citations

  • Data storage management method based on FLASH in MCU and storage medium

    CN121209789A