Data management method integrating real-time analysis and offline analysis

By adopting LSM-Tree structure and change logs in data management, the requirements of real-time analysis and offline analysis are realized, the problem of high operation and maintenance costs in the existing technology is solved, development and operation and maintenance costs are reduced, and data management efficiency is improved.

CN120011363APending Publication Date: 2025-05-16HUBEI RADIO & TELEVISION

Patent Information

Application Number
CN202510094444.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art lacks a data management method that can simultaneously realize real-time analysis and offline analysis requirements and reduce operation and maintenance costs.

Method used

The data is stored using the LSM-Tree structure and its WAL log is replaced with a change log (ChangeLog). By aligning the snapshot with the SST data file, the query that calculates the intermediate data in real time is converted into a query for the LSM-Tree SST data file.

Benefits of technology

It realizes a unified access interface for data batch flow, reduces the development and testing costs of adapting to different data sources, simplifies the data management architecture, reduces operation and maintenance costs, and improves the efficiency of data update, deletion and query under the scale of big data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011363A_ABST
    Figure CN120011363A_ABST
Patent Text Reader

Abstract

The invention discloses a data management method integrating real-time analysis and offline analysis. The method comprises the following steps: storing data by adopting an LSM-Tree structure; a WAL log in the LSM-Tree structure is replaced with a change log; the change log is aligned with the SST data file of the LSM-Tree by using the snapshot, and the message is queried according to the snapshot; positioning a snapshot, positioning a log file, reading log data and recording offset; and the query of real-time calculation of the intermediate data is converted into the query of the SST data file of the LSM-Tree. According to the invention, the problem that the existing data processing framework of the combined and offline data and the real-time data lacks one of the functions of realizing real-time analysis and offline analysis and reducing the operation and maintenance cost at the same time is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data management, and in particular, relates to a data management method integrating real-time analysis and offline analysis. Background Art

[0002] In recent years, big data technology has experienced rapid development. From the perspective of market demand, the scale and complexity of data are rising exponentially, and the demand for real-time computing, storage and analysis is extremely strong, making it difficult for traditional databases to survive. In the early days, the data warehouse production system was mainly based on offline data warehouses, and businesses divided data warehouses into different levels according to their own business needs, such as DWD, DWS, ADS, etc. In the offline data warehouse, business data will enter the data warehouse through offline ETL processing, and data conversion between layers will also be processed using offline ETL. The ADS layer can directly provide query capabilities to the outside world, and the middle layer usually uses Hive to store intermediate data. Some OLAP query capabilities can also be provided based on Hive.

[0003] The advantage of the offline data warehouse production system is that the offline data warehouse production system is very complete, the tool chain is relatively mature, the storage and maintenance costs are relatively low, and the development threshold for users is relatively low. However, the disadvantages are also very obvious:

[0004] First of all, the timeliness of data is very low, usually at the T+1 level, usually at the hourly level, or even the daily level.

[0005] Secondly, the Changelog support is not perfect. Although it is developed for Table, the intermediate storage Hive mainly supports Append (write only, no data update, no data deletion) type of data; at the same time, offline ETL is more suitable for processing full data rather than incremental updates.

[0006] As the amount of data increases, the execution time of offline ETL becomes longer and longer, and the business requirements for data freshness are also getting higher and higher. The business urgently needs a new low-latency data warehouse production system. Therefore, a real-time data warehouse production system has been further evolved based on the offline data warehouse.

[0007] A typical example is the real-time data warehouse production system based on the Lambda architecture. In the real-time data warehouse production system based on the Lambda architecture, the business needs to maintain two links, which divide the production link into a stream processing layer and a batch processing layer:

[0008] The stream processing layer is mainly used to process incremental data in real time. As an acceleration layer for the batch processing layer, this layer usually uses real-time computing engines such as Storm and Flink to process data. The intermediate results are stored in Kafka to provide low-latency streaming consumption capabilities.

[0009] The batch processing layer is the same as the offline data warehouse, completing the T+1 data result output. The service layer will provide external services based on the results of the stream processing layer and the batch processing layer.

[0010] With the continuous development of streaming computing engines, Flink, for example, has achieved the unification of stream and batch computing at the computing layer. In some scenarios, the batch processing layer can be completely removed, and the streaming processing layer can complete the full + incremental computing. In order to provide OLAP query capabilities for intermediate key data, Kafka data still needs to be exported to the data warehouse.

[0011] In this system, the advantages are as follows:

[0012] The data is very timely, usually in seconds. To reduce query latency, a lot of pre-calculations can be done based on the stream processing layer to reduce query latency.

[0013] The disadvantages are as follows: Two sets of technology stacks, high operation and maintenance costs; data warehouse maintenance personnel need to maintain two completely different links from computing to storage, and the development and maintenance costs are relatively high. The storage cost is high. In order to provide low-latency streaming consumption capabilities, Kafka has a higher storage cost than offline storage such as HDFS and S3. At the same time, in order to enable the intermediate data to provide offline query capabilities, an additional copy of the full offline data needs to be stored.

[0014] The data caliber of offline and real-time links is difficult to align. Two completely different sets of technology stacks are used to build the stream processing layer and the batch processing layer. Although the logical abstraction is the same, there are still differences in the specific implementation. In addition, the data of the stream processing layer is constantly processed incrementally, and it is difficult to align the results with the offline processing layer based on a fixed time point.

[0015] The intermediate results of stream computing cannot be queried. The intermediate results in the stream processing link cannot be queried because Kafka only supports streaming sequential consumption and does not have the ability to query points or batches. Although Kafka data can be exported to Hive, the real-time performance is relatively poor.

[0016] In summary, the defect of the prior art is that there is a lack of a data management method that can meet the needs of real-time analysis and offline analysis while reducing operation and maintenance costs. Summary of the invention

[0017] In view of the problem that the existing data processing framework combining offline data and real-time data lacks a framework that can realize the needs of real-time analysis and offline analysis while reducing operation and maintenance costs, the present invention provides a data management method that integrates real-time analysis and offline analysis.

[0018] In order to achieve the above technical objectives, the technical solution adopted by the present invention is as follows:

[0019] A data management method integrating real-time analysis and offline analysis, comprising the steps of:

[0020] S1, using LSM-Tree structure to store data;

[0021] S2. Replace the WAL log in the LSM-Tree structure with the change log;

[0022] S3, the change log (ChangeLog) uses snapshots to align with the SST data file of the LSM-Tree, and the messages are queried according to the snapshots;

[0023] S4, locate the snapshot, locate the log file, read the log data and record the offset;

[0024] S5. The query of real-time calculation intermediate data (ChangeLog) is converted into a query of LSM-Tree SST data file.

[0025] Furthermore, the LSM-Tree structure includes six levels, namely: metadata, snapshot list, table snapshot list Manifest List, table file list Manifest, data partition bucket and LSM-Tree data.

[0026] Furthermore, metadata stores table information and its definition, and points to a snapshot list of the table;

[0027] In the snapshot list, as the latest status record of the table, the snapshot records the latest data file (DataFile) and log file (ChangeLog) of the table;

[0028] The table snapshot list Manifest List is used for the current snapshot execution of the table;

[0029] The table file manifest records which data partitions exist under the table snapshot; points to the data bucket, and which data files and ChangeLogs are contained in the bucket, and marks the newly added and deleted files;

[0030] Data partitioning and bucketing include data partitioning and data bucketing;

[0031] Data partitioning can be used to horizontally split table data according to dimensions such as time;

[0032] Each data partition contains N data buckets. The number of buckets is set dynamically or specified by the user. One data bucket contains a complete LSM Tree.

[0033] An LSM-Tree data consists of two parts: a set of SST data files (DataFile) and a change log (ChangeLog).

[0034] Furthermore, the SST data file, from which the latest records of data can be read, performs memory compression (Compacttion) periodically;

[0035] Furthermore, the change log is a log data file of the message middleware, which is read in a streaming manner according to the offset (Offset) to provide incremental reading services for stream computing. The design of LSM-Tree is based on the design idea of ​​the database engine, and does not provide message access capabilities that support stream computing. Therefore, this solution needs to redesign the WAL log of LSM-Tree so that it can provide message access capabilities for stream computing. In the original design of the WAL Log of LSM-Tree, after the data is written to the SST File, it will be safely deleted, but as a stream message, it needs to be saved for a period of time before it can be cleaned up, otherwise the consumer may miss the message due to abnormal exit and other reasons. In order to distinguish it from the WALLog of LSM-Tree and better represent the meaning of the log file in this proposal, it is named ChangeLog, which is similar to the WAL log of LSM-Tree, which records the insertion, deletion, and update records of the data in its corresponding data file. Through ChangeLog, all operation histories of the table can be completely restored.

[0036] Furthermore, offline real-time data alignment and real-time intermediate result query: the snapshot start mark and snapshot end mark are recorded in the ChangeLog, and the alignment of offline data (LSM-Tree SST data file) and real-time data ChangeLog is achieved based on the snapshot identifier.

[0037] Therefore, querying the intermediate results of real-time calculation can be converted into a query on the LSM-Tree SST data file.

[0038] The SST data files of LSM-Tree are sorted according to the primary key, so better query efficiency can be guaranteed.

[0039] Furthermore, when writing data to the log, insert and delete records can be written as they are, and update records are divided into two cases:

[0040] With the record before the update, there will be two events in the update behavior log, one is the data record before the update, and the other is the data record after the update;

[0041] Without the record before the update, the update behavior log will be reflected as one event, that is, the data record after the update;

[0042] We use +I to indicate inserting a record, -D to indicate deleting a record, +U to indicate an updated record, and -U to indicate a record before the update.

[0043] Data writing includes no log mode, direct log writing mode, log completion before writing mode, and log completion after writing mode.

[0044] Further, offline data reading includes locating snapshots, locating data files, and reading data;

[0045] Positioning snapshot

[0046] Access metadata and obtain a snapshot of the table based on the table name. If no snapshot is specified, the latest snapshot (current snapshot) is used.

[0047] Locating data files

[0048] Find the Manifest List according to the snapshot, then obtain the Manifest file, and then obtain the data partitions and data buckets involved in this reading, and obtain the data bucket list that needs to be read;

[0049] Reading Data

[0050] Get the data bucket to be read from the data bucket list and read the SST data file; if it is concurrent reading, the program that reads the data needs to allocate the concurrent reading order of the data buckets by itself.

[0051] Furthermore, in real-time data reading, the log data is read in an incremental reading manner. When reading log data, first determine whether it is the first time to read the data. The first data reading adopts 4 reading methods:

[0052] Read all historical data from the current snapshot, and use the first log record after the end mark of the current snapshot as the incremental reading start position;

[0053] Read the full amount of historical data from the specified snapshot, and use the first log record after the specified snapshot end marker as the incremental reading start position;

[0054] Do not read historical data. Use the first log record after the end mark of the current snapshot as the incremental reading start position.

[0055] Do not read historical data, and use the first log record after the specified snapshot end mark as the incremental reading start position;

[0056] If it is not the first time to read the data, get the offset of the last read according to the grouping given during reading, and read the data incrementally from the last read offset position.

[0057] Furthermore, the detailed steps of the real-time data reading step include:

[0058] Positioning snapshot

[0059] Access metadata and obtain a snapshot of the table based on the table name. If no snapshot is specified, the latest snapshot (current snapshot) is used.

[0060] Locating log files

[0061] Find the Manifest List according to the snapshot, then obtain the Manifest file, and then obtain the data partitions and data buckets involved in this reading;

[0062] Reading log data

[0063] If it is the first time to read the data, read the data in one of the four ways of reading the data for the first time;

[0064] If it is not the first time to read the data, get the offset of the last read according to the grouping given during reading, and read the data incrementally from the last read offset position;

[0065] Record offset

[0066] According to the grouping, the offset read for each data partition is recorded separately, so that the next time it is read, incremental reading can be performed based on the offset.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] Reduce operation and maintenance costs. A data organization architecture that integrates offline data and real-time data only requires a technology stack to meet real-time analysis and offline analysis needs. There is no need to use Kafka+data warehouse. The architecture is simple and operation and maintenance costs are reduced.

[0069] It reduces development costs and provides a unified access interface for data batches and streams, reducing the development and testing costs of adapting to different data sources. Based on the LSM-Tree SST+ChangeLog approach, it aligns offline data and real-time data, making intermediate data of real-time computing traceable, greatly reducing the debugging costs of real-time computing codes.

[0070] It provides efficient data update and deletion under large data scale. Based on the idea of ​​LSM-Tree, it can achieve better data query, deletion and update costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 This is an overall flow chart of a data management method integrating real-time analysis and offline analysis in an embodiment of the present invention;

[0072] Figure 2 Schematic diagram of the data organization structure of LSM-Tree fusion in an embodiment of the present invention. DETAILED DESCRIPTION

[0073] In order to facilitate the understanding of those skilled in the art, the present invention is further described below in conjunction with embodiments and drawings. The contents mentioned in the implementation modes are not intended to limit the present invention.

[0074] like Figure 1 and 2 As shown, this embodiment provides a data management method integrating real-time analysis and offline analysis, including the steps of:

[0075] S1, using LSM-Tree structure to store data;

[0076] S2. Replace the WAL log in the LSM-Tree structure with the change log;

[0077] S3,ChangeLog uses snapshots to align with the SST data files of the LSM-Tree, and messages are queried according to the snapshots;

[0078] S4, locate the snapshot, locate the log file, read the log data and record the offset;

[0079] S5. The query of real-time calculation intermediate data (ChangeLog) is converted into a query of LSM-Tree SST data file.

[0080] The LSM-Tree structure consists of six levels: metadata, snapshot list, table snapshot list Manifest List, table file list Manifest, data partition buckets and LSM-Tree data.

[0081] The metadata stores table information and its definition, and points to the snapshot list of the table;

[0082] In the snapshot list, as the latest status record of the table, the snapshot records the latest data file (DataFile) and log file (ChangeLog) of the table;

[0083] The table snapshot list Manifest List is used for the current snapshot execution of the table;

[0084] The table file manifest records which data partitions exist under the table snapshot, which data buckets they point to, which data files and ChangeLogs are contained in the buckets, and marks the newly added and deleted files.

[0085] Data partitioning and bucketing include data partitioning and data bucketing;

[0086] Data partitioning can be used to horizontally split table data according to dimensions such as time;

[0087] Each data partition contains N data buckets. The number of buckets is dynamically set or specified by the user. One data bucket contains a complete LSM Tree.

[0088] An LSM-Tree data consists of two parts: a set of SST data files (DataFile) and a change log (ChangeLog).

[0089] SST data file, from which the latest data records can be read, and memory compression (Compacttion) can be performed regularly;

[0090] The change log is a log data file of the message middleware, which is read in a streaming manner according to the offset (Offset) to provide incremental reading services for stream computing. The design of LSM-Tree is based on the design idea of ​​the database engine, and does not provide the message access capability to support stream computing. Therefore, this solution needs to redesign the WAL log of LSM-Tree so that it can provide message access capability for stream computing. In the original design of the WAL Log of LSM-Tree, after the data is written to the SST File, it will be safely deleted, but as a stream message, it needs to be saved for a period of time before it can be cleaned up, otherwise the consumer may miss the message due to abnormal exit and other reasons. In order to distinguish it from the WAL Log of LSM-Tree and better represent the meaning of the log file in this proposal, it is named ChangeLog, which is similar to the WAL log of LSM-Tree, which records the insertion, deletion, and update records of the data in its corresponding data file. Through ChangeLog, all operation history of the table can be completely restored.

[0091] Offline real-time data alignment and real-time intermediate result query: The snapshot start mark and snapshot end mark are recorded in the ChangeLog, and the offline data (LSM-Tree SST data file) and real-time data ChangeLog are aligned based on the snapshot mark.

[0092] Therefore, querying the intermediate results of real-time calculation can be converted into a query on the LSM-Tree SST data file.

[0093] The SST data files of LSM-Tree are sorted according to the primary key, so better query efficiency can be guaranteed.

[0094] When writing data to the log, insert and delete records can be written as they are. There are two cases for updating records:

[0095] With the record before the update, there will be two events in the update behavior log, one is the data record before the update, and the other is the data record after the update;

[0096] Without the record before the update, the update behavior log will be reflected as one event, that is, the data record after the update;

[0097] We use +I to indicate inserting a record, -D to indicate deleting a record, +U to indicate an updated record, and -U to indicate a record before the update.

[0098] Data writing includes no log mode, direct log writing mode, log completion before writing mode, and log completion after writing mode.

[0099] Detailed process of data writing in no-log mode:

[0100] Pre-create snapshot:

[0101] Access metadata, obtain the table, obtain the current snapshot, obtain the manifest file, and then obtain the manifest file to locate the data partition and data bucket;

[0102] Create a new snapshot based on the current snapshot, and update the new snapshot after the write is completed;

[0103] Writing to memory:

[0104] 2.1 Writing to MemTable

[0105] Write data to the MemTable (ordered data structure) in memory; when the data in the MemTable reaches a certain size, it will be converted to Immutable MemTable (read-only), and a new Memtable will be generated for subsequent data writing;

[0106] MemTable is written to disk:

[0107] 3.1 According to the strategy, flush the data of Immutable MemTable to disk and flush MemTable into SSTable (ordered storage file)

[0108] 3.2 Trigger SSTTable merge (optional)

[0109] If the merge trigger condition is met, the SSTables are merged level by level. Each level has a threshold, and when the threshold is reached, the next level of data merging is triggered.

[0110] At the end of the data import, a Flush operation is performed on the MemTable to ensure that all data is written to the disk.

[0111] Commit snapshot:

[0112] (1) When the data is finally imported, a Flush operation is performed on the MemTable to ensure that all data is written to the disk;

[0113] (2) Update the changes (addition, deletion, update) of the corresponding data partitions, data files of the data buckets, and log files in the Manifest file; update the Manifest (addition, deletion, update) in the Manifest List, add the newly created snapshot in the snapshot list, and set it as the current snapshot.

[0114] After all the above operations are completed, the writing process is completed.

[0115] Detailed process of writing data directly into log mode:

[0116] Pre-create snapshot:

[0117] (1) Access metadata, obtain the table, obtain the current snapshot, obtain the manifest file, and then obtain the manifest file to locate the data partition and data bucket;

[0118] (2) Create a new snapshot based on the current snapshot, and update the new snapshot after the writing is completed;

[0119] Writing to memory:

[0120] Write data to the MemTable (ordered data structure) in memory; when the data in the MemTable reaches a certain size, it will be converted to Immutable MemTable (read-only), and a new Memtable will be generated for subsequent data writing;

[0121] Generate log:

[0122] Convert the data into a log format, write the data into a log record, and then write a snapshot identifier into the log;

[0123] MemTable is written to disk:

[0124] (1) According to the strategy, the data of Immutable MemTable is flushed to disk, and MemTable is flushed to SSTable (ordered storage file)

[0125] (2) Trigger SSTTable merge (optional)

[0126] If the merge trigger condition is met, the SSTables are merged level by level. Each level has a threshold, and when the threshold is reached, the next level of data merging is triggered.

[0127] Commit snapshot:

[0128] (1) When the data is finally imported, a Flush operation is performed on the MemTable to ensure that all data is written to the disk;

[0129] (2) Update the changes (addition, deletion, update) of the corresponding data partitions, data files of the data buckets, and log files in the Manifest file; update the Manifest (addition, deletion, update) in the Manifest List, add the newly created snapshot in the snapshot list, and set it as the current snapshot.

[0130] After all the above operations are completed, the writing process is completed.

[0131] Detailed process of writing data in log mode before writing:

[0132] Pre-created snapshots

[0133] (1) Access metadata, obtain the table, obtain the current snapshot, obtain the manifest file, and then obtain the manifest file to locate the data partition and data bucket;

[0134] (2) Create a new snapshot based on the current snapshot, and update the new snapshot after the writing is completed;

[0135] Writing to Memory

[0136] Write data to the MemTable (ordered data structure) in memory; when the data in the MemTable reaches a certain size, it will be converted to Immutable MemTable (read-only), and a new Memtable will be generated for subsequent data writing;

[0137] Generate logs

[0138] Before writing to disk, first access the data file in the disk, find the record before the data update (-U), complete the data before the update, convert the data into log format, write the snapshot start mark in the log, write the data into the log record, and then write the snapshot end mark in the log;

[0139] MemTable written to disk

[0140] 4.1 According to the strategy, flush the data of Immutable MemTable to disk and flush MemTable into SSTable (ordered storage file)

[0141] 4.2 Trigger SSTTable merge (optional)

[0142] If the merge trigger condition is met, the SSTables are merged level by level. Each level has a threshold, and when the threshold is reached, the next level of data merging is triggered.

[0143] Commit Snapshot

[0144] (1) When the data is finally imported, a Flush operation is performed on the MemTable to ensure that all data is written to the disk;

[0145] (2) Update the changes (addition, deletion, update) of the corresponding data partitions, data files of the data buckets, and log files in the Manifest file; update the Manifest (addition, deletion, update) in the Manifest List, add the newly created snapshot in the snapshot list, and set it as the current snapshot.

[0146] After all the above operations are completed, the writing process is completed.

[0147] Detailed process of writing data in post-write log mode:

[0148] Pre-created snapshots

[0149] (1) Access metadata, obtain the table, obtain the current snapshot, obtain the manifest file, and then obtain the manifest file to locate the data partition and data bucket;

[0150] (2) Create a new snapshot based on the current snapshot, and update the new snapshot after the writing is completed;

[0151] Writing to Memory

[0152] Write data to the MemTable (ordered data structure) in memory; when the data in the MemTable reaches a certain size, it will be converted to Immutable MemTable (read-only), and a new Memtable will be generated for subsequent data writing;

[0153] Generate logs

[0154] Before writing to disk, first access the data file in the disk, find the record before the data update (-U), complete the data before the update, convert the data into log format, write the snapshot start mark in the log, write the data into the log record, and then write the snapshot end mark in the log;

[0155] Writing data files

[0156] 4.1 According to the strategy, flush the data of Immutable MemTable to disk and flush MemTable into SSTable (ordered storage file)

[0157] 4.2 Trigger SSTTable merge (optional)

[0158] If the merge trigger condition is met, the SSTables are merged level by level. Each level has a threshold, and when the threshold is reached, the next level of data merging is triggered.

[0159] Commit Snapshot

[0160] (1) When the data is finally imported, a Flush operation is performed on the MemTable to ensure that all data is written to the disk;

[0161] (2) Update the changes (addition, deletion, update) of the corresponding data partitions, data files of the data buckets, and log files in the Manifest file; update the Manifest (addition, deletion, update) in the Manifest List, add the newly created snapshot in the snapshot list, and set it as the current snapshot.

[0162] After all the above operations are completed, the writing process is completed.

[0163] Offline data reading, including locating snapshots, locating data files, and reading data;

[0164] Positioning snapshot

[0165] Access metadata and obtain a snapshot of the table based on the table name. If no snapshot is specified, the latest snapshot (current snapshot) is used.

[0166] Locating data files

[0167] Find the Manifest List according to the snapshot, then obtain the Manifest file, and then obtain the data partitions and data buckets involved in this reading, and obtain the data bucket list that needs to be read;

[0168] Reading Data

[0169] Get the data bucket to be read from the data bucket list and read the SST data file; if it is concurrent reading, the program that reads the data needs to allocate the concurrent reading order of the data buckets by itself.

[0170] In real-time data reading, log data is read in incremental reading mode. When reading log data, first determine whether it is the first time to read the data. The first data reading uses 4 methods:

[0171] Read all historical data from the current snapshot, and use the first log record after the end mark of the current snapshot as the incremental reading start position;

[0172] Read the full amount of historical data from the specified snapshot, and use the first log record after the specified snapshot end marker as the incremental reading start position;

[0173] Do not read historical data. Use the first log record after the end mark of the current snapshot as the incremental reading start position.

[0174] Do not read historical data, and use the first log record after the specified snapshot end mark as the incremental reading start position;

[0175] If it is not the first time to read the data, get the offset of the last read according to the grouping given during reading, and read the data incrementally from the last read offset position.

[0176] The detailed steps of real-time data reading steps include:

[0177] Positioning snapshot

[0178] Access metadata and obtain a snapshot of the table based on the table name. If no snapshot is specified, the latest snapshot (current snapshot) is used.

[0179] Locating log files

[0180] Find the Manifest List according to the snapshot, then obtain the Manifest file, and then obtain the data partitions and data buckets involved in this reading;

[0181] Reading log data

[0182] If it is the first time to read the data, read the data in one of the four ways of reading the data for the first time;

[0183] If it is not the first time to read the data, get the offset of the last read according to the grouping given during reading, and read the data incrementally from the last read offset position;

[0184] Record offset

[0185] According to the grouping, the offset read for each data partition is recorded separately, so that the next time it is read, incremental reading can be performed based on the offset.

[0186] Compared with the prior art, the present invention has the following beneficial effects:

[0187] Reduce operation and maintenance costs. A data organization architecture that integrates offline data and real-time data only requires a technology stack to meet real-time analysis and offline analysis needs. There is no need to use Kafka+data warehouse. The architecture is simple and operation and maintenance costs are reduced.

[0188] It reduces development costs and provides a unified access interface for data batches and streams, reducing the development and testing costs of adapting to different data sources. Based on the LSM-Tree SST+ChangeLog approach, it aligns offline data and real-time data, making intermediate data of real-time computing traceable, greatly reducing the debugging costs of real-time computing codes.

[0189] It provides efficient data update and deletion under large data scale. Based on the idea of ​​LSM-Tree, it can achieve better data query, deletion and update costs.

[0190] Improve LSM-Tree WAL Log to ChangeLog and ChangeLog generation method

[0191] In the data bucketing, LSM-Tree is used as the data management bucket for SST data files and ChangeLog logs. The WAL Log of LSM-Tree is improved to ChangeLog. Four modes are used to generate ChangeLog when writing data, integrating offline data and real-time data.

[0192] A method for aligning offline data and real-time data based on snapshot identification

[0193] In ChangeLog, snapshot identifiers are used to align with the SST data files of LSM-Tree. This allows query capabilities for real-time intermediate data (ChangeLog), which can be converted into queries for LSM-Tree SST data files. There is no need to export data to a data warehouse like Kafka to implement point queries.

[0194] ChangeLog data consumption method based on data alignment

[0195] On the premise that the ChangeLog and LSM-Tree SST data files are aligned, you can flexibly choose whether to read the full amount of historical data stored in the LSM-Tree SST from a certain snapshot identifier, and then start incremental data consumption based on the snapshot identifier.

[0196] The above is a detailed introduction to a data management method that integrates real-time analysis and offline analysis provided by the present application. The description of the specific embodiments is only used to help understand the method and its core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A data management method integrating real-time analysis and offline analysis, characterized in that: Includes steps: S1, using LSM-Tree structure to store data; S2. Replace the WAL log in the LSM-Tree structure with the change log; S3, the change log uses snapshots to align with the SST data file of the LSM-Tree, and the messages are queried according to the snapshots; S4, locate the snapshot, locate the log file, read the log data and record the offset; S5. The query of the intermediate data calculated in real time is converted into the query of the SST data file of the LSM-Tree.

2. The data management method integrating real-time analysis and offline analysis according to claim 1, characterized in that: The LSM-Tree structure consists of six levels: metadata, snapshot list, table snapshot list ManifestList, table file list Manifest, data partition buckets and LSM-Tree data.

3. The data management method integrating real-time analysis and offline analysis according to claim 2, characterized in that: The metadata stores table information and its definition, and points to the snapshot list of the table; In the snapshot list, as the latest status record of the table, the snapshot records the latest data files and log files of the table; The table snapshot list Manifest List is used for the current snapshot execution of the table; The table file manifest records which data partitions are under the table snapshot; which data buckets are pointed to, and which data buckets contain data files and change logs, and marks the newly added and deleted files; Data partitioning and bucketing include data partitioning and data bucketing; Data partitioning divides table data horizontally according to the time dimension; Each data partition contains N data buckets. The number of buckets is set dynamically or specified by the user. One data bucket contains a complete LSM Tree. An LSM-Tree data consists of two parts: a set of SST data files and a change log.

4. The data management method integrating real-time analysis and offline analysis according to claim 3 is characterized in that: The SST data file reads the latest record of data and performs memory compression periodically.

5. The data management method integrating real-time analysis and offline analysis according to claim 4, characterized in that: The change log is a log data file of the message middleware, which is read in a streaming manner according to the offset and provides incremental reading services for stream computing.

6. The data management method integrating real-time analysis and offline analysis according to claim 5, characterized in that: Offline real-time data alignment and real-time intermediate result query: The snapshot start mark and snapshot end mark are recorded in the change log, and the offline data and real-time data change log are aligned based on the snapshot mark. Therefore, the query of the intermediate results of real-time calculation is converted into a query on the SST data file of the LSM-Tree. The SST data files of LSM-Tree are sorted according to the primary key to ensure better query efficiency.

7. The data management method integrating real-time analysis and offline analysis according to claim 6, characterized in that: When writing data to the log, insert and delete records are written as they are, and update records are divided into two cases: With the record before the update, there will be two events in the update behavior log, one is the data record before the update, and the other is the data record after the update; Without the record before the update, the update behavior log will be reflected as one event, that is, the data record after the update; Use +I to insert a record, -D to delete a record, +U to update a record, and -U to update a record.

8. The data management method integrating real-time analysis and offline analysis according to claim 7, characterized in that: Data writing includes no log mode, direct log writing mode, log completion before writing mode, and log completion after writing mode.

9. The data management method integrating real-time analysis and offline analysis according to claim 8, characterized in that: Offline data reading, including locating snapshots, locating data files, and reading data; Positioning snapshot: Access metadata and obtain a snapshot of the table based on the table name. If no snapshot is specified, the latest snapshot is used. Locating data files: Find the Manifest List according to the snapshot, then obtain the Manifest file, and then obtain the data partitions and data buckets involved in this reading, and obtain the data bucket list that needs to be read; Reading data: From the data bucket list, obtain the data bucket to be read and read the SST data file; If it is concurrent reading, the program that reads the data needs to allocate the concurrent reading order of the data buckets by itself.

10. The data management method integrating real-time analysis and offline analysis according to claim 9, characterized in that: In real-time data reading, log data is read in incremental reading mode. When reading log data, first determine whether it is the first time to read the data. The first data reading uses 4 methods: Read all historical data from the current snapshot, and use the first log record after the end mark of the current snapshot as the incremental reading start position; Read the full amount of historical data from the specified snapshot, and use the first log record after the specified snapshot end marker as the incremental reading start position; Do not read historical data. Use the first log record after the end mark of the current snapshot as the incremental reading start position. Do not read historical data, and use the first log record after the specified snapshot end mark as the incremental reading start position; If it is not the first time to read the data, get the offset of the last read according to the grouping given during reading, and read the data incrementally from the last read offset position.

Citation Information

Patent Citations

  • Data storage device and storage control method based on log structured merge tree

    CN116414304A

  • Data integrity validation on LSM tree snapshots

    WO2021061173A1

Cited By

  • Distributed database incremental snapshot method and device and computer equipment

    CN120892259A