Data query method and device of WAL log, time sequence database and medium

By positioning the query target based on the WAL log encoding structure in the timing database, only necessary data page locking is performed, the problem of low cache utilization is solved and query performance is improved.

CN120492490AActive Publication Date: 2025-08-15TSINGHUA UNIVERSITY

Patent Information

Application Number
CN202511000125.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-08-15
Estimated Expiration
2045-07-21

AI Technical Summary

Technical Problem

In the prior art, the timing database performs unnecessary data page locking during data query, resulting in low cache utilization, a large number of useless caches, affecting query performance.

Method used

In the timing database, by receiving log information query instructions, the metadata of the data table is found in the global metadata, the written location information is determined, and the query target is positioned based on the pre-constructed WAL log encoding structure, and only necessary data page locking is performed to improve cache utilization.

Benefits of technology

Improve cache utilization, reduce useless cache, and ensure the stability and efficiency of database query performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492490A_ABST
    Figure CN120492490A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electric digital data processing, in particular to a WAL log data query method and device, a time sequence database and a medium, and the method comprises the steps that a log information query instruction is received to obtain a query target; in response to the log information query instruction, searching metadata of a corresponding data table in global metadata of a time sequence database, and determining write-in position information of a query target based on the metadata; based on the write-in position information and a pre-constructed WAL log coding structure, the query target is positioned, and the WAL log coding structure comprises a plurality of data blocks, a plurality of data pages and a data segment. Therefore, the technical problems that in the related technology, unnecessary data page locking is carried out during data query, the cache utilization rate is low, a large number of useless caches are likely to be generated, and then the database query performance is affected are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic digital data processing, and in particular to a data query method and device for a WAL log, a time series database, and a medium. Background Art

[0002] Time series data is written to the database at high speed, in large quantities, and continuously.

[0003] In related technologies, when reading data, if the data has not yet been flushed to the local disk, the data page remains locked and cannot be swapped out, even if the data being read does not contain the data page. This results in reduced cache utilization, a large amount of useless cache, and other data that needs to be cached cannot be cached, affecting database query performance. This urgently needs improvement. Summary of the Invention

[0004] The present invention provides a data query method and device for a WAL log, a time series database, and a medium to solve the technical problems in related technologies such as unnecessary data page locking during data query, low cache utilization, and the generation of a large amount of useless cache, which in turn affects database query performance.

[0005] A first aspect of the present invention provides a data query method for a WAL log, which is applied to a time series database, wherein the method includes the following steps: receiving a log information query instruction to obtain a query target; in response to the log information query instruction, searching for metadata of a corresponding data table in the global metadata of the time series database, and determining write location information of the query target based on the metadata; locating the query target based on the write location information and a pre-constructed WAL log encoding structure, wherein the WAL log encoding structure includes multiple data blocks, multiple data pages, and data segments.

[0006] Optionally, in one embodiment of the present invention, before receiving the log information query instruction, it also includes: receiving a time series data import instruction, and obtaining the WAL log generated when the time series data is imported; mapping the log information of the WAL log to a corresponding data table, and writing the data table and the log information into a memory composed of the WAL log encoding structure.

[0007] Optionally, in one embodiment of the present invention, each data block stores WAL log information of a data table, each data block stores multiple log information, each log information includes a data header and data, wherein the data header is a timestamp and serial number stored for each log information; each data page includes the multiple data blocks, and the page header of each data page records pre-allocated storage information, a serial number, the number of data blocks, an identifier of the data table corresponding to each data block, a timestamp range, and an offset in the data page; each data segment includes the multiple data pages.

[0008] Optionally, in one embodiment of the present invention, writing the data table and the log information into the memory composed of the pre-built WAL log encoding structure includes: screening the memory that meets the preset storage conditions and initializing each data segment in the memory to obtain the pre-allocated storage information of each data segment; based on the pre-allocated storage information, writing the corresponding number of log information into each data segment in a preset order until all log information is written.

[0009] Optionally, in one embodiment of the present invention, the initialization of each data segment in the memory to obtain the pre-allocated storage information of each data segment includes: storing multiple log information of the corresponding number of log information in the current data block in the current data page of the current data segment until the current data block reaches the pre-allocated number of storage items; storing the remaining log information in the next data block until all data blocks in the current data page reach the pre-allocated number of storage items; storing the remaining log information in the next data page until all data pages in the current data segment reach the pre-allocated storage upper limit, determining that the corresponding number of log information has been written to the current data segment, and updating the storage location of each log information to the metadata of the corresponding data table.

[0010] Optionally, in one embodiment of the present invention, the calculation expression for the actual storage capacity of the current data page is: , in, represents the actual storage capacity, Indicates the number of pre-allocated storage strips, Indicates the corresponding n The size of the data table, Indicates the corresponding n A data table.

[0011] Optionally, in one embodiment of the present invention, locating the query target based on the write position information and the pre-constructed WAL log encoding structure includes: locating at least one data segment to be scanned in the WAL log encoding structure based on the write position information, and traversing the at least one data segment to be scanned using the identifier or timestamp of the at least one data table to find the corresponding data page; traversing the corresponding data page using the timestamp to locate the corresponding data block, and storing the corresponding positioning information in a candidate list; searching for log information that meets the preset query conditions in the candidate list until all data segments to be scanned are scanned.

[0012] A second aspect of the present invention provides a data query device for a WAL log, which is applied to a time series database, wherein the device includes: a receiving module, used to receive a log information query instruction to obtain a query target; a search module, used to search for metadata of a corresponding data table in the global metadata of the time series database in response to the log information query instruction, and determine the write location information of the query target based on the metadata; a positioning module, used to locate the query target based on the write location information and a pre-constructed WAL log encoding structure, wherein the WAL log encoding structure includes multiple data blocks, multiple data pages and data segments.

[0013] Optionally, in one embodiment of the present invention, it further includes: an acquisition module for receiving a time series data import instruction and acquiring a WAL log generated when the time series data is imported; a writing module for mapping the log information of the WAL log to a corresponding data table, and writing the data table and the log information into a memory composed of the WAL log encoding structure.

[0014] Optionally, in one embodiment of the present invention, each data block stores WAL log information of a data table, each data block stores multiple log information, each log information includes a data header and data, wherein the data header is a timestamp and serial number stored for each log information; each data page includes the multiple data blocks, and the page header of each data page records pre-allocated storage information, a serial number, the number of data blocks, an identifier of the data table corresponding to each data block, a timestamp range, and an offset in the data page; each data segment includes the multiple data pages.

[0015] Optionally, in one embodiment of the present invention, the writing module includes: an initialization unit for screening memory that meets preset storage conditions and initializing each data segment in the memory to obtain pre-allocated storage information for each data segment; a writing unit for writing a corresponding number of log information into each data segment in a preset order based on the pre-allocated storage information until all log information is written.

[0016] Optionally, in one embodiment of the present invention, the initialization unit includes: a first storage sub-unit, used to store multiple log information of the corresponding number of log information in the current data block in the current data page of the current data segment, until the current data block reaches the pre-allocated number of storage entries; a second storage sub-unit, used to store the remaining log information in the next data block, until all data blocks in the current data page reach the pre-allocated number of storage entries; a third storage sub-unit, used to store the remaining log information in the next data page, until all data pages in the current data segment reach the pre-allocated storage upper limit, determine that the corresponding number of log information has been written to the current data segment, and update the storage location of each log information to the metadata of the corresponding data table.

[0017] Optionally, in one embodiment of the present invention, the calculation expression for the actual storage capacity of the current data page is: , in, represents the actual storage capacity, Indicates the number of pre-allocated storage strips, Indicates the corresponding n The size of the data table, Indicates the corresponding n A data table.

[0018] Optionally, in one embodiment of the present invention, the positioning module includes: a first positioning unit, used to locate at least one data segment to be scanned in the WAL log encoding structure based on the write position information, and traverse the at least one data segment to be scanned using the identifier or timestamp of the at least one data table to find the corresponding data page; a second positioning unit, used to traverse the corresponding data page using the timestamp to locate the corresponding data block, and store the corresponding positioning information in a candidate list; a scanning unit, used to search the candidate list for log information that meets the preset query conditions until all data segments to be scanned are scanned.

[0019] A third aspect of the present invention provides a time series database, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the WAL log data query method as described in the above embodiment.

[0020] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the WAL log data query method as described in the above embodiment.

[0021] A fifth aspect of the present invention provides a computer program product, including a computer program, which, when executed, is used to implement the above-mentioned WAL log data query method.

[0022] In response to log information query instructions, embodiments of the present invention can search the global metadata of a time series database for metadata related to the corresponding data table. Only necessary data page locks are performed, improving cache utilization. The write location information of the query target is determined based on the metadata. This information is combined with the WAL (Write-Ahead Logging) encoding structure in memory to locate the query target. This ensures query performance while supporting direct database queries of the WAL log. This solves the technical problem in related technologies of unnecessary data page locks during data queries, resulting in low cache utilization and the generation of a large amount of unused cache, which in turn affects database query performance.

[0023] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 A schematic diagram of the structure of a WAL log encoding structure provided according to an embodiment of the present invention; Figure 2 A schematic diagram of the structure of the encoding structure of a WAL log provided according to one embodiment of the present invention; Figure 3 A flowchart of a method for querying WAL log data according to an embodiment of the present invention; Figure 4 A schematic diagram of the writing principle of a data query method for a WAL log provided according to one embodiment of the present invention; Figure 5A schematic diagram of the query principle of a WAL log data query method provided by one embodiment of the present invention; Figure 6 A schematic diagram of the structure of a data query device for a WAL log according to an embodiment of the present invention; Figure 7 A schematic diagram of the structure of a time series database provided according to an embodiment of the present invention.

[0025] Among them, 10-WAL log encoding structure, 101-data block, 102-data page, 103-data segment; 20-WAL log data query device, 201-receiving module, 202-search module, 203-positioning module; 701-memory, 702-processor, 703-communication interface. DETAILED DESCRIPTION

[0026] The following describes embodiments of the present invention in detail, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and are not to be construed as limiting the present invention.

[0027] The following describes, with reference to the accompanying drawings, a data query method, device, time series database, and medium for WAL logs according to embodiments of the present invention. In response to the technical problem in the related art mentioned in the background art above, where unnecessary data page locking is performed during data query, resulting in low cache utilization and a large amount of useless cache, which in turn affects database query performance, the present invention provides a data query method for WAL logs. The method can, in response to a log information query instruction, search the global metadata of the time series database for metadata of the corresponding data table, perform only necessary data page locks, thereby improving cache utilization, and determine the write location information of the query target based on the metadata. The write location information and the WAL log encoding structure in memory are combined to locate the query target, thereby supporting the database to directly query the WAL log while ensuring query performance. This solves the technical problem in the related art, where unnecessary data page locking is performed during data query, resulting in low cache utilization and a large amount of useless cache, thereby affecting database query performance.

[0028] Before describing the data query method for the WAL log according to an embodiment of the present invention, the memory structure involved in the embodiment of the present invention is first described for subsequent understanding.

[0029] In this embodiment of the present invention, the memory may be composed of multiple WAL log encoding structures, such as Figure 1 As shown, the WAL log encoding structure 10 of the embodiment of the present invention may include: a data block 101, a data page 102 and a data segment 103.

[0030] Specifically, multiple data blocks 101, each data block 101 stores WAL log information of a data table, each data block 101 stores multiple log information, each log information includes a data header and data, wherein the data header is a timestamp and sequence number stored for each log information.

[0031] Multiple data pages 102, each data page 102 includes multiple data blocks 101, and the page header of each data page records pre-allocated storage information, a sequence number, the number of data blocks, an identifier of a data table corresponding to each data block 101, a timestamp range, and an offset in the data page.

[0032] Data segments 103 , each data segment 103 includes multiple data pages 102 .

[0033] As you can understand, WAL logging is a technology widely used in databases and other data storage systems, primarily to ensure data persistence and transaction atomicity. Its core principle is to record modifications to a log file before making them to the actual data store. This allows data to be recovered from the log file even in the event of a system failure (such as a crash or power outage), ensuring that data is not lost or corrupted and that transactions are either fully completed or not executed at all.

[0034] like Figure 2 As shown, the WAL log encoding structure of this embodiment of the present invention consists of a three-level data structure. The first-level data structure is the block (data block 101). Each data block 101 contains and only contains WAL log information for the same data table. Each log information (i.e., slot) consists of two parts: a header and data. The header contains the timestamp of the data point plus the corresponding LSN (Log Sequence Number) information (i.e., sequence number).

[0035] The second-level data structure is the page (data page 102). Each data page 102 contains n (n >= 1) data blocks 101, each belonging to a different data table. The page header of each data table contains the pre-allocated slot, LSN, and the number of data blocks 101. Each data block 101 also contains the corresponding data table ID, timestamp range, and offset within the data page 102.

[0036] The third level structure is the segment (data segment 103), which contains n data pages 102. The data segment 103 is the smallest unit for storing and allocating WAL logs in memory and on disk.

[0037] The data segment 103 and data page 102 can be fixed length in this embodiment of the present invention, while the data block 101 can be floating length. The size of each data segment is set to 1 MiB (MB), and the maximum capacity of each data page is 64 KiB. Each page header contains the following content: 1. The data table ID array (i.e., the data table identification array) corresponding to the data stored in the data page 102.

[0038] 2. The data page 102 stores the range of the global LSN of the data.

[0039] The information of each data block 101 stored in the data page 102 includes the identifier of the corresponding data table, the timestamp range, the global LSN range, and the position of the data block 101 in the data page 102 (offset information).

[0040] According to the WAL log encoding structure proposed in the embodiment of the present invention, it can be implemented in the form of block-page-segment. The data block is the smallest storage unit and can store log information and record the data header and data of the log information. The data page can include multiple data blocks, and the contents of multiple data blocks are aggregated to form a page header. The data segment can include multiple data pages. By partitioning and sorting according to the data table, the fragmentation of the same table in the WAL log is reduced, the efficiency of data loading and insertion is improved, and the space occupation is reduced. When importing data, it can be stored layer by layer in the form of data block-data page-data segment, and continuous space is allocated in the memory, reducing the random reading and writing and fragmentation of the memory and disk. When querying or calling data, it can also be checked layer by layer in the form of data segment-data page-data block. Thus, the technical problem in the related art that data needs to be written to the WAL log and cache / disk synchronously when importing data is solved. The storage of two copies of data is not only costly but also affects the overall import performance. Moreover, when importing data, different data tables are stored in different files, causing the memory and disk to be repeatedly randomly read and written, resulting in serious storage fragmentation and affecting subsequent file calls.

[0041] Specifically, Figure 3 A flowchart of a method for querying WAL log data provided by an embodiment of the present invention.

[0042] like Figure 3 As shown in the figure, the data query method of the WAL log is applied to the time series database and includes the following steps: In step S301, a log information query instruction is received to obtain a query target.

[0043] It can be understood that time series data is a set of data points recorded in chronological order, with each data point associated with a specific point in time. This type of data typically reflects the changes in a variable at different moments in time and is an ordered, time-ordered data set. This type of data typically has the following characteristics: Timestamp: Each data point has a timestamp that identifies the time when the data was generated. Continuity: Data is typically generated continuously at fixed or irregular time intervals. High frequency: Data points are generated at a high frequency, such as every second, every minute, or every hour. Multidimensionality: Data can contain multiple dimensions, such as temperature, humidity, pressure, etc.

[0044] A time series database is a database system specifically designed to process time series data. Time series data refers to a series of data points arranged in chronological order, where each data point is typically a measurement value or state record at a specific moment in time.

[0045] The time series database of the embodiment of the present invention may be an Internet of Things (IoT)-based time series database to process and analyze time series data collected from various IoT devices.

[0046] Based on the characteristics of time series data, embodiments of the present invention can achieve high-quality query of log information.

[0047] During the actual execution process, the embodiment of the present invention can receive log information query instructions to clarify the query target in the log information query instruction, that is, the log information and corresponding timestamp corresponding to the time series data stored in the time series database, so as to locate the log information in the subsequent process and return the results to the user.

[0048] In step S302 , in response to the log information query instruction, metadata of the corresponding data table is searched in the global metadata of the time series database, and write location information of the query target is determined based on the metadata.

[0049] Furthermore, an embodiment of the present invention can lock relevant data tables in the global metadata of the time series database according to the log information query instruction, so as to perform metadata query on the locked data table and obtain the write location information of the query target, wherein the global metadata may include the storage location of each log information.

[0050] In step S303, the query target is located based on the write position information and the pre-built WAL log encoding structure, where the WAL log encoding structure includes multiple data blocks, multiple data pages and data segments.

[0051] According to the write position information and the WAL log encoding structure, the embodiment of the present invention can locate the query target to determine the data block where the query target is located, thereby realizing data query.

[0052] Optionally, in one embodiment of the present invention, locating the query target based on the write position information and a pre-constructed WAL log encoding structure includes: locating at least one data segment to be scanned in the WAL log encoding structure based on the write position information, and traversing at least one data segment to be scanned using an identifier or a timestamp of at least one data table to find a corresponding data page; traversing the corresponding data page using the timestamp to locate the corresponding data block, and storing the corresponding positioning information in a candidate list; searching for log information that meets the preset query conditions in the candidate list until all data segments to be scanned are scanned.

[0053] It's understandable that in an IoT-based time series database, data users query stored time series data. After receiving the query, the time series database scans and filters the data and returns the results to the user. Typically, query statements in a time series database are accompanied by a time window, meaning that only data with a timestamp within the time window is selected for further processing. Therefore, optimizing and processing the timestamp selection of time series data has a decisive impact on the performance of the time series database.

[0054] The embodiment of the present invention can simultaneously issue a scan request and return the scanned data as needed. The scanning process is as follows: Step S1. Obtain the identifier of the WAL data segment corresponding to the data of the table through the metadata information of the table; Step S2. Locate the location of the data segment to be scanned; Step S3: Get the header information and make the following judgments: (1) If the data table identifier array does not contain the identifier of the data table to be scanned, skip the page; (2) If the timestamp required by the query is not within the timestamp range of the corresponding data table, skip the page; (3) For the data block information that meets the conditions, record its offset position and store it in the candidate list for scanning; Step S4. Read data block information from the candidate list, read each log information, and if its sequence number meets the database visibility requirements and the corresponding timestamp belongs to the query access, then return the data, otherwise skip; Repeat steps S2-4 until all data segments are scanned.

[0055] Optionally, in one embodiment of the present invention, before receiving the log information query instruction, it also includes: receiving a time series data import instruction, and obtaining the WAL log generated when the time series data is imported; mapping the log information of the WAL log to a corresponding data table, and writing the data table and log information into a memory composed of a WAL log encoding structure.

[0056] It is understandable that the time series database needs to write log information first before it can implement subsequent log information queries. Here, the log information writing process is explained.

[0057] Data ingestion refers to the process of collecting data files from various sources and importing them into a database for storage, processing, and analysis. Data ingestion aims to cleanse data and store it in an accessible and consistent central repository for use within the organization. Data sources include financial systems, third-party data providers, social media platforms, IoT devices, SaaS (Software as a Service) applications, and on-premises business applications such as enterprise resource planning and customer relationship management. These data sources contain both structured and unstructured data. After ingestion, data can be stored in data lakes, data warehouses, integrated data lakes and warehouses, data marts, relational databases, and document storage systems. Organizations ingest data so that it can be used for business intelligence tasks, as well as for machine learning, predictive modeling, and artificial intelligence applications.

[0058] In an embodiment of the present invention, time series data, such as IoT time series data, can be obtained through data ingestion, and the time series data can be mapped to a data table to write the generated log information by importing the data table and the time series data.

[0059] Optionally, in one embodiment of the present invention, the data table and log information are written into a memory composed of a pre-built WAL log encoding structure, including: screening the memory that meets the preset storage conditions, and initializing each data segment in the memory to obtain pre-allocated storage information for each data segment; based on the pre-allocated storage information, writing a corresponding amount of log information into each data segment in a preset order until all log information is written.

[0060] Furthermore, the embodiment of the present invention can first apply for a whole block of memory and store it through the writing structure of the WAL log. For each data segment in the memory, the embodiment of the present invention can initialize the first page and the corresponding data block area.

[0061] During the actual execution process, the embodiment of the present invention can write log information based on the pre-allocated storage information obtained after initialization. For example, according to the pre-allocated storage information, the storage capacity of the current data segment is determined, and the corresponding amount of log information is stored. When the remaining space of the current data segment is insufficient and has been allocated to the new page data structure, a new data segment space is reapplied.

[0062] The preset order may be a chronological order or other orders, and may be specifically set by those skilled in the art according to actual conditions.

[0063] Optionally, in one embodiment of the present invention, each data segment in the memory is initialized to obtain pre-allocated storage information for each data segment, including: storing multiple log information of a corresponding number into a current data block in a current data page of the current data segment until the current data block reaches the pre-allocated number of storage entries; storing the remaining log information into the next data block until all data blocks in the current data page reach the pre-allocated number of storage entries; storing the remaining log information into the next data page until all data pages in the current data segment reach the pre-allocated storage upper limit, determining that the corresponding number of log information has been written into the current data segment, and updating the storage location of each log information into the metadata of the corresponding data table, wherein the calculation expression for the actual storage capacity of the current data page is: , in, Indicates the actual storage capacity, Indicates the number of pre-allocated storage strips. Indicates the corresponding n The size of the data table, Indicates the corresponding n A data table.

[0064] For page initialization, the embodiment of the present invention may use the following formula to calculate the number of slots.

[0065] First, calculate the number of data sources (time series data) being imported n , the frequency of each data source f , and the size of each data m Since IoT time series data is structured data, the size m of each data item is equal to the size of each row of the data table corresponding to the data source. This size can be easily obtained from metadata, and this size is commonly used in the database, so there will be no additional performance loss.

[0066] , in, The maximum capacity of a data page is 64KB, and the denominator is the largest single data item among all data sources. tn Indicates the n data tables, is the number of slots. Then we can find the maximum number of entries that the page can store.

[0067] Afterwards, the embodiment of the present invention can use the following formula to calculate the size of the data block corresponding to each table, that is, the number of entries that each data block can store: , in, tn Indicates the corresponding n Data tables, that is, the overall data page space is divided according to the frequency of each data source, is the acquisition frequency of the nth data table, in Hz, i.e. the number of acquisitions per second. Finally, record the actual size used by the page, i.e. .

[0068] After obtaining the actual size, the embodiment of the present invention can be +Page header size (page header size), allocate actual memory for the page, then the allocator will need to write the identification array of the data table, The value, actual size, and offset of each data block, i.e., the actual position, are written to the page header of the data page. After completion, the state of the data page is changed to writable.

[0069] In combination with the encoding structure of the WAL log, the data writing process of the embodiment of the present invention can be to write the log information directly into the data block data structure of the corresponding table in the allocated page. The process is as follows: Step S1. Get the latest data segment that can be written; Step S2. Update the data segment information to the metadata information of the corresponding table; Step S3: Find the data page that is not fully written in the data segment; Step S4. Read the page header to obtain the address of the corresponding data block; Step S5. Write the latest serial number + the timestamp of the log information + the data line, and retain the address pointer of the current data block; Continue writing until the data block is full, and then go to step S1.

[0070] After the log information is written, the embodiment of the present invention can also track the memory usage status of the WAL log by performing the following operations: Step S1. After the specified time is met, the data blocks of the data segment with log information are appended to the WAL file on the disk; Step S2. After the data block in the data segment is full, the data block is appended to the WAL file on the disk; Step S3: Notify the WAL allocator to reclaim the memory of the data blocks of the data segment written to disk.

[0071] The embodiment of the present invention can convert the data in the corresponding WAL file into a database standard file and create a checkpoint after the database issues a checkpoint instruction, and archive the WAL file that has been converted into a database standard file.

[0072] Combine Figure 4 and Figure 5 , the working principle of the log data processing method of an embodiment of the present invention is described in detail with an embodiment.

[0073] In the actual execution process, the embodiment of the present invention can implement the log information writing and query process based on the WAL writer, WAL distributor, WAL synchronizer, WAL asynchronous writer and WAL log scanner.

[0074] The writing process can be as follows Figure 4 shown.

[0075] The WAL allocator can pre-allocate a whole block of memory for storing the data structure of the data segment. For each data segment, the allocator will initialize the first page and the corresponding data block area.

[0076] The WAL writer can write log information directly into the data structure of the data block of the corresponding table in the allocated page. Each thread of the writer corresponds to a data table imported from a data source.

[0077] The WAL synchronizer and asynchronous writer are responsible for tracking the current WAL log memory usage. The asynchronous writer mainly ensures that after the database issues a checkpoint instruction, the data in the corresponding WAL file is converted to a database standard file and a checkpoint is created. The WAL file that has been converted to a database standard file is archived.

[0078] Based on the above architecture, the time series data imported by the embodiment of the present invention only needs to be written once, that is, it only needs to be written to the WAL log, and then written to the disk through asynchronous technology, which reduces the time for memory copying and waiting. Through the known characteristics of IoT time series data, the allocator can calculate the required memory size in advance, and can apply for the entire block of memory on demand, which reduces the number of memory fragments and avoids memory waste. Each thread can ensure that it is written sequentially within the data block, and there is no random memory I / O. Through the known characteristics of IoT time series data, the WAL writer can ensure that only one thread is writing to the same memory address at the same time, without locking the corresponding data structure, reducing the overhead of concurrent writing and lock waiting. The conversion of WAL logs to database standard files supports mainstream asynchronous writing methods, which improves the performance of data dumps.

[0079] The WAL log data query process can be as follows Figure 5 shown.

[0080] By adding a scanner for WAL logs, the present invention can directly filter out pages that do not need to be scanned based on the query timestamp, eliminating the need for additional indexes, improving query efficiency, and eliminating the extra space occupied by indexes. The WAL scanner only scans data from the required tables, automatically skipping data that does not belong to the scanned tables, eliminating additional overhead.

[0081] In addition, regarding the WAL log data recovery process, the new WAL log generated by the embodiment of the present invention is compatible with the current standard WAL log recovery mechanism, and no additional changes are required for log recovery, thereby reducing database upgrade costs.

[0082] In summary, the encoding method in the embodiment of the present invention partitions and sorts according to the data table, reduces the fragmentation of the same table in the WAL log, improves the efficiency of data loading and insertion, and reduces space occupancy. At the same time, the writer adopts multi-threading, and each thread corresponds to a data block. This ensures that each thread can ensure that it is written sequentially within the data block. While there is no random memory I / O, it can also ensure that through the characteristics of known IoT time series data, the WAL writer can ensure that only one thread is writing to the same memory address at the same time, without locking the corresponding data structure, reducing the overhead and lock waiting of concurrent writing. In addition, compared with the existing technology, the WAL scanner can directly filter pages that do not need to be scanned by querying the timestamp, without the need for additional indexes, improving query efficiency, and at the same time without the additional space occupied by the index. Compared with the existing technology, the WAL scanner only scans the data of the required table, and the data that does not belong to the scanned table will be automatically skipped, and no additional overhead will be generated.

[0083] The new WAL log encoding eliminates the need for additional data storage while ensuring the database's ACID (Atomicity, Consistency, Isolation, Durability) performance remains unchanged. This ensures query performance while supporting direct database queries of the WAL log. Furthermore, WAL logs can be asynchronously archived, allowing WAL log data files to be converted back into a database file format for persistent storage. Compared to existing technologies, imported time-series data only needs to be written once to the WAL log and then flushed to disk using asynchronous technology, reducing memory copying and waiting time.

[0084] The new WAL log is written using space pre-allocation and append mode, which allocates continuous space in the memory in advance and writes it to the disk in append mode for persistent storage, which can reduce random reading and writing of memory and disk and the generation of fragmentation. Through the known characteristics of IoT time series data, the allocator can calculate the required memory size in advance and apply for the entire block of memory on demand, which reduces the number of memory fragments and avoids memory waste. The use of the popular append mode can also ensure that the WAL log supports mainstream asynchronous writing methods, improving the performance of data dumps. In addition, the new WAL log generated by the embodiment of the present invention is compatible with the current standard WAL log recovery mechanism, and there is no need to make additional changes to log recovery, reducing the cost of database upgrades.

[0085] The WAL log data query method proposed in an embodiment of the present invention can respond to log information query instructions by searching the global metadata of the time series database for the metadata of the corresponding data table, performing only necessary data page locks to improve cache utilization. The method also determines the write location information of the query target based on the metadata, and locates the query target by combining the write location information with the WAL log encoding structure in memory. This method supports direct database queries on WAL logs while ensuring query performance. This solves the technical problem in related technologies of unnecessary data page locks during data queries, resulting in low cache utilization and the generation of a large amount of useless cache, which in turn affects database query performance.

[0086] Next, a data query device for a WAL log according to an embodiment of the present invention will be described with reference to the accompanying drawings.

[0087] Figure 6 4 is a block diagram of a data query device for a WAL log according to an embodiment of the present invention.

[0088] like Figure 6 As shown, the data query device 20 for the WAL log includes: a receiving module 201 , a searching module 202 and a positioning module 203 .

[0089] Specifically, the receiving module 201 is configured to receive a log information query instruction to obtain a query target.

[0090] The search module 202 is configured to search for metadata of a corresponding data table in the global metadata of the time series database in response to a log information query instruction, and determine write location information of the query target based on the metadata.

[0091] The positioning module 203 is used to locate the query target based on the write position information and the pre-built WAL log encoding structure, wherein the WAL log encoding structure includes multiple data blocks, multiple data pages and data segments.

[0092] Optionally, in one embodiment of the present invention, the data query device 20 for the WAL log further includes: an acquisition module and a writing module.

[0093] The acquisition module is used to receive the time series data import instruction and obtain the WAL log generated when the time series data is imported.

[0094] The write module is used to map the log information of the WAL log to the corresponding data table, and write the data table and log information to the memory composed of the WAL log encoding structure.

[0095] Optionally, in one embodiment of the present invention, each data block stores WAL log information of a data table, each data block stores multiple log information, each log information includes a data header and data, wherein the data header is a timestamp and serial number stored for each log information.

[0096] Each data page includes multiple data blocks. The page header of each data page records the pre-allocated storage information, sequence number, number of data blocks, identifier of the data table corresponding to each data block, timestamp range and offset in the data page.

[0097] Each data segment includes multiple data pages.

[0098] Optionally, in one embodiment of the present invention, the writing module includes: an initialization unit and a writing unit.

[0099] The initialization unit is used to screen the memory that meets the preset storage conditions and initialize each data segment in the memory to obtain pre-allocated storage information of each data segment.

[0100] The writing unit is used to write a corresponding amount of log information into each data segment in a preset order based on the pre-allocated storage information until all the log information is written.

[0101] Optionally, in one embodiment of the present invention, the initialization unit includes: a first storage subunit, a second storage subunit and a third storage subunit.

[0102] The first storage subunit is configured to store a plurality of log information pieces in a corresponding number of log information pieces in a current data block in a current data page of a current data segment until the current data block reaches a pre-allocated number of storage pieces.

[0103] The second storage subunit is used to store the remaining log information in the next data block until all data blocks in the current data page reach the pre-allocated storage number.

[0104] The third storage sub-unit is used to store the remaining log information in the next data page until all data pages in the current data segment reach the pre-allocated storage limit, determine that the corresponding amount of log information has been written to the current data segment, and update the storage location of each log information in the metadata of the corresponding data table.

[0105] Optionally, in one embodiment of the present invention, the calculation expression for the actual storage capacity of the current data page is: , in, Indicates the actual storage capacity, Indicates the number of pre-allocated storage strips. Indicates the corresponding n The size of the data table, Indicates the corresponding n A data table.

[0106] Optionally, in one embodiment of the present invention, the positioning module 203 includes: a first positioning unit, a second positioning unit and a scanning unit.

[0107] Among them, the first positioning unit is used to locate at least one data segment to be scanned in the WAL log encoding structure based on the write position information, and traverse the at least one data segment to be scanned using the identifier or timestamp of at least one data table to find the corresponding data page.

[0108] The second positioning unit is configured to traverse the corresponding data page using the timestamp to locate the corresponding data block, and store the corresponding positioning information in the candidate list.

[0109] The scanning unit is used to search the candidate list for log information that meets the preset query conditions until all the data segments to be scanned are scanned.

[0110] It should be noted that the aforementioned explanation of the WAL log data query method embodiment is also applicable to the WAL log data query device of this embodiment, and will not be repeated here.

[0111] The WAL log data query device proposed in an embodiment of the present invention can respond to log information query instructions by searching the global metadata of the time series database for the metadata of the corresponding data table, performing only necessary data page locks to improve cache utilization. The device then determines the write location information of the query target based on the metadata, locating the query target by combining the write location information with the WAL log encoding structure in memory. This ensures query performance while supporting direct database queries on the WAL log. This solves the technical problem in related technologies of unnecessary data page locks during data queries, resulting in low cache utilization and the generation of a large amount of useless cache, which in turn affects database query performance.

[0112] Figure 7 A schematic diagram of the structure of a time series database provided in an embodiment of the present invention. The time series database may include: A memory 701 , a processor 702 , and a computer program stored in the memory 701 and executable on the processor 702 .

[0113] When the processor 702 executes the program, the WAL log data query method provided in the above embodiment is implemented.

[0114] Furthermore, the time series database also includes: The communication interface 703 is used for communication between the memory 701 and the processor 702 .

[0115] The memory 701 is used to store computer programs that can be run on the processor 702 .

[0116] The memory 701 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0117] If the memory 701, processor 702, and communication interface 703 are implemented independently, the communication interface 703, memory 701, and processor 702 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0118] Optionally, in a specific implementation, if the memory 701, the processor 702 and the communication interface 703 are integrated on a chip, the memory 701, the processor 702 and the communication interface 703 can communicate with each other through an internal interface.

[0119] The processor 702 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention.

[0120] This embodiment also provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the above-mentioned WAL log data query method is implemented.

[0121] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the WAL log data query method provided by an embodiment of the present invention.

[0122] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0123] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "N" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0124] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or N executable instructions for implementing a custom logical function or step of a process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0125] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" is any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable media include: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0126] It should be understood that various components of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0127] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0128] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0129] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A data query method for WAL logs, characterized in that: Applied to time series databases, the following steps are included: Receive log information query instructions to obtain query targets; In response to the log information query instruction, searching for metadata of a corresponding data table in the global metadata of the time series database, and determining write location information of the query target based on the metadata; The query target is located based on the write position information and a pre-built WAL log encoding structure, wherein the WAL log encoding structure includes multiple data blocks, multiple data pages, and data segments.

2. The data query method of WAL log according to claim 1, characterized in that Before receiving the log information query instruction, it also includes: Receive a time series data import instruction and obtain the WAL log generated when the time series data is imported; The log information of the WAL log is mapped to a corresponding data table, and the data table and the log information are written into a memory composed of the WAL log encoding structure.

3. The data query method of WAL log according to claim 2, characterized in that in, Each data block stores the WAL log information of a data table. Each data block stores multiple log information. Each log information includes a data header and data. The data header is the timestamp and sequence number stored in each log information. Each data page includes the plurality of data blocks, and the header of each data page records pre-allocated storage information, a sequence number, the number of data blocks, an identifier of a data table corresponding to each data block, a timestamp range, and an offset in the data page; Each data segment includes the plurality of data pages.

4. The data query method of WAL log according to claim 3 is characterized in that, Writing the data table and the log information into the memory formed by the pre-built WAL log encoding structure includes: Screening a memory that meets a preset storage condition, and initializing each data segment in the memory to obtain pre-allocated storage information of each data segment; Based on the pre-allocated storage information, a corresponding amount of log information is written into each data segment in a preset order until all log information is written.

5. The data query method of WAL log according to claim 4 is characterized in that, Initializing each data segment in the memory to obtain pre-allocated storage information of each data segment includes: Storing multiple log information pieces from the corresponding number of log information pieces in a current data block in a current data page of a current data segment until the current data block reaches a pre-allocated number of storage pieces; Storing the remaining log information in the next data block until all data blocks in the current data page reach the pre-allocated number of storage entries; The remaining log information is stored in the next data page until all data pages in the current data segment reach the pre-allocated storage limit, and the corresponding amount of log information is determined to have been written to the current data segment, and the storage location of each log information is updated in the metadata of the corresponding data table.

6. The data query method of WAL log according to claim 5, characterized in that: The calculation expression for the actual storage capacity of the current data page is: , in, represents the actual storage capacity, Indicates the number of pre-allocated storage strips, Indicates the corresponding n The size of the data table, Indicates the corresponding n A data table.

7. The data query method of WAL log according to claim 3 is characterized in that, The locating the query target based on the write position information and the pre-built WAL log encoding structure includes: Locating at least one data segment to be scanned in the WAL log encoding structure based on the write position information, and traversing the at least one data segment to be scanned using an identifier or a timestamp of the at least one data table to find a corresponding data page; traversing the corresponding data page using the timestamp to locate the corresponding data block, and storing corresponding location information in a candidate list; The candidate list is searched for log information that meets the preset query conditions until all data segments to be scanned are scanned.

8. A data query device for WAL logs, characterized in that: Applicable to time series databases, including: The receiving module is used to receive the log information query instruction to obtain the query target; A search module, configured to search for metadata of a corresponding data table in the global metadata of the time series database in response to the log information query instruction, and determine write location information of the query target based on the metadata; A positioning module is used to locate the query target based on the write position information and a pre-built WAL log encoding structure, wherein the WAL log encoding structure includes multiple data blocks, multiple data pages and data segments.

9. A time series database, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the WAL log data query method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the WAL log data query method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Oil and gas field time series data storage method and device, oil and gas field time series data query method and device and storage medium

    CN112286867A

  • Time series data storage method and device, computer equipment and storage medium

    CN116149566A

  • Rapid data file merging method and system for time sequence database

    CN116561120A

  • Log query method and device, medium, electronic equipment and program product

    CN119003458A

  • Querying of materialized views for time-series database analytics

    US20200334254A1

Cited By

  • Time series data processing method and device based on edge calculation, equipment and medium

    CN121210112A

  • Edge computing-based time series data processing method, device, equipment and medium

    CN121210112B