Log storage method, database system, server, and storage medium
By directly using the storage engine components in a distributed database system to write transaction logs to online log files and persisting them to block devices, the resource overhead problems caused by the file system are solved, efficient and reliable log storage is achieved, and system performance and availability are improved.
Patent Information
- Application Number
- PCT/IB2024/063178
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-12-26
- Publication Date
- 2025-07-03
AI Technical Summary
In distributed database systems, the existing technology relies on file systems for storing binary logs, resulting in high operational complexity and large resource overhead, affecting system performance and reliability.
By directly writing transaction logs to online log files that support writing anywhere on distributed database nodes using storage engine components and persisting them to block devices, avoiding dependence on file systems and leveraging the recycled use characteristics of online log files to reduce the complexity and resource overhead of log storage operations.
Reduces resource overhead introduced by the file system, improves the efficiency and reliability of log storage, reduces the overhead of metadata synchronization operations, and improves the performance and availability of the system.
Smart Images

Figure IB2024063178_03072025_PF_FP_ABST
Abstract
Description
[0001]Log Storage Method, Database System, Server, and Storage Medium This disclosure claims priority to Chinese patent application number 202311847987.6, filed with the China Patent Office on December 28, 2023, entitled "Log Storage Method, Database System, Server, and Storage Medium," the entire contents of which are incorporated herein by reference. Technical Field This disclosure relates to the field of computer technology, and more particularly to a log storage method, database system, server, and storage medium. Background: In MySQL (a relational database management system), the binary log (Binlog) is used to record internal changes to the database, including insert, update, and delete operations. Binary logs, primarily used for master-slave replication and incremental recovery, are crucial files in distributed database systems. Current distributed database systems primarily rely on the file system to perform operations such as creating, reading, and modifying the binary log. To support a wide range of applications in distributed databases, the file system is designed as a complex, general-purpose component. The complex design of file systems results in high operational complexity in storing binary logs, introducing unnecessary overhead. Therefore, a new solution is needed. SUMMARY OF THE INVENTION Various aspects of the present disclosure provide a log storage method, a database system, a server, and a storage medium to reduce the operational complexity and resource overhead of binary log storage in a distributed database. An embodiment of the present disclosure provides a log storage method applicable to any database node in a distributed database. The method comprises: obtaining a transaction log for a target transaction using a storage engine component; writing the transaction log to a target online log file; the target online log file is an online log file that supports write operations to any location; and persistently storing the transaction log in the target online log file in a corresponding target block device. Optionally, writing the transaction log to the target online log file comprises: determining a first online log file that is in an activated state; determining whether the transaction log data volume is greater than the remaining writable data volume of the first online log file; if so, closing the first online log file and starting a new second online log file as the target online log file; and writing the transaction log to the target online log file.Optionally, writing the transaction log to the target online log file includes: writing the transaction log to a target memory cache block corresponding to the target online log file; and persistently storing the transaction log in the target online log file on the corresponding target block device includes: synchronizing the transaction log written to the target memory cache block to the target block device associated with the target online log file based on log synchronization parameters. Optionally, the method further includes: determining the log synchronization parameters based on at least one of load information of the distributed database, throughput information of the target block device, and backlog information of online log files to be synchronized in the distributed database. Optionally, the storage space size of the target memory cache block and the target block device is the same. Optionally, the method further includes: during writing the transaction log to the target online log file, determining a first block device associated with the target online log file; if the transaction log data volume is larger than the first block device, requesting allocation of a second block device from a storage management component; and associating the target online log file with the second block device to expand the capacity of the target online log file. Optionally, the method further includes: during the process of persisting the transaction log, in response to a commit instruction of the target transaction, writing metadata information of the target online log file into the redo log file of the target transaction; and writing a commit event of the target transaction into the redo log file of the target transaction; and, after completing the persistent storage of the transaction log, completing the commit operation of the target transaction. Optionally, the method further includes: after the target transaction commits, performing a persistence operation on the redo log file of the target transaction to perform a transaction-level persistence operation on the metadata information of the target online log file; or, after the transaction group to which the target transaction belongs commits, performing a persistence operation on the redo log file of the target transaction to perform a transaction-group-level persistence operation on the metadata information of the target online log file. Optionally, the target online log file includes: an initial online log file created during database node initialization or an online log file that has been archived and put into recycling. Optionally, after persistently storing the transaction log in the target online log file in the corresponding target block device, the method further includes: archiving the target online log file to a designated storage space, and updating the target online log file to an archived state, so that the target online log file is put into recycling after being archived.An embodiment of the present disclosure further provides a distributed database system, comprising: a plurality of database nodes, a data block file component, and a storage component; wherein the storage component is configured to provide at least one block device; the database file component is configured to provide at least one online log file; any one of the plurality of database nodes is configured to obtain a transaction log of a target transaction using a storage engine component; the transaction log is written to a target online log file among the at least one online log file; the target online log file is an online log file that supports write operations at any location; and the transaction log in the target online log file is persistently stored on a target block device among the at least one block device that is associated with the target online log file. Optionally, the system further includes: a storage management component; the storage engine component further configured to: during writing the transaction log to the target online log file, determine a first block device associated with the target online log file; if the transaction log data volume exceeds the first block device, send a block device request to the storage management component, so that the storage management component allocates a second block device from the at least one block device to the target online log file; and associate the target online log file with the second block device to expand the capacity of the target online log file. The present disclosure also provides a server, comprising: a memory and a processor; the memory is configured to store one or more computer instructions; the processor is configured to execute the one or more computer instructions to: perform the steps of the method provided in the present disclosure. The present disclosure also provides a computer program product, comprising a computer program, which, when executed by the processor, implements the steps of the method provided in the present disclosure. The present disclosure also provides a computer-readable storage medium storing the computer program, which, when executed by the processor, implements the steps of the method provided in the present disclosure. In the disclosed embodiments, any database node in a distributed database can use a storage engine component to obtain the transaction log of a target transaction and write the transaction log to a target online log file. The database node can use the storage engine to persistently store the transaction log in the target online log file to the corresponding target block device and perform transaction-level persistent storage operations on the metadata of the target online log file. In this embodiment, based on an online log file that supports write operations to any location, the database node can directly store the transaction log using block storage using its own storage engine component without relying on a file system, reducing the resource overhead associated with the introduction of a file system and file storage.BRIEF DESCRIPTION OF THE DRAWINGS The accompanying drawings described herein are intended to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are intended to explain the present disclosure and do not constitute undue limitations thereon. In the accompanying drawings: Figure 1 is a schematic diagram of a single-master, multiple-standby architecture of a shared storage-based database system; Figure 2 is a flow chart of a log storage method provided in an exemplary embodiment of the present disclosure; Figure 3 is a schematic diagram of the architecture of a distributed database system provided in an exemplary embodiment of the present disclosure; Figure 4 is a schematic diagram of the structure of a distributed database system provided in an exemplary embodiment of the present disclosure; and Figure 5 is a schematic diagram of the structure of a server provided in an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION To further clarify the objectives, technical solutions, and advantages of the present disclosure, the technical solutions of the present disclosure will be described clearly and completely below in conjunction with the specific embodiments of the present disclosure and the corresponding drawings. It should be understood that the described embodiments are only some of the embodiments of the present disclosure, and are not intended to be exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort are within the scope of protection of the present disclosure. The terms used in the embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. As used in the embodiments of the present invention and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms, unless the context clearly indicates otherwise. "A plurality" generally includes at least two, but does not exclude the inclusion of at least one. It should be understood that the term "and / or" as used herein merely describes an associative relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent: A alone, A and B simultaneously, or B alone. Furthermore, the character " / " herein generally indicates that the associated objects are in an "or" relationship. It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a list of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such product or system. Without further limitation, the elements specified by the phrase "comprising a..." do not exclude the presence of additional identical elements in the product or system comprising the elements. Shared storage refers to a technology that allows multiple computer systems or processes to share the same physical storage device or storage system.In a shared storage system, multiple computers or processes can access data on the same storage device or storage system and perform operations such as reading, writing, and modifying this data, thereby enabling data sharing and collaborative processing. Shared storage systems can improve data sharing and reliability, thereby supporting larger-scale applications and services. For example, in large-scale database applications, shared storage can improve data access speed and reliability, thereby achieving high performance and high availability. Furthermore, to ensure data consistency and integrity, shared storage systems typically implement mechanisms such as locking, synchronization, and caching. Figure 1 illustrates a typical single-master, multiple-standby architecture for a shared storage-based database system. The architecture shown in Figure 1 includes multiple database nodes, including a master node and multiple standby nodes. The storage engine components running on each of the multiple database nodes form a storage engine cluster. Each database node is implemented as a database instance. The master node can be a write node, and the standby node can be a read node. The write node is responsible for processing data write operations in the database system, while the standby node is responsible for reading data written by the master node. In some embodiments, the master and standby nodes can switch roles. For example, if the original master node fails, a backup node can be switched to become the new master node for data writing. Log data generated locally on the master node can be written to the binary log file in the database file component and then persistently stored from the binary log file to the storage base. Furthermore, metadata about the binary log file can be stored in data tablespaces, such as redo logs or undo logs, to record information about the written log data. As shown in Figure 1, data is read and written between the database file component and the storage base using a file system. In this architecture, the file system (typically a distributed file system) can shield multiple database nodes from shared storage, allowing the database storage engine in the distributed database system to access and modify data just like a standalone database. However, the file system is a complex component used for a wide range of applications and can lead to numerous issues in shared storage scenarios. For example, every storage access operation must go through the file system, which deepens the storage stack. The file system itself consumes a certain amount of computing resources and has high maintenance costs. For example, file systems are designed for general scenarios, and when expanding storage capacity, many factors must be considered and the support flexibility is high, resulting in poor stability in storage expansion.For example, multiple nodes need to synchronize log metadata, but synchronizing metadata through the file system is costly and uncontrollable. Another example is that writing binary logs relies on the file system's cache, which consumes a significant amount of memory resources. In particular, when database instance resources are limited, file system cache cleanup is triggered, resulting in a large number of I / O (Input / Output) operations within a short period of time. In severe cases, this can cause the entire database instance to become unresponsive (hang), rendering it unavailable. To address the aforementioned technical issues, some embodiments of the present disclosure provide a solution that enables binary log file storage without relying on the file system. The technical solutions provided by various embodiments of the present disclosure are described in detail below, in conjunction with the accompanying drawings. Figure 2 is a flow diagram of a log storage method provided by an exemplary embodiment of the present disclosure. The method may include the following steps: Step 201: Obtain a transaction log for a target transaction using a storage engine component. Step 202: Write the transaction log to a target online log file; the target online log file is an online log file that supports write operations to any location. Step 203: Persistently store the transaction log in the target online log file in the corresponding target block device. This embodiment is applied to database nodes in a distributed database system. The architecture of the distributed database system, as shown in Figure 3, may include multiple database nodes, namely, a master node and multiple standby nodes. The storage engine components running on each of the multiple database nodes form a storage engine cluster. Each database node may be implemented as a database instance. The execution entity of this embodiment may be the master node, which is the write node responsible for processing data write operations. In this embodiment, the database node can utilize the storage engine component to directly perform transaction log storage operations without relying on the file system. A transaction is a unit of concurrency control and a user-defined sequence of operations. Transactions have atomicity, consistency, isolation, and persistence. The target transaction can be any transaction processed by the distributed database system. Before a transaction is committed, the transaction log needs to be persistently stored to allow re-execution of transaction modifications in the event of a database failure. The transaction log may include at least one of the following: the transaction ID (Identity document), the transaction operation type (such as insert, operation, or update), the transaction operation object (such as the name and identifier of the database object), the operation timestamp, and the operation content (such as the updated value or deleted record).In this embodiment, without relying on a file system, the storage engine component on the database node can use block storage to persist the target transaction's transaction log to a target block device. A block device is a device with fixed-size storage space. Each read or write operation on a block device is a data block, and it does not support reading or writing to arbitrary locations. Therefore, in this embodiment, after obtaining the target transaction's transaction log, the storage engine component can write the acquired transaction log to a target online log file. When specified conditions are met, the transaction log written to the target online log file is stored in the block device. These specified conditions may include: the target transaction being committed, the target online log file being full, or the target online log file being closed. In this embodiment, the target online log file refers to any online binary log file provided by the distributed database system, referred to herein as an online log file. Online log files have the characteristic of being writeable to arbitrary locations. Furthermore, based on the ability of online log files to be written to any location, write operations to any location on a block device are indirectly enabled, allowing database nodes to directly use block storage to store transaction logs without relying on a file system. In distributed database systems, the number of online log files is typically specified before database instance initialization, and the specified number of online log files is created during database instance initialization. These online log files are maintained by the database file component. Online log files can be recycled, eliminating the need for the database node's storage engine to create online log files. As shown in Figure 3, the database file component maintains multiple online binary files, including online log file OLB1, online log files OLB2, ..., and online log file OLB6. Each online log file has a unique identifier (ID), an index number (Index), a startup status, and an archive status. The startup status describes whether the online log file is writable and can be expressed as 0 (not started) or 1 (started). If an online log file is started, the transaction log generated by a transaction can be written to it. When an online log file is full or a user actively triggers a rotation, the online log file is updated to the closed state and the next online log file is opened. During the time an online log file is being written, its start state remains at 1 (started). Transaction logs are written to online log files sequentially, so the database allows one online log file to be in the started state at a time.The archiving status describes whether the online log file has been archived. The archiving status value can be 0 (not archived) or 1 (archived). After the transaction log in the online log file has been persisted, the online log file can be archived for use by other applications. Optionally, the online log file can be archived to a designated storage space, which can be local to the database or a remote archiving storage node. After the online log file is archived, the archiving status value of the online log file can be updated to 1 to update the online log file to the archived state. Archived online log files can be cyclically overwritten. The index number identifies the sequence number of the online log file generated by the database node. The index number increments in the order in which the online log file is activated. When the same online log file is cyclically used, the index number of the online log file will be different in different cyclic rounds. This is explained below with reference to the accompanying figures. As shown in Figure 3, online log file 0LB1 has an ID value of 1, an index number of 10, a startup status of 0 (i.e., not started), and an archive status of 1 (i.e., archived). Based on this information, we can see that online log file 0LB1 with index number 10 has been archived and is now in circulation. As shown in Figure 3, online log file 0LB2 has an ID value of 2, an index number of 11, a startup status of 1 (i.e., started), and an archive status of 0 (i.e., not archived). Based on this information, we can see that online log file 0LB2 is currently in a writable state. Assume that online log file 0LB3 has an ID value of 3, an index number of 6, a startup status of 0 (i.e., not started), and an archive status of 1 (i.e., archived). During log rotation, if online log file OLB2 is closed, online log file OLB3 can be started. The start status value of online log file OLB3 is updated to 1 (started), the archive status value is updated to 0 (not archived), and a new index number 12 is generated for online log file OLB3. The target online log file described in this embodiment may include: the initial online log file created during database node initialization or an online log file that is archived and recycled. This enables the recycling of online binary log files without relying on the storage engine component to perform online binary log file creation operations, saving resource overhead. After writing the transaction log of the target transaction to the target online log file, the storage engine component of the database node can persistently store the transaction log in the target online log file.Persistently storing the transaction log in the target online log file includes writing the information in the target online log file to a target block device associated with the target online log file using block storage. The target block device can be a physical block device or a virtual block device, which is not a limitation in this embodiment. Block storage can be used to store online binary logs. Compared to file storage, block storage does not rely on an additional file system layer, resulting in more efficient and reliable data reading. Each online log file in the database file component can be constructed on at least one block device. In cloud computing applications, the block device associated with the online log file can be a virtual block device provided by the database storage base. As shown in Figure 3, the storage base provides an elastic block storage service that virtualizes storage space into multiple virtual block devices for block storage. Typically, block devices have a relatively fixed storage capacity. For example, a block device may have a storage capacity of 1G (Gigabyte). Therefore, any online log file has a relatively fixed data size. An online log file can store transaction logs for multiple transactions; a transaction log can be written to a single online log file. Based on this, in some embodiments, when writing the transaction log of a target transaction to a target online log file, the storage engine may determine the first online log file that is in the active state and determine whether the data volume of the transaction log of the target transaction exceeds the remaining writable data volume of the first online log file. If the data volume of the transaction log of the target transaction exceeds the remaining writable data volume of the first online log file, the storage engine may trigger an online log rotation operation. A rotation operation involves closing the currently active online log file and starting the next online log file. Specifically, the storage engine may close the first online log file, start a new second online log file as the target online log file, and write the transaction log of the target transaction to the target online log file. Before closing the first online log file, the transaction log in the first online log file may be archived to prevent log data loss. If the target transaction's transaction log data volume is less than or equal to the remaining writable data volume of the first online log file, the storage engine may use the first online log file as the target online log file and write the target transaction's transaction log to the target online log file. This implementation reduces the risk of writing transaction logs for the same transaction to different online log files by comparing the remaining writable data volume of the opened online log file with the data volume of the transaction's transaction log, thereby improving the reliability of log storage results.In this embodiment, online log files can be recycled. Each online log file has a unique identifier. During recycling, an index number is generated for the online log file each time it is opened to identify the order in which log data was generated. Online log files can be archived and recycled after archiving. When recycled again, a new index number is generated for the online log file to distinguish different log data. In some exemplary embodiments, after persistently storing the transaction log in a target online log file, the storage engine can archive the target online log file to a designated storage space and update the target online log file to an archived state, allowing the target online log file to be recycled after archiving. When the online log file is archived, a target format recognizable by upstream and downstream applications of the database can be obtained, and the transaction log in the online log file can be stored in the target format. Furthermore, even if the persistence format of the online log file differs from the persistence format of ordinary online log files used by the file system, upstream and downstream applications of the database can still use transaction logs based on archived online log files, thereby reducing the risk of transaction log incompatibility with existing upstream and downstream applications. In this embodiment, based on online log files that support write operations to any location, database nodes can directly store transaction logs using block storage based on their own storage engine components without relying on the file system, reducing the resource overhead caused by the introduction of the file system and file storage methods. Furthermore, by writing transaction logs to online log files, the storage engine component can fully utilize the recycling nature of online log files, eliminating the need to create binary logs, further reducing the complexity and resource overhead of log storage operations. In some optional embodiments, before persisting the target online log file, the storage engine can expand the target block device corresponding to the target online log file based on the data volume of the target online log. Optionally, when writing the transaction log of the target transaction to the target online log file, the storage engine can determine the first block device associated with the target online log file. If the transaction log of the target transaction contains more data than the first block device, the storage engine may request allocation of a second block device from the storage management component and associate the target online log file with the second block device to expand the target online log file. The second block device may include one or more block devices. The storage management component may manage the association between the target online log file and the first and second block devices.As shown in Figure 3, the database metafile managed by the storage management component may store the correspondence between online log file 0LB1 and block device 01, the correspondence between online log file 0LB2 and block device 02, and the correspondence between online log file 0LB3 and block device 03. Furthermore, when persisting transaction logs in the target online log file, the storage engine may persist the transaction logs in the target online log file to the first and second block devices based on the associations managed by the storage management component. In this embodiment, the storage engine can expand the capacity of online log files based on the storage management component. Compared to expansion solutions using the file system, the storage engine's expansion operation is less complex and more reliable. In some optional embodiments, when persisting online logs, the database node's storage engine may use memory cache to read and write to any location on the block device. This is explained below as an example. Optionally, in this embodiment, a memory cache segment corresponding to the online binary log file can be set up in the database node's memory. This memory cache segment can be divided into multiple memory cache blocks, each of which can correspond to multiple online binary files. Optionally, the size of any memory cache block is the same as the storage space of the block device. Based on this, when writing the transaction log of the target transaction to the target online log file, the storage engine can first write the transaction log of the target transaction to the target memory cache block corresponding to the target online log file and then synchronize the transaction log written to the target memory cache block to the target block device associated with the target online log file based on log synchronization parameters. Optionally, the log synchronization parameters can be fixed or dynamically determined based on the environment data of the log synchronization operation. The environment data of the log synchronization operation is used to predict the pressure on the database and / or block device caused by the transaction log synchronization operation between the database and the block device. For example, the environment data can include at least one of: load information of the distributed database, throughput information of the target block device, and backlog information of online log files to be synchronized in the distributed database. In some embodiments, the storage engine may determine log synchronization parameters based on the at least one of the aforementioned information. Optionally, the log synchronization parameters may include the log synchronization frequency and / or the amount of data synchronized each time. For example, if database load information indicates that the database is under heavy load, a lower log synchronization frequency and a larger amount of data synchronized may be set to reduce the load on the database caused by log synchronization operations.For example, if the target block device's throughput capacity information indicates a high throughput capacity, a lower log synchronization frequency and a smaller log synchronization data volume can be set to fully utilize hardware resources. For another example, if the backlog information of online log files to be synchronized in the database indicates a high backlog rate, a higher log synchronization frequency and a larger log synchronization data volume can be set to alleviate the backlog and accelerate the rotation of online log files to meet online log file usage requirements. In this implementation, the storage engine can implement write operations to any location in the memory cache blocks, thereby indirectly implementing write operations to any location in the block device. Accordingly, when reading data from the block device, the data in the block device can be read into the memory cache blocks, and the storage engine can implement read operations from any location in the memory cache blocks. Compared to using the file system for read / write operations, having the storage engine directly perform log read / write operations can reduce file system overhead and alleviate the problem of read / write amplification. In addition to recording transaction logs, online log files can also maintain metadata information for the online log files themselves. This metadata information primarily includes the size of the written log data. This metadata information must be persistently stored for subsequent use. Compared to a file system, during the process of writing the target transaction's transaction log to an online log file, the database node's storage engine component can accurately detect the end boundary of the transaction log. Consequently, after the transaction log is completely written, it can update the transaction log's metadata information to record the completion of the transaction log write. Based on this, in some exemplary embodiments, the database node can use the storage engine component to perform transaction-level or transaction-group-level persistent storage operations on the metadata information of the target online log file. Transaction-level persistent storage of metadata information refers to persistent storage of the size of the written log data after the transaction commits. Transaction-group persistent storage of metadata information refers to persistent storage of the size of the written log data after the transaction group to which the transaction belongs commits. This will be explained below in an exemplary manner. In some optional embodiments, after writing the target transaction's transaction log to the target online log file, the storage engine can store the metadata information of the target online log file as an incremental log in the database's redo log. In this embodiment, the storage engine may respond to a commit instruction of the target transaction and write metadata information of the target online log file into the redo log file of the target transaction during the process of persisting the transaction log of the target transaction.While waiting for the transaction log to be persisted, the storage engine can write the target transaction's commit event to the target transaction's redo log file. After the transaction log is persistently stored, the target transaction is committed. This process converts metadata cache operations into write operations to the redo log, facilitating persistent storage of metadata information through redo log persistence, reducing the number of additional block devices required for subsequent metadata persistence. Furthermore, while waiting for the transaction log to be persisted, the storage engine can asynchronously write metadata information and commit events to the redo log file, reducing transaction commit wait time to a certain extent. Optionally, the storage engine can persist the target transaction's redo log file after the target transaction commits, thereby performing transaction-level persistence on the target online log file's metadata. Alternatively, the storage engine can persist the target transaction's redo log file after the transaction group to which the target transaction belongs commits, thereby performing transaction-group-level persistence on the target online log file's metadata. The transaction group to which the target transaction belongs includes the target transaction and one or more transactions with adjacent commit times. This implementation further reduces the frequency of metadata persistence, thereby reducing the overhead associated with metadata synchronization operations. Optionally, the redo log refers to the database's Redo log. A Redo log is a physical log. A log record in the Redo log can be expressed as: modifying a certain data point on a certain data page to a new data point after a certain offset. The data size of a single Redo record is relatively small, thus placing less pressure on the database. When writing metadata information for the target online log file to the Redo log, a log record of a dozen or so bytes can be written to the Redo log. Read nodes in a distributed database can then obtain metadata information for the target online log file from the Redo log file. Compared to using block devices to persist metadata, using incremental logging to persistently store metadata for any online log file achieves a transition from writing to a block device to writing a single log record, further reducing metadata storage costs. In this case, the metadata information persisted is at the transaction level or transaction group level. Compared to performing metadata persistence operations for every log write operation, this approach can, on the one hand, reduce overhead and thus reduce the impact on database performance. On the other hand, binding the persisted metadata to the transaction log boundary can reduce the uncertainty of metadata information and improve database availability.On the other hand, compared to the traditional approach of persisting transaction log metadata before committing the transaction, this embodiment drives the transaction commit within the storage engine before waiting for the transaction log to be persisted. This allows the database to more precisely control the data persistence process, improving database performance. It should be noted that the execution entities of the methods provided in the above embodiments can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 201 to 203 can be device A; another example is that the execution entity of steps 201 and 202 can be device A, and the execution entity of step 203 can be device B; and so on. Furthermore, some processes described in the above embodiments and accompanying figures include multiple operations that appear in a specific order. However, it should be understood that these operations may be executed in a different order than the order in which they appear, or may be executed in parallel. Operation sequence numbers, such as 201 and 202, are merely used to distinguish between different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the terms "first" and "second" herein are used to distinguish between different messages, devices, modules, and the like, and do not represent a sequential order, nor do they limit "first" and "second" to different types. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data, etc.) involved in this disclosure are all authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or deny. In addition to the log storage method described in the aforementioned embodiments, embodiments of the present disclosure also provide a distributed database system, which will be exemplarily described below with reference to the accompanying figures. Figure 4 is a schematic structural diagram of a distributed database system provided in an exemplary embodiment of the present disclosure. As shown in Figure 4, the distributed database system may include: multiple database nodes 401, a database file component 402, and a storage component 403. Multiple database nodes 401 are deployed in a master-slave mode. A storage component 403 is configured to provide at least one block device. Each block device can be a physical block device or a virtual block device, which is not a limitation in this embodiment. A database file component 402 is configured to provide at least one online log file; this at least one online log file can be reused.Each online log file has a unique identifier. During recycling, each time it is opened, an index number is generated for the online log file to identify the order in which log data is generated. Online log files can be archived and recycled after archiving. When recycled again, a new index number is generated for the online log file to distinguish different log data by the index number. Any database node 401 is configured to: obtain a transaction log for a target transaction using a storage engine; write the transaction log to a target online log file in the at least one online log file; the target online log file is an online log file that supports write operations to any location; and persistently store the transaction log in the target online log file on a target block device associated with the target online log file in the at least one block device. In some optional embodiments, when writing the transaction log to the target online log file, the database node 401 is specifically configured to: determine a first online log file that is in an active state using a storage engine; determine whether the data volume of the transaction log is greater than the remaining writable data volume of the first online log file; if so, close the first online log file and start a new second online log file as the target online log file; and write the transaction log to the target online log file. In some optional embodiments, when writing the transaction log to the target online log file, the database node 401 is specifically configured to: write the transaction log to a target memory cache block corresponding to the target online log file using a storage engine; optionally, the target memory cache block has the same storage space size as the target block device. Accordingly, when persistently storing the transaction log in the target online log file on the corresponding target block device, the database node 401 is specifically configured to: synchronize the transaction log written to the target memory cache block to the target block device associated with the target online log file according to log synchronization parameters. In some optional embodiments, database node 401 is further configured to: utilize a storage engine to determine the log synchronization parameters based on at least one of load information of the distributed database, throughput information of the target block device, and backlog information of online log files to be synchronized in the distributed database. In some optional embodiments, the distributed database system further includes a storage management component 404. Database node 401 is further configured to: utilize a storage engine to determine a first block device associated with the target online log file during writing of the transaction log to the target online log file; and if the transaction log data volume exceeds the first block device, send a block device request to storage management component 404.Storage management component 404 is configured to allocate a second block device from the at least one block device to the target online log file. Accordingly, database node 401 is further configured to utilize a storage engine to associate the target online log file with the second block device to expand the capacity of the target online log file. In some optional embodiments, database node 401 is further configured to: during the process of persisting the transaction log, in response to a commit instruction of the target transaction, write metadata information of the target online log file to the redo log file of the target transaction; write a commit event of the target transaction to the redo log file of the target transaction; and, after completing the persistent storage of the transaction log, complete the commit operation of the target transaction. Optionally, database node 401 is further configured to: after the target transaction commits, perform a persistence operation on the redo log file of the target transaction to perform a transaction-level persistence operation on the metadata information of the target online log file; or, after the transaction group to which the target transaction belongs commits, perform a persistence operation on the redo log file of the target transaction to perform a transaction-group-level persistence operation on the metadata information of the target online log file. Optionally, when performing the transaction-level persistent storage operation on the metadata information of the target online log file, database node 401 is specifically configured to: after the target transaction commits, use a storage engine to store the metadata information of the target online log file as an incremental log in the redo log file of the distributed database system; or, after the transaction group to which the target transaction belongs commits, use a storage engine to store the metadata information of the target online log file as an incremental log in the redo log file of the distributed database system. Optionally, the target online log file includes an initial online log file created during database node initialization or an online log file that has been archived and recycled. In some optional embodiments, after persistently storing the transaction log in the target online log file in the corresponding target block device, database node 401 is further configured to: obtain the current index number of the target online log file using a storage engine; add an archive flag to the target online log file corresponding to the current index number; clear the transaction log in the target online log file and update the index number of the target online log file; and restore the target online log file with the updated index number to a usable online log file. In this embodiment, a single database node in a distributed database system can directly store transaction logs using block storage based on its own storage engine component without relying on a file system, thereby reducing resource overhead caused by the introduction of a file system and file storage methods.The storage engine component writes transaction logs to online log files, leveraging the recyclable nature of online log files. This eliminates the need to create binary logs, further reducing the complexity and resource overhead of log storage operations. Furthermore, database nodes can leverage their ability to detect transaction termination to perform transaction-level persistent storage of metadata information in the target online log file. This ensures that the persisted metadata information is complete metadata for the target transaction, improving the controllability and reliability of metadata persistence operations. Figure 5 illustrates a schematic diagram of the structure of a server provided in an exemplary embodiment of the present disclosure, which is applicable to the log storage method provided in the aforementioned embodiments. As shown in Figure 5, the server includes a memory 501, a processor 502, and a communication component 503. Memory 501 is used to store computer programs and can be configured to store various other data to support operations on the server. Examples of such data include instructions for any application or method operating on the server. The processor 502 is coupled to the memory 501 and configured to execute a computer program in the memory 501, configured to: obtain a transaction log of a target transaction using a storage engine component; write the transaction log to a target online log file; the target online log file is an online log file that supports write operations to any location; and persistently store the transaction log in the target online log file in a corresponding target block device. Optionally, when writing the transaction log to the target online log file, the processor 502 is configured to: determine, using the storage engine component, a first online log file that is in an activated state; determine whether the data volume of the transaction log is greater than the remaining writable data volume of the first online log file; if so, close the first online log file and activate a new second online log file as the target online log file; and write the transaction log to the target online log file. Optionally, when writing the transaction log to the target online log file, the processor 502 is specifically configured to: use a storage engine component to write the transaction log to a target memory cache block corresponding to the target online log file; correspondingly, when persistently storing the transaction log in the target online log file in the corresponding target block device, the processor 502 is specifically configured to: use a storage engine component to synchronize, according to a log synchronization parameter, the transaction log written in the target memory cache block to a target block device associated with the target online log file.Optionally, the processor 502 is further configured to: determine the log synchronization parameters using a storage engine component based on at least one of load information of the distributed database, throughput information of the target block device, and backlog information of online log files to be synchronized in the distributed database. Optionally, the processor 502 configures the target memory cache block and the target block device to have the same storage space size. Optionally, the processor 502 is further configured to: during writing the transaction log to the target online log file, determine using the storage engine component a first block device associated with the target online log file; if the transaction log data volume is larger than the first block device, request allocation of a second block device from the storage management component; and associate the target online log file with the second block device to expand the capacity of the target online log file. Optionally, the processor 502 is further configured to: during the process of persisting the transaction log, in response to a commit instruction of the target transaction, write the metadata information of the target online log file into the redo log file of the target transaction; and write the commit event of the target transaction into the redo log file of the target transaction; and, after completing the persistent storage of the transaction log, complete the commit operation of the target transaction. Optionally, the processor 502 is further configured to: after the target transaction commits, perform a persistence operation on the redo log file of the target transaction to perform a transaction-level persistence operation on the metadata information of the target online log file; or, after the transaction group to which the target transaction belongs commits, perform a persistence operation on the redo log file of the target transaction to perform a transaction-group-level persistence operation on the metadata information of the target online log file. Optionally, the target online log file includes: an initial online log file created during database node initialization or an online log file that is archived and recycled. Optionally, after persistently storing the transaction log in the target online log file in the corresponding target block device, the processor 502 is further configured to archive the target online log file to a designated storage space and update the target online log file to an archived state, so that the target online log file is recycled after archiving. Furthermore, as shown in FIG5 , the server also includes other components, such as a power supply component 504. FIG5 schematically illustrates only some components and does not imply that the server includes only the components shown in FIG5 .The memory 501 may be implemented by any type of volatile or non-volatile memory device, or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The communication component 503 is configured to facilitate wired or wireless communication between the device in which the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as Wi-Fi (wireless network communication technology), 2G (such as Global System for Mobile Communications (GSM)), 3G (such as Wideband Code Division Multiple Access (WCDMA), 4G (such as Long Term Evolution (ETE)), 4G+ (such as ETE-Advanced (ETE-A)), or 5G (Fifth Generation Mobile Communication Technology), or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IRDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, and other technologies. Among them, The power supply component 504 is used to provide power to various components of the device where the power supply component is located.The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component resides. In this embodiment, the database node can store transaction logs directly using block storage based on its own storage engine component, without relying on a file system. This reduces the resource overhead associated with the introduction of a file system and file storage. The storage engine component writes transaction logs to online log files, leveraging the cyclical nature of online log files and eliminating the need to create binary logs, further reducing the complexity and resource overhead of log storage operations. Furthermore, the database node can leverage its ability to detect transaction termination to perform transaction-level persistent storage operations on the metadata information of the target online log file. This ensures that the persisted metadata information is complete metadata for the target transaction, improving the controllability and reliability of the metadata persistence operation. The present disclosure also provides a computer program product, including a computer program. When executed by a processor, the computer program implements the method of the aforementioned method embodiment. Accordingly, the present disclosure also provides a computer-readable storage medium storing the computer program. When executed, the computer program implements the steps of the aforementioned method embodiment that can be performed by the server. Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM (Compact Disc Read-Only Memory), optical storage, etc.) containing computer-usable program code. The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, may be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more blocks in the block diagram.These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in one or more flows in a flowchart and / or one or more blocks in a block diagram. These computer program instructions may also be loaded onto a computer or other programmable data processing device, causing the computer or other programmable device to execute a series of operational steps to produce a computer-implemented process, such that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows in a flowchart and / or one or more blocks in a block diagram. In a typical configuration, a computing device includes one or more processors (Central Processing Units, CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium. Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, Parallel Random Access Machine (PRAM), Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), other types of Random Access Memory (RAM), Read-Only Memory (ROM), Electrically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other memory technologies, Compact Disc Read-Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.It should also be noted that the terms "comprise," "include," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, product, or apparatus comprising a list of elements may include not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, product, or apparatus. Without further limitation, the phrase "comprising a..." does not preclude the presence of additional, identical elements in the process, method, product, or apparatus comprising the recited elements. The foregoing description is merely an example of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations of the present disclosure are possible. Any modifications, equivalent substitutions, improvements, and the like made within the spirit and principles of the present disclosure are intended to be encompassed by the claims of the present disclosure.
Claims
Claims 1. A log storage method, applicable to any database node in a distributed database, wherein, Including: Obtaining the transaction log of the target transaction by using a storage engine component; Writing the transaction log into a target online log file; The target online log file is an online log file that supports write operations at any position; Persistently storing the transaction log in the target online log file into a corresponding target block device.
2. The method according to claim 1, wherein, Writing the transaction log into the target online log file includes: Determining a first online log file in a startup state; Judging whether the data volume of the transaction log is greater than the remaining writable data volume of the first online log file; If so, closing the first online log file and starting a new second online log file as the target online log file; Writing the transaction log into the target online log file.
3. The method according to claim 1 or 2, wherein Writing the transaction log into the target online log file includes: Writing the transaction log into a target memory cache block corresponding to the target online log file; Persistently storing the transaction log in the target online log file into a corresponding target block device includes: Synchronizing the transaction log written in the target memory cache block to a target block device associated with the target online log file according to a log synchronization parameter.
4. The method according to claim 3, wherein Further including: Determining the log synchronization parameter according to at least one of the load information of the distributed database, the throughput capacity information of the target block device, and the backlog information of the online log file to be synchronized in the distributed database.
5. The method according to claim 3 or 4, wherein, The storage space sizes of the target memory cache block and the target block device are the same.
6. According to the method according to any one of claims 1-5, wherein Further including: During the process of writing the transaction log into the target online log file, determining a first block device already associated with the target online log file; If the data volume of the transaction log is greater than the first block device, applying to a storage management component for allocating a second block device; Associating the target online log file with the second block device to expand the target online log file.
7. The method according to any one of claims 1-6, wherein, Further including: During the process of persisting the transaction log, in response to a commit instruction of the target transaction, writing metadata information of the target online log file into a redo log file of the target transaction; And writing a commit event of the target transaction into the redo log file of the target transaction; And after completing the persistent storage of the transaction log, completing a commit operation of the target transaction.
8. The method according to claim 7, wherein Further including: After the target transaction is committed, performing a persistent operation on the redo log file of the target transaction to perform a transaction-level persistent operation on the metadata information of the target online log file; Or, After a transaction group to which the target transaction belongs is committed, performing a persistent operation on the redo log file of the target transaction to perform a transaction-group-level persistent operation on the metadata information of the target online log file.
9. The method according to any one of claims 1-8, wherein The target online log file includes: An initial online log file created during database node initialization or an online log file that is put into circular use after being archived.
10. The method according to any one of claims 1-8, wherein, After persistently storing the transaction log in the target online log file into the corresponding target block device, the method further includes: archiving the target online log file to a specified storage space, and updating the target online log file to an archived state, so that the target online log file can be recycled after archiving.
11. - A distributed database system, wherein, The method includes: a plurality of database nodes, a database file component, and a storage component; wherein, the storage component is configured to provide at least one block device; the database file component is configured to provide at least one online log file; any one of the plurality of database nodes is configured to: obtain the transaction log of the target transaction by using a storage engine component; write the transaction log into a target online log file among the at least one online log file; the target online log file is an online log file that supports write operations at any position; persistently store the transaction log in the target online log file onto a target block device associated with the target online log file among the at least one block device.
12. The system according to claim 11, wherein The system further includes: a storage management component; the storage engine component is further configured to: during the process of writing the transaction log into the target online log file, determine a first block device associated with the target online log file; if the data volume of the transaction log is greater than the first block device, send a block device application to the storage management component, so that the storage management component allocates a second block device for the target online log file from the at least one block device; associate the target online log file with the second block device to expand the target online log file.
13. A server, wherein, The method includes: a memory and a processor; the memory is configured to store one or more computer instructions; the processor is configured to execute the one or more computer instructions to: execute the steps in the method according to any one of claims 1-10.
14. A computer program product, wherein, The method includes a computer program, which when executed by a processor, implements the log storage method according to any one of claims 1-10.
15. A computer-readable storage medium storing a computer program, wherein, A computer program, when executed by a processor, can implement the log storage method according to any one of claims 1-10.
Citation Information
Patent Citations
Data file processing method, device and system and storage medium
CN110908589A
Database real-time backup method and device, computer equipment and storage medium
CN115454717A
Data segment writing method and device and data reading method and device
CN116701387A
Cited By
Multi-data source acquisition and storage method
CN121262088A
Database log storage method and device, equipment and medium
CN121807793A