Data writing method and device, equipment and storage medium
By generating and directly sending physical logs to multiple log nodes in a distributed system, the majority threshold confirmation mechanism is adopted, and the network overhead and delay problems of data writing in a distributed system are solved, improving write reliability and system performance.
Patent Information
- Application Number
- CN202510528263.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-26
AI Technical Summary
During the data writing process, the existing distributed systems require all nodes to be completed simultaneously, resulting in high network overhead and delay, affecting system performance, and complex abnormal and fault recovery mechanisms, affecting data consistency and recovery efficiency.
The calculation node generates a physical log containing the physical change information and transaction semantic information of the data page, and directly sends it to at least 3 log nodes. The majority threshold confirmation mechanism is used to determine that the write operation is successful, reducing network communication and simplifying the synchronization process.
It improves the reliability and system performance of data writing, reduces network communication overhead between nodes, and enhances the fault tolerance and robustness of the system.
Smart Images

Figure CN120540581A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and in particular to the field of database and distributed storage technology. Background Art
[0002] Current distributed systems typically ensure data availability by replicating data across multiple nodes. Common strategies include two-phase commit and consensus algorithms like the Replicated and Fault Tolerant (Raft) protocol to ensure consensus across multiple replicas. Data replication typically requires all or most nodes to confirm receipt of the data before marking the write as successful, which can result in high network overhead in some cases. Summary of the Invention
[0003] The present disclosure provides a data writing method, apparatus, device, and storage medium.
[0004] According to one aspect of the present disclosure, a data writing method is provided, comprising:
[0005] In response to a write operation, the computing node generates a physical log containing physical change information of the data page and transaction semantic information based on the data to be written;
[0006] The computing node directly sends the physical log to at least three log nodes;
[0007] If the number of successful confirmations received by the computing node reaches a predetermined majority threshold, the write operation is determined to be successful.
[0008] According to another aspect of the present disclosure, there is provided a data writing device, comprising:
[0009] A log generation module is used to generate a physical log containing physical change information of the data page and transaction semantic information based on the data to be written in response to a write operation by a computing node;
[0010] A sending module, configured to send the physical log directly to at least three log nodes by the computing node;
[0011] The determination module is configured to determine that the write operation is successful if the number of successful confirmations received by the computing node reaches a predetermined majority threshold.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.
[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.
[0017] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.
[0018] The present disclosure improves the reliability of data writing and the fault tolerance of the system, reduces the network communication overhead between nodes, and improves system performance.
[0019] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0021] Figure 1 is an exemplary architecture diagram of a distributed storage system according to an embodiment of the present disclosure;
[0022] Figure 2 is a flowchart of a data writing method provided according to an embodiment of the present disclosure;
[0023] Figure 3 This is a schematic diagram of the structure of log records provided according to an embodiment of the present disclosure;
[0024] Figure 4 is a structural diagram of a data writing device provided according to an embodiment of the present disclosure;
[0025] Figure 5 is a block diagram of an electronic device used to implement the data writing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0026] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0027] In related technologies, distributed storage systems typically require all nodes to be synchronized before proceeding to the next step. This can result in high latency, especially in scenarios with poor network conditions or low performance on certain nodes. Data synchronization confirmation across all nodes can lead to high network traffic and latency, impacting system write performance. Exception and failure recovery mechanisms are complex, requiring additional coordination and communication to ensure data consistency and recovery.
[0028] In order to at least partially solve one or more of the above-mentioned problems and other potential problems, the embodiments of the present disclosure provide a data writing method. By utilizing the technical solutions of the embodiments of the present disclosure, the reliability and fault tolerance problems of data writing operations in distributed systems can be solved, while reducing the network communication overhead caused by data synchronization between nodes.
[0029] This technology is applicable to various distributed database systems, distributed storage systems, distributed file systems, and other distributed applications that require high availability and high reliability. These systems usually require data replication on multiple nodes to ensure data persistence and availability.
[0030] In order to make the purpose, technical solutions and advantages of this disclosure clearer, Figure 1 The embodiments of the present disclosure are described in detail. Figure 1 An exemplary architecture of a distributed storage system according to an embodiment of the present disclosure is shown. The system is designed to provide high availability, scalability, and data consistency.
[0031] In one embodiment, the distributed storage system may include: one or more proxy nodes (Proxy), a computing layer, a log service (LogService), a metadata management service (Page Server Manager) and a storage layer (composed of multiple Page Servers).
[0032] The client interacts with the system through the proxy node, which is responsible for routing client requests to the computing layer.
[0033] The computing layer may include a primary compute node and multiple secondary compute nodes (including special hot standby nodes, called Hot_Secondary). The primary compute node is responsible for processing read and write requests, including parsing SQL, generating and executing query plans. Standby compute nodes are typically responsible for processing read-only requests and maintaining state consistency by synchronizing with the primary node so that they can take over the primary node's functions when needed (for example, through hot or cold switching mechanisms). Hot standby nodes can continuously synchronize more comprehensive state information (such as metadata and buffer pool contents) to achieve faster primary-standby switching.
[0034] The metadata management service (Page Server Manager), for example, uses the Raft protocol to implement high-availability cluster deployment, and is responsible for managing the system's global metadata, such as cluster topology, tablespace definition, and the mapping relationship between data fragments (such as segments) and physical storage locations.
[0035] The LogService can also deploy multiple replicas using consensus protocols such as Raft to ensure high availability and data persistence. It is responsible for receiving and persisting the physical redo logs (RedoLogs) generated by the primary compute node. These redo logs record physical modifications to the data. The LogService uses protocols such as Quorum to ensure the reliability of log writes and provides log reading services for other system components, serving as a key hub for data synchronization. Log records typically include a unique, monotonically increasing Log Sequence Number (LSN).
[0036] The storage layer consists of multiple page servers, also known as storage nodes. Each page server is responsible for managing the storage resources of the physical node where it is located. The system's data is organized into logical units (for example, segments), and each segment has multiple replicas (replicas). These replicas are distributed across different page servers to achieve data redundancy and high availability. Each segment replica independently pulls (or receives) Redo logs from the log service and replays (applies) physical changes to the data pages of its local storage (such as PageStore files) based on the log content, thereby maintaining data synchronization with the master node.
[0037] An exemplary data synchronization process includes the following steps:
[0038] 1. The client initiates a Data Definition Language (DDL) operation through the proxy node, such as a table creation operation.
[0039] 2. The proxy node forwards the request to the main computing node.
[0040] 3. The primary computing node processes the request:
[0041] a. Interact with the metadata management service (Page Server Manager) to perform operations such as creating tablespaces, registering metadata, and establishing related mappings.
[0042] b. Based on the above operations, one or more corresponding physical redo log records are generated, for example, labeled "create_tablespace_sync redo log". This log record contains the physical layer change information required to perform the DDL operation.
[0043] c. The primary computing node sends the generated Redo log records to the log service cluster.
[0044] 4. The log service cluster persists the Redo log record based on its consensus protocol (such as Raft) and write protocol (such as Quorum) and assigns it an LSN to ensure reliable storage of the log.
[0045] 5. The data replicas in the storage layer (Segment Replicas, located on different Page Servers) independently and continuously pull or subscribe to new Redo log records from the log service. Upon receiving the "create_tablespace_syncredolog" record, each relevant replica will replay the log. This may involve creating corresponding metadata structures or initializing data pages in its local storage space to reflect the state of the created table. During the replay process, the replica can verify the continuity of the log stream and the correctness of the application based on the sequence information in the log record (such as LSN or the enhanced checksum field proposed in this invention).
[0046] At the same time or later, a standby compute node (especially a hot standby node) also obtains the same Redo log records from the log service. The standby compute node replays the log records to update its internal state, such as cached metadata, to synchronize with the primary node and prepare for possible read-only query services or fast master-slave failover.
[0047] Figure 2 FIG. 1 is a flow chart of a data writing method according to an embodiment of the present disclosure. Figure 2 As shown, the method comprises at least the following steps:
[0048] S210 . In response to a write operation, the computing node generates a physical log including physical change information of the data page and transaction semantic information according to the data to be written.
[0049] In the disclosed embodiment, when the storage system receives a write operation that needs to be persisted, such as an insert or update request from a client, a computing node (such as the master computing node in the aforementioned architecture) is responsible for processing the operation. The computing node generates a corresponding physical log record based on the data to be written. This physical log record contains both physical change information for the underlying data page and the necessary transaction semantic information (such as the transaction ID to which it belongs and possible transaction boundary markers).
[0050] S220: The computing node directly sends the physical log to at least three log nodes.
[0051] After generating a physical log record (or batch of records) containing the above information, the compute node sends the physical log directly, typically in parallel, to at least three log nodes in the system. These log nodes correspond to the LogService replicas in the aforementioned architecture. This direct transmission eliminates the intermediate forwarding steps that may be present in traditional master-slave replication models (for example, eliminating the need to first send the data to a log master node and then distribute it to slave nodes), thereby reducing network latency. The selection of at least three nodes establishes a foundation for majority confirmation, thereby improving fault tolerance.
[0052] S230: When the number of successful confirmations received by the computing node reaches a predetermined majority threshold, determine that the write operation is successful.
[0053] After sending the physical log, the compute node listens for successful confirmations from each log node, which indicates that the corresponding log node has successfully received and persisted the physical log. The compute node counts the number of valid confirmations received. When the number of successful confirmations received reaches the predetermined majority threshold (for example, for 3 log nodes, the threshold is 2), the compute node can determine that the write operation is successful. This quorum-based confirmation mechanism ensures that even if a few log nodes fail or respond slowly, as long as the majority of nodes confirm persistence, the write operation can be considered to be completed reliably, thereby ensuring high availability and durability of data writes.
[0054] According to the solution of the embodiment of the present disclosure, the end-to-end data writing process optimization is achieved by generating a unified physical log containing physical changes and transaction semantics by the computing node and sending it directly to multiple log nodes for Quorum confirmation. It avoids the conversion complexity and potential inconsistency risks between multiple log formats, while reducing the number of network hops in the write path (no need for intermediate master node forwarding), significantly reducing write latency and improving system throughput. The Quorum-based confirmation mechanism enhances the fault tolerance and data reliability of write operations, simplifies the distributed database system architecture as a whole, and improves performance and robustness.
[0055] In one possible implementation, in response to a write operation, a computing node generates a physical log containing physical change information and transaction semantic information of the data page according to the data to be written, including:
[0056] S211 . In response to a write operation, the computing node determines a transaction identifier and physical change information of a target data page according to the data to be written.
[0057] S212: Obtain a physical log according to the transaction identifier, the physical change information, and the log sequence number.
[0058] In the embodiment of the present disclosure, first, in response to a write operation, the computing node needs to determine the core content required for the log record. Based on the data to be written and the context of the current operation, the computing node will determine the transaction identifier (Transaction Identifier, TxID) of the transaction to which it belongs, which is a mark that uniquely identifies the transaction. At the same time, the computing node will analyze which target data page or pages will be modified by the write operation, and determine the physical change information describing these modifications. This includes the new data that needs to be written, the specific location (offset) in the page, and the original data that may need to be recorded.
[0059] After determining the transaction identifier and physical change information, the compute node combines this information with a log sequence number (LSN) assigned to the log record. An LSN is a unique, typically monotonically increasing number that identifies the order and position of the log record within the log stream. By combining the transaction identifier, physical change information, and LSN (possibly with other metadata such as timestamps and checksums), a structured physical log record is ultimately obtained.
[0060] According to the solution of the embodiment of the present disclosure, the basic steps for generating unified physical log records are clarified. By first determining the transaction identifier and core physical change information, and then combining them with the assigned log sequence number, it is ensured that each basic log record structurally contains the necessary context (transaction ownership) and content (physical modification) as well as its sequential identification in the log stream, which helps to ensure the integrity of the log record.
[0061] In one possible implementation, in response to a write operation, the computing node determines, based on the data to be written, a transaction identifier and physical change information of a target data page, including:
[0062] S211a. In response to the write operation, the computing node determines a target transaction associated with the write operation and obtains a transaction identifier.
[0063] In the disclosed embodiments, when a compute node processes a write operation, the operation typically belongs to an active transaction. The compute node needs to determine the target transaction associated with the write operation and obtain or assign a unique transaction identifier (TxID) for the transaction. This identifier associates the log record generated by the write operation with the entire lifecycle of the transaction (start, commit, or abort).
[0064] When a new transaction starts, a new unique transaction identifier is assigned. For operations within the transaction (data writing, committing, aborting), the transaction identifier assigned to the transaction is used.
[0065] S211b. Determine the current change information and historical change information of the target data page according to the data to be written.
[0066] Next, the computing node needs to comprehensively determine the change information of the target data page based on the data to be written. This includes not only the current change information directly caused by the current write (for example, the specific content and write location of the new data), but also the historical change information related to the modification history of the page. In a preferred embodiment of the present disclosure, this historical change information is an identifier pointing to the log record of the last modification of the target data page, which is used to subsequently implement page-level modification chain verification.
[0067] S211c. Obtain physical change information based on the current change information and historical change information.
[0068] Finally, the compute node combines or packages the current change information determined in the previous step with the historical change information (if any) to obtain the final, complete physical change information for logging. This packaged information will contain all the necessary details so that the receiving node can correctly apply the changes and perform relevant verification.
[0069] According to the solution of the embodiment of the present disclosure, when determining physical change information, not only the current change information is considered, but also the historical change information (i.e., the identifier pointing to the log record of the previous modification of the page). By actively including link information of the page modification history when generating log records, the necessary data foundation is provided for downstream storage nodes to implement accurate page-level modification chain integrity verification. This enhances the accuracy of data synchronization, allowing the storage layer to discover potential data inconsistencies earlier and more accurately during log playback.
[0070] In a possible implementation, S211b determines the current change information and historical change information of the target data page based on the data to be written, including:
[0071] The target data page is determined based on the data page modification caused by the data to be written.
[0072] Based on the table space identifier and physical page number of the target data page, query the previous log record of the last modification of the target data page.
[0073] Get the change information based on the data to be written, table space identifier and physical page number.
[0074] Obtain historical change information based on the identification sequence number of the previous log record.
[0075] In the embodiment of the present disclosure, the computing node first locates the target data page that needs to be operated based on the data page modification caused by the data to be written.
[0076] To retrieve historical change information, a compute node uses the physical identifier of the target data page, namely the tablespace ID and the physical page number, to query an internally maintained map (e.g., primary_page_lsn_map). This query aims to locate relevant information about the previous log record that recorded the previous modification of the target data page, specifically its identifying sequence number.
[0077] The computing node generates the change information describing the specific modification content according to the data to be written, the table space identifier and the physical page number of the target data page.
[0078] The compute node uses the identification sequence number of the previous log record (i.e., the start or end LSN of the previous log record that modified the page) as historical change information. This historical change information will be included in the log record to be generated, and may be stored in the second verification field (page_last_rec_end_lsn field).
[0079] According to the solution of the embodiment of the present disclosure, historical change information can be obtained specifically and efficiently. By using the physical identifier of the target data page (tablespace ID, page number) to directly query the previous log record information maintained inside the computing node, the accuracy and timeliness of the historical change information obtained are ensured, and the page chain pointer contained in the generated physical log record is guaranteed to be correct and reliable, thereby effectively supporting the downstream storage nodes to perform strict page modification chain consistency verification.
[0080] In a possible implementation, S212 obtains the physical log according to the transaction identifier, the physical change information, and the log sequence number, including:
[0081] S212a: Combine the transaction identifier, physical change information, and transaction boundary record to generate multiple log records.
[0082] In the disclosed embodiments, a computing node generates multiple types of log information during the processing of a transaction. This includes log records containing TxID and physical change information generated by data write operations, as well as transaction boundary records that mark the transaction lifecycle status, such as records marking the start (Begin), successful end (Commit), or failed end (Abort) of a transaction. The computing node combines these two types of records (physical change records and transaction boundary records) according to their logical order during the transaction execution process to form a sequence of multiple log records.
[0083] S212b. Assign a log sequence number to the log stream formed by the multiple log records to obtain a physical log.
[0084] Next, the compute node assigns a log sequence number to the log stream formed by these multiple log records. This LSN assignment ensures the global order of the entire log stream. After LSN assignment and possible packaging (for example, forming a log block), the physical log is ultimately ready for transmission. This physical log contains complete transaction information (through boundary records and embedded TxIDs) as well as all physical change details.
[0085] According to the solution of the embodiment of the present disclosure, by combining log records describing physical changes with transaction boundary records marking the start and end (commit / abort) of transactions, and assigning sequence numbers to the resulting complete log stream, the generated physical log itself is ensured to contain complete transaction context information. This allows downstream systems (such as storage nodes or standby computing nodes) to understand the transaction scope, final state, and apply physical changes directly by consuming this single physical log stream, completely eliminating the need for independent logical logs (such as binlogs) and significantly simplifying replication or recovery logic that requires transaction awareness.
[0086] In one possible implementation, S212b assigns a log sequence number to a log stream formed by multiple log records to obtain a physical log, including:
[0087] For a log stream formed by multiple log records, a unique and continuously increasing log sequence number is assigned to each unit in the log stream.
[0088] Writing the identification sequence number of a first log record in the log stream into a first verification field of a second log record, wherein the first log record is a previous log record adjacent to the second log record.
[0089] In the disclosed embodiment, for a log stream formed by multiple log records, a unique and continuously increasing log sequence number (LSN) is assigned to each unit in the log stream (i.e., a single log record). This ensures the strict order and identifiability of the log records.
[0090] To support subsequent log stream continuity verification, when constructing a log stream, the system writes the identification sequence number (which can be the start LSN or end LSN of the first log record in the log stream (i.e., the physically previous log record)) into a specific field of the second log record (the current record) that follows it. This field is called the first verification field (for example, the last_rec_end_lsn field). In this way, each log record contains a pointer to its immediate predecessor, allowing the recipient to verify that the log stream is intact. The resulting sequence of log records (or the packaged blocks) with LSNs and link pointers is the resulting physical log.
[0091] According to the solution of the embodiment of the present disclosure, a mechanism for enhancing log stream integrity verification is provided. By assigning a unique and continuously increasing LSN to each record unit in the log stream and embedding a first verification field pointing to the end LSN of its physical predecessor record in each log record, an internal self-verifying log chain is constructed. This enables the log receiver to efficiently verify the physical continuity of the entire log stream when processing each record, and promptly identify log record loss or disorder issues that may occur due to generation, transmission, and other links, thereby improving the robustness of data synchronization and the accuracy of problem location.
[0092] In a possible implementation, the method further includes the steps of:
[0093] The modification chain identifier of the target data page is updated in the computing node according to the sequence number of the log record used to modify the target data page in the physical log.
[0094] The modification chain identifier is used to indicate the location of the previous log record of the previous modification of the target data page.
[0095] In the embodiment of the present disclosure, a modification chain identifier associated with the target data page is updated in a computing node according to a sequence number of a log record in a physical log used to modify the target data page.
[0096] The modification chain identifier here is a state information maintained internally by the computing node (for example, the value stored in the memory map primary_page_lsn_map). It is used to indicate (record) the LSN of the modification log record that has just been generated. This LSN also becomes the location of the previous log record that needs to be queried and filled into the page_last_rec_end_lsn field (the second verification field) of the new log record when the same target data page is modified next time. This update action usually occurs after the LSN of the log record is determined, or after the log record is successfully sent and may have been confirmed by the Quorum (depending on the specific consistency model implemented), ensuring that the computing node always holds the latest modification history information for each page. This step is crucial to ensure that the correct page history change information is contained in the subsequently generated log records.
[0097] According to the solution of the embodiment of the present disclosure, this step ensures that the computing node can continuously and accurately generate physical log records containing historical change information (page chain pointers). After generating a log record for modifying a data page, the modification chain identifier of the page maintained internally by the computing node is promptly updated (i.e., the LSN just assigned is recorded as the latest modified version of the page), ensuring that when the page is modified again in the future, the correct and latest preceding log record LSN can be queried. This is a key prerequisite for achieving the correct transmission of page modification chain information in the log, and ensures the accuracy of the data source of the entire page-level consistency verification mechanism.
[0098] In a possible implementation, the log service node is used to store a physical log, which is used to record operations on data pages stored in the storage node.
[0099] In the disclosed embodiment, the log service (LogService) node can deploy multiple copies using consensus protocols such as Raft to ensure high availability and data persistence. It is responsible for receiving and persisting the redo logs (Redo Log) generated by the main computing node. These redo logs record the modification operations on the physical level of the data. The log service node uses protocols such as Quorum to ensure the reliability of log writing and provides log reading services to other components of the system for data synchronization. Log records usually contain a unique, monotonically increasing log sequence number (Log Sequence Number, LSN). For example, the physical log block of the physical log has a common header (Common Header). The header contains key metadata information, such as the starting position of the sequence number of the content contained in the log block (for example, the block_start_lsn field) and the total data length of the log block (for example, the block_total_length field, indicating the expected number of bytes of the block from beginning to end).
[0100] In one possible implementation, the serial number of the physical log is used to verify the log stream continuity of each log record contained in the physical log; the historical change information of the log record and the current version identifier of the target data page to be modified by the log record are used to verify the integrity of the modification chain; wherein the historical change information is used to indicate the preceding log record that last modified the target data page.
[0101] In the embodiment of the present disclosure, the log record of the physical log includes at least one of the following information: table space identifier, data page number, data length, real data, end sequence number of the previous record, end sequence number of the previous record on the page, etc. For example, a log record of a physical log may include the following fields:
[0102] 1. Single record flag+type (record flag and type):
[0103] Structure: For example, it may occupy 1 byte (1B).
[0104] Function: Used to identify the specific operation type and possible control flags of the log record. The operation type indicates what kind of physical operation is performed by the log record, for example Figure 3 MLOG_INIT_FILE_PAGE (initialization file page) or MLOG_WRITE_STRING (write data string) in the log file. This field is the basis for parsing and applying log records.
[0105] 2. spaceid (tablespace identifier):
[0106] Structure: Its size is determined by the system design, for example, it can be a 4-byte or 8-byte integer.
[0107] Function: It is used to uniquely identify the logical table space or equivalent data organization unit to which the data page operated by the log record belongs. Figure 3 As shown, the spaceid of both records is 100.
[0108] 3. page_no (physical page number):
[0109] Structure: Its size is determined by the system design, for example, it can be a 4-byte or 8-byte integer.
[0110] Function: Used to uniquely identify the specific physical data page number operated by the log record within the specified spaceid. Figure 3 As shown, the page_no of both records is 0. The combination of space id and page_no uniquely identifies the target data page.
[0111] 4. bodylen (data body length):
[0112] Structure: Its size is variable, for example, it can be 3 bytes (3B), or it can be designed as other fixed or variable length representations as needed.
[0113] Purpose: Used to indicate the actual byte length of the data field that follows. This enables the parser to accurately read variable-length physical change data.
[0114] 5. data (real data):
[0115] Structure: variable-length byte sequence, whose length is specified by the bodylen field.
[0116] Function: Contains the actual physical changes, which can be a new piece of data starting at a certain offset within the data page, differential data, or a specific pattern used to initialize the page.
[0117] 6. last_rec_end_lsn (last record end LSN / first verification field):
[0118] Structure: Usually an 8-byte (8B) integer used to store the log sequence number (LSN).
[0119] Function: It stores the end LSN of the previous log record immediately before the current log record in the entire log stream. Figure 3In the example, the last_rec_end_lsn of the second record (MLOG_WRITE_STRING) is 9020, which points to the end LSN of the first record (MLOG_INIT_FILE_PAGE).
[0120] 7. page_last_rec_end_lsn (end LSN of the previous record on the page / second verification field):
[0121] Structure: Usually an 8-byte (8B) integer.
[0122] Function: A core field used to implement page-level modification chain integrity verification. It stores the end LSN of the last log record that modified the same data page (i.e., the same space ID and page_no) before the current log record.
[0123] like Figure 3 As shown, the first record (MLOG_INIT_FILE_PAGE) modifies Page (100,0), and its page_last_rec_end_lsn is 0, which usually means that this is the first record modification of the page, or its predecessor state corresponds to LSN0.
[0124] The second record (MLOG_WRITE_STRING) also modifies Page (100, 0). Because the first record is the immediately preceding record and modifies the same page, the second record's page_last_rec_end_lsn value is 9020, precisely pointing to the end LSN of the first record. When processing the second record, the storage node can compare 9020 with the current version identifier of Page (100, 0) in its local record to complete the modification chain verification.
[0125] In the embodiment of the present disclosure, after the storage node successfully receives the physical log, it can perform an integrity check of the physical log. An example of an integrity check process includes: the storage node pulls the physical log to be applied from a log source, such as a log service node, reads the block_start_lsn field in the log block header of the physical log, and verifies whether its value is equal to the next LSN (i.e., Lsn_N+1) that the storage node expects to receive. If they are not equal, it may indicate that an erroneous or disordered log block has been received, and the check has failed. Read the block_total_length field in the log block header. Compare the value of the block_total_length field with the number of bytes (received_bytes) of the log block actually received. If received_bytes is equal to block_total_length, and block_start_lsn is also as expected, it indicates that the received physical log block has not been truncated or lost data during the transmission process, and it is complete as a physical unit, and the integrity check passes. If received_bytes is not equal to block_total_length, it means that the log block is incomplete or has data errors. It may have been truncated during network transmission, causing verification failure.
[0126] After passing the physical log block integrity check, the storage node can further perform more fine-grained log stream continuity and modification chain integrity verification on each log record contained in the block.
[0127] When the log integrity check passes, the storage node begins to parse the log records contained in the physical log block one by one according to the order of the serial numbers (LSN from small to large). The parsing process usually involves reading the header information of each log record. After each log record is parsed, the storage node can perform log stream continuity verification on the log record. Only if the log stream continuity verification passes will the log record be further applied. After processing the log record, the next log record will be processed in the order of the serial numbers. Before processing the log record, the storage node uses the starting LSN and total length information recorded in the log block header to quickly and effectively detect problems with physical log blocks that are damaged or incomplete due to network or other reasons, avoiding subsequent parsing and application based on erroneous data and improving the robustness of data synchronization.
[0128] In the disclosed embodiment, the storage node can perform log stream continuity verification on each log record contained in the physical log in sequence according to the sequence number order of the physical log. For example, the storage node determines the first unprocessed log record in the physical log according to the sequence number order of the physical log to be applied; and performs log stream continuity verification based on the first identifier contained in the log record and the completion status sequence number of the current storage node. The first identifier is used to indicate the position of the log record in the log stream relative to the previous log record.
[0129] In the disclosed embodiment, the storage node will process the log records in the physical log in the order of the serial numbers (i.e., the order of LSN from small to large). For the first log record to be processed, Record_{N+1} (whose LSN is Lsn_{N+1}), the node will read the last_rec_end_lsn field contained in the record, i.e., the first identifier. This field indicates the end LSN of the log record immediately before Record_{N+1} from the perspective of the master node. The node performs log stream continuity verification, i.e., compares whether Record_{N+1}.last_rec_end_lsn is equal to the node_applied_lsn (i.e., Lsn_N) currently maintained by the node. When the two are equal, it indicates that the received log stream and the log stream currently processed by the node are continuous and without gaps, and the verification is passed. The storage node can verify the continuity of the global log stream before applying each physical log record, which helps to ensure the accuracy and reliability of data synchronization.
[0130] In the embodiment of the present disclosure, when the log stream continuity verification is passed, the storage node performs modification chain integrity verification based on the historical change information of the log record (the second identifier, stored in the second verification field) and the current version identifier of the target data page to be modified by the log record. For example, when the log stream continuity verification is passed, the storage node determines the target data page to be modified based on the header information of the log record; and performs modification chain integrity verification based on the second identifier of the log record and the version identifier of the target data page. Through page-level modification chain verification, a more fine-grained verification of the consistency of the data page state is provided, which is crucial for handling node internal failures, recovery, and potential application logic errors, and provides a strong guarantee for the accuracy and reliability of data synchronization.
[0131] In the disclosed embodiment, the storage node performs modification chain integrity verification based on the second identifier of the log record and the version identifier of the target data page. For example: the storage node performs modification chain integrity verification based on whether the version identifier and the second identifier of the target data page point to the same physical log. The version identifier is determined based on the third physical log corresponding to the last time the target data page was successfully modified. In order to make a comparison, the storage node needs to know the current "version status" of the target data page. The node reads the version identifier associated with the target data page that it maintains internally. The version identifier here, in a preferred embodiment, is the end LSN (page_current_lsn) of the third physical log (i.e., the last log record successfully applied to the page) recorded by the storage node when the target data page was last successfully modified.
[0132] The storage node obtains the second identifier in the current log record header (i.e., the value of the page_last_rec_end_lsn field, which represents the LSN of the log record immediately preceding the modification of this page from the perspective of the master node). The storage node then verifies the integrity of the modification chain: it compares this "second identifier" (page_last_rec_end_lsn) read from the log record with the "version identifier" (page_current_lsn) of the target data page maintained locally on the slave node to see if they point to the same physical log (i.e., if their LSN values are equal). If they are equal, the current log record is the immediate successor of the last successful modification applied to the page, and the page modification chain is continuous and complete. If they are not equal, it indicates a page-level inconsistency. This may be because a previous log record modifying the page was omitted or skipped, or the node's state was incorrect after recovery from a failure. By maintaining and reading accurate version identifiers for each responsible page, the node has a local benchmark for determining the continuity of subsequent modifications, providing fine-grained consistency verification at the page level. It goes beyond the continuity check of the global log stream and can capture and prevent specific data page content errors caused by internal errors or state synchronization of replicas. This greatly enhances the accuracy and reliability of data synchronization.
[0133] In the embodiment of the present disclosure, when the integrity verification of the modification chain is passed, the storage node modifies the target data page according to the log record. For example, when the version identifier is the same as the second identifier or points to the same physical log, the storage node overwrites the current data of the target data page in the memory according to the real data and offset position in the log record; and updates the version identifier of the target data page according to the identifier of the physical log or the identifier of the log record. Since it has been verified that the current log record Lsn_{N+1} is the next expected operation applied to the target data page Page_X, the storage node believes that it can be modified safely. The node reads the real data contained in the log record body (data field) (i.e., the byte changes at the physical level), as well as the offset position (offset) and data length (body_len) indicated by the record header. The node then directly in the memory (for example, on the copy of the target data page Page_X cached in the Buffer Pool), starting from the specified offset, uses the data in the log record to overwrite the current data of body_len length. After successfully applying the changes of the physical log to the data page in the memory, the status information of the page is updated to reflect this modification. The storage node needs to update the version identifier (i.e., page_current_lsn) maintained internally for the target data page Page_X. The node updates the value of page_current_lsn to the identifier of the log record or physical log that has just been successfully applied. Specifically, it can be the starting LSN or ending LSN of the log record or physical log (e.g., Lsn_{N+1}).
[0134] This embodiment aims to provide a highly reliable, low-latency physical log writing method for use in the aforementioned distributed storage system architecture. This method involves a primary compute node interacting directly with multiple replica nodes of the LogService to achieve quorum-based log persistence and include a mechanism for detecting and repairing data loss between log nodes.
[0135] System Prerequisites:
[0136] The LogService consists of N replica nodes (e.g., N=3) that collectively store the physical Redo log. These nodes act as peer receivers in this write method and do not rely on internal leader-follower replication for persistence of the write operation.
[0137] The primary compute node is responsible for generating physical log records (or log batches) with unique, monotonically increasing sequence numbers (LSNs, which can be considered version numbers).
[0138] The quorum is defined as floor(N / 2) + 1. For N=3, the quorum=2.
[0139] Complete steps:
[0140] 1. Log generation and distribution:
[0141] Step 1.1: When executing a data modification operation (DML) or DDL within a transaction, the primary compute node generates a corresponding unified physical redo log record (or batch) based on its integrated transaction management and consistency logic. This log record contains both transaction semantics and a precise description of the physical changes to the data page, eliminating the need for subsequent conversion. The node assigns a unique, monotonically increasing LSN to this record (or batch).
[0142] Step 1.2: The primary compute node sends this unified physical Redo log record directly and in parallel to all N LogService replica nodes.
[0143] 2. Log reception and persistence:
[0144] Step 2.1: Each LogService node independently receives unified physical Redo log data from the master compute node.
[0145] Step 2.2: Perform LSN check, data packet integrity and other verifications.
[0146] Step 2.3: Write the verified log data to local persistent storage.
[0147] Step 2.4: After the persistence is successful, a success confirmation (ACK) is sent to the primary computing node, carrying the LSN (Lsn_X) of the successful persistence.
[0148] 3. Quorum confirmation and transaction advancement:
[0149] Step 3.1: The master compute node collects ACKs from the LogService node.
[0150] Step 3.2: Count the number of valid ACKs for Lsn_X.
[0151] Step 3.3:
[0152] Success scenario: After receiving ACKs equal to or exceeding the quorum, the primary compute node determines that the unified physical log corresponding to LSN Lsn_X has been reliably persisted. Based on this confirmation, the primary compute node can safely advance the transaction status (for example, commit the transaction) and ultimately return a success message to the client. Due to the unified and direct write nature of the log, this confirmation ensures both transaction durability and consistency across multiple replicas, providing a highly integrated process.
[0153] Failure scenario: If the quorum is not reached within the timeout period, the write is deemed to have failed and a retry is performed or the transaction is aborted.
[0154] According to the solution of the embodiment of the present disclosure, by deeply integrating the transaction management and multi-copy consistency logic of the database master into the master computing node, and unifying the database-level write-ahead log (WAL) and the storage-level physical replication log into a single physical Redo log stream, the master computing node is directly managed by the log's Quorum write, significantly simplifying the system architecture and improving performance and robustness. Specific improvements and beneficial effects include but are not limited to:
[0155] The primary compute node is responsible for not only SQL parsing, query plan generation, and execution, but also manages transaction status and coordinates multi-replica storage consistency, similar to traditional database master databases. The write-ahead log (WAL) and physical page change information generated by transactions are merged into a single physical redo log stream.
[0156] Based on this unification, the primary compute node can directly and concurrently write this unified physical Redo log to the Log Service's N replica nodes, and confirm the write success based on the Quorum, without relying on the LogService's internal Leader or a separate storage replication master node for relay.
[0157] The LogService cluster serves as a reliable persistent storage backend, receiving and persisting unified physical Redo logs sent directly from the master compute node and providing read services. Data loss detection and repair capabilities are available between nodes based on LSN and continuity detection.
[0158] The master compute node writes directly to N log nodes. Compared to the two-hop model of "compute node -> log master node -> log slave node", this eliminates one network hop and significantly reduces network latency along the write path. Reducing network hops and simplifying internal replication logic help improve overall write performance.
[0159] Based on the quorum mechanism, writes succeed as long as a majority of log nodes successfully persist. The system can tolerate temporary failures or slow responses from a few log nodes, ensuring high availability of write operations.
[0160] When processing writes from compute nodes, the log service layer does not require complex leader election and replication logic, reducing system complexity.
[0161] Through the LSN and continuity detection mechanism between nodes, it is possible to automatically detect and pull data from other healthy nodes to repair lost or damaged logs, ensuring the eventual consistency of each replica data and improving the system's self-healing ability.
[0162] Figure 4 Schematic diagram of the structure of a data writing device provided according to an embodiment of the present disclosure. Figure 3 As shown, the device includes:
[0163] The log generation module 401 is configured to generate, in response to a write operation, a physical log containing physical change information of the data page and transaction semantic information according to the data to be written by the computing node.
[0164] The sending module 402 is configured to send the physical log directly to at least three log nodes from the computing node.
[0165] The determination module 403 is configured to determine that the write operation is successful when the number of successful confirmations received by the computing node reaches a predetermined majority threshold.
[0166] In one possible implementation, the log generation module 401 is used to:
[0167] In response to a write operation, the computing node determines the transaction identifier and the physical change information of the target data page according to the data to be written.
[0168] Obtain the physical log based on the transaction identifier, physical change information, and log sequence number.
[0169] In one possible implementation, the log generation module 401 is used to:
[0170] In response to the write operation, the computing node determines a target transaction associated with the write operation and obtains a transaction identifier.
[0171] Based on the data to be written, the current change information and historical change information of the target data page are determined.
[0172] Get the physical change information based on the current change information and historical change information.
[0173] In one possible implementation, the log generation module 401 is used to:
[0174] The target data page is determined based on the data page modification caused by the data to be written.
[0175] Based on the table space identifier and physical page number of the target data page, query the previous log record of the last modification of the target data page.
[0176] Get the change information based on the data to be written, table space identifier and physical page number.
[0177] Obtain historical change information based on the identification sequence number of the previous log record.
[0178] In a possible implementation, the log generation module 401 is configured to combine the transaction identifier, the physical change information, and the transaction boundary record to generate multiple log records.
[0179] Assign a log sequence number to the log stream formed by multiple log records to obtain a physical log.
[0180] In a possible implementation, the log generation module 401 is configured to: for a log stream formed by a plurality of log records, assign a unique and continuously increasing log sequence number to each unit in the log stream.
[0181] Writing the identification sequence number of a first log record in the log stream into a first verification field of a second log record, wherein the first log record is a previous log record adjacent to the second log record.
[0182] In a possible implementation, an update module is further included, configured to:
[0183] The modification chain identifier of the target data page is updated in the computing node according to the sequence number of the log record used to modify the target data page in the physical log.
[0184] The modification chain identifier is used to indicate the location of the previous log record of the previous modification of the target data page.
[0185] For the description of specific functions and examples of each module and submodule of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.
[0186] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0187] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0188] Figure 5A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0189] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0190] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0191] The computing unit 501 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the data writing method. For example, in some embodiments, the data writing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the data writing method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the data writing method in any other appropriate manner (e.g., by means of firmware).
[0192] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0193] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0194] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0195] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0196] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0197] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0198] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.
[0199] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A data writing method, comprising: In response to a write operation, the computing node generates a physical log containing physical change information of the data page and transaction semantic information based on the data to be written; The computing node directly sends the physical log to at least three log nodes; If the number of successful confirmations received by the computing node reaches a predetermined majority threshold, the write operation is determined to be successful.
2. The method according to claim 1, wherein In response to the write operation, the computing node generates a physical log containing physical change information and transaction semantic information of the data page according to the data to be written, including: In response to a write operation, the computing node determines the transaction identifier and the physical change information of the target data page based on the data to be written; A physical log is obtained according to the transaction identifier, the physical change information, and the log sequence number.
3. The method according to claim 2, wherein: In response to the write operation, the computing node determines the transaction identifier and the physical change information of the target data page according to the data to be written, including: In response to the write operation, the computing node determines a target transaction associated with the write operation and obtains a transaction identifier; Determining current change information and historical change information of a target data page based on the data to be written; Physical change information is obtained based on the current change information and historical change information.
4. The method according to claim 3, wherein: The determining, based on the data to be written, current change information and historical change information of the target data page includes: Determine the target data page based on the data page modification caused by the data to be written; According to the table space identifier and physical page number of the target data page, query the previous log record of the last modification of the target data page; Obtaining current change information according to the data to be written, the table space identifier, and the physical page number; The historical change information is obtained according to the identification serial number of the preceding log record.
5. The method according to claim 2, wherein: Obtaining the physical log according to the transaction identifier, the physical change information, and the log sequence number includes: Combining the transaction identifier, the physical change information, and the transaction boundary record to generate a plurality of log records; A log sequence number is assigned to the log stream formed by the multiple log records to obtain a physical log.
6. The method according to claim 5, wherein: The step of assigning a log sequence number to the log stream formed by the plurality of log records to obtain a physical log includes: For a log stream formed by the plurality of log records, assigning a unique and continuously increasing log sequence number to each unit in the log stream; Writing the identification serial number of the first log record in the log stream into the first verification field of the second log record; wherein the first log record is the previous log record adjacent to the second log record.
7. The method according to claim 6, further comprising: updating a modification chain identifier of the target data page in the computing node according to a sequence number of a log record in the physical log used to modify the target data page; The modification chain identifier is used to indicate the location of the preceding log record of the previous modification of the target data page.
8. The method according to claim 1, wherein The log service node is used to store a physical log, and the physical log is used to record operations on data pages. The data pages are stored in a storage node.
9. The method according to any one of claims 1 to 8, wherein The serial number of the physical log is used to verify the log stream continuity of each log record contained in the physical log; the historical change information of the log record and the current version identifier of the target data page to be modified by the log record are used to verify the integrity of the modification chain; wherein, the historical change information is used to indicate the previous log record that last modified the target data page.
10. A data writing device, comprising: A log generation module is used to generate a physical log containing physical change information of the data page and transaction semantic information based on the data to be written in response to a write operation by a computing node; A sending module, configured to send the physical log directly to at least three log nodes by the computing node; The determination module is configured to determine that the write operation is successful if the number of successful confirmations received by the computing node reaches a predetermined majority threshold.
11. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.