Separated memory transaction system based on layered block version chain
Through the layered block version chain and RDMA batch processing mechanism, the access delay and capacity limitation problems of the separated memory transaction system are solved, low-latency, high-concurrency multi-version data management is achieved, the system performance and scalability are improved, and transaction consistency is guaranteed.
Patent Information
- Application Number
- CN202510850702.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-14
AI Technical Summary
The traditional chain-structured separated memory transaction system has the problem of high access latency, while the array-structured system has the problems of limited capacity and false conflicts, and cannot simultaneously meet the requirements of low latency and high concurrency scalability.
It adopts a hierarchical block version chain structure, combined with the RDMA batch access mechanism and asynchronous garbage collection strategy. Through the hierarchical block version chain, multiple version data are organized in fixed-size block units and connected through a linked list to form a version chain that can be expanded on demand. It combines the block-level index table and notification queue for asynchronous index updates, prioritizes the reuse of recyclable outdated version slots, and adopts a hybrid garbage collection strategy and version integrity verification mechanism to ensure transaction consistency.
It achieves low-latency and high-concurrency multi-version data management, improves the performance and scalability of the separated memory transaction system, reduces the overhead of index synchronization and garbage collection, and ensures transaction consistency.
Smart Images

Figure CN120780412A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of distributed transaction systems, and more particularly relates to a separated memory transaction system based on a hierarchical block version chain. BACKGROUND
[0002] With the continuous development of cloud infrastructure, resource decoupling, horizontal expansion and high elasticity have become important demands of modern database systems. The separated memory architecture, as a new system structure that separates computing resources from memory resources, is gradually applied to cloud databases, distributed transaction processing and high-performance computing fields.
[0003] However, under this architecture, since the data is distributed in remote memory, the access of the transaction system to multi-version data usually requires cross-node communication, and the access delay becomes a performance bottleneck. The existing multi-version data management structure mainly adopts a chain structure or an array structure. The chain version structure stores multiple versions of a data item in the form of a linked list, has dynamic expansion capability, but needs to traverse along the pointer hop by hop each time the target version is accessed, and each traversal needs remote access, resulting in frequent network round trips and high access delay. Under high concurrency and hot write load, the version chain of the chain structure will quickly become long, thus exacerbating the delay and performance bottleneck. In contrast, the array version structure stores multiple versions of the same data item in a continuous fixed-size memory block, and the computing node can obtain the entire version array through one remote access, thereby significantly reducing the number of pointer jumps and access delay. However, the array structure has limitations due to fixed capacity: when the concurrent write transactions of hot data come frequently, the version array is easy to be filled, at which time the system needs to adopt the strategy of covering the oldest version, resulting in the phenomenon of "pseudo conflict", i.e. read transactions are forced to abort due to the required version being overwritten by a new version.
[0004] Therefore, the traditional chain structure has the problem of high access delay, and the array structure has the problems of limited capacity and pseudo conflict, which cannot simultaneously meet the requirements of low delay and high concurrency scalability of the separated memory transaction system. Therefore, a new version management structure and mechanism are needed to balance access efficiency and structural scalability, while solving the additional overhead caused by index synchronization and garbage collection, so as to improve the overall performance of the system. SUMMARY
[0005] In view of the above defects or improvement needs of the prior art, the present application provides a separated memory transaction system based on a hierarchical block version chain, which aims to realize a low-delay remote read-write, high-throughput transaction processing mechanism with multi-version expansion through structure optimization and protocol coordination.
[0006] To achieve the above object, according to a first aspect of the present application, a separated memory transaction system based on hierarchical block version chain is provided, comprising a plurality of remote memory nodes and a plurality of computing nodes, the remote memory nodes store a plurality of version data of each data item in the form of fixed-size block units, forming version blocks corresponding to each data item, and the adjacent version blocks are connected by a linked list, forming a block version chain; the head of the block version chain comprises the address of the latest version block, version number and lock flag;
[0007] The computing nodes comprise a block-level index table and a block-level index notification queue; the block-level index table is used to store the addresses of the 2nd to nth version blocks of each data item; the block-level index table is updated by an asynchronous index updating mechanism: when a new version block is generated by a transaction submission, the coordinator sends the metadata information of the new version block to the block-level index notification queue of each computing node, and each computing node asynchronously reads the information from the block-level index notification queue and updates the local block-level index table.
[0008] Preferably, the system further comprises a hybrid garbage collection module, which is used to sequentially execute a node-level transaction state tracking and synchronization process, an obsolete version memory reuse strategy and a background batch recycling strategy.
[0009] The node-level transaction state tracking and synchronization process comprises: maintaining the time stamp of active transactions locally at each computing node, and periodically synchronizing with other nodes to obtain the global earliest active transaction time stamp, so as to determine the range of obsolete versions that can be safely recycled.
[0010] The obsolete version memory reuse strategy is: in the process of a write transaction, the obsolete version slots that can be recycled are preferentially reused.
[0011] The background batch recycling strategy is: when the number of records in the block-level index table reaches a preset threshold, the background garbage collection thread identifies the obsolete versions of hot data in combination with the block-level index table, and re-integrates these versions in time stamp order into new version blocks, while releasing the corresponding old version blocks.
[0012] Preferably, the obsolete version memory reuse strategy in the process of a write transaction comprises:
[0013] In the existing block version chain, check the obsolete version slots that can be reused, if there is a available slot, write the new version data into the slot; otherwise, allocate a new version block according to the on-demand expansion strategy and write the new version data into the new version block.
[0014] Preferably, the start and end positions of each version block are provided with check counters; the transaction consistency guarantee mechanism adopted by the system includes: a read transaction checks whether the check counters at the start and end of a version are consistent when reading the version, and if yes, it is determined that the version data is complete, otherwise it is determined that the version is in a write-incomplete state and the version is skipped.
[0015] Preferably, the computing node further includes a local bitmap management unit for managing the allocation state of the pre-allocated block-shaped version units.
[0016] Preferably, the address of the first version block of each data item is obtained through hash index calculation.
[0017] Preferably, the system adopts a transaction execution protocol based on RDMA one-sided primitives.
[0018] Overall, compared with the prior art, the above technical solutions conceived by the present application can achieve the following beneficial effects:
[0019] 1. The MiT-DM system based on the layered block version chain is suitable for a separated memory architecture scene with resource decoupling and high remote access delay. The system uses a version management structure, i.e., a layered block version chain, which has low-delay access and dynamic expansion capability, to solve the problems of remote access jump amplification of the existing chain structure and pseudo conflict caused by the rigidity of the array structure capacity. A plurality of version data is organized in a fixed-size block unit, and a plurality of version blocks are connected through a linked list to form a version chain that can be expanded on demand. While the local advantage of array structure access is retained, the pseudo conflict problem caused by the limitation of array capacity is avoided. Meanwhile, a lightweight block-level index table is constructed on the computing node to index the version blocks of hot data, and the RDMA Doorbell batch processing technology is combined to enable the computing node to read all versions of a data item through a single network round trip, further improving the remote access performance. In order to fully exert the version access acceleration advantage of the layered block version chain under block-level indexing and RDMA batch processing, the system uses the chain traversable and block characteristics to add a block-level index notification queue to each computing node, and realizes low-overhead asynchronous index updating by decoupling version submission and index synchronization, avoiding the high synchronization cost and submission blocking caused by strong consistency synchronization. Based on the traversable and block characteristics of the layered block version chain, an asynchronous index synchronization scheme based on the notification queue is adopted: when a new version block is generated by transaction submission, the coordinator sends the block-level index update information (data item identifier and new version block address) of the version block to the block-level index notification queue of each computing node; each computing node asynchronously reads the information from the notification queue and updates the local block-level index table; by decoupling the version submission and index updating processes, low-overhead index synchronization is realized, avoiding the blocking waiting and high communication overhead caused by the traditional strong consistency synchronization protocol, thereby maintaining high concurrency performance in a multi-node environment.
[0020] Further, to suppress the performance degradation caused by the expansion of the version chain, the system provided by the application adopts a hybrid garbage collection strategy based on transaction state awareness to compress the version chain and recycle obsolete versions: the earliest start timestamp of the local active transaction is maintained on the computing node, and the global earliest active transaction timestamp is periodically synchronized to determine the safe recyclable version range. During the write transaction process, the coordinator preferentially reuses obsolete version slots that have been determined to be recyclable, thereby delaying the expansion of the chain structure and delaying the expansion of the block version chain structure. The garbage collection thread in the background of the computing node is triggered after the block-level index record reaches a threshold, locates hot data based on the block-level index, and cleans up the corresponding obsolete versions, and reorganizes these versions in timestamp order and writes them into new version blocks, while releasing the corresponding old version blocks; by cleaning up the invalid version data, the size of the version chain and the index is effectively reduced, and the access performance of the system in a high-concurrency scenario is improved.
[0021] Further, to solve the non-atomic problem of the layered block version chain in the non-lock concurrent read and write, the system provided by the application introduces a version integrity check mechanism to guarantee transaction consistency, increases check counts at the beginning and end of the version tuple and version data, and only when the head and tail counts of the read metadata and data are consistent, it is determined that the version data is complete, effectively avoiding the inconsistent read data problem caused by concurrent writing. On this basis, the transaction processing based on the RDMA single side communication primitive is designed, which supports two isolation levels of serializable and snapshot isolation, guarantees the ACID transaction semantics, and reduces the remote interaction overhead.
[0022] To further improve the system efficiency, the system provided by the application also adopts a version memory reuse strategy and an optimization strategy of local resource management: when writing a version, by checking the free slot in the current block version chain, the obsolete version slot is preferentially reused, avoiding frequent allocation of new version blocks on the critical path, reducing the need for frequent allocation of new blocks in the transaction submission process, thereby reducing the synchronization lock competition and network communication overhead caused by version block allocation, improving the write efficiency and throughput of the system; a bitmap structure is maintained locally on the computing node to manage the pre-allocated block version units, which can quickly find and allocate available slots, improve the version block allocation efficiency, and reduce the lock competition and communication overhead of global memory allocation. These optimization methods improve the memory resource utilization and reduce the metadata management overhead.
[0023] In summary, the application realizes low-latency access and high scalability of the multi-version data management structure of the separate memory transaction system, can effectively reduce the overhead of version management, and improve the overall performance of the separate memory transaction system; at the same time, it provides an efficient and reliable solution for index synchronization, garbage collection and consistency guarantee, which can play a synergistic effect in high-concurrency transaction processing scenarios. BRIEF DESCRIPTION OF DRAWINGS
[0024] Figure 1 The overall architecture of the separate memory transaction system provided by the embodiment of the application is shown in the figure;
[0025] Figure 2 The layered block version chain structure provided by the embodiment of the application is shown in the figure;
[0026] Figure 3 The access flow based on the layered block version chain provided by the embodiment of the application is shown in the figure;
[0027] Figure 4 The version tuple memory reuse strategy provided by the embodiment of the application is shown in the figure;
[0028] Figure 5 The block version memory allocation flow provided by the embodiment of the application is shown in the figure;
[0029] Figure 6 A block-level index asynchronous update mechanism based on a notification queue provided for an embodiment of the present application is shown in the schematic diagram.
[0030] Figure 7 A transaction state tracking and synchronization design schematic diagram provided for an embodiment of the present application is shown in the schematic diagram.
[0031] Figure 8 A background thread garbage collection process schematic diagram provided for an embodiment of the present application is shown in the schematic diagram.
[0032] Figure 9 A transaction processing schematic diagram based on a hierarchical block version chain provided for an embodiment of the present application is shown in the schematic diagram. DETAILED DESCRIPTION
[0033] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application. In addition, the technical features involved in the various embodiments of the present application described below can be combined with each other as long as they do not conflict with each other.
[0034] With the wide deployment of Disaggregated Memory (DM) architecture in cloud infrastructure, the traditional transaction system exposes serious performance bottlenecks in this environment. Based on this, the present application proposes a Disaggregated Memory transaction system based on a hierarchical block version chain. In view of the problems of remote access jump multiple of the chain version management structure and the pseudo conflict of the array type version management structure, the hierarchical block version chain structure is adopted, which is combined with the RDMA batch access mechanism and the asynchronous garbage collection strategy, so as to effectively improve the transaction processing performance and system sustainability under the disaggregated architecture.
[0035] The Disaggregated Memory transaction system based on a hierarchical block version chain provided by the embodiment of the present application comprises a plurality of remote memory nodes and a plurality of computing nodes.
[0036] The remote memory nodes store a plurality of version data of each data item in the form of a block unit with a fixed size, form a version block (also referred to as a block version or a block version tuple) corresponding to each data item, connect adjacent version blocks through a linked list, and form a block version chain. The head of the block version chain comprises a latest version block address, a version number and a lock flag.
[0037] The computing node comprises a block-level index table and a block-level index notification queue; the block-level index table is used to store addresses of second to nth version blocks of each data item; the block-level index table is updated through an asynchronous index updating mechanism: when a new version block is generated by transaction submission, the coordinator sends metadata information of the version block to the block-level index notification queue of each computing node, and each computing node asynchronously reads information from the block-level index notification queue and updates the local block-level index table.
[0038] The overall architecture of the system provided by the embodiment of the application is shown in Figure 1 The system comprises a plurality of remote memory nodes and a plurality of computing nodes; the computing node is responsible for executing transaction processing logic, and the remote memory node is used to store multi-version data of a plurality of data items.
[0039] In the remote memory node, the multi-version data of each data item is organized into fixed-size block units to form version blocks corresponding to each data item, and the blocks are connected through a linked list to form a version chain that can be expanded on demand; that is, each remote memory node stores version data organized in a hierarchical block version chain in a memory pool thereof; adjacent version blocks are connected through pointers and constitute a linked list, and the head of the version chain comprises metadata such as the address value_ptr of the latest version block, the version number major_version, and the lock flag lock. As shown in Figure 1 The n version blocks (i.e., block version tuples) are connected through a linked list to form a version chain, and each version block comprises k version tuple slots.
[0040] The computing node maintains a lightweight block-level index table for each data item, which is used to record the identification and location of each version block of the data item.
[0041] With the help of the RDMA Doorbell batch processing mechanism, the computing node can batch acquire all versions of a data item in a single network round trip. The hierarchical block version chain structure combines the dynamic expansion capability of the chain structure and the local access advantage of the array structure, effectively reduces the remote access delay, avoids the pseudo conflict caused by the fixed capacity, and significantly improves the access efficiency and system throughput.
[0042] Based on the traversable and blockable characteristics of the hierarchical block version chain, the system provided by the embodiment of the application adopts an asynchronous index synchronization scheme based on a notification queue:
[0043] When a new version block is generated by transaction submission, the transaction coordinator pushes the update information of the block where the version is located to the block-level index notification queue of each computing node; each computing node asynchronously reads information from the notification queue and updates the local block-level index table.
[0044] Through the asynchronous index updating mechanism, communication and blocking overheads caused by using a strong consistency protocol for index synchronization among multiple nodes are avoided.
[0045] To suppress performance degradation caused by version chain expansion, further, the system provided by the embodiment of the application adopts a hybrid garbage collection module combining transaction state awareness, for sequentially executing a node-level transaction state tracking and synchronization process, an obsolete version memory reuse strategy and a background batch collection strategy;
[0046] The execution of the node-level transaction state tracking and synchronization process comprises: maintaining a timestamp of an active transaction locally at each computing node, and periodically synchronizing with other nodes to obtain a global earliest active transaction timestamp (Global-OAT) to determine a range of obsolete versions that can be safely collected;
[0047] The obsolete version memory reuse strategy is: in the process of a write transaction, a reusable obsolete version slot is reused preferentially, and new versions are written into an existing block, thereby delaying the expansion of a chain structure; that is, in an existing block version chain, a reusable obsolete version slot is checked, if there is a usable slot, new version data is written into the slot; otherwise, a new version block is allocated according to an on-demand expansion strategy, and new version data is written into the new version block.
[0048] The background batch collection strategy is: when the number of records of a block-level index table reaches a preset threshold, a background garbage collection thread identifies obsolete versions of hot data in combination with the block-level index table, and writes these versions into a new version block in timestamp order, while releasing the corresponding old version block.
[0049] Through the above hybrid garbage collection strategy, the version chain and index scale can be effectively compressed, obsolete version data in access can be reduced, and low access latency under high concurrency can be ensured.
[0050] In view of the non-atomicity problem of lock-free concurrent read and write in a separate memory architecture, the application further proposes a version integrity verification mechanism to ensure transaction consistency: a verification counter is maintained at the start and end positions of each version tuple data, a write transaction updates the counter after writing version data, and a read transaction checks whether the start and end counters of the version tuple are consistent when accessing; only when the counters match, the version data is considered to be a complete and valid version, otherwise, the version is considered to be still in writing and the version is skipped. In combination with the integrity verification mechanism, safe concurrent read and write of version data can be realized in a lock-free environment.
[0051] On this basis, the application adopts a transaction execution protocol based on an RDMA single-sided primitive, supports serializable and snapshot isolation levels, reduces the number of remote interactions and delay overheads while ensuring ACID transaction semantics.
[0052] Further, the computing node further comprises a local bitmap management unit for managing the allocation state of the pre-allocated block version units.
[0053] When allocating version slots, the local bitmap is queried to quickly locate available slots, thereby avoiding frequent requests to the global memory allocator.
[0054] Through the local bitmap management mechanism, the system improves the concurrency efficiency of version block allocation and eliminates unnecessary cross-node communication delay.
[0055] The system workflow provided by the embodiment of the application comprises the following stages:
[0056] 1. Initialization stage: a unique block-level version chain starting node is allocated for each logical primary key. Each version chain is composed of a series of physical blocks, and each block contains a plurality of slots for storing different version tuples; the blocks are connected through chain pointers. The computing node pre-constructs a block-level index structure locally, records the version block address corresponding to each primary key, and the index structure supports dynamic expansion and cache replacement.
[0057] 2. Data access stage: when reading data, the transaction obtains the remote version block address according to the primary key through the local block-level index, initiates a multi-block read request in batches through the RDMA Doorbell mechanism, and returns the block data containing version information remotely. The system parses the effective slot according to the bitmap structure, judges the visible version according to the transaction timestamp and isolation level, and returns the result to the upper application, and the process does not need remote logical participation. The write operation adopts an incremental appending strategy, splices the new version tuple at the tail of the block, and only involves the target field and necessary metadata; if there is no free slot in the current block, a new block is allocated and connected to the tail of the version chain.
[0058] 3. Commitment and index synchronization stage: in the transaction commitment stage, the system writes the new block and version information into the local change buffer, and sends the index update message to other replica nodes through the asynchronous notification queue mechanism, completes the distributed block-level index synchronization, and avoids the synchronization operation into the transaction execution main path. The index synchronization adopts a bidirectional channel, supports version confirmation and redundancy invalidation marking, and ensures the index consistency and query correctness.
[0059] 4. Garbage collection phase: The background thread traverses the local block-level index structure periodically, and determines whether each version slot is still dependent on transactions based on the transaction visibility list. For the blocks that are not visible and have stable version sequence, the compression process is performed: the active version is rewritten to a new block, the old block is discarded, and the chain pointer and local index are updated. The version recycling of the present application adopts a hybrid strategy of obsolete version memory reuse and background thread garbage collection. The compression operation is performed in block units, and the unused slots can be recycled while rewriting the version, thereby reducing the memory occupancy and improving the subsequent appending efficiency.
[0060] The above phases will be further described below in conjunction with the accompanying drawings.
[0061] Multi-version data layout: As shown in Figure 2 , the data layout based on the hierarchical block version chain includes the block version chain in the memory pool and the block-level index table in the computing pool. Figure 2 Take n = 2 and k = 5 as an example, that is, a total of 2 block version tuples, each of which includes 5 block units.
[0062] Block version chain: composed of a Header structure and multiple block version tuples (BlockVersionCell) linked by a next_block pointer. The Header stores data item meta information, such as the primary version data pointer value_ptr, the delta data area pointer delta_ptr, the write lock lock, the primary version number major_version, and the read reference count read_count. Each BlockVersionCell contains a fixed-size version tuple array block_node (for example, 4) and a next_block pointer.
[0063] Version tuple (VersionCell): stores the metadata of a single version, and the key fields include version start / version end, valid flag, stale flag, version number (timestamp), delta data offset, and delta attribute bitmap.
[0064] Block-level index table (BlockIndexTable): stored in the local memory of the computing node, and adopts a sparse record strategy. Each entry contains the key of the data item and an address array quick_index, which stores the address of the data item from the second block version tuple. The address of the first block version tuple is obtained by hash index calculation.
[0065] Target version data access: as shown in Figure 3As shown, when the computing node (coordinator) needs to read the target version of the data item, the following steps are performed:
[0066] 1. Read all chunk version tuples of the data: The target version data reading first relies on the index information in the chunk version tuple, in which the Header structure contains the main version data address of the data, and the version tuple contains the corresponding incremental attribute data offset of the version.
[0067] 2. Determine whether the transaction read version is the main version: The determination is based on the transaction start timestamp, and the transaction start timestamp is compared with the main version submission timestamp to determine the target version to be read. If the transaction start timestamp is greater than the main version submission timestamp, the main version data is directly read, and if it is less than the main version submission timestamp, the version tuple with the largest submission timestamp but still less than the transaction start timestamp is found.
[0068] 3. Read the target version data: If the target version to be read is the main version, the main version data is directly read based on the value_ptr in the Header; if the target version is not the main version, the main version data and the incremental attribute value from the main version to the target version are reconstructed to construct the complete target version data value. This process needs to sequentially apply the incremental data in the version tuple to restore the target version to be read by the transaction. In addition, during the transaction operation, the read reference count of the data needs to be maintained to avoid the interference of the garbage collection operation on the normal operation of the transaction. After the transaction is submitted or aborted, the read reference count of the read version needs to be reduced accordingly to release the reference to the version.
[0069] 4. Perform version integrity check: The main version data read and the head and tail check counts of the main version tuple are checked. If the four are consistent, it means that no other transaction has performed write during the transaction reading, and the target version data can be constructed based on the main version data; if any check count is different from the other check counts, the read process is ended and the read reference to the version is released.
[0070] 5. Return the target version data read and end the read data process.
[0071] Version metadata memory management: as shown in Figure 4 and Figure 5 When the coordinator needs to write a new version, the following steps are performed:
[0072] 1. Find useful slot: First, read all the block version tuples of the data item locally (same as read procedure step 1-5). Then, get the current global earliest active timestamp Global-OAT. Traverse the local block version tuples, find if there is an idle slot (never used) or a reusable stale slot (stale flag is true and version < Global-OAT).
[0073] 2. Reuse or allocate: If a useful slot useful_tuple_addr is found, the new version tuple will be written to this address. Setting stale flag of the stale version tuple to true is part of the reuse policy;
[0074] If no useful slot is found, a new block version tuple needs to be allocated. The coordinator calls the local block version memory allocator BlockCellAllocator. The allocator locks the mutex of the reserved memory segment AllocatorReserveRegion corresponding to the memory node ID where the first block of the data item is located. Find the first 0 bit in the bitmap and set it to 1, calculate the corresponding remote memory address new_block_addr, and return this address after unlocking. The new version tuple will be written to the first slot of new_block_addr.
[0075] 3. Write data and metadata:
[0076] 1) Calculate the delta data delta_data of the new version relative to the current major version;
[0077] 2) Construct the new version tuple new_version_tuple, including the new version number (commit timestamp TxCommitTS), delta data offset, bitmap, etc.
[0078] 3) Execute in order by batch RDMA operations:
[0079] 4) Write delta_data to the delta data area;
[0080] 5) Write the new major version data major_data to the major version data area (update the position pointed by value_ptr);
[0081] 6) Write the new version tuple new_version_tuple to the found useful slot useful_tuple_addr or the slot of the newly allocated block;
[0082] 7) Write the updated Header (update major_version to TxCommitTS, set lock to 0 to release the lock);
[0083] 8) If a new block is allocated, the next_block pointer of the previous block also needs to be updated to point to new_block_addr.
[0084] 4. Update index and notify: If a new block version tuple new_block_addr is allocated, insert an entry corresponding to the data item in the local block-level index table, add new_block_addr to the quick_index array, and write an update message of <TableID, Key, new_block_addr> to the notification queue of other computing nodes through RDMA WRITE.
[0085] The system provided by the present application adopts an asynchronous index update mechanism based on a notification queue, with reference to Figure 6 , the asynchronous index update mechanism is as follows:
[0086] 1. Each computing node C_i maintains a ring notification queue InfoQueue_{i->j} for each other computing node C_j (j!= i) in the memory of C_j. The queue includes capacity, head pointer head, tail pointer tail, and message array message;
[0087] 2. When C_i allocates a new block new_block_addr and updates the local index, it sends a notification to all other nodes C_j. C_i atomically increases the tail pointer of InfoQueue_{i->j} using RDMACAS, and then writes the update message to the position pointed to by the tail using RDMAWRITE;
[0088] 3. Each node C_j has a background synchronization thread that periodically polls all notification queues pointing to itself (InfoQueue_{k->j} for all k!= j);
[0089] 4. If the queue is found to be non-empty (head!= tail), the background thread reads all messages in the queue and calls the block_index_insert function to update the local block-level index table according to the message content (TableID, Key, Addr);
[0090] 5. Since the block version chain itself can be traversed through the next_block pointer, even if the index of a certain node is temporarily behind, the read operation can find all versions through the slow path to ensure correctness, so strong consistency is not necessary.
[0091] The system provided by the embodiment of the application adopts a garbage collection strategy based on transaction state awareness, including transaction state tracking and synchronization, an obsolete version memory reuse strategy, and a background thread garbage collection strategy.
[0092] As shown in FIG. 1, the transaction state tracking and synchronization process is executed according to the following steps: Figure 7
[0093] 1. Each computing node maintains a local active transaction table (LATT, a small top heap) to record the start timestamp of the active transaction of the node, and the top of the heap is Local-OAT.
[0094] 2. The nodes exchange their respective Local-OAT periodically (for example, after a certain period of time or a certain number of transactions).
[0095] 3. The coordinator calculates the minimum value of all node Local-OATs to obtain Global-OAT.
[0096] The obsolete version tuple memory reuse strategy is executed according to the following steps:
[0097] 1. Referring to step 1 of the embodiment 1, when searching for an available slot, if the version of a version tuple is less than Global-OAT, the stale flag of the version tuple is set to true.
[0098] 2. If a new version tuple needs to be written but the current block is full, a slot with a true stale flag is searched for to perform overwrite writing, instead of immediately allocating a new block.
[0099] As shown in FIG. 3, the background thread garbage collection mechanism is executed according to the following steps: Figure 8
[0100] 1. Check whether the size of the block-level index table exceeds a threshold value: during transaction execution, it is detected whether the size of the block-level index table exceeds a threshold value, and if the threshold value is exceeded, a wake-up signal is sent to the background garbage collection thread to notify it to start executing the garbage collection task.
[0101] 2. Wake up the background garbage collection thread: after receiving the garbage collection wake-up signal, the background thread starts the garbage collection thread and completes the initialization of related resources.
[0102] 3. Traverse the block-level index table to read the block-shaped version tuples of the target data: the garbage collection thread scans the block-level index table, and based on the block-level index information, batch processes reading of all block-shaped version tuples of the corresponding data and the Header structure of the version chain.
[0103] 4. Detect the state of the data, ensure that the recycling operation can be performed: detect whether the lock variable in the Header structure is 0 and whether read_cnt is 0, and only when the two states are 0, the metadata of the block version chain of the data can be recycled. The obsolete version data page of the incremental version data area can be directly recycled without additional state detection, because the obsolete version data will not be referenced by any transaction.
[0104] 5. Modify the lock state to GC_STATE, mark that the garbage collection is in progress: prevent other transactions from accessing the block version chain of the data during the recycling process, and ensure the data consistency of the transaction execution.
[0105] 6. Release the idle block version tuple, recycle the memory: migrate the version tuple that is not identified as obsolete to the front of the version chain, release the block version tuple space that can be recycled, clear the data in the block memory, and modify the corresponding bitmap bit in the computing node and update the block-level index table.
[0106] 7. End the garbage collection process: release the lock state, modify the lock to 0, allow transactions to read and write the data, and put the garbage collection thread to sleep.
[0107] The transaction processing flow of the system provided by the embodiment of the application is as shown in Figure 9 , mainly including three stages of a transaction execution stage, a verification stage and a submission stage.
[0108] 1. Transaction execution stage
[0109] The transaction execution stage mainly has two steps, and each step needs one network round trip.
[0110] 1) Read the block version tuple of the data, and find the target version locally. When the T0 transaction starts, the coordinator first obtains the start timestamp TStart from the timestamp service, and finds the block version tuple address of the data participating in the transaction in the block-level index table. Then, the coordinator performs Doorbell batch processing RDMA READ based on the block version tuple address, reads all block version tuples of the data from the memory node of the primary copy of the storage data. After obtaining the block version tuple information of the data, the coordinator verifies whether the index address in the block-level index table is wrong or not based on the chain index of the first block version tuple. If the verification fails, the index information of the block-level index table is corrected, and the block version tuple is read from the memory pool based on the chain index of the block version chain. If the verification succeeds, the coordinator selects the target version Vtarget locally, which is the maximum submission timestamp version in the version tuple whose transaction submission timestamp is less than TStart.
[0111] Under SR isolation level, the coordinator can abort the transaction early by version tuple information when it selects the target version of data locally, avoiding the waste of computing resources. If the coordinator finds a commit version greater than TStart when it selects the target version, it means that another transaction T1 completed commit after T0 started. In this case, the coordinator can abort transaction T0 early to ensure serialization. The reason is that even if the transaction is executed based on Vtarget selected by TStart, T0 will be aborted in the validation phase. Under SI isolation level, this early abort detection is not needed to be executed because the transaction under SI isolation level only sees the data version at its own start time, even if other transactions have committed new data, it will not be immediately visible.
[0112] After version selection, the coordinator uses batched RDMA READ to read the primary version data and the incremental data to build the target version value. For read-write data {A, B} in T0 transaction, the coordinator batches the RDMA CAS instruction and the RDMA READ instruction to achieve the lock on the write data first and then read the data, solving the write-write conflict of concurrent transactions. For read-only data {C}, the coordinator batches the RDMA FAA instruction and the RDMA READ instruction to achieve the read reference count plus one of the read-only data, so as to achieve the perception of the transaction state during garbage collection. In addition, version integrity verification is required at this link to check whether the primary version data and the head and tail check count in the version tuple are consistent to ensure data consistency during concurrent read and write.
[0113] 2. Validation phase
[0114] After successfully locking the remote block version tuple of the write data, the coordinator will obtain the commit timestamp TCommit from the timestamp service. If the transaction does not contain read-only data, the following operation can be skipped to reduce the delay, because all write data has been locked in the previous stage. If the transaction contains read-only data, such as the read set {C} shown in Figure 9 the coordinator needs to verify that the primary version of the read-only data has not changed during TStart to TCommit to ensure serialization. Under SI isolation level, the transaction processing flow of MiT-DM does not need the validation phase because SI is not visible to the commit of other transactions.
[0115] 3. Commit phase
[0116] When the verification phase is successful, the coordinator will update the data and write it to all remote memory nodes in parallel, and reduce the read reference count of read-only data by one. The coordinator uses batch processing RDMA WRITE and RDMACAS to write the update data into the memory pool while releasing the lock of the write data. The batch write data contains the Header structure of the block version chain structure, the version tuple of the write data, the main version data, and the incremental version data. In order to ensure the consistency of read-write concurrency, the batch write operation initiated by the computing node needs to be written in a certain order. First, the difference version data of the old version relative to the current commit version needs to be written, which is used as the undo-log required for transaction rollback, and then the new main version data, the Header structure and the version tuple are written in turn. In addition, according to the global minimum active transaction timestamp, the Stale attribute of the obsolete version tuple in the block version chain is changed to 1, which identifies that the version can be recycled.
[0117] In the commit phase, if the version tuple slot in the current block version chain is full, the first tuple with the Stale attribute of 1 is found in the order of the block version tuple chain, and the write of the new version tuple is completed in the manner of memory reuse. If there is no tuple in the current block version chain that is identified as an obsolete version, the memory allocator needs to be used to allocate the corresponding block version tuple for the current write.
[0118] Those skilled in the art will readily understand that the above description is only the preferred embodiment of the present application, and is not intended to limit the present application. Any modification, equivalent replacement and improvement made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A separated memory transaction system based on a hierarchical block version chain, comprising a plurality of remote memory nodes and a plurality of computing nodes, characterized in that: The remote memory node stores multiple versions of each data item in the form of fixed-size block units, forming version blocks corresponding to each data item. Adjacent version blocks are connected by a linked list to form a block version chain. The head of the block version chain includes the latest version block address, version number, and lock flag. The computing node includes a block-level index table and a block-level index notification queue; the block-level index table is used to store the addresses of the 2nd to nth version blocks of each data item; the block-level index table is updated through an asynchronous index update mechanism: when a transaction is committed to generate a new version block, the coordinator sends the metadata information of the new version block to the block-level index notification queue of each computing node, and each computing node asynchronously reads the information from the block-level index notification queue and updates the local block-level index table.
2. The system according to claim 1, wherein It also includes a hybrid garbage collection module, which is used to sequentially execute the node-level transaction status tracking and synchronization process, the outdated version memory reuse strategy, and the background batch recycling strategy; The execution node-level transaction status tracking and synchronization process includes: maintaining the timestamp of active transactions locally on each computing node, and periodically synchronizing with other nodes to obtain the global earliest active transaction timestamp to determine the range of outdated versions that can be safely recycled; The outdated version memory reuse strategy is: in the write transaction process, reclaimable outdated version slots are reused first; The background batch recycling strategy is: when the number of records in the block-level index table reaches a preset threshold, the background garbage collection thread combines the block-level index table to identify outdated versions of hot data, and reintegrates these versions into new version blocks in timestamp order, while releasing the corresponding old version blocks.
3. The system according to claim 2, wherein: The outdated version memory reuse strategy in the write transaction process includes: Check the existing block version chain for reusable outdated version slots. If there is an available slot, write the new version data to the slot; otherwise, allocate a new version block according to the on-demand expansion strategy and write the new version data to the new version block.
4. The system according to claim 1, wherein: A checksum counter is set at the start and end positions of each version block; the transaction consistency guarantee mechanism adopted by the system includes: when reading a version, the read transaction checks whether the checksum counts at the beginning and end of the version are consistent. If so, the version data is considered complete; otherwise, the version is considered to be in an incomplete write state and is skipped.
5. The system according to claim 1, wherein: The computing node further includes a local bitmap management unit for managing the allocation status of pre-allocated block version units.
6. The system according to claim 1, wherein: The address of the first version block of each data item is obtained through hash index calculation.
7. The system according to claim 1, wherein: The system adopts a transaction execution protocol based on RDMA unilateral primitives.
Citation Information
Cited By
Big data knowledge construction and management system
CN121901189A