A memory dynamic resource allocation and scheduling method for big data system integration

CN122653849BActive Publication Date: 2026-09-25CHONGQING SHENGYAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611107262.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-24
Publication Date
2026-09-25
Estimated Expiration
2046-07-24

AI Technical Summary

Technical Problem

由于磁盘存储的读写速度远低于内存,大量数据的磁盘换入操作会引发较长的等待时间,从而降低计算任务的整体执行效率

Benefits of technology

[0053]1、本发明通过跟踪分布式计算作业中数据块的下游访问状态,在确定目标数据完成当前阶段读取后,选定远端计算节点并获取远端已注册内存区域对应的远端内存访问元数据,利用RDMA技术将本地物理页帧中的数据转移至远端计算节点的远端已注册内存区域中。系统在内核态构建远端内存索引表,生成整数索引,并将整数索引与远端内存访问元数据建立绑定关系;在确认远端写入完成后,通过清除对应本地页表项的存在位,并在处理器架构允许软件使用的位域内记录远端重定向标记和整数索引,随后将承载原数据的物理页帧释放回操作系统。当上层应用因容错重算、推测执行回滚或缓存复用等机制再次访问该部分逻辑地址时,由硬件地址转换失败触发缺页异常。操作系统在缺页异常处理上下文中识别重定向标记,通过整数索引查询维护在内核态的远端内存索引表,获取远端寻址信息和权限参数,执行RDMA读操作将数据自动恢复至新分配的本地物理页帧。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653849B_ABST
    Figure CN122653849B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of distributed computing and memory management, and discloses a memory dynamic resource allocation and scheduling method for big data system integration, which comprises the following steps: analyzing a distributed computing job to establish a mapping of a data block logical address to a local page table entry; generating an event notification when a downstream operator completes reading of the data block; writing data in a local physical page frame into a registered memory area of a remote computing node through RDMA, and constructing a remote memory index table in a kernel state; writing a remote redirection mark and an integer index, and releasing the corresponding local physical page frame; when a target page table entry triggers a page fault exception and contains the redirection mark, the system acquires a remote data address according to the integer index by querying the index table, pulls the remote data to update the page table entry, and restores process execution. The application effectively releases local memory resources, realizes on-demand recovery of remote data through a page fault exception mechanism of an operating system, and reduces system recalculation overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed computing and memory management technology, specifically a method for dynamic memory resource allocation and scheduling for big data system integration. Background Technology

[0002] In distributed big data computing systems, the computing engine generates numerous intermediate data blocks during execution. These intermediate data blocks typically reside in the local physical memory of the computing nodes, for downstream operators to read or for subsequent fault-tolerant recalculation. As the job execution progresses and the data flow graph expands, the continuous accumulation of intermediate data blocks consumes local memory space, gradually reducing the available memory of the current computing node.

[0003] When local memory is insufficient, existing memory management mechanisms typically swap some data out to the local disk. However, when distributed computing jobs face node failure recovery or speculative execution rollback, the computing engine needs to re-access this swapped-out data. Since disk storage read / write speeds are much slower than memory, large-scale disk swapping operations can cause lengthy waiting times, thereby reducing the overall execution efficiency of the computing task.

[0004] To reduce reliance on local disk read / write operations, some solutions choose to move intermediate data to independent external storage nodes. However, this approach typically requires modifications to the upper-layer computing engine code, demanding that applications explicitly call network communication interfaces to initiate data read / write requests. This not only alters the application's original direct access mode based on virtual memory addresses but also requires the application layer to maintain network addressing and state distribution information for the data, increasing the development cost of system integration. Therefore, how to move data no longer used in the current computing phase out of local memory while maintaining the upper-layer application's virtual memory access logic, and automatically restore the data through underlying mechanisms when needed later, is a problem that current memory scheduling and management needs to solve. Summary of the Invention

[0005] The technical problem solved by this invention is that intermediate data blocks generated during the execution of existing distributed computing jobs occupy the memory resources of local computing nodes for a long time, resulting in limited overall system available memory. Directly cleaning up these data blocks will trigger a full recalculation or a large number of disk swap-in operations when fault tolerance recalculation or cache reuse occurs, thereby generating high system overhead and affecting the execution efficiency of computing tasks.

[0006] To address the above problems, the present invention provides the following technical solution:

[0007] The first aspect of this invention provides a method for dynamic memory resource allocation and scheduling for big data system integration, comprising:

[0008] Receive distributed computing jobs, parse the physical execution plan, and establish a mapping relationship between the logical address space of data blocks and local page table entries;

[0009] Monitor the execution status of the distributed computing job, and generate a completion event notification when it is determined that the downstream operator has completed the reading and processing of the data block;

[0010] In response to the completion event notification, a remote computing node is selected and the remote memory access metadata corresponding to the remote registered memory region in the remote computing node is obtained. The data in the local physical page frame is written into the remote registered memory region, and a remote memory index table is built in the kernel mode. An integer index is generated and a binding relationship is established between the integer index and the remote memory access metadata.

[0011] After confirming that the data in the local physical page frame has been written to the remote registered memory region, modify the target local page table entry, write the remote redirection flag and the integer index, and release the corresponding local physical page frame to the reclaimable memory pool.

[0012] When receiving and processing a page fault triggered by accessing the target logical address corresponding to the target local page table entry and the target local page table entry being invalid, and confirming that the target local page table entry contains the remote redirection flag, the remote memory index table is queried according to the integer index to obtain the data address in the remote registered memory region, the remote data is pulled through an RDMA read operation, the target local page table entry is updated, and process execution is resumed.

[0013] Furthermore, the process of receiving distributed computing jobs, parsing physical execution plans, and establishing mapping relationships between data block logical address spaces and local page table entries specifically includes:

[0014] The distributed computing job is subjected to directed acyclic graph structure parsing to extract the operator node set and data dependency edge set;

[0015] Determine the intermediate data block output by the operator node during execution, and allocate a contiguous logical address space corresponding to the intermediate data block in the virtual memory space;

[0016] In local memory, allocate corresponding physical page frames for the contiguous logical address space, map the contiguous logical address space to the allocated physical page frames, and form a set of local page table entries corresponding to the contiguous logical address space;

[0017] A mapping registry is maintained in memory, which uses data block identifiers as indexes and records the binding relationships between each intermediate data block and the local page table entry set.

[0018] Furthermore, the monitoring of the distributed computing job execution status, and the generation of a completion event notification when it is determined that the downstream operator has completed the reading and processing of the data block, specifically includes:

[0019] Based on the data dependency edges in the directed acyclic graph structure obtained by parsing the physical execution plan, determine the set of direct downstream operators for the target data block;

[0020] Register a status callback interface in the data input / output scheduling component, receive the read completion confirmation flag of the downstream operator through the status callback interface, and update the read completion status of the downstream operator for the target data block.

[0021] When all downstream operators in the set of direct downstream operators submit read completion confirmation flags, the event triggering condition is determined to be met, and a completion event notification for the target data block is generated.

[0022] Furthermore, the step of responding to the completion event notification, selecting a remote computing node and obtaining the remote memory access metadata corresponding to the remote registered memory region in the remote computing node, and writing the data in the local physical page frame into the remote registered memory region specifically includes:

[0023] In response to the completion event notification, the cluster resource status table is queried, and a computing node with free memory capacity not less than the target data length is selected as the remote computing node, wherein the target data length is the data length to be written to the remote registered memory area;

[0024] Send a remote memory request to the remote computing node to allocate a remote memory region on the remote computing node, perform access registration on the remote memory region to form the remote registered memory region, and return the remote memory access metadata.

[0025] A remote direct memory access communication connection is established between the first network card of the first computing node carrying the local physical page frame and the remote computing node. An RDMA write command is issued to write the data in the local physical page frame into the remote registered memory area.

[0026] Furthermore, the step of constructing a remote memory index table in kernel mode, generating an integer index, and binding the integer index with the remote memory access metadata specifically includes:

[0027] The remote memory index table is constructed in the kernel address space, and the integer index is generated for the remote memory access metadata.

[0028] The binding relationship between the integer index and the remote memory access metadata is written into and maintained in the remote memory index table, wherein the remote memory access metadata includes at least the remote network address, the remote registration key, the remote base address, the offset of the data in the remote registered memory area, and the data length.

[0029] Furthermore, the modification of the target local page table entry, writing the remote redirection flag and the integer index, specifically includes:

[0030] After setting the local physical page frame corresponding to the target local page table entry to the unloading ready state and confirming that no new writes have occurred during the unloading ready state, the presence bit in the target local page table entry is cleared to indicate that the physical page frame corresponding to the logical address is not in a valid mapping state.

[0031] In the target local page table entry, the processor architecture allows the software to divide the bit field into a first bit field and a second bit field, write the remote redirection flag in the first bit field, and write the integer index in the second bit field.

[0032] Furthermore, the step of releasing the corresponding local physical page frame to the reclaimable memory pool specifically includes:

[0033] After the target local page table entry is modified, an address translation backup buffer refresh operation is performed on the logical address range where the target logical address corresponding to the target local page table entry is located.

[0034] The kernel physical page release interface is invoked to clear the page lock state of the local physical page frame, reduce or clear the corresponding reference count, and return the local physical page frame to the operating system's reclaimable memory pool.

[0035] Furthermore, when receiving and processing a page fault triggered by accessing the target logical address corresponding to the target local page table entry and the target local page table entry being invalid, and confirming that the target local page table entry contains the remote redirection flag, querying the remote memory index table according to the integer index to obtain the data address in the remote registered memory region specifically includes:

[0036] Obtain the target logical address that triggered this page fault, and locate the target local page table entry that triggered the page fault in the page table structure;

[0037] Extract the integer index from the bit fields that the processor architecture allows the software to use in the target local page table entry, query the remote memory index table in the kernel address space, and obtain the remote memory access metadata corresponding to the integer index;

[0038] Based on the offset of the target logical address relative to the starting position of the data block logical address space, calculate the remote read start address and read length corresponding to this page fault.

[0039] Furthermore, the step of retrieving remote data via RDMA read operations specifically includes:

[0040] Submit the remote data recovery request to the page fault recovery execution context for processing, and suspend the application thread that triggered this page fault.

[0041] Allocate free physical page frames from the reclaimable memory pool in local memory;

[0042] The remote memory access metadata is used as a parameter to issue an RDMA read instruction, which reads data within the corresponding range from the remote registered memory region of the remote computing node, and writes the read data into the newly requested free physical page frame.

[0043] Furthermore, updating the target local page table entry and resuming process execution specifically includes:

[0044] After the remote data reading is completed and passes the verification, the physical page frame address field of the target local page table entry is set to the newly requested free physical page frame, the presence bit is set to valid again, and the original remote redirection flag is cleared.

[0045] An address translation back buffer refresh operation is performed on the target logical address to release the suspended application thread, so that the resumed process can re-execute the interrupted memory read instruction.

[0046] A second aspect of the present invention provides a memory dynamic resource allocation and scheduling system for big data system integration, comprising:

[0047] The network communication connects a first computing node and a second computing node. The first computing node is equipped with a first network interface card (NIC), and the second computing node is equipped with a second NIC. The second computing node is equipped with a remote memory service unit, which is used to respond to remote memory request requests, allocate remote memory regions, perform access registration to form remote registered memory regions, and return remote memory access metadata.

[0048] The operating system module runs on each computing node and includes a page table management unit that maintains page table mapping relationships and a page fault handling unit that receives and processes page faults.

[0049] The computing engine module includes a status tracking unit for monitoring the completion status of downstream operator data reading and generating completion event notifications;

[0050] The scheduling management module includes an event receiving unit and an unloading control unit. It is used to respond to completion event notifications, write data in the local physical page frame to the remote registered memory area of ​​the remote computing node through the first network card of the first computing node, and notify the page table management unit to modify the target local page table entry to the remote redirection state.

[0051] The operating system module is further configured to construct a remote memory index table in kernel mode, generate an integer index, and establish a binding relationship between the integer index and the remote memory access metadata. The page table management unit is configured to, after confirming that the data in the local physical page frame has been written to the remote registered memory region, write the remote redirection flag and the integer index to the target local page table entry, and release the corresponding local physical page frame to the reclaimable memory pool. The page fault handling unit is configured to, when receiving and processing a page fault triggered by accessing the target logical address corresponding to the target local page table entry and the target local page table entry being invalid, query the remote memory index table according to the integer index, obtain the data address in the remote registered memory region, and pull the remote data through an RDMA read operation.

[0052] This invention provides a method for dynamic memory resource allocation and scheduling for big data system integration. It has the following advantages:

[0053] 1. This invention tracks the downstream access status of data blocks in a distributed computing job. After determining that the target data has completed reading in the current stage, it selects a remote computing node and obtains the remote memory access metadata corresponding to the remote registered memory region. Using RDMA technology, it transfers the data in the local physical page frame to the remote registered memory region of the remote computing node. The system constructs a remote memory index table in kernel mode, generates an integer index, and establishes a binding relationship between the integer index and the remote memory access metadata. After confirming that the remote write is complete, it clears the existence bit of the corresponding local page table entry and records the remote redirection flag and integer index in the bit field allowed by the processor architecture. Then, it releases the physical page frame carrying the original data back to the operating system. When the upper-layer application accesses this logical address again due to fault-tolerant recalculation, speculative execution rollback, or cache reuse mechanisms, a page fault is triggered by hardware address translation failure. The operating system identifies the redirection flag in the page fault handling context, queries the remote memory index table maintained in kernel mode using the integer index, obtains the remote addressing information and permission parameters, and performs an RDMA read operation to automatically restore the data to the newly allocated local physical page frame.

[0054] 2. The present invention releases local computing node memory space while maintaining the validity of application layer logical addresses. It adopts an indirect index structure of kernel tables to establish a correspondence between the integer index recorded in the page table entries and the remote memory access metadata in the remote memory index table, solving the problem that the bit width of a single page table entry cannot directly store complete remote network and permission metadata. At the same time, the entire data migration and on-demand recovery process remains transparent to the user-space computing engine, eliminating the need for upper-layer applications to re-initiate read requests and effectively reducing system fault tolerance overhead and improving memory resource utilization efficiency. Attached Figure Description

[0055] Figure 1 This is a schematic diagram of a memory dynamic resource allocation and scheduling system architecture for big data system integration according to an embodiment of the present invention;

[0056] Figure 2 This is a flowchart of a memory dynamic resource allocation and scheduling method for big data system integration according to an embodiment of the present invention;

[0057] Figure 3 This is a timing diagram of cross-level interaction during the data block logical address allocation and page table mapping stage according to an embodiment of the present invention;

[0058] Figure 4 This is a logic diagram of the state transition and unloading condition determination of the target data block during the downstream operator reading process according to an embodiment of the present invention;

[0059] Figure 5 This is a multi-component collaborative swimlane diagram for remote memory offloading and index table construction based on RDMA, according to an embodiment of the present invention.

[0060] Figure 6 This is a comparison diagram of the data structure of a target local page table entry before and after performing a remote redirection operation, according to an embodiment of the present invention.

[0061] Figure 7 This is a flowchart illustrating the determination and execution process for transparent remote data recovery based on a page fault interception mechanism, according to an embodiment of the present invention.

[0062] Figure 8 This is a sequence diagram showing the global lifecycle evolution of an intermediate data block from generation and unloading to on-demand recovery, according to an embodiment of the present invention.

[0063] Among them, 100 is the first computing node; 101 is the first network interface card (NIC); 110 is the second computing node; 111 is the second NIC; 200 is the operating system module; 210 is the page table management unit; 220 is the page fault handling unit; 300 is the computing engine module; 310 is the status tracking unit; 400 is the scheduling management module; 410 is the event receiving unit; and 420 is the unloading control unit. Detailed Implementation

[0064] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0065] See attached document Figure 1 The present invention provides a memory dynamic resource allocation and scheduling system for big data system integration, which may include: a first computing node 100, a second computing node 110, an operating system module 200, a computing engine module 300, and a scheduling management module 400.

[0066] The first computing node 100 and the second computing node 110 are connected via a network communication link; the first computing node 100 is equipped with a first network interface card 101, and the second computing node 110 is equipped with a second network interface card 111. In this embodiment, the first computing node 100 is a local computing node that carries distributed computing jobs and stores local physical page frames, and the second computing node 110 is a remote computing node that provides remotely registered memory regions.

[0067] The second computing node 110 deploys a remote memory service unit. This unit responds to remote memory request requests, allocates a remote memory region from the memory resources of the second computing node 110, performs access registration on the remote memory region to form a remote registered memory region, and returns remote memory access metadata. The remote memory access metadata includes at least the remote node identifier, remote network address, remote registration key, remote base address, and data length, and may also include at least one of the following: remote memory region identifier, offset within the remote registered memory region, access permission information, and verification information.

[0068] The first network interface card 101 and the second network interface card 111 support cross-node memory data transfer operations based on Remote Direct Memory Access (RDMA). During the data transfer phase, the central processing unit of the first computing node 100 does not need to perform memory copy operations on the data in the local physical page frame. The central processing unit is used to issue RDMA transfer control instructions, complete event processing, and perform page fault recovery scheduling operations.

[0069] The operating system module 200 runs on the first computing node 100 and the second computing node 110. The operating system module 200 includes a page table management unit 210 and a page fault handling unit 220. The page table management unit 210 is used to maintain the page table mapping relationship between the logical address space and the physical page frame. The page fault handling unit 220 is used to receive and process page faults triggered by invalid target local page table entries, and to perform remote data recovery processing when it is confirmed that there is a remote redirection flag in the target local page table entry.

[0070] The computing engine module 300 runs in user space above the operating system module 200; the computing engine module 300 is used to receive and parse distributed computing jobs from external input. The computing engine module 300 has an embedded state tracking unit 310. The state tracking unit 310 is used to monitor the completion status of downstream operators reading data blocks in the distributed computing job.

[0071] The scheduling management module 400 runs independently in user space; the scheduling management module 400 includes an event receiving unit 410 and an unloading control unit 420. The scheduling management module 400 connects to the operating system module 200 via a system call interface and to the computing engine module 300 via an inter-process communication mechanism.

[0072] After determining that the target data block meets the unloading triggering conditions, the status tracking unit 310 in the computing engine module 300 generates a completion event notification. The event receiving unit 410 in the scheduling management module 400 receives the completion event notification and transmits it to the unloading control unit 420.

[0073] Upon receiving a completion event notification, the unloading control unit 420 determines the data in the local physical page frame to be unloaded based on the logical address range corresponding to the target data block and the local page table entry set. It then writes the data from the local physical page frame to the remote registered memory region of the second computing node 110 via the first network interface card 101. After writing, the unloading control unit 420 notifies the page table management unit 210 to modify the page table mapping relationship, so that the target local page table entry changes from the local physical page frame mapping state to the remote redirection state.

[0074] Page fault handling unit 220 receives and processes page faults triggered by the computing engine module 300 accessing a target logical address. When the target local page table entry that triggered the page fault contains a remote redirection flag, page fault handling unit 220 locates the data address in the remote registered memory region through the remote memory index table in operating system module 200, and performs an RDMA read operation through the first network card 101 to pull the remote data into the local physical page frame. After the remote data is recovered, page table management unit 210 restores the valid mapping state of the corresponding local page table entry, enabling the computing engine module 300 to continue accessing the recovered data.

[0075] See attached document Figure 2 Based on the aforementioned dynamic memory resource allocation and scheduling system, this embodiment of the invention also provides a dynamic memory resource allocation and scheduling method for big data system integration. The method includes the following steps:

[0076] S10 receives distributed computing jobs, parses the physical execution plan, and establishes a mapping relationship between the logical address space of data blocks and local page table entries;

[0077] S20, monitors the execution status of distributed computing jobs, and generates a completion event notification when it is determined that the downstream operator has completed the reading and processing of the data block;

[0078] S30, responding to the completion event notification, writes the data in the local physical page frame to the pre-allocated remote registered memory region in the remote computing node, and builds the remote memory index table in kernel mode;

[0079] S40, Modify the target local page table entry, write the remote redirection flag and integer index, and release the corresponding local physical page frame to the reclaimable memory pool;

[0080] S50: Upon receiving and processing a page fault exception triggered by a target local page table entry, and confirming that the target local page table entry contains a remote redirection flag, the system queries the remote memory index table based on the integer index to obtain the data address in the remote registered memory region, retrieves the remote data through an RDMA read operation, updates the target local page table entry, and resumes process execution.

[0081] The following sections describe steps S10 to S50, including job parsing, reading completion status determination, remote memory writing, page table entry redirection, and page fault recovery.

[0082] See attached document Figure 3 In some specific implementations, step S10 is used to complete the physical execution plan parsing, data block logical address space allocation, and local page table entry mapping establishment before the distributed computing job enters the execution phase. Specifically, it includes the following steps:

[0083] S110, the computing engine module 300 receives the distributed computing job from external input and performs directed acyclic graph (DAG) structure parsing on the distributed computing job. The DAG structure is used to represent the data dependencies between operator nodes, enabling the computing engine module 300 to determine the execution order of each operator node and the location of intermediate data blocks.

[0084] Specifically, the computing engine module 300 extracts the operator node set and the data dependency edge set from the physical execution plan of the distributed computing job, and constructs a directed acyclic graph structure accordingly. The operator node set includes multiple computing operator nodes; the data dependency edge set is used to represent the predecessor and successor dependencies between different operator nodes, as well as the correspondence between the output data of the upstream operator node and the reading of the downstream operator node.

[0085] For parsing the physical execution plan of distributed computing jobs, the computing engine module 300 can call the existing job parsing component of the distributed computing framework to obtain operator nodes, data dependency edges, and data input / output relationships. This job parsing component can be an execution plan parsing component from Spark, Flink, or other distributed computing frameworks.

[0086] S120, the computing engine module 300 determines the intermediate data blocks output by the operator node during execution and allocates logical address space for the intermediate data blocks.

[0087] For any operator node capable of generating intermediate data, the computation engine module 300 determines one or more intermediate data blocks output by that operator node. For each intermediate data block, the computation engine module 300 allocates a contiguous logical address space corresponding to that intermediate data block in the virtual memory space. The contiguous logical address space is used to carry the access entry of the intermediate data block in the local node and is used to subsequently establish the mapping relationship between logical addresses and local page table entries.

[0088] In some embodiments, different intermediate data blocks correspond to different logical address ranges, thereby distinguishing intermediate data blocks at the virtual address level. A logical address range may include the starting position of the logical address, the length of the logical address, and the corresponding page size.

[0089] S130, the operating system module 200 responds to the logical address space allocation request and establishes a mapping relationship between the logical address space and local page table entries.

[0090] The computing engine module 300 triggers the page table management unit 210 in the operating system module 200 through a system call. The page table management unit 210 allocates corresponding physical page frames for the logical address space in local memory and maps the logical address space to the allocated physical page frames, forming a set of local page table entries corresponding to the logical address space.

[0091] Local page table entries are used to record the correspondence between logical memory pages and local physical page frames, as well as access permission information. The specific page table hierarchy can be determined according to the processor architecture, such as using a single-level page table or a multi-level page table structure; this invention does not limit this.

[0092] After establishing the page table mapping relationship, the computing engine module 300 maintains a mapping registry in memory. The mapping registry uses data block identifiers as indexes to record the binding relationship between each intermediate data block and the local page table entry set. In some embodiments, the mapping registry also records the logical address start position, logical address length, page size, logical page sequence number range, and data block length corresponding to the data block.

[0093] When a data block spans multiple local page table entries, the system can calculate the page offset or byte offset within the data block based on the specific logical address that triggered the access, combined with the starting position of the logical address and the page size recorded in the mapping registry. Therefore, when tracking the state of a specific data block, the computing engine module 300 can locate the set of local page table entries carrying that data block based on the mapping registry; the operating system module 200 can also determine the range of local physical page frames that need to be processed when subsequently performing remote memory unloading or page fault recovery.

[0094] See attached document Figure 4 In some specific implementations, step S20 is used to monitor the completion status of downstream operators reading the target data block based on the data dependencies in the physical execution plan, and generate a completion event notification when the unload trigger condition is met. Specifically, this includes the following steps:

[0095] S210, the state tracking unit 310 in the computing engine module 300 determines the set of direct downstream operators of the target data block based on the data dependency edges in the directed acyclic graph structure.

[0096] During the execution of a distributed computing job, for any target data block output by an operator node that generates intermediate data, the state tracking unit 310 searches for downstream operator nodes with that operator node as the upstream node based on the data dependency edges recorded in the directed acyclic graph structure, and further determines whether the target data block belongs to the input data of the corresponding downstream operator node.

[0097] Therefore, the state tracking unit 310 determines the set of direct downstream operators corresponding to the target data block. The set of direct downstream operators includes one or more downstream operators that need to read the target data block in the current physical execution plan. Through the determination of the above data dependencies, the computing engine module 300 can obtain the scope of the target data block to be read in the current distributed computing job.

[0098] S220, the state tracking unit 310 registers a state callback interface in the data input / output scheduling component of the computing engine module 300 to update the read completion status of the downstream operator on the target data block.

[0099] In some embodiments, when any downstream operator in the direct downstream operator set initiates a read request for the target data block, the computing engine module 300 records the start state of the read process through a status callback interface. After the downstream operator completes the read processing of the target data block, it submits a read completion confirmation flag to the computing engine module 300. The status tracking unit 310 updates the read status corresponding to the downstream operator to "completed" based on the read completion confirmation flag.

[0100] For downstream operators that have not yet completed the reading process, the state tracking unit 310 keeps their corresponding reading status as incomplete. Thus, the reading status of each downstream operator corresponding to the target data block can be continuously updated during the execution of the distributed computing job.

[0101] For the registration of underlying data input / output scheduling and status callback interfaces, existing data flow scheduling interfaces or task status callback interfaces in the distributed computing framework can be called. For example, the computing engine module 300 can register status callback interfaces in the data reading operator, input stream processing component, or task completion callback component to receive the start status and completion confirmation flag of the target data block.

[0102] S230, the status tracking unit 310 determines whether the event triggering condition is met based on the reading completion status of each downstream operator, and sends a completion event notification to the scheduling management module 400 when the event triggering condition is met.

[0103] Specifically, during the execution of the distributed computing job, the state tracking unit 310 continuously checks the set of direct downstream operators corresponding to the target data block. If all downstream operators in the set have submitted read completion confirmation flags, the state tracking unit 310 determines that the target data block has completed downstream read processing in the current normal computing process and confirms that the event triggering conditions are met.

[0104] In some specific implementations, if the target data block does not have a direct downstream operator, in order to avoid the situation where the event triggering conditions cannot be met, the state tracking unit 310 further determines whether the target data block belongs to the final output data of the distributed computing job.

[0105] If the target data block does not belong to the final output data, or if the target data block belongs to the final output data but has already completed external result submission, persistent writing, or result confirmation operations, then the state tracking unit 310 determines that the target data block meets the event triggering conditions.

[0106] When the event trigger condition is met, it means that the target data block will no longer be read by downstream operators in the current normal calculation process. It should be noted that this state does not mean that the target data block has been deleted or permanently invalidated; the system will create and retain a remote data copy of the target data block and its corresponding index information in the subsequent unloading process to support data recovery in scenarios such as fault-tolerant recalculation, speculative execution rollback, or cache reuse.

[0107] The status tracking unit 310 generates a completion event notification for the target data block and sends the completion event notification to the event receiving unit 410 in the scheduling management module 400 through the cross-process communication channel to trigger the subsequent local physical page frame data unloading process.

[0108] See attached document Figure 5 In some specific implementations, step S30 is used to determine the remote computing node after the target data block meets the unload trigger condition, complete the writing of data in the local physical page frame to the remote registered memory region, and establish a remote memory index table in the operating system module 200 for subsequent data recovery. Specifically, it includes the following steps:

[0109] S310, the scheduling management module 400 responds to the completion event notification, queries the cluster resource status table to select the remote computing node, and calls the first network card 101 to write the data in the local physical page frame into the remote registered memory area in the selected remote computing node.

[0110] After receiving a completion event notification for the target data block, the event receiving unit 410 in the scheduling management module 400 sends the notification to the unloading control unit 420. The unloading control unit 420 determines the logical address range, local page table entry set, and target data length corresponding to the target data block based on the mapping registry. The target data length is the length of the data to be written to the remote registered memory region. The target data length can be determined based on the logical address length, number of pages, page size, or actual byte length corresponding to the target data block. When the target data block spans multiple local physical page frames, the target data length is the total length of the data to be written to the remote registered memory region in those multiple local physical page frames. The unloading control unit 420 requests the page table management unit 210 to set the target logical address range to an unloading preparation state via a system call.

[0111] Page table management unit 210 sets an unloading preparation state for the local physical page frame corresponding to the target local page table entry. The unloading preparation state indicates that the corresponding local physical page frame is undergoing remote unloading processing. During this state, page table management unit 210 prevents the local physical page frame from being newly written to, migrated, or prematurely reclaimed by at least one of the following methods: page locking, reference count increment, read-only setting, or inaccessible setting. In the following text, the local physical page frame corresponding to the target local page table entry and carrying the target data block is also referred to as the target physical page frame. Simultaneously, page table management unit 210 performs page table permission updates and address translation backstop refresh operations on the corresponding logical address range to prevent the target physical page frame from being written to, migrated, or prematurely reclaimed during remote data transmission.

[0112] In some specific implementations, the cluster manager maintains the memory resource status across the cluster and forms a cluster resource status table. The unloading control unit 420 queries this cluster resource status table to obtain the current free memory capacity of each compute node, and selects a compute node with a free memory capacity not less than the target data length as a remote compute node. In this embodiment, the second compute node 110 serves as a remote compute node, used to provide remote registered memory regions and return remote memory access metadata.

[0113] If there are no compute nodes in the current cluster that meet the free memory capacity requirements, the system can temporarily keep the target data block in local memory or use a local disk swap-out mechanism to save the corresponding data in order to avoid data loss.

[0114] After selecting the second computing node 110, the unloading control unit 420 sends a remote memory request to the remote memory service unit in the second computing node 110. The remote memory service unit allocates a remote memory region in the free memory pool or pre-registered memory pool of the second computing node 110, and performs access registration on the remote memory region through the second network card 111 or the driver of the second network card 111, generating a remote access handle, a remote registration key, a remote base address, a data length, and a remote memory region identifier, and returns the above information to the unloading control unit 420.

[0115] The unloading control unit 420 sends an RDMA write command to the first network card 101 through the network card driver interface. An RDMA communication connection is established between the first network card 101 and the second network card 111. During the data transmission phase, the central processing unit of the first computing node 100 does not need to perform a memory copy operation on the data in the local physical page frame, and writes the data in the target local physical page frame into the remote registered memory area of ​​the second computing node 110.

[0116] For the connection establishment, authorization verification, data fragmentation, and transmission confirmation processes of the RDMA transmission protocol, those skilled in the art can use existing InfiniBand or RoCE technologies.

[0117] S320, after confirming that the data transmission is completed, the scheduling and management module 400 obtains and records the remote memory access metadata that carries the data in the second computing node 110.

[0118] In some embodiments, in order to locate and read the unloaded data during subsequent fault-tolerant recalculation, speculative execution rollback, or cache reuse, the unloading control unit 420 collects the network addressing parameters and memory access parameters corresponding to the remote registered memory region and records them as remote memory access metadata.

[0119] Specifically, the remote memory access metadata includes the network address of the second computing node 110, the remote registration key, the remote base address, the remote memory region identifier, the offset of the data within the remote registered memory region, and the data length. In some embodiments, the remote memory access metadata may also include the remote node identifier, page size, logical address start position, page sequence number offset, version number, and verification information.

[0120] Remote memory access metadata provides the addressing and permission verification information required for cross-node reading of a target remote registered memory region. Specifically, the network address identifies the remote compute node, the remote registration key verifies remote memory access permissions, the remote base address and offset within the remote registered memory region determine the data's location within that region, and the data length determines the read range. For data blocks spanning multiple local page table entries, the page size and logical address start position are used to calculate the corresponding remote read range based on the logical address that triggered the page fault.

[0121] S330, the operating system module 200 constructs a remote memory index table in the kernel address space, stores remote memory access metadata in the remote memory index table, and generates the corresponding integer index.

[0122] In common processor architectures, the bit width of a single page table entry is limited, making it typically impossible to directly write complete remote memory access metadata. Therefore, the scheduling management module 400 triggers the operating system module 200 via a system call, which then allocates a storage region in the kernel address space to maintain the remote memory index table.

[0123] A remote memory index table includes at least an index field, a logical address range field, a remote node identifier field, a remote registration key field, a remote base address field, a data length field, and a validity status field. Depending on the actual implementation requirements, the remote memory index table may also include at least one of the following: a data block identifier field, a remote memory region identifier field, a page offset field, an access permission field, a version number field, and a checksum field.

[0124] In some embodiments, when the system supports backup remote data copy recovery, the remote memory index table further includes at least one of the following: backup remote node field, backup remote registration key field, backup remote base address field, backup remote memory region identifier field, and backup copy validity status field, so as to select a backup remote node to perform data recovery when the primary remote node is unreachable.

[0125] Operating system module 200 generates an integer index for the remote memory access metadata newly written to the remote memory index table. The integer index is unique within the remote memory index table or within the first compute node 100, and is bound to the corresponding remote memory access metadata.

[0126] When a data block corresponds to multiple local page table entries, and the data block is stored contiguously in a remote registered memory region, the operating system module 200 can generate the same integer index for the data block and record the logical address start position, remote base address, page size, and data length of the data block in the remote memory index table. When any subsequent logical page triggers a page fault, the operating system module 200 calculates the corresponding remote read address and read length based on the offset of the logical address that triggered the page fault relative to the logical address start position.

[0127] When the remote registered memory region corresponding to a data block is not contiguous, the operating system module 200 can generate an independent integer index for each logical page or each contiguous page segment, and record the corresponding remote memory access metadata.

[0128] In some embodiments, the range of integer index values ​​is determined by the bit field width allowed by the processor architecture in the local page table entry. For example, when the bit field width is W bits, the range of integer index values ​​can be set to an integer range that can be represented by W bits, to ensure that the integer index can be written to the target local page table entry and can uniquely correspond to a set of remote memory access metadata in the remote memory index table or in the first compute node 100.

[0129] Operating system module 200 writes and maintains the binding relationship between the integer index and the remote memory access metadata in the remote memory index table. Through this indirect index structure, the target local page table only needs to record the integer index. During subsequent page fault recovery, the corresponding remote registered memory region, remote read address, and read length can be determined based on the remote memory index table, thereby avoiding the problem of insufficient page table entry storage space caused by directly writing the complete remote memory access metadata into the page table entry.

[0130] Meanwhile, for data blocks spanning multiple local physical page frames, the operating system module 200 can combine the logical address offset that triggered the page fault to determine the specific remote read location from the remote memory index table to support subsequent data recovery.

[0131] See attached document Figure 6 In some specific implementations, step S40 is used to modify the target local page table entry to a remote redirection state after the data in the local physical page frame has been written remotely, and to release the corresponding local physical page frame after the page table mapping update is completed. Specifically, it includes the following steps:

[0132] S410, the scheduling management module 400 triggers the page table management unit 210 in the operating system module 200 through a system call to perform a validity modification operation on the target local page table entry.

[0133] After confirming that the data in the local physical page frame corresponding to the target data block has been written to the remote registered memory region, the page table management unit 210 further confirms that no new writes have occurred to the target physical page frame during the unloading preparation state, or confirms that the dirty page state of the target physical page frame has been synchronized. Once the above conditions are met, the page table management unit 210 locates the corresponding target local page table entry in the page table structure based on the logical address of the target data block.

[0134] Page table management unit 210 clears the presence bit in the target local page table entry. The presence bit is used to indicate whether the physical page frame corresponding to the logical address is in a valid mapping state; for example, in the x86 architecture, the presence bit can be the P bit in the page table entry. After the presence bit is cleared, the processor memory management unit can no longer obtain a valid local physical address through this page table entry when performing address translation for this logical address. Therefore, subsequent memory accesses to this logical address will trigger a page fault, and the operating system module 200 will enter the corresponding data recovery process.

[0135] S420, Page Table Management Unit 210 uses the bit fields in the target local page table entry that the processor architecture allows the software to use to write the remote redirection flag and the integer index.

[0136] After the presence bit of the target local page table entry is cleared, the page table management unit 210 writes redirection information into the bit fields that are allowed to be used by the software in the processor architecture. The bit fields that are allowed to be used by the software in the processor architecture refer to the bit fields that can be used by the software to record status information when the page table entry is invalid; the page table management unit 210 does not set bits that are reserved for hardware and prohibited from being set by the processor architecture.

[0137] Specifically, the page table management unit 210 divides the aforementioned bit field into a first bit field and a second bit field. The first bit field is used to write the remote redirection flag. This remote redirection flag indicates that the data corresponding to the current logical address has not been deleted or swapped out to the local disk, but has been written to the remote registered memory area in the remote computing node.

[0138] The second bit field is used to write the integer index; the integer index points to the remote memory access metadata corresponding to the target data block in the remote memory index table. Subsequently, when this logical address triggers a page fault, the operating system module 200 can query the remote memory index table based on the integer index in the page table entry, thereby obtaining information such as the remote node, remote registration key, remote base address, offset, and read length.

[0139] When a target data block corresponds to multiple local page table entries, the page table management unit 210 writes a remote redirection flag and an integer index to each target local page table entry. For data blocks stored contiguously in a remotely registered memory region, multiple local page table entries can be written with the same integer index; during subsequent page fault recovery, the operating system module 200 determines the specific remote read location based on the offset of the logical address that triggered the page fault relative to the starting position of the logical address of the target data block. For data blocks that are not contiguously stored remotely, the page table management unit 210 can write different integer indices for different logical pages or contiguous page segments.

[0140] After modifying the target local page table entry, the page table management unit 210 performs an address translation back buffer (TLB) refresh operation on the target logical address range. In a multi-processor core scenario, the page table management unit 210 also sends a TLB synchronization failure request to other processor cores so that each processor core no longer uses the old address mapping before the modification.

[0141] Through the above processing, the target local page table entry no longer points to the original local physical page frame, but instead records the remote redirection flag and integer index. Therefore, subsequent accesses to this logical address can trigger a page fault, and remote data recovery can be completed through the remote memory index table.

[0142] S430, the page table management unit 210 calls the kernel memory release interface to transfer the local physical page frames that have been unmapped from the page table to the operating system's reclaimable memory pool.

[0143] After confirming that the RDMA write operation is complete, the target local page table entry has been modified to a remote redirection state, and the TLB for the corresponding logical address range has been refreshed, the local physical page frame that originally carried the target data block is no longer directly mapped by that logical address. At this point, the page table management unit 210 calls the kernel physical page release interface to clear the page lock state of the physical page frame, reduce or clear the corresponding reference count, and return the physical page frame to the reclaimable memory pool.

[0144] Physical page frames returned to the reclaimable memory pool can be reallocated to other processes or computing tasks by the operating system's memory allocator. The release of physical page frames, maintenance of the free page list, and the management methods of the memory allocator can be implemented using the memory management mechanisms in standard operating systems.

[0145] See attached document Figure 7 In some specific implementations, step S50 is used to identify the remote redirection flag in the target local page table entry after a page fault is triggered at the target logical address, query the remote memory index table based on the integer index, and restore the data in the remote registered memory region to the local physical page frame. Specifically, this includes the following steps:

[0146] S510, when the computing engine module 300 accesses the unloaded logical address again due to fault-tolerant recalculation, the page fault exception handling unit 220 in the operating system module 200 receives and processes the page fault exception triggered by the processor.

[0147] In scenarios involving fault-tolerant recalculation, speculative execution rollback, or data reuse in distributed computing, the computing engine module 300 may again initiate a read instruction to the logical address corresponding to the unloaded intermediate data block. Since the local page table entry corresponding to this logical address has had its existence bit cleared during the aforementioned unloading process, the processor memory management unit cannot obtain a valid local physical address during address translation, thus triggering a page fault.

[0148] The page fault handling unit 220 in the operating system module 200 enters the kernel-mode page fault handling process and obtains the target logical address that triggered the current page fault. The target logical address is the specific logical address that triggered the current page fault, and the target logical address range is the address range formed by one or more logical addresses corresponding to the target data block or the target local page table entry set. When a target data block corresponds to multiple local page table entries, the target logical address range includes the logical addresses corresponding to each of the multiple local page table entries; when the current page fault corresponds to only one local page table entry, the target logical address is used to locate the target local page table entry that triggered the current page fault. In some embodiments, when the processor adopts an x86 architecture, the page fault handling unit 220 can read the target logical address from the CR2 register.

[0149] The page fault is recognized by the operating system module 200 as a remote data redirection recovery trigger event, rather than a direct condition for the abnormal termination of the application process. Through the above processing, the computing engine module 300 can recover the unloaded data in the page fault handling process without explicitly initiating a remote data read request.

[0150] S520, the page fault handling unit 220 parses the target local page table entry that triggered the page fault, and after extracting the integer index, queries the remote memory index table to obtain remote memory access metadata.

[0151] The page fault handling unit 220 locates the target local page table entry that triggered the page fault in the page table structure based on the target logical address. Subsequently, the page fault handling unit 220 reads the data in the target local page table entry and parses the bit fields that are allowed to be used by the software by the processor architecture.

[0152] If the target local page table entry does not contain a remote redirection flag, the operating system module 200 performs local physical page frame allocation, disk swapping, or exception reporting operations according to the standard page fault handling procedure.

[0153] If the target local page table entry contains a remote redirection flag, the page fault handling unit 220 extracts an integer index from the predefined field of the target local page table entry. The integer index corresponds to a set of remote memory access metadata in the remote memory index table.

[0154] After obtaining the integer index, the page fault handling unit 220 queries the remote memory index table in the kernel address space to obtain the remote memory access metadata corresponding to the integer index. The remote memory access metadata includes the network address of the remote computing node, the remote registration key, the remote base address, the offset within the remote registered memory region, the remote memory region identifier, and the read length.

[0155] When the remote memory access metadata record is a contiguous storage area at the data block level, the page fault handling unit 220 calculates the remote read start address and read length corresponding to this page fault based on the offset of the target logical address relative to the start position of the data block logical address. Thus, the operating system module 200 can determine the specific remote read location based on the integer index in the target local page table entry.

[0156] S530, the operating system module 200 allocates free physical page frames locally, performs RDMA read operation through the first network card 101 to pull remote data based on remote memory access metadata, updates the target local page table entry and resumes process execution.

[0157] After identifying the remote redirection flag and extracting the integer index, the page fault handling unit 220 submits the remote data recovery request to a sleepable kernel recovery thread, kernel work queue, or equivalent page fault recovery execution context for processing, and suspends the application thread that triggered the page fault until the remote data recovery is completed or the recovery failure result is returned.

[0158] The operating system module 200 requests new free physical page frames from the reclaimable memory pool in local memory. If no available free physical page frames are available in the current reclaimable memory pool, the operating system module 200 can trigger a standard memory reclamation mechanism, such as reclaiming some inactive pages based on the least recently used strategy, to obtain local physical page frames for data recovery.

[0159] Subsequently, the operating system module 200 uses the remote memory access metadata as parameters and issues an RDMA read command to the first network card 101 through the network card driver interface. The first network card 101 reads the data within the corresponding range from the remote registered memory region of the remote computing node based on the network address, remote memory region identifier, remote base address, remote registration key, remote read offset, and read length, and writes the read data into a newly allocated local physical page frame.

[0160] When the read length exceeds the capacity of a single local physical page frame, the operating system module 200 can allocate multiple consecutive or non-consecutive local physical page frames and perform segmented writing of the read data according to the page size.

[0161] If the remote compute node is unreachable, the remote registration key is invalid, data verification fails, or the remote read times out, the operating system module 200 returns a recovery failure event to the compute engine module 300, and keeps the target local page table entry that triggered the page fault in an invalid or pre-recovery state. If the remote memory index table records backup remote data copy information, the operating system module 200 can select a backup remote node to re-execute the data read based on the backup remote node, backup remote registration key, backup remote base address, and backup remote memory region identifier. If no available backup remote data copy exists, the compute engine module 300 triggers upper-layer fault-tolerant recalculation or re-reads the corresponding data from persistent storage.

[0162] For connection maintenance, permission verification, data verification, and transmission confirmation processes in RDMA read operations, existing InfiniBand or RoCE technologies can be used.

[0163] After the remote data reading is completed and passes verification, the operating system module 200 modifies the target local page table entry corresponding to the target logical address. Verification includes at least one of length verification, version number verification, checksum verification, or hash verification. The operating system module 200 can perform a consistency check on the data returned by the RDMA read operation based on the data length, version number, verification information, or remote memory access metadata recorded in the remote memory index table. Specifically, the page table management unit 210 points the physical page frame address field of the target local page table entry to the newly allocated local physical page frame, resets the exist bit to a valid state, and clears the original remote redirection flag.

[0164] After the target local page table entry is restored, the page table management unit 210 performs an address translation lookup buffer (TLB) refresh operation on the target logical address and releases the application thread that was suspended due to the remote data recovery request. Subsequently, the page fault handling process returns, and the operating system module 200 resumes the execution of the computing engine module 300's process. The computing engine module 300 re-executes the memory read instructions that were previously interrupted by the page fault, thereby reading the restored data.

[0165] In some embodiments, when a distributed computing job ends, the target data block is confirmed to no longer be used for fault-tolerant recalculation, speculative execution rollback or cache reuse, or the remote data copy has been replaced by a new remote data copy, the scheduling management module 400 deletes the corresponding entry in the remote memory index table via a system call, or marks the corresponding entry in the remote memory index table as invalid, and sends a remote memory release request to the remote memory service unit to release the remote registered memory region in the second computing node 110. After deleting or invalidating the corresponding entry, the operating system module 200 marks the integer index corresponding to the entry as idle, so that it can be rebound or reused for subsequent remote memory access metadata.

[0166] See attached document Figure 8 The following description, using an intermediate data block processing scenario in a distributed computing job as an example, further illustrates the memory dynamic resource allocation and scheduling system and method provided in this embodiment of the invention. It should be noted that this embodiment is only used to illustrate one application of the invention in a distributed computing job and does not constitute a limitation on the application scenarios of the invention.

[0167] In this embodiment, the first computing node 100 serves as one of the execution nodes of the distributed computing job, used to execute the upstream operator tasks in the physical execution plan. The second computing node 110 serves as a remote memory bearer node, used to provide a remote registered memory region. The computing engine module 300 executes the distributed computing job on the first computing node 100 and determines the data dependencies between the upstream operators, the target intermediate data block, and the direct downstream operator set according to the physical execution plan.

[0168] Specifically, when the computing engine module 300 executes the upstream operator task, it generates a target intermediate data block in the local memory of the first computing node 100. The page table management unit 210 in the operating system module 200 allocates a corresponding local physical page frame for the target intermediate data block and establishes a mapping relationship between the logical address range of the target intermediate data block and the local page table entry set. The computing engine module 300 records the data block identifier, logical address start position, logical address length, page size, and corresponding local page table entry set of the target intermediate data block in the mapping registry.

[0169] After the target intermediate data block is generated, multiple direct downstream operators in the physical execution plan initiate read requests for the target intermediate data block. The status tracking unit 310 determines the set of direct downstream operators corresponding to the target intermediate data block based on the data dependencies in the physical execution plan, and tracks the read completion status of each direct downstream operator on the target intermediate data block through the status callback interface. When any direct downstream operator completes the read processing of the target intermediate data block, the status tracking unit 310 updates the read status of the direct downstream operator to "completed".

[0170] When all downstream operators in the direct downstream operator set have completed reading and processing of the target intermediate data block, the state tracking unit 310 determines that the target intermediate data block will no longer be read by downstream operators in the current normal calculation process, and generates a completion event notification for the target intermediate data block. The state tracking unit 310 sends the completion event notification to the event receiving unit 410 in the scheduling management module 400 through the inter-process communication channel. After receiving the completion event notification, the event receiving unit 410 transmits the completion event notification to the unloading control unit 420.

[0171] The unloading control unit 420 determines the logical address range, local page table entry set, and target data length corresponding to the target intermediate data block based on the mapping registry, and requests the page table management unit 210 to set the unloading preparation state for the corresponding logical address range via a system call. The page table management unit 210 sets the page lock state or increases the reference count for the target physical page frame, and temporarily sets the corresponding logical address range to a read-only state or an inaccessible state. At the same time, it performs page table permission update and address translation backup buffer refresh operations to prevent the target intermediate data block from being written to, migrated, or prematurely reclaimed during remote data transmission.

[0172] After the unloading control unit 420 queries the cluster resource status table and determines that the second compute node 110 has sufficient free memory capacity to meet the target data length requirement, it sends a remote memory request to the remote memory service unit in the second compute node 110. The remote memory service unit allocates a remote memory region in the second compute node 110 and performs access registration on the remote memory region to form a remote registered memory region. The remote memory service unit then returns remote memory access metadata to the unloading control unit 420. The remote memory access metadata includes at least one of the following: remote node identifier, remote network address, remote registration key, remote base address, remote memory region identifier, offset within the remote registered memory region, and data length.

[0173] After obtaining the remote memory access metadata, the unloading control unit 420 issues an RDMA write command to the first network interface card 101 via the network interface card driver interface. Based on the remote memory access metadata, the first network interface card 101 establishes an RDMA communication connection with the second network interface card 111 in the second computing node 110, and writes the data in the local physical page frame corresponding to the target intermediate data block into the remote registered memory region of the second computing node 110. During this data transfer phase, the central processing unit of the first computing node 100 does not need to perform a memory copy operation on the data in the local physical page frame.

[0174] After the RDMA write operation is completed, the operating system module 200 constructs or updates the remote memory index table in the kernel address space. The remote memory index table records the integer index, logical address range, remote node identifier, remote network address, remote registration key, remote base address, offset within the remote registered memory region, data length, and validity status corresponding to the target intermediate data block. The page table management unit 210 then modifies the target local page table entry corresponding to the target intermediate data block, clears the presence bit in the target local page table entry, and writes the remote redirection flag and integer index into the bit field that the processor architecture allows software to use. After completing the modification of the target local page table entry, the page table management unit 210 performs an address translation back buffer refresh operation and releases the local physical page frame that originally carried the target intermediate data block to the reclaimable memory pool.

[0175] Therefore, after the target intermediate data block completes downstream reading processing, the local physical page frame corresponding to the target intermediate data block can be released and reused by other processes or computing tasks; at the same time, the data copy of the target intermediate data block is retained in the remote registered memory area of ​​the second computing node 110, and the target local page table entry is associated with the remote memory index table through the remote redirection flag and integer index.

[0176] In subsequent scenarios involving rollback, fault-tolerant recalculation, or cache reuse, the computing engine module 300 may access the target logical address corresponding to the target intermediate data block again. Since the presence bit in the target local page table entry corresponding to this target logical address has been cleared, the processor memory management unit cannot obtain a valid local physical address during address translation, thus triggering a page fault. The page fault handling unit 220 receives and processes the page fault triggered by the processor, obtains the target logical address, and locates the target local page table entry that triggered the page fault.

[0177] The page fault handling unit 220 parses the bit fields in the target local page table entry that the processor architecture allows the software to use. When it is confirmed that the target local page table entry contains a remote redirection flag, the page fault handling unit 220 extracts the integer index from the target local page table entry and queries the remote memory index table based on the integer index to obtain the remote memory access metadata corresponding to the target intermediate data block. If the target intermediate data block is stored contiguously in the remote registered memory region, the page fault handling unit 220 determines the remote read start address and read length corresponding to this page fault based on the offset of the target logical address relative to the starting position of the logical address of the target intermediate data block.

[0178] The operating system module 200 requests a new free physical page frame in local memory and performs an RDMA read operation through the first network interface card 101. The first network interface card 101 reads the corresponding data of the target intermediate data block from the remote registered memory area of ​​the second computing node 110 based on the remote network address, remote registration key, remote base address, offset in the remote registered memory area, and read length, and writes the read data into the newly allocated local physical page frame.

[0179] After the remote data is read and verified, the page table management unit 210 updates the target local page table entry, pointing the physical page frame address field of the target local page table entry to the newly allocated local physical page frame, resetting the exist bit to a valid state, and clearing the original remote redirection flag. After the target local page table entry is restored, the page table management unit 210 performs an address translation backstop buffer refresh operation on the target logical address and releases the application threads suspended due to the remote data recovery request. Subsequently, the computing engine module 300 re-executes the memory read instructions that were previously interrupted by the page fault exception, thereby reading the restored target intermediate data block.

[0180] Through the above application embodiments, the present invention can unload intermediate data blocks in a distributed computing job from the local physical page frame of the first computing node 100 to the remote registered memory region of the second computing node 110 after downstream reading processing is completed, and maintain the remote recovery path through the remote redirection flag and integer index in the target local page table entry. When subsequent computing processes access the target intermediate data block again, the system can automatically trigger remote data recovery based on the page fault exception handling mechanism, thereby releasing local memory resources while maintaining data support for fault-tolerant recalculation, speculative execution rollback, and cache reuse scenarios.

[0181] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for dynamic memory resource allocation and scheduling for big data system integration, characterized in that, Includes the following steps: Receive distributed computing jobs, parse the physical execution plan, and establish a mapping relationship between the logical address space of data blocks and local page table entries; Monitor the execution status of the distributed computing job, and generate a completion event notification when it is determined that the downstream operator has completed the reading and processing of the data block; In response to the completion event notification, a remote computing node is selected and the remote memory access metadata corresponding to the remote registered memory region in the remote computing node is obtained. The data in the local physical page frame is written into the remote registered memory region, and a remote memory index table is built in the kernel mode. An integer index is generated and a binding relationship is established between the integer index and the remote memory access metadata. After confirming that the data in the local physical page frame has been written to the remote registered memory region, modify the target local page table entry, write the remote redirection flag and the integer index, and release the corresponding local physical page frame to the reclaimable memory pool. When receiving and processing a page fault triggered by accessing the target logical address corresponding to the target local page table entry and the target local page table entry being invalid, and confirming that the target local page table entry contains the remote redirection flag, the remote memory index table is queried according to the integer index to obtain the data address in the remote registered memory region, the remote data is pulled through the RDMA read operation, the target local page table entry is updated, and process execution is resumed.

2. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, The process of receiving distributed computing jobs, parsing physical execution plans, and establishing mappings from data block logical address spaces to local page table entries specifically includes: The distributed computing job is subjected to directed acyclic graph structure parsing to extract the operator node set and data dependency edge set; Determine the intermediate data block output by the operator node during execution, and allocate a contiguous logical address space corresponding to the intermediate data block in the virtual memory space; In local memory, allocate corresponding physical page frames for the contiguous logical address space, map the contiguous logical address space to the allocated physical page frames, and form a set of local page table entries corresponding to the contiguous logical address space; A mapping registry is maintained in memory, which uses data block identifiers as indexes and records the binding relationships between each intermediate data block and the local page table entry set.

3. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, The monitoring of the distributed computing job execution status, and the generation of a completion event notification when it is determined that the downstream operator has completed the reading and processing of the data block, specifically includes: Based on the data dependency edges in the directed acyclic graph structure obtained by parsing the physical execution plan, determine the set of direct downstream operators for the target data block; Register a status callback interface in the data input / output scheduling component, receive the read completion confirmation flag of the downstream operator through the status callback interface, and update the read completion status of the downstream operator for the target data block. When all downstream operators in the set of direct downstream operators submit read completion confirmation flags, the event triggering condition is determined to be met, and a completion event notification for the target data block is generated.

4. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, The process of responding to the completion event notification, selecting a remote computing node and obtaining the remote memory access metadata corresponding to the remote registered memory region in the remote computing node, and writing the data in the local physical page frame into the remote registered memory region specifically includes: In response to the completion event notification, the cluster resource status table is queried, and a computing node with free memory capacity not less than the target data length is selected as the remote computing node, wherein the target data length is the data length to be written to the remote registered memory area; Send a remote memory request to the remote computing node to allocate a remote memory region on the remote computing node, perform access registration on the remote memory region to form the remote registered memory region, and return the remote memory access metadata. A remote direct memory access communication connection is established between the first network card of the first computing node carrying the local physical page frame and the remote computing node. An RDMA write command is issued to write the data in the local physical page frame into the remote registered memory area.

5. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, The step of constructing a remote memory index table in kernel mode, generating an integer index, and binding the integer index with the remote memory access metadata specifically includes: The remote memory index table is constructed in the kernel address space, and the integer index is generated for the remote memory access metadata. The binding relationship between the integer index and the remote memory access metadata is written into and maintained in the remote memory index table, wherein the remote memory access metadata includes at least the remote network address, the remote registration key, the remote base address, the offset of the data in the remote registered memory area, and the data length.

6. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, The modification of the target local page table entry, including writing the remote redirection flag and the integer index, specifically includes: After setting the local physical page frame corresponding to the target local page table entry to the unloading ready state and confirming that no new writes have occurred during the unloading ready state, the presence bit in the target local page table entry is cleared to indicate that the physical page frame corresponding to the logical address is not in a valid mapping state. In the target local page table entry, the processor architecture allows the software to divide the bit field into a first bit field and a second bit field, write the remote redirection flag in the first bit field, and write the integer index in the second bit field.

7. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 6, characterized in that, The step of releasing the corresponding local physical page frame to the reclaimable memory pool specifically includes: After the target local page table entry is modified, an address translation backup buffer refresh operation is performed on the logical address range where the target logical address corresponding to the target local page table entry is located. The kernel physical page release interface is invoked to clear the page lock state of the local physical page frame, reduce or clear the corresponding reference count, and return the local physical page frame to the operating system's reclaimable memory pool.

8. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 1, characterized in that, When receiving and processing a page fault triggered by accessing the target logical address corresponding to the target local page table entry and the target local page table entry being invalid, and confirming that the target local page table entry contains the remote redirection flag, the method of querying the remote memory index table according to the integer index to obtain the data address in the remote registered memory region specifically includes: Obtain the target logical address that triggered this page fault, and locate the target local page table entry that triggered the page fault in the page table structure; Extract the integer index from the bit fields that the processor architecture allows the software to use in the target local page table entry, query the remote memory index table in the kernel address space, and obtain the remote memory access metadata corresponding to the integer index; Based on the offset of the target logical address relative to the starting position of the data block logical address space, calculate the remote read start address and read length corresponding to this page fault.

9. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 8, characterized in that, The process of retrieving remote data via RDMA read operations specifically includes: Submit the remote data recovery request to the page fault recovery execution context for processing, and suspend the application thread that triggered this page fault. Allocate free physical page frames from the reclaimable memory pool in local memory; The remote memory access metadata is used as a parameter to issue an RDMA read instruction, which reads data within the corresponding range from the remote registered memory region of the remote computing node, and writes the read data into the newly requested free physical page frame.

10. The memory dynamic resource allocation and scheduling method for big data system integration according to claim 9, characterized in that, The step of updating the target local page table entry and resuming process execution specifically includes: After the remote data reading is completed and passes the verification, the physical page frame address field of the target local page table entry is set to the newly requested free physical page frame, the presence bit is set to valid again, and the original remote redirection flag is cleared. An address translation back buffer refresh operation is performed on the target logical address to release the suspended application thread, so that the resumed process can re-execute the interrupted memory read instruction.

Citation Information

Patent Citations

  • Operation method of distributed shared memory protocol

    CN116680229A

  • Spark cache optimization method and system based on RDD reusability

    CN117093369A