Task processing method, computing device, storage medium, and computer program product
Patent Information
- Application Number
- CN202610636610.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-11
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-05-11
AI Technical Summary
这种较高的延迟使得基于CXL.io的通信方式难以高效支撑不同粒度、尤其是细粒度的数据处理任务,无法满足近场数据处理对低通信延迟的要求
[0010]As can be seen from the above technical solutions, the task processing method provided in this specification is applied to the CXL memory extender. The CXL memory extender includes a memory mapping region. Based on this memory mapping region, the target program in the processor can simplify the task message sending process into a simple memory storage operation via the CXL.mem protocol, eliminating the need to call complex device driver interfaces and removing the overhead caused by operating system kernel context switching and software protocol stack parsing. Compared to sending task messages to the CXL memory extender via the CXL.io protocol, the latency of task message transmission is reduced from microseconds to nanoseconds, lowering communication latency and improving communication efficiency. Simultaneously, the task message carries the target virtual address. The CXL memory extender can query the TLB based on the address information including the target virtual address to obtain the target device physical address corresponding to the target virtual address, and execute near-field tasks based on the target device physical address, thus achieving the goal of reducing communication latency during near-field task processing. In summary, the task processing method provided in this specification is based on the CXL.mem protocol and memory mapping area to achieve rapid delivery of task messages. The CXL memory extender is based on TLB to achieve the address translation and task processing required for near-field task processing, without the need for repeated communication with the host (i.e., the processor), thereby reducing the communication latency of near-field task processing.
Smart Images

Figure CN122173417B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, specifically to memory expansion technology in the field of computer technology, and more specifically to a task processing method, computing device, storage medium, and computer program product. Background Technology
[0002] With the rapid development of computationally intensive applications such as artificial intelligence (AI) training and big data analytics, the demand for computer system memory capacity has increased dramatically. In order to expand the local memory capacity of processors at a lower cost, memory expansion technology based on Compute Express Link (CXL) has emerged.
[0003] However, in practical applications, it has been found that the data processing functions within the CXL memory expander are typically implemented by the CXL controller, and these solutions are mostly customized for specific application scenarios. This customized design lacks versatility and flexibility, making it difficult to adapt to diverse practical application needs and limiting the broad application potential of the CXL memory system.
[0004] Furthermore, data communication and control between the main processor and the CXL controller typically rely on the CXL.io (Compute Express Link Input / Output) protocol (which is based on a similar protocol stack to the PCIe protocol stack). This communication path involves software protocol stack processing, resulting in microsecond (µs) level latency overhead. This high latency makes CXL.io-based communication methods difficult to efficiently support data processing tasks of different granularities, especially fine-grained ones, and cannot meet the low communication latency requirements of near-field data processing. Summary of the Invention
[0005] This specification provides a task processing method, computing device, storage medium, and computer program product to reduce communication latency during near-field data processing.
[0006] To achieve the above technical objectives, the embodiments of this specification provide the following technical solutions: In a first aspect, one embodiment of this specification provides a task processing method applied to a compute fast link CXL memory extender, wherein the CXL memory extender establishes a communication connection with a processor, and the CXL memory extender includes a memory-mapped region; the task processing method includes: Receive a task message from the processor, the task message being written by the target program based on the host physical address corresponding to the memory-mapped region via the CXL.mem protocol; the task message includes the target virtual address; In response to the task message, the TLB is queried using address information to obtain the target device physical address corresponding to the target virtual address. The address information includes the target virtual address. The TLB is used to store address mapping relationships, which include the correspondence between virtual addresses and device physical addresses. Access the target data based on the target device's physical address to perform the processing operation indicated by the task message.
[0007] Secondly, one embodiment of this specification also provides a computing device, including: a processor, a CXL memory expander, and memory; wherein, The processor establishes a communication connection with the CXL memory expander, and the CXL memory expander establishes a communication connection with at least one of the memory modules. The CXL memory expander includes a memory mapping region. The processor is configured to write to the memory-mapped region via the CXL.mem protocol using the target program based on the host physical address corresponding to the memory-mapped region, in order to send task messages to the CXL memory extender; The CXL memory extender is configured to, in response to the task message, query the TLB using address information to obtain the target device physical address corresponding to the target virtual address, wherein the address information includes the target virtual address; the TLB is used to store address mapping relationships, wherein the address mapping relationships include the correspondence between virtual addresses and device physical addresses; Access the target data based on the target device's physical address to perform the processing operation indicated by the task message.
[0008] Thirdly, one embodiment of this specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the task processing method described above.
[0009] Fourthly, embodiments of this specification provide a computer program product or computer program, the computer program product including a computer program stored in a computer-readable storage medium; the processor of the computer device reads the computer program from the computer-readable storage medium, and when the processor executes the computer program, it implements the steps of the task processing method described above. Optionally, the computer program may be stored in a computer-readable storage medium or in the cloud; the processor of the computer device reads the computer program from the readable storage medium or in the cloud.
[0010] As can be seen from the above technical solutions, the task processing method provided in this specification is applied to the CXL memory extender. The CXL memory extender includes a memory mapping region. Based on this memory mapping region, the target program in the processor can simplify the task message sending process into a simple memory storage operation via the CXL.mem protocol, eliminating the need to call complex device driver interfaces and removing the overhead caused by operating system kernel context switching and software protocol stack parsing. Compared to sending task messages to the CXL memory extender via the CXL.io protocol, the latency of task message transmission is reduced from microseconds to nanoseconds, lowering communication latency and improving communication efficiency. Simultaneously, the task message carries the target virtual address. The CXL memory extender can query the TLB based on the address information including the target virtual address to obtain the target device physical address corresponding to the target virtual address, and execute near-field tasks based on the target device physical address, thus achieving the goal of reducing communication latency during near-field task processing. In summary, the task processing method provided in this specification is based on the CXL.mem protocol and memory mapping area to achieve rapid delivery of task messages. The CXL memory extender is based on TLB to achieve the address translation and task processing required for near-field task processing, without the need for repeated communication with the host (i.e., the processor), thereby reducing the communication latency of near-field task processing. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this specification. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of the structure of a computing device provided for one embodiment of this specification.
[0013] Figure 2 This is a flowchart illustrating a task processing method provided for one embodiment of this specification.
[0014] Figure 3 This is a schematic diagram of a storage area in a CXL memory expander provided for one embodiment of this specification.
[0015] Figure 4 This is a schematic diagram illustrating an address translation process provided for one embodiment of this specification.
[0016] Figure 5 This is a schematic diagram of a state machine for an NFDU provided as one embodiment of this specification. Detailed Implementation
[0017] Unless otherwise defined, the technical or scientific terms used in the embodiments of this specification shall have the ordinary meaning understood by one of ordinary skill in the art to which this specification pertains. The terms "first," "second," and similar terms used in the embodiments of this specification do not indicate any order, quantity, or importance, but are merely used to avoid confusion of constituent elements.
[0018] Unless the context otherwise requires, throughout this specification, "a plurality of" means "at least two," and "including" is interpreted as open-ended or encompassing, that is, "including, but not limited to." In the description of this specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this specification. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example.
[0019] The technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0020] Before introducing the discovery process of the technical problem to be solved by the task processing method provided in the embodiments of this specification, a brief explanation of the technologies that may be involved will be given first: Near-Field Task Processing (NLP) can refer to a task processing mode where a computational task initiated by the processor is offloaded to the CXL memory extender (i.e., the "near field") where the data resides. The CXL memory extender then directly accesses and processes the local data, and finally returns the result to the host (i.e., the processor).
[0021] NFDU (Near-Field Data Processing Unit) can refer to a hardware module integrated inside or near the CXL memory expander controller for performing data processing tasks (such as filtering, transformation, calculation, etc.).
[0022] CXL (Compute Express Link) is a high-performance, low-latency open standard for interconnecting processors and devices based on the PCIe physical layer. Its core objective is to achieve cache coherency between processors and devices such as accelerators and memory expanders.
[0023] CXL.io (Compute Express Link Input / Output) protocol: This can refer to the base layer of the CXL protocol stack. Essentially, it inherits and is compatible with PCIe I / O semantics, and can be responsible for device discovery, configuration, register and I / O space access, and other initialization and management functions.
[0024] CXL.mem (Compute Express Link Memory Protocol): A protocol layer in the CXL protocol stack used for memory expansion and access. It allows the host CPU to access memory on the CXL device directly, as if it were local memory, using load / store instructions, with latency in the hundreds of nanoseconds.
[0025] CXL.cache (Compute Express Link Cache Protocol) is an agent coherency protocol whose main function is to support devices caching the memory space managed by the host (usually the CPU) and maintain cache consistency.
[0026] CXL Switch (Compute Express Link Switch): This can refer to a network switching device used to connect multiple CXL ports and route data packets between them. It allows hosts and multiple CXL devices (such as memory extenders and accelerators) to build complex network topologies, enabling resource pooling and flexible sharing.
[0027] TLB (Translation Lookaside Buffer): In this specification, it can refer to a high-speed cache used to store the translation results (or correspondences) of recently used virtual addresses to physical addresses. It can be used to speed up the address translation process and avoid looking up the page table in memory every time it is accessed.
[0028] ASID (Address Space Identifier): This identifier identifies which process or virtual machine the address translation belongs to. During context switching, comparing the ASID avoids flushing the entire TLB, thus distinguishing between virtual addresses that may be shared by different processes.
[0029] DIMM (Dual In-line Memory Module): A standard module that assembles multiple DRAM chips on a single printed circuit board and connects to the computer motherboard via gold finger slots.
[0030] DRAM (Dynamic Random Access Memory) is a type of volatile semiconductor memory that requires periodic refreshing to retain data but allows for high-speed random read and write operations. It is a major component of computer memory.
[0031] Of the three protocols defined in the CXL standard (CXL.io, CXL.mem, and CXL.cache), CXL.io possesses complete and flexible device management and control capabilities. Originating from and compatible with the PCIe protocol stack, it naturally assumes the responsibilities of device enumeration, configuration, interrupt management, and general I / O communication. Therefore, in related technologies, when the processor issues near-field tasks to the CXL memory extender, it typically uses the CXL.mem protocol to build a data channel for large-volume data reads and writes, enjoying its near-local memory low latency and high bandwidth; and uses the CXL.io protocol to build a control channel for task processing, state synchronization, and command issuance. This is because a task descriptor is essentially a structured control information, not a raw data stream to be processed. The original CXL.mem protocol does not support the transmission of task descriptors; therefore, only the CXL.io protocol can be relied upon for issuing near-field tasks.
[0032] However, during application, it was found that when issuing near-field tasks based on the CXL.io protocol, the communication process requires traversing the entire PCIe / CXL.io software protocol stack. This involves multiple context switches between user mode and kernel mode, driver intervention, and possible memory copies. The latency introduced by this series of operations is typically on the order of microseconds (µs). In fine-grained task issuance scenarios, if each fine-grained computation task (e.g., processing a 4KB page) incurs a latency of up to microseconds, the latency during near-field task processing becomes unacceptable.
[0033] To address this issue, research revealed a solution: the conventional approach of requiring the CXL.io protocol for task assignment can be broken, transforming this control process into a single, efficient memory write operation. To achieve this, a memory-mapped region can be configured within the CXL memory extender. This allows target programs on the host machine to write task messages, carrying the target virtual address of the data to be processed, into this memory-mapped region using the CXL.mem protocol and based on the host's physical address. This single memory write operation based on the CXL.mem protocol achieves task message assignment, reducing latency from microseconds to nanoseconds, lowering communication delays, and improving the efficiency of near-field task processing. Once the task message is sent to the CXL memory extender, the extender can use the address information to query the TLB, obtain the target device's physical address corresponding to the target virtual address, and further process the target data based on that physical address, thus enabling the execution of near-field tasks locally on the CXL memory extender.
[0034] In summary, this specification provides a task processing method applied to the CXL memory extender, which includes a memory-mapped region. Based on this region, the target program in the processor can simplify the task message delivery process into a simple memory storage operation via the CXL.mem protocol, eliminating the need to call complex device driver interfaces and removing the overhead of operating system kernel context switching and software protocol stack parsing. Compared to delivering task messages to the CXL memory extender via the CXL.io protocol, the latency of task message transmission is reduced from microseconds to nanoseconds, lowering communication latency and improving communication efficiency. Simultaneously, the task message carries the target virtual address. The CXL memory extender can query the TLB based on the address information including the target virtual address to obtain the physical address of the target device corresponding to the target virtual address, and then execute the near-field task based on the target device's physical address, thus reducing communication latency during near-field task processing. In summary, the task processing method provided in this specification is based on the CXL.mem protocol and memory mapping area to achieve rapid delivery of task messages. The CXL memory extender is based on TLB to achieve the address translation and task processing required for near-field task processing, without the need for repeated communication with the host (i.e., the processor), thereby reducing the communication latency of near-field task processing.
[0035] Based on the above concept, the embodiments of this specification provide a task processing method. The task processing method provided by the embodiments of this specification will be described exemplarily below with reference to the accompanying drawings.
[0036] To be applied to, for example Figure 1Taking the CXL memory expander in the computing device shown as an example, this specification provides a task processing method, such as... Figure 2 As shown, the task processing method includes: S201: Receives a task message from a processor connected to the CXL memory expander. The task message is written by the target program based on the host physical address corresponding to the memory-mapped region, via the CXL.mem protocol. The task message includes the target virtual address. S202: In response to the task message, query the TLB using the address information to obtain the target device physical address corresponding to the target virtual address. The address information includes the target virtual address. The TLB is used to store address mapping relationships, which include the correspondence between virtual addresses and device physical addresses. S203: Access target data based on the target device's physical address to perform the processing operation indicated by the task message.
[0037] Figure 1 In this context, the computing device can be a server or other device with computing power and requirements for memory expansion and near-field task processing. Figure 1 In this context, the computing device may include a DIMM, a processor, and a computing fast link switch (in... Figure 1 The CXL memory expander (referred to as CXL SW) and CXL memory extender are used to connect multiple DRAMs as processor memory. The CXL memory expander can include an NFDU controller, NFDU, address interleaving module, and memory controller. DIMMs are the main memory directly connected to the processor. The processor is the computing and control center of the computing device. A fast link switch is optional hardware for the computing device and can be used to expand the system topology, such as... Figure 1 As shown, it allows a single processor to connect to multiple CXL memory expanders, thereby constructing a memory resource pool. The NFDU controller is the control center of the CXL memory expander, responsible for parsing task messages read from the memory-mapped area and coordinating the collaborative work of hardware such as the NFDU. The NFDU is the hardware engine that executes computing tasks, capable of receiving the device physical address translated by the NDFU controller and directly initiating data load / store requests to the memory controller. The address interleaving module is an optional module of the computing device, which can interleave the continuous device physical address space at the granularity of cache lines or pages onto the DRAM corresponding to the multiple memory channels it manages. The memory controller can manage multiple DRMAs connected to the CXL memory expander. It can receive memory access requests from the NFDU and convert them into electrical signals that meet the timing requirements of DRAM to complete the actual data read and write operations. It is understandable that, depending on the needs of the actual application scenario, the computing device may include more than Figure 1This manual does not limit the number of hardware modules that may be included.
[0038] Before implementing the task processing method, such as Figure 3 As shown, the CXL memory expander can reserve a memory-mapped area. This area is used by target programs in the processor to write task messages to this area via the CXL.mem protocol using memory read / write operations, achieving low-latency, high-efficiency task dispatching. In addition, the CXL memory expander can also include a system storage area and a TLB area. The memory-mapped area and system storage area can be storage areas visible to the operating system in the processor.
[0039] When implementing the task processing method, based on the memory mapping region, the target program in the processor can simplify the task message delivery process into a simple memory storage operation through the CXL.mem protocol, eliminating the need to call complex device driver interfaces and removing the overhead caused by operating system kernel context switching and software protocol stack parsing. Compared to delivering task messages to the CXL memory extender via the CXL.io protocol, the latency of task message transmission is reduced from microseconds to nanoseconds, lowering communication latency and improving communication efficiency. Simultaneously, the task message carries the target virtual address. The CXL memory extender can query the TLB based on the address information including the target virtual address to obtain the target device physical address corresponding to the target virtual address, and then execute the near-field task based on the target device physical address, achieving the goal of reducing communication latency during near-field task processing. In summary, the task processing method provided in this specification achieves rapid task message delivery based on the CXL.mem protocol and memory mapping region. The CXL memory extender performs the address translation and task processing required for near-field task processing based on the TLB, eliminating the need for repeated communication with the host (i.e., the processor), thereby reducing communication latency in near-field task processing.
[0040] In one implementation, the address mapping relationship specifically includes the correspondence between virtual addresses, device physical addresses, and address space identifiers; the address space identifier is used to identify the process of the target program; the task message further includes: the target address space identifier; the address information further includes the target address space identifier; The step of querying the TLB using address information to obtain the target device physical address corresponding to the target virtual address includes: Using the target address space identifier and the target virtual address, the TLB is queried to obtain the physical address of the target device.
[0041] In this embodiment, to correctly handle near-field tasks initiated by different processes in a multi-process concurrent execution environment, an address space identifier (ASID) is introduced. Assume two independent application processes are running simultaneously on the processor. Process A has an address space identifier of 0x0010 and possesses a virtual address of 0x4000_1000, which is mapped to a physical address on a CXL memory extender. Process B has an address space identifier of 0x0020 and also uses the virtual address 0x4000_1000, which is mapped to a different physical address on the CXL memory extender. In this case, to correctly identify near-field tasks initiated by different processes and obtain correct address translation results, an address space identifier representing the process can be introduced into the address mapping relationship and address information to ensure correct handling of near-field tasks in a multi-process environment.
[0042] In one implementation, the address mapping relationship is stored in a TLB entry, the TLB entry including a tag array and a data array corresponding to the tag array, wherein the tag array is used to store the high-order data of the address space identifier and at least a portion of the data of the virtual address; The data array is used to store the low-order data and mapping result of the address space identifier; the high-order data is the high N-order data of the address space identifier, and the low-order data is the low M-order data of the address space identifier. The step of querying the TLB using the target address space identifier and the target virtual address to obtain the physical address of the target device includes: The tag array of the TLB entry is queried using the high N bits of the target address space identifier and the target virtual address to obtain a target tag array that matches both the high N bits of the target address space identifier and the target virtual address. If the low M bits of the target address space identifier match the low bits of the data array corresponding to the target tag array, the physical address of the target device is determined to include the mapping result in the data array corresponding to the target tag array.
[0043] Since the size of the tag array in the TLB does not support the complete storage of the address space identifier, in this embodiment, the high N bits of the address space identifier are stored in the tag array for indexing, and the low M bits are stored in the data array for secondary matching. Furthermore, storing the high N bits in the tag array also improves query efficiency. During a query, the high bits of the target virtual address and the target address space identifier can be used for parallel matching in the tag array. Because the high bits of the target address space identifier have already participated in the query in the tag array, a large number of entries with mismatched address space identifiers can be quickly eliminated in the first comparison, improving matching and query efficiency.
[0044] Furthermore, to ensure the correctness of the matching results and avoid the risk of false hits, specifically, after querying the tag array and determining the target tag array, a second query is performed using the low M bits of the target address space identifier. Only when the low M bits of the target address space identifier match the low bits of the data array corresponding to the target tag array is it determined that the target device physical address includes the mapping result in the data array corresponding to the target tag array. This ensures the correctness of the matching results.
[0045] In one implementation, to improve the efficiency of address translation, a query system of first TLB - page cache - second TLB is constructed. Specifically, refer to... Figure 4 ,exist Figure 4 In this context, Miss indicates a query miss, and Hit indicates a query hit. The CXL memory extender is connected to at least one memory. The CXL memory extender also includes a page cache. The TLB includes a first TLB and a second TLB. The first TLB is stored in the CXL memory extender, and the second TLB is stored in the memory. The page cache is used to cache the address mapping relationship read from the second TLB. The method of querying the TLB using address information includes: The address information is used to query the first TLB and / or the page cache. If a hit occurs, the physical address of the target device corresponding to the target virtual address is obtained. If a miss occurs, the second TLB is queried to obtain the physical address of the target device corresponding to the target virtual address.
[0046] Specifically, when querying the second TLB, if the query hits (Hit), the physical address of the target device corresponding to the target virtual address is obtained and sent to the memory controller; if the query misses (Miss), a page table traversal and TLB filling operation can be triggered. This page table traversal and TLB filling operation can specifically include: triggering the page table traversal operation and writing the newly obtained entry into the second TLB.
[0047] In this embodiment, the second TLB in memory is used to expand the TLB hierarchy and increase the coverage of the TLB. This helps to improve the hit rate when querying the TLB, reduce the probability of traversing the page table, and thus improve query efficiency.
[0048] In one implementation, querying the second TLB includes: Data is read using a standard data exchange unit of the memory interface as a cache line, and the address mapping relationship obtained from the reading is cached in the page cache area. The standard data exchange unit includes 32 bytes or 64 bytes.
[0049] In this embodiment, all lookup operations for the second TLB are performed on the DRAM data bus at the DRAM interface granularity. Since each TLB entry is designed to be 16 bytes, this means that a single read will force the fetching of 2 or 4 consecutive TLB entries. Thus, the passive hardware limitation can be actively transformed into a system optimization strategy: by caching multiple entries read at once in the page cache, the mapping relationship between adjacent virtual addresses is essentially prefetched. Due to the spatial locality of program access, subsequent requests are likely to hit these prefetched entries, thereby transforming a single memory access into multiple subsequent fast cache hits, significantly reducing the average lookup latency.
[0050] In one embodiment, the second TLB is a directly mapped TLB, and the task processing method further includes: In response to an update request for a target TLB entry in the second TLB, a partial write operation is performed on the target TLB entry via the memory interface to update the target TLB entry.
[0051] In this embodiment, by leveraging the partial write capability of DRAM, updating the target TLB entry does not require reading and rewriting other irrelevant entries in the entire cache line. This eliminates redundant data transfer during the update process, thereby significantly reducing the memory bandwidth and read / write latency consumed by the update operation. For systems that frequently undergo mapping changes (such as process switching and page migration), the update method provided in this embodiment can significantly reduce the performance overhead of maintaining the second TLB.
[0052] In one implementation, the CXL memory extender further includes a base address register; The process of querying the second TLB includes: Based on the base address stored in the base address register, the second TLB index corresponding to the target virtual address, and the size of the second TLB entry, the physical address of the target TLB entry in the memory is calculated. Based on the physical address of the target TLB entry in the memory, the target TLB entry is read from the memory and a query process is performed.
[0053] In this embodiment, a feasible method for calculating the physical address of the second TLB entry in memory is provided. By introducing a base address register, the physical address can be calculated using the formula "base address (which can be stored in the base address register) + second TLB index × correlation × second TLB entry size". Compared with addressing through software calculation or complex memory management units, this method helps to reduce the latency of locating the entry position.
[0054] Specifically, since the second TLB is stored in the memory, the physical address of the target TLB entry in the memory must be calculated. The CXL memory extender stores the base address of the second TLB through a base address register. During a query, the physical address of the target TLB entry is calculated using combinational logic based on the base address, the second TLB index corresponding to the target virtual address, and the size of the second TLB entry.
[0055] Once the target TLB entry is read from the memory, the query process compares address information to determine whether it is a hit or a miss. If a hit occurs, the physical address of the target device is retrieved from the entry. If a miss occurs in the second TLB, a page table traversal operation is triggered to obtain the correct address mapping, and this mapping is written as a new TLB entry to the second TLB.
[0056] For example, consider a directly mapped second TLB (corresponding to a 4KB memory page) with 1M entries and its base address configured as 0. For the target virtual address 0xffff12345000, its second TLB index can be calculated as 0x12345. Assuming the second TLB entry size is 16 bytes, the physical address of the target TLB entry in memory is 0x123450 (i.e., 0 + 0x12345 × 1 × 16, with associativity = 1). The lookup process compares the key bits in the address information (e.g., 0xffff) with the information stored in the entry to ultimately determine whether the lookup has been successful or not.
[0057] In one implementation, such as Figure 5 As shown, the query process for the target TLB entry includes: Determine whether the target virtual address matches the target TLB entry; if so, obtain the target device physical address from the target TLB entry. If not, a page table traversal operation is triggered to obtain the physical address of the target device, and the obtained address mapping relationship is written as a new TLB entry into the second TLB.
[0058] Figure 5 The NFDU state machine is shown. Upon receiving the target virtual address, a translation process is executed. First, the TLB is queried to determine if there is a hit or miss. If a hit occurs, the query result (i.e., the physical address of the target device) is output. If a miss occurs, the next-level TLB is queried, and the above judgment process is repeated. If all TLBs are missed, a page table traversal operation is performed, and the TLBs are filled. In this embodiment, if a miss still occurs after querying the second TLB, a page table traversal operation is triggered, and the newly obtained entry is written to the second TLB. In this way, the CXL memory extender can autonomously supplement and update the second TLB during the address translation process, reducing the processor's burden in maintaining the second TLB.
[0059] In one implementation, the address mapping relationship is stored in a TLB entry, and the task processing method further includes: When a TLB invalidation instruction is received from the processor, in response to the TLB invalidation instruction, the TLB entry storing the address mapping relationship targeted by the TLB invalidation instruction is set to an invalid state.
[0060] A TLB invalidation instruction is an instruction used to notify one or more TLB entries in the CXL memory extender that they have become invalid. When executed by the CXL memory extender, this instruction sets the TLB entry it points to to an invalid state, thereby ensuring the correctness of the TLB entries stored in the CXL memory extender.
[0061] In this embodiment, a method is provided for handling invalid TLB entries when page table entries are changed (such as pages being swapped out, permissions being changed, or processes exiting). This method ensures that the address translation cache inside the CXL memory extender is always strictly synchronized with the page table maintained by the host CPU, preventing the NFDU from using invalid or incorrect address mappings to access data, thereby avoiding calculation errors, system crashes, or security vulnerabilities caused by data corruption.
[0062] In summary, the task processing method based on a fast computation link provided in this specification, in a memory extender supporting the CXL.mem protocol, achieves data communication and access between the host and the device by collaboratively utilizing memory mapping and memory binding mechanisms. The memory mapping mechanism refers to directly mapping the storage resources of the CXL device (i.e., the memory mapping region) in the host system address space, establishing a communication path based on the native CXL.mem protocol. The memory binding mechanism, through a hardware management unit, dynamically maintains and updates address translation entries in the device's DRAM-TLB to respond to host access. The resource management of the near-field data processing unit is associated with the memory binding mechanism, enabling flexible configuration of its computing resources based on the DRAM-TLB status or host instructions. Compared to other methods, the task processing method provided in this specification can significantly improve the speed of various applications requiring large amounts of memory, including in-memory OLAP (Online Analytical Processing), KVStore (Key-Value Store), LLM (Large Language Model), DLRM (Deep-Learning Recommendation Model), and graph analytics applications.
[0063] In one exemplary embodiment of this specification, a task processing apparatus is also provided, applied to a compute fast link CXL memory extender, the CXL memory extender establishing a communication connection with a processor, the CXL memory extender including a memory-mapped region, and the task processing apparatus comprising: The first module is configured to receive a task message from the processor, the task message being written by the target program based on the host physical address corresponding to the memory-mapped region via the CXL.mem protocol; the task message includes a target virtual address. The second module is used to respond to the task message by querying the TLB using address information to obtain the target device physical address corresponding to the target virtual address. The address information includes the target virtual address. The TLB is used to store address mapping relationships, which include the correspondence between virtual addresses and device physical addresses. The third module is used to access target data based on the physical address of the target device in order to perform the processing operation indicated by the task message.
[0064] For specific limitations regarding the task processing device, please refer to the limitations regarding the task processing method above, which will not be repeated here. Each module in the aforementioned task processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the operations corresponding to each module.
[0065] An exemplary embodiment of this specification also provides a computing device, including: a processor, a CXL memory expander, and memory; wherein, The processor establishes a communication connection with the CXL memory expander, and the CXL memory expander establishes a communication connection with at least one of the memory modules. The CXL memory expander includes a memory mapping region. The processor is configured to write to the memory-mapped region via the CXL.mem protocol using the target program based on the host physical address corresponding to the memory-mapped region, in order to send task messages to the CXL memory extender; The CXL memory extender is configured to, in response to the task message, query the TLB using address information to obtain the target device physical address corresponding to the target virtual address, wherein the address information includes the target virtual address; the TLB is used to store address mapping relationships, wherein the address mapping relationships include the correspondence between virtual addresses and device physical addresses; Access the target data based on the target device's physical address to perform the processing operation indicated by the task message.
[0066] Optionally, the address mapping relationship specifically includes the correspondence between virtual address, device physical address, and address space identifier; the address space identifier is used to identify the process of the target program; the task message further includes: the target address space identifier; the address information further includes the target address space identifier; The step of querying the TLB using address information to obtain the target device physical address corresponding to the target virtual address includes: Using the target address space identifier and the target virtual address, the TLB is queried to obtain the physical address of the target device.
[0067] Optionally, the address mapping relationship is stored in a TLB entry, the TLB entry including a tag array and a data array corresponding to the tag array, wherein the tag array is used to store the high-order data of the address space identifier and at least a portion of the data of the virtual address; The data array is used to store the low-order data and mapping result of the address space identifier; the high-order data is the high N-order data of the address space identifier, and the low-order data is the low M-order data of the address space identifier. The step of querying the TLB using the target address space identifier and the target virtual address to obtain the physical address of the target device includes: The tag array of the TLB entry is queried using the high N bits of the target address space identifier and the target virtual address to obtain a target tag array that matches both the high N bits of the target address space identifier and the target virtual address. If the low M bits of the target address space identifier match the low bits of the data array corresponding to the target tag array, the physical address of the target device is determined to include the mapping result in the data array corresponding to the target tag array.
[0068] Optionally, the CXL memory extender is connected to at least one memory, the CXL memory extender further includes a page cache, the TLB includes a first TLB and a second TLB, the first TLB is stored in the CXL memory extender, the second TLB is stored in the memory, and the page cache is used to cache the address mapping relationship read from the second TLB; The method of querying the TLB using address information includes: The address information is used to query the first TLB and / or the page cache. If a hit occurs, the physical address of the target device corresponding to the target virtual address is obtained. If a miss occurs, the second TLB is queried to obtain the physical address of the target device corresponding to the target virtual address.
[0069] Optionally, querying the second TLB includes: Data is read using a standard data exchange unit of the memory interface as a cache line, and the address mapping relationship obtained from the reading is cached in the page cache area. The standard data exchange unit includes 32 bytes or 64 bytes.
[0070] Optionally, the second TLB is a directly mapped TLB, and the task processing method further includes: In response to an update request for a target TLB entry in the second TLB, a partial write operation is performed on the target TLB entry via the memory interface to update the target TLB entry.
[0071] Optionally, the CXL memory expander further includes a base address register; The process of querying the second TLB includes: Based on the base address stored in the base address register, the second TLB index corresponding to the target virtual address, and the size of the second TLB entry, the physical address of the target TLB entry in the memory is calculated. Based on the physical address of the target TLB entry in the memory, the target TLB entry is read from the memory and a query process is performed.
[0072] Optionally, the query process for the target TLB entry includes: Determine whether the target virtual address matches the target TLB entry; if so, obtain the target device physical address from the target TLB entry. If not, a page table traversal operation is triggered to obtain the physical address of the target device, and the obtained address mapping relationship is written as a new TLB entry into the second TLB.
[0073] Optionally, the address mapping relationship is stored in a TLB entry, and the task processing method further includes: When a TLB invalidation instruction is received from the processor, in response to the TLB invalidation instruction, the TLB entry storing the address mapping relationship targeted by the TLB invalidation instruction is set to an invalid state.
[0074] In addition to the methods and devices described above, the task processing methods provided in the embodiments of this specification can also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the task processing methods according to various embodiments of this specification as described in the above-described task processing method section.
[0075] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0076] The computer program product described herein can be written in any combination of one or more programming languages to perform the operations of the embodiments described herein. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0077] Furthermore, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of the steps in the task processing methods according to various embodiments of this specification as described in the above-described task processing method section.
[0078] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this specification can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0079] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0080] The embodiments described above are merely illustrative of several implementation methods outlined in this specification. While the descriptions are specific and detailed, they should not be construed as limiting the scope of the solutions provided in this specification. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this specification, and these all fall within the scope of protection of this specification. Therefore, the scope of protection for this patent should be determined by the appended claims.
Claims
1. A task processing method, characterized in that, The CXL memory extender is applied to computing fast-link memory expansion. The CXL memory extender establishes a communication connection with a processor, which runs a target program. The CXL memory extender includes a memory-mapped region corresponding to a host physical address. The task processing method includes: The processor receives a task message, which is written by the target program based on the host physical address corresponding to the memory-mapped region, via the CXL.mem protocol; the task message includes the target virtual address. In response to the task message, the TLB is queried using address information to obtain the target device physical address corresponding to the target virtual address, wherein the address information includes the target virtual address; the TLB is used to store address mapping relationships, wherein the address mapping relationships include the correspondence between virtual addresses and device physical addresses; the TLB includes a first TLB, which is stored in the CXL memory expander; Access the target data based on the target device's physical address to perform the processing operation indicated by the task message.
2. The method according to claim 1, characterized in that, The address mapping relationship specifically includes the correspondence between virtual addresses, device physical addresses, and address space identifiers; the address space identifiers are used to identify the process of the target program. The task message further includes: a target address space identifier; the address information further includes the target address space identifier; The step of querying the TLB using address information to obtain the target device physical address corresponding to the target virtual address includes: Using the target address space identifier and the target virtual address, the TLB is queried to obtain the physical address of the target device.
3. The method according to claim 2, characterized in that, The address mapping relationship is stored in a TLB entry, which includes a tag array and a data array corresponding to the tag array. The tag array is used to store the high-order data of the address space identifier and at least a portion of the data of the virtual address. The data array is used to store the low-order data and mapping result of the address space identifier; the high-order data is the high N-order data of the address space identifier, and the low-order data is the low M-order data of the address space identifier. The step of querying the TLB using the target address space identifier and the target virtual address to obtain the target device physical address includes: The tag array of the TLB entry is queried using the high N bits of the target address space identifier and the target virtual address to obtain a target tag array that matches both the high N bits of the target address space identifier and the target virtual address. If the low M bits of the target address space identifier match the low bits of the data array corresponding to the target tag array, the physical address of the target device is determined to include the mapping result in the data array corresponding to the target tag array.
4. The method according to claim 1, characterized in that, The CXL memory extender is connected to at least one memory, and the CXL memory extender also includes a page cache. The TLB also includes a second TLB, which is stored in the memory. The page cache is used to cache the address mapping relationship read from the second TLB. The method of querying the TLB using address information includes: The address information is used to query the first TLB and / or the page cache. If a hit is found, the target device physical address corresponding to the target virtual address is obtained. If a match is not found, the second TLB is queried to obtain the target device physical address corresponding to the target virtual address.
5. The method according to claim 4, characterized in that, The query of the second TLB includes: Data is read using a standard data exchange unit of the memory interface as a cache line, and the address mapping relationship obtained from the reading is cached in the page cache area. The standard data exchange unit includes 32 bytes or 64 bytes.
6. The method according to claim 4, characterized in that, The second TLB is a directly mapped TLB, and the task processing method further includes: In response to an update request for a target TLB entry in the second TLB, a partial write operation is performed on the target TLB entry via the memory interface to update the target TLB entry.
7. The method according to claim 4, characterized in that, The CXL memory expander also includes a base address register; The process of querying the second TLB includes: Based on the base address stored in the base address register, the second TLB index corresponding to the target virtual address, and the size of the second TLB entry, calculate the physical address of the target TLB entry in the memory; Based on the physical address of the target TLB entry in the memory, the target TLB entry is read from the memory and a query process is performed.
8. The method according to claim 7, characterized in that, The query process for the target TLB entry includes: Determine whether the target virtual address matches the target TLB entry; if so, obtain the target device physical address from the target TLB entry. If not, a page table traversal operation is triggered to obtain the physical address of the target device, and the obtained address mapping relationship is written as a new TLB entry into the second TLB.
9. The method according to claim 1, characterized in that, The address mapping relationship is stored in a TLB entry, and the task processing method further includes: When a TLB invalidation instruction is received from the processor, in response to the TLB invalidation instruction, the TLB entry storing the address mapping relationship targeted by the TLB invalidation instruction is set to an invalid state.
10. A computing device, characterized in that, include: Processor, CXL memory expander, and memory; among which, The processor establishes a communication connection with the CXL memory expander, and the CXL memory expander establishes a communication connection with at least one of the memory modules. The CXL memory expander includes a memory mapping region, and the memory mapping region corresponds to a host physical address. The processor is configured to send a task message to the CXL memory extender by writing to the memory-mapped region via the CXL.mem protocol using a target program based on the host physical address corresponding to the memory-mapped region; the task message includes the target virtual address. The CXL memory extender is configured to respond to the task message by querying the TLB using address information to obtain the target device physical address corresponding to the target virtual address, wherein the address information includes the target virtual address; the TLB is used to store address mapping relationships, wherein the address mapping relationships include the correspondence between virtual addresses and device physical addresses; the TLB includes a first TLB, which is stored in the CXL memory extender; Access the target data based on the target device's physical address to perform the processing operation indicated by the task message.
11. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the task processing method according to any one of claims 1 to 9.
12. A computer program product, characterized in that, The computer program product includes a computer program, which, when executed by a processor, implements the task processing method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Memory access method and device
CN114840445A
Data processing equipment and method
CN117370228A