Polling method, system and host

By setting up the polling bitmap in the device memory module of the CXL device and updating the cache module using the reverse failure mechanism, the problem of CPU frequent access to host memory is solved, and more efficient processor processing is achieved.

CN120469882APending Publication Date: 2025-08-12DAPUSTOR CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342743.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

In the prior art, the polling mechanism causes the CPU to frequently access the host memory, resulting in interrupts and waste of resources, and reducing processing efficiency.

Method used

Set up a polling bitmap in the device memory module of the CXL device, and update the polling bitmap in the processor's cache module in real time through the reverse failure mechanism, so that the processor handles events without interruption. If the event misses the cache module, data will be obtained from the CXL device.

Benefits of technology

It reduces the latency of accessing host memory, improves the processor's processing efficiency, and reduces the overhead of invalid polling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469882A_ABST
    Figure CN120469882A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a polling method, a polling system and a host. A polling bitmap is set in a device memory module of a CXL device, and the polling bitmap in a cache module of a processor is updated in real time according to a processing result of an event; the processor of the host can perform event processing under the condition that the processor is not interrupted; moreover, during polling, if the event does not hit the cache module of the processor, the data corresponding to the event is acquired from the CXL equipment, so that the processor does not need to access the host memory during polling, the time delay caused by access to the host memory can be reduced, and the processing efficiency of the processor of the host is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and in particular to a polling method, system, and host. Background Art

[0002] Polling is a mechanism whereby the host CPU actively queries the device for changes in status or task completion. The CPU periodically or cyclically checks the device's status registers or data structures to determine whether the device's task progress or results are ready.

[0003] Currently, the polling area is usually set in the host memory. The CXL device can transfer status information to the host memory through the DMA mechanism. This allows the host to periodically check the status information in the host memory (the result of the device update via DMA) without directly accessing the device registers. When the host detects a status update, it takes corresponding actions.

[0004] However, this method requires the CPU to occupy CPU time every time it accesses the host memory. Polling with a short check interval will occupy a large amount of CPU time. Invalid polling (polling when the event is not completed) will lead to resource waste, thereby reducing the CPU processing efficiency. Summary of the Invention

[0005] Embodiments of the present application provide a polling method, system, and host. These methods set a polling bitmap in the device memory module of a CXL device and update the polling bitmap in the processor's cache module in real time based on event processing results, enabling the host's processor to process events without interruption. Furthermore, during polling, if an event misses the processor's cache module, the data corresponding to the event is retrieved from the CXL device. This eliminates the need for the processor to access the host memory during polling, thereby reducing latency caused by accessing the host memory and improving the processing efficiency of the host's processor.

[0006] The embodiments of this application provide the following technical solutions:

[0007] In a first aspect, an embodiment of the present application provides a polling method, which is applied to a host, wherein the host includes a processor, the processor includes a cache module, the host is connected to a CXL device, the CXL device includes a device memory module, and the device memory module is used to store a polling bitmap;

[0008] Methods include:

[0009] When the host polls the CXL device for the first time, the polling bitmap is read from the device memory module of the CXL device and loaded into the cache module;

[0010] According to the processing results of the event, the polling bitmap in the cache module is updated in real time;

[0011] During the next polling, if an event hits the cache module, the data corresponding to the event is obtained from the cache module;

[0012] If the event does not hit the cache module, the data corresponding to the event is obtained from the CXL device.

[0013] In some embodiments, the polling bitmap is updated by the CXL device through a reverse invalidation mechanism, the reverse invalidation mechanism being configured to update the polling bitmap in the cache module after the polling bitmap in the device memory module is updated; the polling bitmap includes a plurality of bits, each bit corresponding to a one-to-one event;

[0014] Based on the event processing results, the polling bitmap in the cache module is updated in real time, including:

[0015] If the processing result of a certain event is completed, based on the reverse invalidation mechanism, the bit corresponding to the event is set to one in the polling bitmap of the cache module to update the polling bitmap in the cache module.

[0016] In some embodiments, the reverse invalidation mechanism is specifically used to: after the bit corresponding to the event in the polling bitmap of the device memory module is set to one, the bit corresponding to the event in the polling bitmap of the cache module is set to one, and the data corresponding to the event in the cache module is marked from valid data to invalid data, wherein the valid data is used to indicate that the data corresponding to the event in the current cache module is consistent with the data corresponding to the event in the device memory module, and the invalid data is used to indicate that the data corresponding to the event in the current cache module is inconsistent with the data corresponding to the event in the device memory module.

[0017] In some embodiments, the method further comprises:

[0018] Determine whether an event hits the cache module, specifically including:

[0019] Determine whether the data corresponding to the event in the cache module is valid data;

[0020] If so, it is determined that the event hits the cache module;

[0021] If not, it is determined that the event does not hit the cache module, wherein when the data corresponding to the event is valid data, the value of the bit corresponding to the event is a preset value.

[0022] In some embodiments, the method further comprises:

[0023] Sending a first request to the CXL device, wherein the first request corresponds to a first event;

[0024] querying a polling bitmap in the cache module according to the first request, determining a first index position in the polling bitmap, and sending a first event number corresponding to the first index position to the CXL device, wherein the first index position is a position where the first bit in the polling bitmap is zero;

[0025] After the CXL device receives the first event number, a processing result of the first event corresponding to the first event number sent by the CXL device is obtained.

[0026] In some embodiments, the method further comprises:

[0027] If the processing result of the first event is completed, the bit corresponding to the first event is set to 1 in the polling bitmap of the cache module through the reverse invalidation mechanism to update the polling bitmap in the cache module;

[0028] If the processing result of the first event is incomplete, a timer task is generated to periodically check the processing result of the first event until the processing result of the first event is completed;

[0029] After the processing result of the first event is completed, the bit corresponding to the first event in the polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module, and a check command is sent to the CXL device, so that the CXL device sets the bit corresponding to the first event to 0 after receiving the check command.

[0030] In some embodiments, the check command includes a polling bitmap, the polling bitmap includes at least two bits, and each bit corresponds to a first event;

[0031] The method also includes:

[0032] After the processing results of at least two first events are completed, the bit corresponding to each first event in the polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module.

[0033] In some embodiments, the cache module includes a first bitmap unit, wherein the first bitmap unit is used to cache the polling bitmap, and the first bitmap unit is a fixed area in the cache module, wherein the polling bitmap in the first bitmap unit is not evicted during polling;

[0034] The device memory module of a CXL device is read-only.

[0035] In a second aspect, an embodiment of the present application provides a host, including:

[0036] at least one processor; and

[0037] a memory communicatively connected to at least one processor; wherein,

[0038] The memory stores instructions that can be executed by at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the polling method according to the first aspect.

[0039] In a third aspect, an embodiment of the present application provides a polling system, including:

[0040] The host according to the second aspect, wherein the host comprises at least two processors, and the processors comprise cache modules;

[0041] A CXL device is connected to a host, wherein the CXL device includes:

[0042] a device memory module, the device memory module being configured to store at least two polling bitmaps, wherein the polling bitmaps correspond one to one with the processors;

[0043] The event manager is configured to set the bit corresponding to the first event corresponding to the check command to zero after the CXL device receives the check command sent by the host;

[0044] The CXL controller is configured to send a reverse invalidation message to the host through a reverse invalidation mechanism after the bits of the polling bitmap of the device memory module are updated, so as to update the polling bitmap in the cache module of the host.

[0045] In a fourth aspect, an embodiment of the present application provides a polling method, which is applied to the polling system of the third aspect, and the method includes:

[0046] When the host writes data to the CXL device, the host generates a first request and generates a first command based on the first request, wherein the first request corresponds to a first event;

[0047] The host sends a first command to the CXL device. After the CXL device receives the first command, the event manager establishes a mapping relationship between the first command and the event number.

[0048] After the CXL device completes the first command, the event manager sets the bit corresponding to the first event in the polling bitmap of the device memory module to one;

[0049] After the bit corresponding to the first event is updated, a reverse invalidation message is sent to the host through the reverse invalidation mechanism to update the polling bitmap in the cache module of the host;

[0050] The host polls the processing result of the first event, and if the processing result of the first event is completed, sends an event bitmap to the CXL device, wherein the event bitmap is used to determine the first event whose processing result is completed;

[0051] The event manager sets to zero the bit corresponding to the first event with a completed processing result in the polling bitmap of the device memory module according to the event bitmap.

[0052] In some embodiments, the method further comprises:

[0053] When a CXL device establishes communication with a host, the CXL device creates a send queue and a completion queue, where the number of bits in the polling bitmap is equal to the queue depth of the completion queue;

[0054] After the host generates the first command according to the first request, the host adds the first command to the sending queue.

[0055] The advantageous effects of the embodiments of the present application are as follows: Different from the prior art, the embodiments of the present application provide a polling method applied to a host, the host including a processor including a cache module, the host connected to a CXL device, the CXL device including a device memory module, the device memory module being used to store a polling bitmap; the method comprising: when the host initially polls the CXL device, reading the polling bitmap from the device memory module of the CXL device and loading the polling bitmap into the cache module; updating the polling bitmap in the cache module in real time based on the result of event processing; and during the next polling, if an event hits the cache module, obtaining data corresponding to the event from the cache module; if the event does not hit the cache module, obtaining data corresponding to the event from the CXL device.

[0056] By setting a polling bitmap in the device memory module of the CXL device and updating the polling bitmap in the processor's cache module in real time based on the event processing results, the host processor can process events without being interrupted. Moreover, during polling, if the event does not hit the processor's cache module, the data corresponding to the event is obtained from the CXL device, so that the processor does not need to access the host memory during polling, thereby reducing the delay caused by accessing the host memory and improving the processing efficiency of the host processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] One or more embodiments are exemplarily described by corresponding drawings, which do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0058] Figure 1This is a schematic diagram of polling interaction between a host and a CXL device provided in an embodiment of the present application;

[0059] Figure 2 This is a flowchart of a polling method provided in an embodiment of the present application;

[0060] Figure 3 yes Figure 2 A detailed flow chart of step S202 in FIG.

[0061] Figure 4 yes Figure 2 A detailed flow chart of step S203 in FIG.

[0062] Figure 5 This is a flowchart of obtaining a processing result of a first event provided by an embodiment of the present application;

[0063] Figure 6 This is a flowchart of determining whether the processing result of a first event is completed, provided by an embodiment of the present application;

[0064] Figure 7 This is a schematic diagram of a process for updating a polling bitmap of a cache module provided by an embodiment of the present application;

[0065] Figure 8 This is a schematic diagram of the structure of a host provided in an embodiment of the present application;

[0066] Figure 9 This is a schematic diagram of the structure of a polling system provided in an embodiment of the present application;

[0067] Figure 10 This is a flowchart of another polling method provided in an embodiment of the present application;

[0068] Figure 11 This is a flow chart of adding a first command to a sending queue provided by an embodiment of the present application;

[0069] Figure 12 This is a schematic diagram of the overall flow of data processing between a host and a CXL device provided in an embodiment of the present application.

[0070] Description of Figure Numbers:

[0071] Label name Label name 90 Polling system 901 Host 911 processor 9111 Cache module 912 Memory 902 CXL equipment 921 Device memory module 922 CXL Controller 923 Event Manager DETAILED DESCRIPTION

[0072] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0073] In addition, the technical features involved in each embodiment of the present application described below can be combined with each other as long as they do not conflict with each other.

[0074] Unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification and in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.

[0075] Currently, the polling area is typically set in the host memory. CXL devices can transfer status information to the host memory through the DMA mechanism. This eliminates the need for the host to directly access the device registers. Instead, the host periodically checks the status information in the host memory (the result of the device updating via DMA). When a status update is detected, the host takes appropriate action. This solution has the following two disadvantages:

[0076] (1) Interrupts interrupt the current task of the CPU, generate additional context switching overhead, and reduce system performance. When the device requires a response from the host (such as when data transmission is completed or new data is available), it actively sends an interrupt signal to the host. After receiving the interrupt signal, the host suspends the current task and jumps to the interrupt handler to handle the device's needs. In this process, the data in the cache of the original task may be evicted, causing the CPU to re-access the main memory, reducing overall performance. When an interrupt occurs, the operating system needs to save the state of the current process, switch to the interrupt handler, and then restore the state of the process after handling the interrupt. This process will cause context switching overhead, especially when interrupts occur frequently, which may significantly affect system performance.

[0077] (2) Traditional polling repeatedly checks memory / device registers, which is slow and inefficient. Polling with a short check interval will take up a lot of CPU time, and invalid checks will lead to resource waste. Although resources can be saved by reasonably setting a fixed time interval and performing checks periodically, each access to host memory or device register will take hundreds of nanoseconds or even several microseconds, which leads to inefficient checks. For example, SPDK (Storage Performance Development Kit) is a development toolkit specifically designed to provide high performance and low latency for NVMe storage devices. It supports polling mode to process I / O requests. In this mode, the application continuously checks the memory area of the completion queue (CQ) exposed by the device. Depending on the host configuration, these areas are usually placed in the host's DMA buffer (referring to the memory used for DMA access by the device, usually set to non-cacheable), or they can be placed inside the device and mapped to the host via PCIe. When the CQ is placed in the host memory, each poll will access the memory, and the memory latency is usually at the level of hundreds of nanoseconds. If the CQ is placed in the device memory, each polling will involve accessing the PCIe bus, which usually takes microseconds, resulting in high query overhead.

[0078] To address the shortcomings of the above-mentioned solution, the present application utilizes a reverse invalidation mechanism to update the polling bitmap in the processor's cache module in real time, allowing the host processor to process events without being interrupted. Furthermore, during polling, if an event does not hit the processor's cache module, the data corresponding to the event is obtained from the CXL device, so that the processor does not need to access the host memory during polling, thereby reducing the delay caused by accessing the host memory and improving the processing efficiency of the host processor.

[0079] The technical solution of this application is described in detail below with reference to the accompanying drawings:

[0080] See also Figure 1 , Figure 1 This is a schematic diagram of polling interaction between a host and a CXL device provided in an embodiment of the present application;

[0081] like Figure 1 As shown, a host is communicatively connected to a CXL device. The host includes at least two processors (CPUs), such as CPU0 and CPU1. The host also includes an application and a cache module. The cache module includes multiple levels of cache, such as a first-level cache, a second-level cache, and a third-level cache. The CXL device includes an event manager and a device memory module.

[0082] This application proposes a cached polling mechanism for CXL devices and hosts. This mechanism leverages the reverse invalidation mechanism in CXL cache coherence to redesign the polled memory from a traditional DMA buffer to a cacheable polling bitmap. This effectively reduces CPU overhead for invalid polling (polling when an event is not completed), while also rapidly expanding the number of events and supporting multi-event polling. The cached polling process primarily consists of three parts: the host's kernel polling service, the CXL device's event manager, and the CXL device's polling bitmap.

[0083] The kernel polling service is mainly used to expose an interface for registering events and polling events to drivers and user programs, complete periodic access to the polling bitmap, and support blocked waiting interfaces.

[0084] The event manager is used to track events within the device, recycle event numbers, and establish event-related data structures when the host requests resources.

[0085] The polling bitmap is a special memory area set aside for each CPU in the CXL device. This memory area supports the CXL.mem protocol and includes a reverse invalidation mechanism. When accessing this memory, the CPU reads and writes at a 64-byte cache line granularity, corresponding to the size of each polling bitmap. The reverse invalidation mechanism allows the CPU to update the cache without interruption, and the latest polling bitmap will be read from the CXL device on the next poll.

[0086] In the embodiment of the present application, the host is used to send a request and event number to the CXL device; or obtain a polling result sent by the CXL device; or send a query command to the CXL device to notify the CXL device that the host has completed the query of the processing result of a certain event.

[0087] Specifically, a user program or driver calls a kernel polling service interface, generates a first request, and encapsulates the request into a CXL.io transaction layer packet (TLP). The host sends the request and the event number corresponding to the event to the CXL device via the CXL.io protocol or other PCIe protocols. The request is used to obtain the processing result of the first event from the CXL device. Alternatively, the host requests the CXL device via the kernel polling service to obtain the processing result of the first event corresponding to the first event number. The status of the processing result is stored in a polling bitmap in the memory of the CXL device, where each event number corresponds to a bit. The host obtains the processing result of the first event corresponding to the first event number sent by the CXL device by querying the bit corresponding to the first event number.

[0088] Among them, the processing result of the first event includes but is not limited to DMA transfer completion and computing task readiness. The request can be initiated by the user program (such as calling the poll interface or io_uring interface) or actively queried by the driver. This application does not limit this; or, after querying that the processing result of the first event is completed, the bit corresponding to the first event is set in the polling bitmap of the cache module, that is, the bit corresponding to the first event is set from 0 to 1 to update the polling bitmap of the cache module, and the data corresponding to the first event is set to invalid data.

[0089] Furthermore, after the host updates the polling bitmap in the cache module, the host sends the bitmap of the read event to the event manager via a check command through the kernel polling service, wherein the check command is used to indicate that the host has completed checking the processing result of the first event, so that after receiving the check command, the CXL device sets the bit corresponding to the first event to 0 through the event manager and recovers the first event number corresponding to the first event.

[0090] In the embodiment of the present application, the device memory module is used to store at least two polling bitmaps, wherein the polling bitmaps correspond to the processors one by one, that is, each polling bitmap Figure 1 One corresponds to one processor; specifically, the device memory module of the CXL device is a high-performance memory expansion hardware based on the Compute Express Link (CXL) protocol, which realizes efficient interconnection between the processor and memory resources through a standardized interface; the CXL memory module is usually composed of DDR5 / DRAM chips, CXL memory controllers and protocol interfaces, and supports cache coherence communication with the host processor. The device memory module is specifically used to store at least two polling bitmaps, where each polling bitmap Figure 1 One corresponds to a host processor, for example:

[0091] Figure 1 The CPU0 polling bitmap corresponds to CPU0, and the CPU1 polling bitmap corresponds to CPU1. The polling bitmap is a fixed 64-byte bitmap, where each byte is 8 bits. The polling bitmap has a total of 512 (64*8) bits, and each bit corresponds to an entry. There are a total of 512 entries, each entry represents an event, and the value of each bit is used to indicate the completion status of the event. For example, when the value of the bit corresponding to an event is 1, it indicates that the processing result of the event is incomplete. When the value of the bit corresponding to an event is 0, it indicates that the processing result of the event is completed.

[0092] In an embodiment of the present application, the event manager is configured to, after the CXL device receives a check command issued by the host, reset a bit corresponding to a first event corresponding to the check command to zero. Specifically, after the host inquires that the processing result of the first event is completed, the event manager resets the bit corresponding to the first event in the polling bitmap of the cache module to one, i.e., sets the bit corresponding to the first event from 0 to 1, thereby updating the polling bitmap of the cache module and setting the data corresponding to the first event to invalid data. Furthermore, after the host updates the polling bitmap in the cache module, the host sends a bitmap of read events to the event manager via a check command through the kernel polling service. The check command indicates that the host has completed the processing result of the first event, so that after receiving the check command, the CXL device, through the event manager, resets the bit corresponding to the first event corresponding to the check command to zero and reclaims the first event number corresponding to the first event.

[0093] In this embodiment of the present application, after updating the bits in the polling bitmap of the device's memory module, the CXL controller is configured to send a reverse invalidation message to the host via a reverse invalidation mechanism to update the polling bitmap in the host's cache module. Specifically, the reverse invalidation mechanism involves the CXL device proactively sending an invalidation request to the host via the CXL protocol, invalidating old data in the host's cache, thereby maintaining data consistency between the CXL device and the host. This mechanism allows the device to more efficiently manage cache status, reduces host intervention, and thus improves overall performance.

[0094] In an embodiment of the present application, a reverse invalidation mechanism is a mechanism used by CXL devices in the CXL protocol to proactively maintain cache coherence. It notifies the host via a reverse invalidation message to update the completion status of an event in the cache module, ensuring data consistency between the host and the CXL device while reducing latency and protocol overhead. The polling bitmap is updated by the CXL device via the reverse invalidation mechanism. The reverse invalidation mechanism is used to update the polling bitmap in the cache module after the polling bitmap in the device memory module is updated. Furthermore, the reverse invalidation mechanism is specifically configured to: after setting the bit corresponding to the event in the polling bitmap in the device memory module to 1, the CXL controller sends a reverse invalidation message to the host, causing the host to set the bit corresponding to the event in the polling bitmap in the cache module to 1 upon receipt of the reverse invalidation message, and to mark the data corresponding to the event in the cache module from valid data to invalid data. Valid data indicates that the data corresponding to the event in the current cache module is consistent with the data corresponding to the event in the device memory module, while invalid data indicates that the data corresponding to the event in the current cache module is inconsistent with the data corresponding to the event in the device memory module. For example:

[0095] Assuming that the current processing result of an event is incomplete, the bit corresponding to the event is zero. When the processing result of the event is completed, the CXL device will set the bit corresponding to the event to one according to the reverse invalidation mechanism, that is, the bit corresponding to the event in the polling bitmap of the cache module is set from 0 to 1, and the data corresponding to the event in the host cache module is marked from valid data to invalid data through the reverse invalidation mechanism. For example:

[0096] Assuming this event is a write operation, after the host modifies the data in the CXL device corresponding to the write operation event, the event processing result is completed. The CXL device will set the bit corresponding to the event in the host's cache module from 0 to 1 according to the back-invalidation mechanism. Furthermore, because the data corresponding to the event has changed, the CXL device also uses the back-invalidate mechanism to mark the old data corresponding to the event in the host's cache module from valid data to invalid data. In this embodiment of the application, combined with the CXL back-invalidate mechanism, the polling area can be converted from a traditional non-cacheable memory area to a cacheable CXL memory area.

[0097] See also Figure 2 , Figure 2 This is a flowchart of a polling method provided in an embodiment of the present application;

[0098] The polling method is applied to a host, the host includes a processor, the processor includes a cache module, the host is connected to a CXL device, the CXL device includes a device memory module, and the device memory module is used to store a polling bitmap.

[0099] like Figure 2 As shown, the polling method includes:

[0100] Step S201 : When the host polls a CXL (Compute Express Link) device for the first time, the host reads a polling bitmap from a device memory module of the CXL device and loads the polling bitmap into a cache module.

[0101] Specifically, the polling bitmap is a bitmap with a fixed 64-byte group, where each byte is 8 bits. The polling bitmap has a total of 512 (64*8) bits, and each bit corresponds to an entry, so there are a total of 512 entries. Each entry represents an event, and the value of each bit is used to indicate the completion status of the event. For example, when the value of the bit corresponding to an event is 1, it indicates that the processing result of the event is in an unfinished state. When the value of the bit corresponding to an event is 0, it indicates that the processing result of the event is in a completed state.

[0102] It is understandable that the number of events supported simultaneously can be linearly expanded by increasing the number of polling bitmaps. For example, assuming that a polling bitmap currently supports 256 events, if the system needs to expand the number of concurrent processes, the number of polling bitmaps needs to be increased, and the number of events supported by the polling bitmap can be expanded to 512. The polling bitmap compacts event flags into a single cache line, which can minimize cache overhead.

[0103] The polling bitmap is set in the device memory module of the CXL device. The CXL device must support the CXL.mem protocol and set the polling bitmap's memory area to read-only. The CXL.mem protocol is part of the CXL protocol and is primarily used for communication between the CPU and memory. It is a transaction layer protocol that uses CXL's physical layer and link layer for cross-chip communication and is responsible for handling the transaction interface between the CPU and memory. With the support of the CXL protocol, the host CPU's access to the polling bitmap enters the CPU's cache module. After the first poll, when the host first polls the CXL device, it reads the polling bitmap from the CXL device's device memory module and loads it into the host's cache module.

[0104] It can be understood that the CPU's cache module refers to the cache space inside the CPU used to temporarily store data. The cache module includes at least two levels of cache. The cache size and cache levels are the core parameters of the processor. Different processors have different cache sizes and levels. For example, assuming that the cache level number of a processor's cache module is 3, the cache module includes a first-level cache (L1 cache), a second-level cache (L2 cache) and a third-level cache (L3 cache). The CPU's cache will automatically adjust the position of the data according to the access heat of the data. After the bitmap is locked, it will remain in at least the last layer. For example, data with high access heat will be placed in the first-level cache.

[0105] It is understandable that cache query is a hardware behavior of the CPU. When an event does not hit the data in the first-level cache, the CPU will automatically query downwards level by level. That is, when the host accesses the cache, it will first access the first-level cache. When the first-level cache is not hit, it will then access the second-level cache. When both the first-level cache and the second-level cache are not hit, the third-level cache will be accessed last. After the polling bitmap is loaded from the CXL device to the CPU's cache module, since the latency of polling to access memory is usually 100ns and the latency of accessing device registers is usually 1us, the next time the event is polled, if it hits the first-level cache, the result can be queried in about 2ns. By polling event information within the host's cache module, the polling overhead can be reduced by up to 50 times or 500 times, saving the host's access time and thereby improving the host's access efficiency.

[0106] Step S202: updating the polling bitmap in the cache module in real time;

[0107] For details, please refer to Figure 3 , Figure 3 yes Figure 2 A detailed flow chart of step S202 in FIG.

[0108] like Figure 3 As shown, step S202: real-time updating of the polling bitmap in the cache module includes:

[0109] Step S221: If the processing result of an event is completed, based on the reverse invalidation mechanism, the bit corresponding to the event in the polling bitmap of the cache module is set to 1 to update the polling bitmap in the cache module;

[0110] Specifically, the reverse invalidation mechanism refers to the CXL device proactively sending invalidation requests to the host via the CXL protocol, invalidating old data in the host cache, thereby maintaining data consistency between the CXL device and the host. This mechanism allows the device to more efficiently manage cache status, reduce host intervention, and improve overall performance. In the embodiments of the present application, the reverse invalidation mechanism is a mechanism by which the CXL device proactively maintains cache consistency within the CXL protocol. It notifies the host via reverse invalidation messages to update the completion status of events in the cache module, ensuring data consistency between the host and the CXL device while reducing latency and protocol overhead. The polling bitmap is updated by the CXL device via the reverse invalidation mechanism. The reverse invalidation mechanism is used to update the polling bitmap in the cache module after the polling bitmap in the device memory module is updated. Furthermore, the reverse invalidation mechanism is specifically used to: after the bit corresponding to the event in the polling bitmap in the device memory module is set to 1, the bit corresponding to the event in the polling bitmap in the cache module is set to 1, and the data corresponding to the event is marked from valid data to invalid data. For example:

[0111] If the current processing result of an event is incomplete, the bit corresponding to that event is zero. Once the processing result of the event is complete, the CXL device sets the corresponding bit in the cache module's polling bitmap from 0 to 1 according to the reverse invalidation mechanism. This means that the bit corresponding to the event in the cache module's polling bitmap is set from 0 to 1, and the reverse invalidation mechanism marks the data corresponding to the event in the host cache module from valid to invalid. For example, if the event is a write operation, and the host modifies the data in the CXL device corresponding to the write operation, indicating that the processing result of the event is complete, the CXL device sets the corresponding bit in the host cache module from 0 to 1 according to the reverse invalidation mechanism. Furthermore, since the data corresponding to the event has changed, the CXL device also marks the old data corresponding to the event in the host cache module as invalid data through the reverse invalidation mechanism.

[0112] It's understandable that because the data corresponding to the event in the host's cache module has been marked as invalid, the host will not be able to find the data corresponding to the event in the cache module during the next poll. Therefore, the host will access the device memory module of the CXL device. The latency of the host accessing the CXL device is typically 250ns. Through this access, the host can obtain the data corresponding to the completed event and pass it to the corresponding application or driver.

[0113] By placing the polling bitmap in a memory area that supports reverse invalidation, all invalid polling overhead can be significantly reduced. Only valid polling will enter the device memory module of the CXL device to obtain data. Invalid polling refers to the situation where the processing results of multiple query events are all incomplete. Valid polling refers to the situation where the processing results of the queried events are all completed. It is understandable that if the CPU repeatedly queries the event processing results as incomplete, it will lead to CPU resource waste. For example, in the traditional technical solution without the polling bitmap and reverse invalidation mechanism, if the processing result of a certain event is incomplete, the CPU needs to occupy CPU time each time it accesses the host memory. Polling with a short check interval will occupy a large amount of CPU time. Invalid polling (polling when the event is not completed) will lead to resource waste, thereby reducing the CPU processing efficiency. This application is based on a polling bitmap and reverse invalidation mechanism. All memory accesses by the CPU will first be queried in the polling bitmap in the cache module. If the event does not hit the cache module, the corresponding CXL device memory module will be accessed according to the address to obtain the data corresponding to the event from the CXL device. This makes it unnecessary for the processor to access the host memory during polling, thereby reducing the delay caused by accessing the host memory and improving the processing efficiency of the host processor.

[0114] Step S203: During the next polling, determine whether a certain event hits the cache module;

[0115] For details, please refer to Figure 4 , Figure 4 yes Figure 2 A detailed flow chart of step S203 in FIG.

[0116] like Figure 4 As shown, step S203: in the next polling, determining whether a certain event hits the cache module includes:

[0117] Step S231: Obtain data corresponding to the event in the cache module;

[0118] Specifically, based on the above-mentioned reverse invalidation mechanism, when the device updates the bitmap, the host's cache will be marked as invalid, forcing the CPU to obtain the latest data from the device during the next access, thereby avoiding invalid polling. During the next polling, the CPU will first obtain the data corresponding to the event in the cache module. It can be understood that the data corresponding to the event is valid data or invalid data. If the data corresponding to the event is valid data, it means that the event is not completed (the bit corresponding to the event in the polling bitmap is 0); if the data corresponding to the event is invalid data, it means that the event has been completed (the bit corresponding to the event in the polling bitmap is 1).

[0119] It can be understood that valid data refers to the bitmap data in the host's cache module being consistent with the bitmap data in the device memory module of the CXL device; while invalid data refers to the data in the host's cache module being outdated and needing to be updated. That is, valid data is used to indicate that the data corresponding to the event in the current cache module is consistent with the data corresponding to the event in the device memory module, and invalid data is used to indicate that the data corresponding to the event in the current cache module is inconsistent with the data corresponding to the event in the device memory module. When the CPU polls, if the data in the cache module is valid (not invalidated), it can be used directly; otherwise, it is necessary to access the device memory module of the CXL device to obtain the data corresponding to the event. That is, valid data means that the bitmap data in the cache is the latest data and its status can be trusted, while invalid data requires rechecking the device memory. Therefore, when the data in the cache module is valid data, the CPU does not need to frequently poll the device memory of the CXL device, thereby saving polling overhead.

[0120] Step S232: determining whether the data corresponding to the event in the cache module is valid data;

[0121] Specifically, when obtaining the data corresponding to the event in the cache module, the host determines whether the data corresponding to the event in the cache module is valid data through the kernel polling service. If the data corresponding to the event in the cache module is valid data, step S233 is entered; if the data corresponding to the event in the cache module is not valid data, step S234 is entered.

[0122] Step S233: Determine whether the event hits the cache module;

[0123] Specifically, if the data corresponding to the event in the cache module is valid, it indicates that the data corresponding to the event has not been invalidated and can be used directly. In other words, valid data means that the bitmap data in the cache is current and its state can be trusted. Therefore, the event is determined to have hit the cache module. It can be understood that when the data in the cache module is valid, the CPU does not need to frequently poll the device memory of the CXL device. Instead, it directly obtains the data corresponding to the event from the cache module, thereby reducing polling overhead.

[0124] Step S234: determining that the event does not hit the cache module;

[0125] Specifically, if the data corresponding to the event in the cache module is invalid data, it means that the data corresponding to the event has been invalidated, then it is determined that the event did not hit the cache module, and it is necessary to access the device memory module of the CXL device to obtain the data corresponding to the event. That is, when the data corresponding to the event is invalid data, it is necessary to recheck the device memory module of the CXL device, and then obtain the data corresponding to the event from the device memory module of the CXL device.

[0126] Step S204: Obtain data corresponding to the event from the CXL device;

[0127] Specifically, during the next host poll, if the event does not hit the cache, the host accesses the CXL device's device memory. The latency for host access to the device memory is typically 250ns. This access to the device memory allows the host to obtain the data for the completed event and pass it to the corresponding application or driver.

[0128] Step S205: Obtaining data corresponding to the event from the cache module;

[0129] Specifically, when the host polls next time, since the event hits the cache module, the host will directly obtain the data corresponding to the event from the cache module. Through this access to the cache module, the host can obtain the data corresponding to the event in the cache module and pass the information to the corresponding application or driver.

[0130] In the embodiment of the present application, by placing the polling bitmap in a memory area that supports reverse invalidation, all invalid polling overheads can be significantly reduced, and only valid polling will enter the device to obtain data.

[0131] Please refer to Figure 5 , Figure 5 This is a flowchart of obtaining a processing result of a first event provided by an embodiment of the present application;

[0132] like Figure 5 As shown, the process of obtaining the processing result of the first event includes:

[0133] Step S501: Send a first request to the CXL device;

[0134] Specifically, the CXL.io protocol is a physical layer interface defined in the Compute Express Link (CXL) specification that can provide lower latency, higher bandwidth, and better scalability than traditional PCIe. CXL.io uses SerDes technology (a technology that converts serial data to parallel data and vice versa) to simultaneously transmit multiple different data streams on a single physical channel. These data streams can include bandwidth-intensive data streams, low-latency command and control information, and configuration registers and status information.

[0135] In an embodiment of the present application, a user program or driver calls a kernel polling service interface to generate a first request and encapsulates the first request into a CXL.io transaction layer packet (TLP). The host sends the first request to the CXL device via the CXL.io protocol or other PCIe protocols. The first request is used to obtain a processing result of a first event from the CXL device. The processing result of the first event includes, but is not limited to, accelerator calculation completion, DMA transfer status, and the like. The information content of the first request includes, but is not limited to, the CPU ID of the first request, the requested page address, the flash memory read request ID sent to the backend, and the like.

[0136] Step S502: according to the first request, query the polling bitmap in the cache module, determine the first index position in the polling bitmap, and send the first event number corresponding to the first index position to the CXL device;

[0137] Specifically, for the first request that needs to wait for resources, the host will query the free event number in the polling bitmap and send it as the event number of the request to the device. In an embodiment of the present application, the host also maintains an event number bitmap, which is used to record the event numbers currently in use in the process of each CPU, wherein the event number bitmap corresponding to the event number in use is 1. The host first performs an AND operation on the event number bitmap and the polling bitmap, and then searches from the beginning to the end for the first position with a value of 0 to find the first unused first index position in the polling bitmap, wherein the first index position corresponds to a first event number, and the first index position is the position where the first bit in the polling bitmap is zero. For example, assuming that the host kernel polling service scans the first bit with a value of 0 in the bitmap (indicating an unused event number) and records its index position as N, the host uses index N as the first event number and marks the event number bitmap corresponding to the index position as "in use" to prevent duplicate allocation for other requests. The host then sends event number N to the event manager of the CXL device via the CXL.io protocol.

[0138] In some embodiments, the host may also maintain an event number usage table (e.g., a bit mask or hash table) to record currently allocated event numbers and their associated request contexts (e.g., process ID, callback function address). If all event numbers are used, the kernel polling service may trigger a blocked wait or return an error code (e.g., EBUSY).

[0139] Step S503: After the CXL device receives the first event number, obtain a processing result of the first event corresponding to the first event number sent by the CXL device;

[0140] Specifically, a CXL device includes an event manager and a CXL controller. Upon receiving an event number, the event manager performs a validity check, including a range check and a status check. The range check verifies whether the first event number is within the maximum event number range supported by the CXL device, and the status check verifies whether the first event number has been marked as "in use" by the host (to prevent duplicate processing or illegal access). After the validity check passes, the event manager determines that the received first request is valid and creates relevant information about the first request. The relevant information includes the CPU ID of the first request, the requested page address, and the flash memory read request ID sent to the backend. This information constitutes the data structure of an event, and the event element corresponding to the first request is written to the device memory.

[0141] When the task is completed or the resources are ready, the event manager will proactively set the bit corresponding to the event corresponding to the first request to 1. At this point, the CXL device will send the processing result of the first event corresponding to the first event number to the host. Furthermore, the CXL controller will pass a back-invalidation message to the host, that is, a back-invalidation message sent to the host via the CXL.cache protocol, invalidating the corresponding cache line in the host CPU cache. It will be appreciated that after the host receives the processing result of the first event corresponding to the first event number and the back-invalidation message, the CPU cache automatically marks the cache line, without CPU intervention.

[0142] Please refer to Figure 6 , Figure 6 This is a flowchart of determining whether the processing result of a first event is completed, provided by an embodiment of the present application;

[0143] like Figure 6 As shown, the process of determining whether the processing result of the first event is completed includes:

[0144] Step S601: Obtaining a processing result of a first event corresponding to a first event number sent by a CXL device;

[0145] Specifically, the host requests the CXL device through the kernel polling service to obtain the processing result of the first event corresponding to the first event number. The status of the processing result is stored in the polling bitmap in the memory of the CXL device. Each event number corresponds to a bit. The host obtains the processing result of the first event corresponding to the first event number sent by the CXL device by querying the bit corresponding to the first event number. The processing result of the first event includes but is not limited to DMA transfer completion and computing task readiness. The request can be initiated by the user program (such as calling the poll interface or io_uring interface) or by the driver's active query, which is not limited in this application.

[0146] Step S602: Determine whether the processing result of the first event is completed;

[0147] Specifically, after the host receives the processing result of the first event corresponding to the first event number sent by the CXL device, it determines whether the processing result of the first event is completed, that is, it determines whether the bit corresponding to the first event number in the polling bitmap of the device memory module is 1. If the bit corresponding to the first event number is 1, it is determined that the processing result of the first event is completed, and the process goes to step S603; if the bit corresponding to the first event number is 0, it is determined that the processing result of the first event is not completed, and the process goes to step S604.

[0148] Step S603: Using the reverse invalidation mechanism, the bit corresponding to the first event is set to 1 in the polling bitmap of the cache module to update the polling bitmap in the cache module;

[0149] Specifically, if the processing result of the first event is completed, the CXL device sets the bit corresponding to the first event in the polling bitmap of the cache module to 1 through the reverse invalidation mechanism, thereby updating the polling bitmap in the cache module. It can be understood that the reverse invalidation mechanism immediately invalidates the data corresponding to the first event in the host cache module, ensuring that the host reads the latest valid data corresponding to the first event from the device memory module of the CXL device during the next polling.

[0150] Step S604: Generate a timer task to periodically check the processing result of the first event until the processing result of the first event is completed;

[0151] Specifically, when the kernel polling service receives the returned first event number, it will register and generate a timer task for the process that initiated the blocking call. The timer task is used to periodically check whether the first event is completed, that is, to check the processing result of the first event at fixed intervals until the processing result of the first event is completed.

[0152] Step S605: After the processing result of the first event is completed, the bit corresponding to the first event in the polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module, and a check command is sent to the CXL device, so that the CXL device sets the bit corresponding to the first event to zero after receiving the check command.

[0153] Specifically, after querying that the processing result of the first event is completed, the bit corresponding to the first event in the polling bitmap of the cache module is set to one, that is, the bit corresponding to the first event is set from 0 to 1 to update the polling bitmap of the cache module, and the data corresponding to the first event is set to invalid data.

[0154] Furthermore, after the host updates the polling bitmap in the cache module, the host sends the bitmap of the read event to the event manager via a check command through the kernel polling service, wherein the check command is used to indicate that the host has completed checking the processing result of the first event, so that after receiving the check command, the CXL device sets the bit corresponding to the first event to 0 through the event manager and recovers the first event number corresponding to the first event.

[0155] Please refer to Figure 7 , Figure 7 This is a schematic diagram of a process for updating a polling bitmap of a cache module provided by an embodiment of the present application;

[0156] like Figure 7 As shown in FIG, the process of updating the polling bitmap of the cache module includes:

[0157] Step S701: After the processing results of at least two first events are completed, the bit corresponding to each first event in the polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module;

[0158] Specifically, the host can simultaneously receive the processing results of at least two first events. After the processing results of at least two first events are completed, the host sets the bit corresponding to each first event in the cache module's polling bitmap to 1, thereby updating the cache module's polling bitmap. For example, if the CXL device sends the processing results of three first events to the host, and all three first events are completed, the host can simultaneously receive the processing results of the three first events and set the bit corresponding to each first event in the cache module's polling bitmap to 1, thereby updating the cache module's polling bitmap. In embodiments of the present application, the host can also simultaneously check for multiple events and issue multiple check commands to the CXL device. Before the host checks, the CXL device's event manager can also update the polling bitmap multiple times and provide feedback to the host on the processing results of multiple events, thereby improving system data processing efficiency.

[0159] In an embodiment of the present application, a cache locking mechanism is provided. This mechanism locks the polling bitmap in the cache to prevent cache evictions due to excessively long polling intervals. Specifically, the cache module includes a first bitmap unit, which is used to cache the polling bitmap. The first bitmap unit is a fixed area within the cache module, and the polling bitmap in the first bitmap unit is not evicted during polling. The device memory module of the CXL device is read-only.

[0160] It is understandable that cache lock is a processor technology that allows users to select a memory area and keep it fixed in the cache so that it will not be evicted. In combination with this cache lock technology, the polling effect of the present application can be enhanced, that is, locking the cache line of the polling bitmap can avoid the situation where the bitmap is evicted during low-frequency polling. Caches often use strategies such as least recently used (LRU) or least frequently used (LFU) to evict cache lines. When an application uses a low-frequency polling method, the interval between two polls may be too long, and the polling bitmap has been evicted from the cache, so that the second polling needs to obtain data from the CXL device again. Based on this, the present application uses cache lock technology to lock the polling bitmap in the cache when the kernel polling service is initialized, so as to avoid cache eviction due to excessively large polling intervals.

[0161] In an embodiment of the present application, a polling method is provided. The method is applied to a host, wherein the host includes a processor, the processor includes a cache module, and the host is connected to a CXL device. The CXL device includes a device memory module, and the device memory module is used to store a polling bitmap. The method includes: when the host initially polls the CXL device, reading the polling bitmap from the device memory module of the CXL device and loading the polling bitmap into the cache module; updating the polling bitmap in the cache module in real time based on the result of event processing; and during the next polling, if an event hits the cache module, obtaining data corresponding to the event from the cache module; if the event does not hit the cache module, obtaining data corresponding to the event from the CXL device.

[0162] By setting a polling bitmap in the device memory module of the CXL device and using the reverse invalidation mechanism and event processing results to update the polling bitmap in the processor's cache module in real time, the host processor can process events without being interrupted. Moreover, during polling, if the event does not hit the processor's cache module, the data corresponding to the event is obtained from the CXL device, so that the processor does not need to access the host memory during polling, thereby reducing the delay caused by accessing the host memory and improving the processing efficiency of the host processor.

[0163] Please refer to Figure 8 , Figure 8 This is a schematic diagram of the structure of a host provided in an embodiment of the present application;

[0164] like Figure 8 As shown, the host 901 includes one or more processors 911 and a memory 912. Figure 8 A processor 911 is taken as an example.

[0165] The processor 911 and the memory 912 may be connected via a bus or other means. Figure 8 The bus connection is taken as an example.

[0166] Processor 911 is configured to provide computing and control capabilities to control host 901 to perform corresponding tasks. Processor 911 includes a cache module 9111, wherein cache module 9111 is configured to store a polling bitmap. For example, cache module 9111 controls host 901 to perform the polling method described in any of the above method embodiments. The polling method includes: when the host initially polls a CXL device, reading a polling bitmap from a device memory module of the CXL device and loading the polling bitmap into the cache module; updating the polling bitmap in the cache module in real time based on event processing results; and during the next polling, if an event hits the cache module, obtaining data corresponding to the event from the cache module; and if the event does not hit the cache module, obtaining data corresponding to the event from the CXL device.

[0167] By setting a polling bitmap in the device memory module of the CXL device and updating the polling bitmap in the processor's cache module in real time based on the event processing results, the host processor can process events without being interrupted. Furthermore, during polling, if the event does not hit the processor's cache module, the data corresponding to the event is obtained from the CXL device, so that the processor does not need to access the host memory during polling. This application can reduce the delay caused by accessing the host memory and improve the processing efficiency of the host processor.

[0168] Processor 911 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or any combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0169] The memory 912 is a non-transient computer-readable storage medium that can be used to store non-transient software programs, non-transient computer executable programs and modules, such as the program instructions / modules corresponding to the polling method in the embodiment of the present application. The processor 911 can implement the polling method in any of the above method embodiments by running the non-transient software programs, instructions and modules stored in the memory 912. Specifically, the memory 912 may include a volatile memory (VM), such as a random access memory (RAM); the memory 912 may also include a non-volatile memory (NVM), such as a read-only memory (ROM), a flash memory (flash memory), a hard disk drive (HDD) or a solid-state drive (SSD) or other non-transient solid-state storage device; the memory 912 may also include a combination of the above types of memories.

[0170] The memory 912 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 912 may optionally include a memory remotely located relative to the processor 911, and such remote memory may be connected to the processor 911 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0171] One or more modules are stored in the memory 912, and when executed by one or more processors 911, execute the polling method in any of the above method embodiments, for example, execute the above described Figure 2 The steps shown.

[0172] In the embodiment of the present application, the host 901 may also have components such as a wired or wireless network interface, a keyboard, and an input and output interface for input and output. The host 901 may also include other components for realizing device functions, which will not be described in detail here.

[0173] Please refer to Figure 9 , Figure 9 This is a schematic diagram of the structure of a polling system provided in an embodiment of the present application;

[0174] like Figure 9 As shown, the polling system 90 includes a host 901 and a CXL device 902, wherein the host 901 includes at least two processors, and the processors include a cache module;

[0175] The CXL device 902 is communicatively connected to the host 901, wherein the CXL device 902 includes:

[0176] Device memory module 921, the device memory module is used to store at least two polling bitmaps, wherein the polling bitmaps correspond one to one with the processor; specifically, the device memory module of the CXL device is a high-performance memory expansion hardware based on the Compute Express Link (CXL) protocol, which realizes efficient interconnection between the processor and memory resources through a standardized interface; the CXL memory module is usually composed of a DDR5 / DRAM chip, a CXL memory controller and a protocol interface, and supports cache coherence communication with the host processor. The device memory module is specifically used to store at least two polling bitmaps, wherein each polling bitmap has a one-to-one correspondence with the processor. Figure 1 One corresponds to a processor of a host. For example, the CPU0 polling bitmap corresponds to CPU0, and the CPU1 polling bitmap corresponds to CPU1. The polling bitmap is a bitmap of a fixed 64-byte group, where each byte is 8 bits. The polling bitmap has a total of 512 (64*8) bits, and each bit corresponds to an entry. There are a total of 512 entries, each entry represents an event, and the value of each bit is used to indicate the completion status of the event. For example, when the value of the bit corresponding to an event is 1, it indicates that the processing result of the event is in an unfinished state. When the value of the bit corresponding to an event is 0, it indicates that the processing result of the event is completed.

[0177] CXL controller 922 is configured to send a reverse invalidation message to the host via a reverse invalidation mechanism after updating bits in the polling bitmap of the device's memory module, thereby updating the polling bitmap in the host's cache module. Specifically, the reverse invalidation mechanism involves the CXL device proactively sending an invalidation request to the host via the CXL protocol to invalidate old data in the host's cache, thereby maintaining data consistency between the CXL device and the host. This mechanism allows the device to more efficiently manage cache status, reduce host intervention, and improve overall performance. In this embodiment of the present application, the reverse invalidation mechanism is a mechanism in the CXL protocol by which the CXL device proactively maintains cache consistency. It notifies the host of the completion status of events in the cache module via a reverse invalidation message, ensuring data consistency between the host and the CXL device while reducing latency and protocol overhead. The polling bitmap is updated by the CXL device via the reverse invalidation mechanism. The reverse invalidation mechanism is used to update the polling bitmap in the cache module after the polling bitmap in the device's memory module is updated.

[0178] Furthermore, the reverse invalidation mechanism is specifically configured to: after setting the bit corresponding to the event in the polling bitmap of the device memory module to 1, the CXL controller sends a reverse invalidation message to the host, so that upon receiving the reverse invalidation message, the host sets the bit corresponding to the event in the polling bitmap of the cache module to 1, and marks the data corresponding to the event from valid data to invalid data. For example, assuming that the current processing result of a certain event is incomplete, the bit corresponding to the event is zero. When the processing result of the event is completed, the CXL device will set the bit corresponding to the event to 1 according to the reverse invalidation mechanism, that is, set the bit corresponding to the event in the polling bitmap of the cache module from 0 to 1, and mark the data corresponding to the event in the host cache module from valid data to invalid data through the reverse invalidation mechanism. For example, assuming that this event is a write operation event, after the host modifies the data in the CXL device corresponding to the write operation event, the processing result of this event is completed. The CXL device will set the bit corresponding to this event in the host's cache module from 0 to 1 according to the reverse invalidation mechanism. In addition, since the data corresponding to this event has been changed, the CXL device will also mark the old data corresponding to this event in the host's cache module as invalid data through the reverse invalidation mechanism.

[0179] The event manager 923 is configured to, after the CXL device receives a check command issued by the host, reset a bit corresponding to a first event corresponding to the check command to zero. Specifically, after the host inquires that the processing result of the first event is completed, the event manager 923 resets the bit corresponding to the first event in the polling bitmap of the cache module to one, i.e., sets the bit corresponding to the first event from 0 to 1, thereby updating the polling bitmap of the cache module and setting the data corresponding to the first event to invalid data. Furthermore, after the host updates the polling bitmap in the cache module, the host sends a bitmap of read events to the event manager via a check command through the kernel polling service. The check command indicates that the host has completed the processing result of the first event, so that after receiving the check command, the CXL device, through the event manager, resets the bit corresponding to the first event corresponding to the check command to zero and reclaims the first event number corresponding to the first event.

[0180] Please refer to Figure 10 , Figure 10 This is a flowchart of another polling method provided in an embodiment of the present application;

[0181] In an embodiment of the present application, the host is communicatively connected to a CXL device, which includes a CXL SSD. The CXL SSD is a storage device that supports both CXL.io and CXL.mem, and can also interact with the host through the NVMe protocol.

[0182] This polling method is used for Figure 9 The polling system includes a host and a CXL device.

[0183] like Figure 10 As shown, the polling method includes:

[0184] Step S1001: When a host writes data to a CXL device, the host generates a first request and, based on the first request, generates a first command;

[0185] Specifically, the host is communicatively connected to the CXL device. When the host writes data to the CXL device, the user program or driver calls the kernel polling service interface to generate a first request. The host sends the first request to the CXL device through the CXL.io protocol or other PCIe protocols. The first request corresponds to a first event and is used to obtain a processing result of the first event from the CXL device. The processing result of the first event includes but is not limited to accelerator calculation completion, DMA transfer status, etc. The information content of the first request includes but is not limited to the CPU ID of the first request, the requested page address, the flash memory read request ID sent to the backend, and other information. Then, the NVMe driver encapsulates the first request into an NVMe command (transaction layer data packet TLP) according to the first request to generate a first command, wherein the first command includes an NVMe command.

[0186] Step S1002: The host sends a first command to the CXL device. After the CXL device receives the first command, the event manager establishes a mapping relationship between the first command and the event number.

[0187] Specifically, the CXL device includes an event manager. A host sends a first command to the CXL device via a PCIe physical link. After receiving the first command, the CXL device establishes a mapping relationship between the first command and an event number through the event manager. The mapping relationship is established by performing a modulo operation on the mapping relationship between the first command and the event number when the range of the command ID (65535) is greater than the range of a polling bitmap (512).

[0188] Step S1003: After the CXL device completes the first command, the event manager sets the bit corresponding to the first event in the polling bitmap of the device memory module to 1;

[0189] Specifically, after the CXL device completes the first command, the event manager in the CXL device sets the bit corresponding to the first event. Specifically, the event manager in the device's memory module sets the bit corresponding to the event in the polling bitmap from 0 to 1. For example, assuming the event is a write operation, after the host modifies the data in the CXL device corresponding to the write operation, the event is considered complete. The CXL device then sets the bit corresponding to the write operation in the polling bitmap in the device's memory module from 0 to 1.

[0190] Step S1004: After the bit corresponding to the first event is updated, a reverse invalidation message is sent to the host through the reverse invalidation mechanism to update the polling bitmap in the cache module of the host;

[0191] Specifically, after the event manager updates the bit corresponding to the first event, since the data corresponding to the event has changed, the CXL device also needs to mark the old data corresponding to the event in the host cache module as invalid data through the reverse invalidation mechanism. The CXL controller sends a reverse invalidation message to the host according to the reverse invalidation mechanism to set the bit corresponding to the first event in the host cache module from 0 to 1, thereby updating the polling bitmap in the host cache module.

[0192] Step S1005: polling the processing result of the first event on the host, and if the processing result of the first event is completed, sending the event bitmap to the CXL device;

[0193] Specifically, when the host polls the processing result of the first event, the kernel polling service in the host periodically polls the entire bitmap. If the processing result of the first event is completed, that is, the first command is completed, the relevant processing thread is awakened, and the event number of the completed request is sent to the event manager in bitmap form, that is, the event bitmap is sent to the CXL device, wherein the event bitmap is used to determine the first event whose processing result is completed.

[0194] Step S1006: The event manager sets the bit corresponding to the first event with a completed processing result to zero in the polling bitmap of the device memory module according to the event bitmap;

[0195] Specifically, the event manager parses the event bitmap sent by the host, and traverses all bits set to 1 (such as event number N) in the polling bitmap of the device memory module, releases the associated command context resources (such as memory buffer, task queue slot) for each completed event and recycles the data structure of the relevant event number, and sets the bit corresponding to the first event whose processing result is completed to zero.

[0196] In the embodiment of the present application, the polling bitmap and the reverse invalidation mechanism are used to improve the polling efficiency of the host, avoid the host processor spending a lot of time in invalid polling, and indirectly improve the CPU utilization of the entire polling system.

[0197] Please refer to Figure 11 , Figure 11 This is a flow chart of adding a first command to a sending queue provided by an embodiment of the present application;

[0198] like Figure 11 As shown, the process of adding the first command to the sending queue includes:

[0199] Step S1101: When the CXL device establishes communication with the host, the CXL device creates a send queue and a completion queue;

[0200] Specifically, when a CXL device establishes communication with a host, it creates a SubmissionQueue (SQ) and a CompletionQueue (CQ). The SubmissionQueue is used to send the first command to the CXL device. The host places the first command in the SubmissionQueue and writes to the CXL SSD's Doorbell Register via the CXL.io protocol to notify the CXL device of the command arrival. The CompletionQueue is a structure used by the CXL device to feedback command processing results to the host. The host confirms whether the command has been completed by reading the CQ entry in the CompletionQueue. The device reads the unprocessed commands in the SQ (from the current head pointer to the new tail pointer) and writes the results to the CompletionQueue (CQ) after execution is complete. The number of bits in the polling bitmap is equal to the queue depth of the CompletionQueue.

[0201] Step S1102: After the host generates a first command according to the first request, the host adds the first command to a sending queue;

[0202] Specifically, the host encapsulates the first request into an NVMe command (transaction layer data packet TLP) according to the first request through the NVMe driver to generate a first command, wherein the first command includes the NVMe command. After generating the first command, the host writes the first command to the next free entry of the send queue (based on the current tail pointer) through the CXL.mem protocol, wherein the write operation needs to ensure atomicity (such as 64-byte alignment) to avoid the device reading partially updated commands.

[0203] Please refer to Figure 12 , Figure 12 This is a schematic diagram of the overall process of data processing between a host and a CXL device provided by an embodiment of the present application;

[0204] like Figure 12 As shown in Figure 1, the overall process of data processing between the host and the CXL device includes:

[0205] Step S1201: Initialize the queue and polling bitmap;

[0206] Specifically, CXL devices include CXL SSDs. Taking the CXL SSD as an example, when the CXL SSD shakes hands with the host, it establishes a submission queue (SQ) and a completion queue (CQ). The host initializes the submission queue and completion queue and allocates a polling bitmap with the same depth as the completion queue to complete cache locking. The polling bitmap is stored in the device memory module of the CXL device and the host's cache module, and consistency must be maintained through CXL.cache.

[0207] Step S1202: the application writes data to the CXL device;

[0208] Specifically, the application in the host writes data to the CXL SSD. The application calls the write() system call, and the data is passed to the NVMe driver through the VFS layer, so that the driver allocates a physically continuous or scatter-gather (SGL) buffer and triggers a DMA write operation to the CXL device memory, thereby writing data to the CXL device.

[0209] Step S1203: The NVMe driver generates a first command and writes the first command into a sending queue;

[0210] Specifically, the host is communicatively connected to the CXL device. When the host writes data to the CXL device, the user program or driver calls the kernel polling service interface to generate a first request. The host sends the first request to the CXL device through the CXL.io protocol or other PCIe protocols. The first request corresponds to a first event and is used to obtain a processing result of the first event from the CXL device. The processing result of the first event includes but is not limited to accelerator calculation completion, DMA transfer status, etc. The information content of the first request includes but is not limited to the CPU ID of the first request, the requested page address, the flash memory read request ID sent to the backend, and other information. Then, the NVMe driver encapsulates the first request into an NVMe command (transaction layer data packet TLP) according to the first request to generate a first command, and writes the first command into a send queue (SQ). The first command includes an NVMe command.

[0211] Step S1204: The host writes the doorbell register via the CXL protocol;

[0212] Specifically, the doorbell register is used in NVMe to notify the CXL device that a new command has been submitted. The doorbell register is created by registering the MMIO address of the device doorbell (such as the HDM decoder mapping of CXL 2.0). The host writes a new SQ tail pointer value to the device doorbell address through the CXL.io protocol (based on the PCIe 5.0 PHY layer) so that when the first command is sent to the CXL device, the doorbell register notifies the CXL device that a command has arrived.

[0213] Step S1205: The event manager creates information related to the command and event number;

[0214] Specifically, the cache polling mechanism can be combined with the queue mechanism in the NVMe protocol. The event manager can assign an event number (Event ID) according to the command ID (Command ID) of the completion queue, that is, the event manager can assign a unique event number to each command, wherein the allocation method includes but is not limited to generation through an incremental counter or a hash function. The event manager creates an association between the event number and the command based on the command ID. Since the command ID continues to increase, there will be no overlap of event numbers between commands. The relevant information of the command and the event number includes the association information of the Event ID and the SQ entry, CQ status, callback function, etc. In the embodiment of the present application, a hash table or array is created to record the association information of the Event ID and the SQ entry, CQ status, callback function, etc.

[0215] Step S1206: the event manager creates a mapping relationship between the first command and the event number;

[0216] Specifically, the related information between the first command and the event number includes a mapping relationship. The CXL device includes an event manager. The host sends the first command to the CXL device via a PCIe physical link. After receiving the first command, the CXL device establishes a mapping relationship between the first command and the event number through the event manager. The mapping relationship is established by performing a modulo operation on the mapping relationship between the first command and the event number when a range of a command ID (65535) is larger than a range of a polling bitmap (512).

[0217] Step S1207: The event manager sets the bit to 1 after the command is completed;

[0218] Specifically, after the CXL device completes the first command, the event manager in the CXL device sets the bit corresponding to the first event. Specifically, the event manager in the device's memory module sets the bit corresponding to the event in the polling bitmap from 0 to 1. For example, assuming the event is a write operation, after the host modifies the data in the CXL device corresponding to the write operation, the event is considered complete. The CXL device then sets the bit corresponding to the write operation in the polling bitmap in the device's memory module from 0 to 1.

[0219] Step S1208: The kernel polls the service and finds that the event is completed, waking up the processing thread;

[0220] Specifically, the kernel polling service periodically polls the entire bitmap (e.g., once every 10 μs). When the NVMe command is completed, that is, when the bit corresponding to a certain event is detected to be 1, the relevant processing thread is woken up by the wake_up_interruptible() command, and the event number of the completed request is sent to the event manager in the form of a bitmap.

[0221] Step S1209: the host sends the processed event bitmap to the CXL device;

[0222] Specifically, the host writes the event bitmap into the device memory module of the CXL device through the CXL.mem protocol, so that the CXL device reads the event bitmap, wherein the event bitmap includes event IDs, and releases resources (such as DMA buffers) associated with these event IDs.

[0223] Step S1210: The event manager sets the bit corresponding to the first event to zero;

[0224] Specifically, the event manager recycles the data structure of the relevant event number, executes the atomic_clear_bit(Event_ID) command to set the bit corresponding to the above event number to 0, and deletes the Event ID entry from the hash table.

[0225] In the embodiments of the present application, through data interaction between the host and the CXL device, combined with the high bandwidth and low latency characteristics of the CXL protocol, the present application can achieve efficient data transmission and event management.

[0226] The present application also provides a non-volatile computer-readable storage medium, such as a memory including program code, wherein the program code can be executed by a processor to perform the polling method in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0227] The present application also provides a non-volatile computer-readable storage medium, such as a memory including program code, wherein the program code can be executed by a processor to perform the polling method in the above embodiment. For example, the non-volatile computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CDROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0228] The present application also provides a computer program product comprising one or more program codes stored in a non-volatile computer-readable storage medium. A processor of a flash memory device reads the program code from the non-volatile computer-readable storage medium and executes the program code to perform the steps of the polling method provided in the above embodiment.

[0229] Those skilled in the art will understand that all or part of the steps for implementing the above embodiments may be accomplished by hardware, or by hardware related to program code, and the program may be stored in a non-volatile computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk, or an optical disk, etc.

[0230] Through the description of the above embodiments, it can be clearly understood by those skilled in the art that each embodiment can be implemented by means of software plus a general hardware platform, or of course by hardware. It can be understood by those skilled in the art that all or part of the processes in the above embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0231] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Based on the idea of the present application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the present application as mentioned above. For the sake of simplicity, they are not provided in detail. Although the present application has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A polling method, characterized in that: The device is applied to a host, the host including a processor including a cache module, the host being connected to a CXL device, the CXL device including a device memory module, and the device memory module being used to store a polling bitmap; The method comprises: When the host polls the CXL device for the first time, reading the polling bitmap from the device memory module of the CXL device and loading the polling bitmap into the cache module; According to the processing result of the event, the polling bitmap in the cache module is updated in real time; During the next polling, if an event hits the cache module, data corresponding to the event is obtained from the cache module; If the event does not hit the cache module, data corresponding to the event is obtained from the CXL device.

2. The method according to claim 1, characterized in that The polling bitmap is updated by the CXL device through a reverse invalidation mechanism, wherein the reverse invalidation mechanism is used to update the polling bitmap in the cache module after the polling bitmap in the device memory module is updated; The polling bitmap includes a plurality of bits, each bit corresponding to an event; The updating of the polling bitmap in the cache module in real time according to the event processing result includes: If the processing result of a certain event is completed, based on the reverse invalidation mechanism, the bit corresponding to the event is set to one in the polling bitmap of the cache module to update the polling bitmap in the cache module.

3. The method according to claim 2, characterized in that The reverse invalidation mechanism is specifically used to: after the bit corresponding to the event in the polling bitmap of the device memory module is set to one, the bit corresponding to the event in the polling bitmap of the cache module is set to one, and the data corresponding to the event in the cache module is marked from valid data to invalid data, wherein the valid data is used to indicate that the data corresponding to the event in the current cache module is consistent with the data corresponding to the event in the device memory module, and the invalid data is used to indicate that the data corresponding to the event in the current cache module is inconsistent with the data corresponding to the event in the device memory module.

4. The method according to claim 1, wherein The method further comprises: Determining whether an event hits the cache module specifically includes: Determining whether the data corresponding to the event in the cache module is valid data; If so, determining that the event hits the cache module; If not, it is determined that the event does not hit the cache module, wherein when the data corresponding to the event is valid data, the value of the bit corresponding to the event is a preset value.

5. The method according to claim 1, wherein The method further comprises: sending a first request to the CXL device, wherein the first request corresponds to a first event; querying a polling bitmap in the cache module based on the first request to determine a first index position in the polling bitmap, and sending a first event number corresponding to the first index position to the CXL device, wherein the first index position is a position in the polling bitmap where the first bit is zero; After the CXL device receives the first event number, a processing result of the first event corresponding to the first event number sent by the CXL device is obtained.

6. The method according to claim 5, characterized in that The method further comprises: If the processing result of the first event is completed, setting a bit corresponding to the first event in the polling bitmap of the cache module to 1 through a reverse invalidation mechanism to update the polling bitmap in the cache module; If the processing result of the first event is incomplete, generating a timer task to periodically check the processing result of the first event until the processing result of the first event is completed; After the processing result of the first event is completed, a bit corresponding to the first event in a polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module, and a check command is sent to the CXL device, so that the CXL device sets the bit corresponding to the first event to 0 after receiving the check command.

7. The method according to claim 6, characterized in that The check command includes a polling bitmap, the polling bitmap includes at least two bits, and each bit corresponds to a first event; The method further comprises: After the processing results of at least two first events are completed, the bit corresponding to each first event in the polling bitmap of the cache module is set to 1 to update the polling bitmap of the cache module.

8. The method according to any one of claims 1 to 7, characterized in that The cache module includes a first bitmap unit, wherein the first bitmap unit is used to cache the polling bitmap, and the first bitmap unit is a fixed area in the cache module, wherein the polling bitmap in the first bitmap unit is not evicted during polling; The device memory module of the CXL device is read-only.

9. A host, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

10. A polling system, characterized in that: include: The host of claim 9, wherein the host comprises at least two processors, the processors comprising a cache module; A CXL device is connected to the host, wherein the CXL device includes: a device memory module, the device memory module being configured to store at least two polling bitmaps, wherein the polling bitmaps correspond one-to-one to the processors; an event manager configured to, after the CXL device receives a check command sent by the host, set a bit corresponding to a first event corresponding to the check command to zero; The CXL controller is configured to send a reverse invalidation message to the host through a reverse invalidation mechanism after the bits of the polling bitmap of the device memory module are updated, so as to update the polling bitmap in the cache module of the host.

11. A polling method, characterized in that: Applied to the polling system according to claim 10, the method comprises: When the host writes data to the CXL device, the host generates a first request and generates a first command according to the first request, wherein the first request corresponds to a first event; The host sends the first command to the CXL device. After the CXL device receives the first command, the event manager establishes a mapping relationship between the first command and the event number. After the CXL device completes the first command, the event manager sets a bit corresponding to the first event in a polling bitmap of a memory module of the device to one; After the bit corresponding to the first event is updated, sending a reverse invalidation message to the host through a reverse invalidation mechanism to update the polling bitmap in the cache module of the host; The host polls the processing result of the first event, and if the processing result of the first event is completed, sends an event bitmap to the CXL device, wherein the event bitmap is used to determine the first event whose processing result is completed; The event manager sets to zero the bit corresponding to the first event with a completed processing result in the polling bitmap of the device memory module according to the event bitmap.

12. The polling method according to claim 11, wherein: The method further comprises: When the CXL device establishes communication with the host, the CXL device creates a send queue and a completion queue, wherein the number of bits in the polling bitmap is equal to the queue depth of the completion queue; After the host generates a first command according to the first request, the host adds the first command to the sending queue.