A prefetch control method and device for CXL memory devices
Patent Information
- Application Number
- CN202610201661.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-11
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-02-11
AI Technical Summary
[0007]本发明旨在解决现有CXL内存设备侧预取方案中,因缺乏统一的在途请求状态跟踪与去重机制,导致冗余预取请求生成;且因无法感知系统负载而盲目预取,缺乏自适应节流机制,进而造成高负载下正常访存延迟加剧、系统性能下降的技术问题,该问题存在于CXL内存设备及具备请求代理能力的中间节点中
[0030] 1. This invention maintains the register unit in a missing state to track and manage all in-transit requests in a unified manner, effectively realizing the merging and deduplication of memory access requests for the same cache line, fundamentally eliminating the waste of backend storage bandwidth caused by redundant prefetch requests, and reducing unnecessary queuing delays.
Smart Images

Figure CN122195875B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer storage, and more particularly to a prefetch control method and apparatus for CXL memory devices. Background Technology
[0002] As drones and other aerial platforms become increasingly intelligent and complex in their missions, the demand for memory capacity and bandwidth in their onboard computing systems is growing dramatically, while simultaneously facing stringent energy constraints and real-time requirements. CXL (ComputeExpress Link), a high-speed coherent interconnect protocol built on the PCIe physical layer, offers a new solution to these challenges. CXL provides efficient memory expansion and resource pooling solutions for memory-intensive applications. Furthermore, CXL supports the expansion of non-volatile memory devices and enables rapid saving and restoration of the entire system's operating state, potentially helping drone platforms achieve "zero-power" deep hibernation of their computing systems during mission intervals, thereby effectively extending their flight endurance.
[0003] However, CXL memory devices, especially those based on non-volatile media, typically have significantly higher access latency than processor-local DRAM (Dynamic Random Access Memory). This characteristic drastically reduces the performance of traditional DRAM-optimized hardware prefetching mechanisms on CXL memory devices, failing to effectively hide data access latency and thus becoming a key bottleneck restricting the performance potential of CXL technology in high-energy-efficiency, high-real-time over-the-air scenarios.
[0004] To solve the above problems, one feasible technical approach is to adopt a device-side prefetching scheme, which involves pushing the prefetching logic down to the controller of the CXL memory device and loading data into the buffer of the CXL memory device in advance through the prefetcher, thereby shortening the data access path of the processor. However, although this scheme can improve the access latency of the CXL memory device, it may lead to limited access bandwidth of the CXL memory device in actual application scenarios. There are two main influencing factors: (1) The device-side prefetcher may generate redundant prefetch requests. Without unified in-transit status tracking and deduplication, the prefetcher may generate prefetch requests with the same access address as prefetch requests that are already in transit or requests issued by the processor, repeatedly fetching the same cache line, which wastes the access bandwidth of the memory device's back-end storage medium and increases queuing latency; (2) The prefetching strategy is not flexible. The prefetcher usually adopts a static strategy, which makes it difficult to perceive the system load status such as CXL link utilization and prefetching demand queue depth in real time. Under high load, it still aggressively issues prefetch requests, preempting the available bandwidth of the processor's normal memory access requests, which may increase the overall latency of application access data and reduce system performance.
[0005] Therefore, there is an urgent need in the field for a prefetch control method and apparatus for CXL memory devices that can effectively sense the global system status and real-time load, dynamically optimize prefetch decisions, thereby improving data access latency while minimizing negative competition for the available bandwidth of CXL memory devices. Summary of the Invention
[0006] (a) Technical problems to be solved
[0007] This invention aims to solve the technical problems in the existing CXL memory device-side prefetching scheme, which are caused by the lack of a unified in-transit request status tracking and deduplication mechanism, resulting in the generation of redundant prefetch requests; and the blind prefetching due to the inability to perceive system load and the lack of an adaptive throttling mechanism, which leads to increased normal memory access latency and degraded system performance under high load. This problem exists in CXL memory devices and intermediate nodes with request proxy capabilities.
[0008] (II) Technical Solution
[0009] To address the aforementioned issues, this invention proposes a prefetch control method and apparatus for CXL memory devices. The core idea is to deploy integrated functional modules in the CXL memory device controller or a CXL switch controller with memory proxy functionality to achieve unified tracking and deduplication of in-transit requests. Simultaneously, based on multi-dimensional system state parameters, the prefetch strategy and scheduling weights are dynamically adjusted to balance prefetch efficiency and bandwidth usage, thereby ensuring overall system performance.
[0010] First, this invention proposes a prefetch control method for CXL memory devices. This method is executed in a CXL memory device controller or a CXL switch with memory proxy functionality, and specifically includes the following steps:
[0011] S1 receives memory access requests, i.e., demand requests, initiated by the host processor through the CXL memory device response port.
[0012] S2, upon receiving a memory access request, the system initiates multiple processes in parallel: the link bandwidth statistics unit records the statistical information of the request for bandwidth and load assessment. This unit adopts a window mechanism based on the number of memory access requests. When the cumulative number of memory access requests reaches the preset window size, it sends a window end notification to the prefetch throttling control unit and resets the counter; the prefetcher generates a prefetch request based on the address of the memory access request; the prefetch buffer queries the address to determine if a fast path has been hit; the missing status holding register unit detects the in-transit status of the same address to avoid subsequent duplicate memory access requests.
[0013] S3, the arbitration decision unit performs a fast path decision based on the query results of the prefetch buffer and the missing state holding register unit.
[0014] S4, to be handled according to the judgment result:
[0015] If the prefetch buffer hit is determined, the hit data is directed to the return queue and quickly returned to the host processor, and the missing status holding register unit is notified that no request entry needs to be allocated for this request.
[0016] If the prefetch buffer is missed but the prefetch request entry corresponding to the address exists in the missing status holding register, then it is further determined whether the prefetch request has been scheduled by the weighted scheduling unit: if it has not been scheduled, the missing status holding register unit sends the demand request to the demand queue and notifies the weighted scheduling unit to cancel the scheduling of the original prefetch request; if it has been scheduled, the missing status holding register unit marks the prefetch entry as a demand triggered state.
[0017] If the decision indicates a prefetch buffer miss and there is no relevant entry in the missing state holding register unit, then the request will create an entry from the missing state holding register unit and submit it to the request queue.
[0018] S5, for requests entering the slow path, the missing state holding register unit performs merging and deduplication of the demand and prefetch requests to ensure that no duplicate memory access operations are issued for the same cache line, and then the unique request is distributed to the demand queue or the prefetch queue.
[0019] S6, the weighted scheduling unit selects the next request to be served from the demand queue and the prefetch queue according to the scheduling weight set by the prefetch throttling control unit; the weighted scheduling unit maintains a list of canceled scheduling flags, identifies and discards invalid prefetch requests, and sends valid requests to the backend access port for actual data access; if the scheduled request comes from the prefetch queue, the weighted scheduling unit sends a scheduled notification to the missing status holding register unit.
[0020] S7 After the data is returned from the backend access port, the missing status holding register unit accurately writes the data into the prefetch buffer or the guided return queue according to the type of the original request.
[0021] S8, the prefetch throttling control unit receives the window end notification and collects the current system status parameters, including link bandwidth utilization, prefetch accuracy, prefetch coverage, prefetch timeliness, missing status hold register occupancy rate, and demand queue depth; when deployed on a CXL switch, the parameters additionally include the remote device response latency statistically analyzed by the backend access port.
[0022] S9, the prefetching throttling control unit calculates a new prefetching intensity level and scheduling weight based on the collected parameters, generates control signals and sends them in parallel to the prefetcher and the weighted scheduling unit, so as to dynamically adjust the prefetching aggressiveness of the prefetcher and adjust the scheduling weight of the weighted scheduling unit. When the system load is high, priority is given to ensuring access to the demand queue and suppressing the competition of the prefetching queue.
[0023] Secondly, based on the aforementioned method, this invention proposes a prefetch control device for CXL memory devices. This device is a configurable hardware system integrated into a CXL memory expansion card (CXL Type 3 device) or a CXL switch with memory proxy functionality, used to implement the aforementioned prefetch control method. The device mainly includes four functional modules: a prefetch processing module, a state management module, a scheduling and control module, and a data path module.
[0024] The prefetching module is responsible for capturing memory access streams from the host processor and generating prefetch requests. Internally, this module contains two sub-units: a prefetcher 104 and a prefetch buffer 105. The prefetcher 104 analyzes access patterns and predicts future memory access addresses; its behavior is dynamically adjusted by control signals from the prefetch throttling control unit 108. The prefetch buffer 105 acts as a fast buffer, storing the data loaded by the prefetch requests issued by the prefetcher 104 and providing address lookup functionality for a fast response to subsequent requests.
[0025] The state management module is responsible for the unified tracking and management of all in-transit requests. This module consists of an arbitration decision unit 106 and a missing state holding register unit 107, which work together to manage request states. The arbitration decision unit 106 receives query results from the prefetch buffer 105 and state information from the missing state holding register unit 107, performs fast path decision, and issues instructions to the missing state holding register unit 107 to cancel redundant entries, trigger scheduling queries, or submit new requests. The missing state holding register unit 107, as the core state recording unit, is responsible for merging and deduplicating requests and queue distribution and data return routing for both request and prefetch requests, ensuring that duplicate memory access operations are not issued for the same cache line.
[0026] The scheduling and control module is responsible for dynamically adjusting resource allocation strategies based on the current operating status of the system. This module integrates a weighted scheduling unit 111, a prefetch throttling control unit 108, and a link bandwidth statistics unit 102. The weighted scheduling unit 111 selects the next request to be served from the demand queue 109 and the prefetch queue 110 according to the weights set by the prefetch throttling control unit 108. The prefetch throttling control unit 108, as a control unit, receives multi-dimensional feedback signals from the link bandwidth statistics unit 102, the prefetch buffer 105, the missing status holding register unit 107, and the demand queue 109; when deployed on a CXL switch, it additionally receives remote device response latency parameters statistically collected from the backend access port 112; and outputs control commands accordingly to dynamically adjust the behavior of the prefetcher 104 and the weights of the weighted scheduling unit 111. The link bandwidth statistics unit 102 continuously monitors the bandwidth usage and memory access frequency of the front-end CXL link. It adopts a window mechanism based on the number of demand requests. When the cumulative number of requests reaches the preset window size, it sends a window end notification to the prefetch throttling control unit 108 to provide a time reference for adaptive control.
[0027] The data path module is responsible for the input, output, and temporary storage of data between various modules. This module includes a CXL memory device response port 101, a backend access port 112, a return queue 103, a demand queue 109, and a prefetch queue 110. The CXL memory device response port 101 is the interface between the system and the host processor, responsible for receiving all memory access requests. The backend access port 112 is the system's backend communication interface, and its physical connection method is dynamically configured according to the deployment node: in the CXL memory device, it connects to the DRAM controller; in the CXL switch, it connects to the CXL link controller, used to access remote CXL memory devices and enable their integrated counters to collect remote device response latency parameters and feed them back to the prefetch throttling control unit 108. The return queue 103 stores all data to be returned to the host processor. The demand queue 109 and the prefetch queue 110 are used to temporarily store raw demand requests and prefetch requests requiring backend access, respectively, awaiting scheduling by the weighted scheduling unit 111.
[0028] (III) Beneficial Effects
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] 1. This invention maintains the register unit in a missing state to track and manage all in-transit requests in a unified manner, effectively realizing the merging and deduplication of memory access requests for the same cache line, fundamentally eliminating the waste of backend storage bandwidth caused by redundant prefetch requests, and reducing unnecessary queuing delays.
[0031] 2. This invention uses a prefetch throttling control unit to collect multi-dimensional system load parameters in real time, such as link bandwidth utilization, queue depth, prefetch accuracy, prefetch coverage, and prefetch timeliness. Based on these parameters, it dynamically adjusts the aggressiveness of the prefetcher and the weight of the weighted scheduling unit, enabling prefetching behavior to adapt to changes in system load. Under high load, it automatically suppresses prefetching, prioritizing bandwidth requests from the processor, avoiding vicious competition between prefetch requests and normal memory access, and ensuring the overall service quality and performance stability of the system.
[0032] 3. The device and method proposed in this invention can be flexibly deployed in CXL memory device controllers or CXL switches, exhibiting strong architectural versatility. Especially when deployed on CXL switches, by introducing remote device response latency as a control parameter, it can better address latency uncertainties in multi-level interconnection scenarios, resulting in more significant optimization effects.
[0033] 4. This invention achieves a fast path response to prefetch hit requests through the cooperation of the arbitration decision unit and the prefetch buffer, maximizing the latency-hiding benefits of prefetching and improving data access efficiency. Attached Figure Description
[0034] Figure 1 This is a diagram of the prefetch throttling control device on the CXL memory device side;
[0035] Figure 2 This is a request processing flowchart;
[0036] Figure 3 This is a flowchart of the arbitration decision-making process;
[0037] Figure 4 This is a flowchart of the scheduling request process;
[0038] Figure 5 This is a flow chart for throttling control. Detailed Implementation
[0039] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0040] Example 1: System Architecture Description
[0041] like Figure 1As shown, this embodiment provides a prefetch control device for CXL memory devices, which can be configured in any of the following nodes: (a) a CXL memory device controller, directly connected to and managing the physical DRAM media; or (b) a CXL switch controller with memory proxy functionality, wherein the CXL switch integrates a local cache and a request proxy engine. The system includes the following functional modules:
[0042] CXL memory device response port 101 serves as the interface for interaction between the system and the host processor. It is responsible for receiving memory access requests initiated by the host processor and distributing the requests in parallel to the link bandwidth statistics unit 102, prefetcher 104, prefetch buffer 105, and missing status holding register unit 107. This parallel processing mechanism ensures that the system can simultaneously perform traffic statistics, prefetch prediction, cache lookup, and status tracking, providing a basis for subsequent fast path decision-making.
[0043] The link bandwidth statistics unit 102 receives request information from the CXL memory device response port 101, records request traffic data, and calculates the bandwidth utilization of the CXL link in real time. Simultaneously, the link bandwidth statistics unit 102 employs a window mechanism based on the number of demand requests. When the cumulative number of memory access requests reaches a preset window size (e.g., 4096 times), it sends a window end notification to the prefetch throttling control unit 108 and resets the counter, providing a time reference for adaptive control.
[0044] Prefetcher 104 analyzes access patterns and predicts addresses that may be accessed in the future based on received request addresses, and generates prefetch requests. Prefetcher 104 can be configured as different types of prefetchers, such as Stream, SPP, BOP, SandBox, MLOP, etc. Prefetcher 104 receives control signals from prefetch throttling control unit 108 and can dynamically adjust prefetch behavior, including enabling / disabling prefetching function and adjusting prefetch aggressiveness parameters such as prefetch step size, to adapt to different system load conditions.
[0045] The prefetch buffer 105 serves as a high-speed cache unit, storing data pre-loaded by the prefetcher 104 but not yet hit, and provides an address lookup function to determine whether the current request has been hit in the buffer. The hit result of the prefetch buffer 105 directly affects the fast path decision of the arbitration decision unit 106, and is a key component for achieving low-latency response.
[0046] The missing status holding register unit 107, as the core status management unit of the system, is responsible for tracking the lifecycle of all in-transit requests (including original demand requests and prefetch requests), performing request deduplication, and providing status query services to the arbitration decision unit 106. The missing status holding register unit 107 can receive instructions from the arbitration decision unit 106 to cancel entries for specific requests, check if a prefetch request entry exists at a specific address, and query whether a prefetch request has been scheduled by the weighted scheduling unit 111. When a request needs to be submitted to the backend, the missing status holding register unit 107 distributes the request to the demand queue 109 or the prefetch queue 110 according to the arbitration decision result, and can notify the weighted scheduling unit 111 not to schedule prefetch requests associated with a specific address. When data is returned, the missing status holding register unit 107 writes the data to the prefetch buffer 105 or directs it to the return queue 103 according to the original request type.
[0047] Arbitration decision unit 106 receives the query results from prefetch buffer 105 and missing status register unit 107, performs fast path decision, and determines the processing path of the request. Specifically, arbitration decision unit 106 first determines whether prefetch buffer 105 is hit; if hit, data is returned directly; if miss, it checks whether there is a prefetch request entry for that address in missing status register unit 107; if so, it further determines whether the prefetch request has been scheduled. Based on these determinations, arbitration decision unit 106 coordinates the direct return of data, waiting for data in transit, or submitting a new request.
[0048] The prefetch throttling control unit 108 acts as a global adaptive controller, receiving status information from the link bandwidth statistics unit 102, prefetch buffer 105, missing status holding register unit 107, and demand queue 109. This information includes parameters such as link bandwidth utilization, prefetch accuracy, prefetch coverage, prefetch timeliness, missing status holding register occupancy rate, and queue depth. When deployed on a CXL switch, it additionally receives remote device response latency parameters from the backend access ports. At the end of each statistical window, the prefetch throttling control unit 108 calculates a new prefetch strength level and scheduling weight based on the collected multidimensional parameters and historical decisions, generating control signals that are sent in parallel to the prefetcher 104 and the weighted scheduling unit 111.
[0049] The weighted scheduling unit 111 selects requests from the demand queue 109 and the prefetch queue 110 based on the scheduling weights provided by the prefetch throttling control unit 108, and sends them to the backend access port 112. At the beginning of each scheduling window, the weighted scheduling unit 111 updates the scheduling weights and selects candidate requests based on the weights and queue status. Simultaneously, the weighted scheduling unit 111 maintains a list of canceled scheduling flags, enabling it to identify and discard invalid prefetch requests. When scheduling a prefetch request, the weighted scheduling unit 111 sends a scheduled notification to the missing status holding register unit 107, updating the status of in-transit requests.
[0050] Backend access port 112 is the communication interface between the system and the backend storage entity. The specific physical object it connects to depends on the system deployment node: when the system is deployed on a CXL memory device, this port connects to the DRAM controller, responsible for the actual data read and write operations; when the system is deployed on a CXL switch, this port connects to the CXL link controller, used to forward requests to remote CXL memory devices and receive their returned data, as well as to calculate the remote device response latency parameters and feed them back to the prefetch throttling control unit. The returned data is returned to the missing status holding register unit 107 via backend access port 112, where it is distributed and processed according to the request type.
[0051] Demand queue 109 and prefetch queue 110 are used to temporarily store original demand requests and prefetch requests that require backend access, respectively, awaiting scheduling by weighted scheduling unit 111. Return queue 103 is used to temporarily store data to be returned to the host processor.
[0052] Example 2: Request Processing Flow
[0053] like Figure 2 As shown, this embodiment details the complete processing flow of host processor memory access requests in a CXL memory device or a CXL switch with memory proxy functionality. The specific steps are as follows:
[0054] 201. The system receives the host processor's request through the CXL memory device response port 101.
[0055] 202. The system initiates multiple processes in parallel: the link bandwidth statistics unit 102 records traffic information; the prefetcher 104 generates a prefetch request based on the requested address; the prefetch buffer 105 performs cache lookup on the address; and the missing status holding register unit 107 performs in-transit status lookup on the requested address.
[0056] 203. The arbitration decision unit 106 summarizes the query results of the prefetch buffer 105 and the missing status holding register unit 107, and the execution path returns the arbitration decision.
[0057] 204. Determine the arbitration decision result. If it is a pre-fetch hit, proceed to step 205; if it is a hit in transit or a complete miss, proceed to step 206.
[0058] 205, the request is served by the fast path, and the data is returned to the host processor from the return queue 103 via the CXL memory device response port 101, and the process ends.
[0059] 206. Requests to enter the slow path service are managed uniformly by the missing state holding register unit 107, including deduplication and distributing the unique prefetch request or demand request to the prefetch queue 110 or the demand queue 109.
[0060] 207. The system executes the request scheduling process and sends the request to the backend access port 112.
[0061] 208, Data is returned from the backend access port to the missing status holding register unit 107.
[0062] 209. The missing status holding register unit distributes data according to the request type: if it is prefetch data, it is written to the prefetch buffer 105; if it is demand data, it is sent to the return queue 103.
[0063] 210. If data enters the return queue 103, it is returned to the host processor via the CXL memory device response port 101.
[0064] Example 3: Arbitration Decision-Making Process
[0065] like Figure 3 As shown, this embodiment describes in detail the workflow of the arbitration decision unit 106, and the specific steps are as follows:
[0066] 301, the arbitration decision unit 106 receives the query results from the prefetch buffer 105 and the status information from the missing status holding register unit 107.
[0067] 302. Determine if the prefetch buffer 105 is hit. If it is hit, proceed to step 303; otherwise, proceed to step 304.
[0068] 303, redirect the corresponding data to the return queue 103, and notify the missing status holding register unit 107 that it is not an allocation request entry for this request.
[0069] 304. Determine whether there is a prefetch request entry for the address in the missing status holding register unit 107. If it exists, proceed to step 305; otherwise, proceed to step 307.
[0070] 305. Check whether the prefetch request has been scheduled by the weighted scheduling unit 111. If it has been scheduled, proceed to step 306; otherwise, proceed to step 307.
[0071] 306, waiting for the data from the prefetch request to be returned.
[0072] 307, the missing status holding register unit 107 submits the request as a new demand request to the demand queue 109 and notifies the weighted scheduling unit 111 not to schedule a prefetch request with the same address.
[0073] Example 4: Weighted Scheduling Process
[0074] like Figure 4 As shown, this embodiment describes in detail the workflow of the weighted scheduling unit 111, and the specific steps are as follows:
[0075] 401, When the new window begins, the weighted scheduling unit 111 updates the scheduling weights from the prefetch throttling control unit 108.
[0076] 402, The weighted scheduling unit 111 selects a candidate request from the demand queue 109 or the prefetch queue 110 according to the current scheduling weight. If the selected queue is empty, another queue is selected.
[0077] 403, The weighted scheduling unit 111 checks whether the selected candidate request needs to be invalidated. If the request needs to be invalidated, the request is discarded; otherwise, the request is confirmed as the service object for this cycle.
[0078] 404, the weighted scheduling unit 111 sends the confirmed valid request to the backend access port 112. If the scheduled request comes from the prefetch queue 110, the weighted scheduling unit 111 sends a "scheduled notification" to the missing status holding register unit 107.
[0079] Example 5: Prefetch Throttling Control Process
[0080] like Figure 5 As shown, this embodiment describes in detail the prefetch throttling control process of the prefetch throttling control unit 108. The specific steps are as follows:
[0081] 501, Link bandwidth statistics unit 102 records the number of memory access requests passing through CXL memory device response port 101.
[0082] 502. Determine whether the number of memory access requests has reached the preset window size (e.g., 4096 times). If it has, proceed to step 503; otherwise, continue counting.
[0083] 503, the link bandwidth statistics unit 102 notifies the prefetch throttling control unit 108 that the current window has ended and resets the counter.
[0084] 504. After receiving the notification, the prefetch throttling control unit 108 collects the current system status parameters, including link bandwidth utilization, prefetch accuracy, prefetch coverage, prefetch timeliness, missing status holding register unit 107 occupancy rate, and demand queue 109 depth. If this device is deployed on a CXL switch, the parameters also include remote device response latency.
[0085] The prefetch throttling control unit 108 determines the state value or adjusts the action based on the state parameters using a weighted calculation strategy or a reinforcement learning strategy.
[0086] Method 1 (when using a weighted calculation strategy):
[0087] Calculate the state value using the following formula. :
[0088] If deployed on a CXL memory device:
[0089]
[0090] When deployed on a CXL switch:
[0091]
[0092] in These represent the prefetch accuracy score, prefetch coverage score, and prefetch timeliness score, respectively. The missing state holds the register occupancy rate score; This indicates a score for link bandwidth utilization. Indicates the depth score of the demand queue; Indicates the response latency score of remote devices; This indicates the prefetched intensity level of the previous window; These are the weighting coefficients for each performance indicator.
[0093] Method 2 (when using reinforcement learning strategies):
[0094] 1. State Construction: The prefetch throttling control unit 108 discretizes and encodes the collected parameters to construct the state feature vector of the current window. Specifically, link bandwidth utilization, missing state holding register occupancy, and demand queue depth are divided into several congestion levels (if this device is deployed in a switch, remote device response latency should also be included in the congestion level classification); prefetch accuracy, prefetch coverage, and prefetch timeliness are divided into several performance levels; and the prefetch strength level of the previous window is also considered. Incorporate the feature vector.
[0095] 2. Decision Model: The pre-fetching throttling control unit 108 internally maintains a state-action value table (Q-Table) or a neural network model. The model is based on the current state feature vector. Output the corresponding action command The action command The space includes {-1, 0, +1}, which correspond to the adjustment values for decreasing, maintaining, or increasing the pre-selected intensity level, respectively.
[0096] 3. Reward Mechanism: The system calculates reward values based on feedback after an action is executed. Positive rewards are given when prefetch accuracy, coverage, or timeliness improves; negative penalties are given when link bandwidth utilization, missing state hold register occupancy, or demand queue depth exceeds a preset threshold. The model outputs the optimal action instruction as an adjustment value by maximizing the cumulative reward value.
[0097] 505, the pre-fetch throttling control unit 108 based on the status value Determine the prefetch intensity level of the current window. ,in Defined by a three-digit saturation counter, its value ranges from 0 to 7. Specifically, the adjustment value is first obtained:
[0098] If method one is used, then the state value calculated in step 504 will be... Adjustment values mapped to -1, 0, or +1;
[0099] If method two is adopted, the adjustment value output by the reinforcement learning model in step 504 is used directly.
[0100] Subsequently, the adjustment value is applied to a two-digit cyclic counter. When the two-digit cyclic counter reaches state 11 and the adjustment value is +1, When the two-bit cyclic counter reaches 00 and the adjustment value is -1, ;otherwise .
[0101] 506, the pre-fetch throttling control unit 108 converts the calculated parameters into control signals.
[0102] 507, the prefetch throttling control unit 108 sends control signals in parallel to the prefetcher 104 and the weighted scheduling unit 111.
[0103] The above description is merely an example embodiment of the present invention and is not intended to limit the present invention in any way. Any simple modifications, equivalent changes or alterations made by those skilled in the art using the disclosed technical content shall fall within the protection scope of the present invention.
Claims
1. A prefetch control method for CXL memory devices, characterized in that, The method, executed in a CXL memory device controller or a CXL switch with memory proxy functionality, includes: S1 receives memory access requests, i.e., demand requests, initiated by the host processor through the CXL memory device response port; S2, the following processes are initiated in parallel: the link bandwidth statistics unit records the traffic information of the demand request, the prefetcher generates a prefetch request based on the address of the demand request, the prefetch buffer performs cache lookup on the address, and the missing status holding register unit performs in-transit status lookup on the address. S3, the arbitration decision unit performs a fast path decision based on the query results of the prefetch buffer and the query status information of the missing status holding register unit; S4, if the prefetch buffer is hit, the hit data is directed to the return queue and returned; if the prefetch buffer is not hit but there is a prefetch request entry with the corresponding address in the missing status holding register unit, i.e., the prefetch request is hit in transit, then it is further determined whether the prefetch request has been scheduled by the weighted scheduling unit, and the request processing is coordinated according to the determination result; if the prefetch is completely missed, the missing status holding register unit creates a new entry and submits the request to the request queue. S5. For requests entering the slow path, the missing state holding register unit performs merging and deduplication before distributing the requests to the demand queue or the prefetch queue. S6, the weighted scheduling unit selects a request from the demand queue and the prefetch queue and sends it to the backend access port according to the scheduling weight set by the prefetch throttling control unit; S7. After the data is returned from the backend access port, the missing state holding register unit writes the data into the prefetch buffer or directs it to the return queue according to the original request type. S8, the pre-fetching throttling control unit collects system status parameters; S9, the prefetching throttling control unit dynamically adjusts the prefetching intensity level of the prefetcher and the scheduling weight of the weighted scheduling unit based on the collected parameters.
2. The prefetch control method according to claim 1, characterized in that, The link bandwidth statistics unit adopts a window mechanism based on request counting. When the cumulative number of requests reaches a preset threshold, it sends a window trigger signal to the prefetch throttling control unit.
3. The prefetch control method according to claim 1, characterized in that, The system status parameters collected by the prefetch throttling control unit include link bandwidth utilization, prefetch accuracy, prefetch coverage, prefetch timeliness, missing state hold register unit occupancy rate, and demand queue depth. When the method is executed in a CXL switch, the collected parameters also include remote device response latency statistically analyzed by the back-end access port.
4. The prefetch control method according to claim 2, characterized in that, The prefetching throttling control unit responds to the window trigger signal, calculates the prefetching intensity level and scheduling weight based on the collected multidimensional parameters, and generates control signals that are sent in parallel to the prefetcher and the weighted scheduling unit.
5. The prefetch control method according to claim 1, characterized in that, In step S4, when it is found that there is an already en route prefetch request entry at the address of the demand request and the prefetch request has not been scheduled by the weighted scheduling unit, the missing status holding register unit notifies the weighted scheduling unit to cancel the scheduling of the prefetch request and submits the demand request as a new request to the demand queue.
6. The prefetch control method according to claim 1, characterized in that, The weighted scheduling unit maintains a list of canceled scheduling flags to identify and discard invalid prefetch requests; when scheduling a request from the prefetch queue, it sends a scheduled notification to the missing status holding register unit to update the status of the in-transit request.
7. A prefetch control device for CXL memory devices, characterized in that, An integrated CXL memory device controller or a CXL switch with memory proxy function is used to implement the prefetch control method according to any one of claims 1 to 6, wherein the device includes a prefetch processing module, a state management module, a scheduling and control module, and a data path module; The prefetching processing module includes a prefetcher and a prefetching buffer; The state management module includes an arbitration decision unit and a missing state holding register unit; The scheduling and control module includes a weighted scheduling unit, a prefetch throttling control unit, and a link bandwidth statistics unit; The data path module includes a CXL memory device response port, a backend access port, a return queue, a demand queue, and a prefetch queue.
8. The prefetch control device according to claim 7, characterized in that, The prefetch throttling control unit is connected to the link bandwidth statistics unit, the prefetch buffer, the missing status holding register unit, and the demand queue, and is used to receive system status parameters and output control commands; the prefetch throttling control unit is also connected to the prefetcher and the weighted scheduling unit, and is used to send the control commands.
9. The prefetch control device according to claim 7, characterized in that, The missing state holding register unit is used to perform merging and deduplication, queue distribution and data return routing for both demand and prefetch requests; the arbitration decision unit is used to perform fast path decision based on the query results of the prefetch buffer and the query state information of the missing state holding register unit.
10. The prefetch control device according to claim 7, characterized in that, The back-end access port connects to the DRAM controller when the device is deployed on a CXL memory device, and to the CXL link controller when deployed on a CXL switch to access remote CXL memory devices. When deployed on a switch, the port also collects remote device response latency parameters and feeds them back to the prefetch throttling control unit.
Citation Information
Patent Citations
Prefetching control method and device, electronic equipment and readable storage medium
CN118245512A
Intelligent network card and distributed object access method based on intelligent network card
CN120111107A