Memory access method and device, storage medium and computer equipment
By designing a dual-queue interface in a separate memory architecture, access requests are prioritized based on critical section contention, thus resolving the performance bottleneck when the number of clients expands and improving the application's throughput and scalability.
Patent Information
- Application Number
- CN202511045625.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-11-11
AI Technical Summary
In a split memory architecture, as the number of clients increases, the execution time of critical sections lengthens, leading to a decrease in application throughput and limiting the scalability of the number of clients.
By designing a dual-queue interface on the memory node, a high-priority or low-priority queue is allocated according to whether the access request is in a contention critical section, and requests in the contention critical section are executed first. This leverages the QoS characteristics of the RDMA network to implement a strict priority access policy.
It improved the application's throughput, enhanced application performance when the number of clients increased, and improved the application's scalability.
Smart Images

Figure CN120929407A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and more specifically, to a memory access method, apparatus, storage medium, and computer device. Background Technology
[0002] In a split-memory architecture, computing and memory resources within a cluster are divided into independent compute pools and memory pools. The compute pool consists of compute nodes, responsible for running application client processes; the memory pool consists of memory nodes, responsible for storing the data accessed by the application. Compute nodes can access the data stored in memory nodes through the memory semantics access interface provided by the memory nodes. Locks, as a common concurrency control method for applications using split-memory architectures, can create critical sections when clients running on compute nodes access data in memory nodes, thereby ensuring that only one client can access the data in a memory node at a time, improving data access security.
[0003] As the number of clients running an application increases, more sporadic memory access requests will occur. However, since the hardware resources within the cluster for executing sporadic memory access requests are limited, the execution time of critical sections will increase with the number of sporadic memory access requests. If multiple applications compete for the same critical section at this time, it will lead to a decrease in application throughput, affecting application performance when the number of clients expands, and limiting the scalability of the application's client base. Summary of the Invention
[0004] In view of this, this application provides a memory access method, apparatus, storage medium, and computer device to improve the performance of applications using a discrete memory architecture as the number of clients expands.
[0005] Specifically, this application is implemented through the following technical solution:
[0006] In a first aspect, embodiments of this disclosure provide a memory access method applied to an application client running on a computing node, comprising:
[0007] After generating a separate memory access request, if the separate memory access request is located in a critical section, it is determined whether the critical section is a contested critical section; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that is contested by multiple application clients.
[0008] Based on the judgment result, the target queue is determined from two initial queues pre-designed for the first memory node; the two initial queues have different queue priorities.
[0009] Using the queue interface provided by the first memory node for the target queue, the separate memory access request is submitted to the target queue; the first memory node is used to execute the separate memory access request in the two initial queues according to the queue priority.
[0010] In one possible implementation, determining whether the critical region belongs to a competing critical region includes:
[0011] Obtain the currently stored hot set of objects; the hot set of objects includes a first data object whose access frequency from the application client is higher than a set frequency; the first data object is located in a memory node in the memory pool;
[0012] Based on whether the critical section where the split memory access request is located is related to any of the first data objects, it is determined whether the critical section is a contested critical section.
[0013] In one possible implementation, based on the determination result, the target queue is determined from two initial queues pre-designed for the first memory node, including:
[0014] If the determination result is yes, the first queue with the higher queue priority among the two initial queues is determined as the target queue;
[0015] If the determination result is negative, the second queue with the lower priority among the two initial queues is determined as the target queue.
[0016] In one possible implementation, the queue interface for the two initial queues of the first memory node is provided through the following steps:
[0017] Enable Quality of Service (QoS) functions for each network device in the cluster consisting of each compute node and the memory node;
[0018] By configuring a virtual channel arbitration table for the quality of service function, two virtual channels are started, and a correspondence between the two virtual channels and the two initial queues is constructed according to the different queue priorities set for the two virtual channels.
[0019] Two queue pairs are allocated to the application client as queue interfaces for two initial queues, and different service levels are set for the two queue pairs respectively;
[0020] Based on the queue priority and the service level, establish a mapping relationship between the queue pairs and the virtual channels;
[0021] The step of using the queue interface provided by the first memory node for the target queue to submit the separate memory access request to the target queue includes:
[0022] The separate memory access request is submitted to the target queue pair corresponding to the target queue. The network device, based on the mapping relationship between the target queue pair and the virtual channel, and the correspondence between the virtual channel and the initial queue, submits the separate memory access request in the target queue pair to the target queue.
[0023] In one possible implementation, the first memory node is configured to execute separate memory access requests in the two initial queues according to the queue priority, following the steps described below:
[0024] Use a multiplexer to determine whether there are separate memory access requests in the first queue with high queue priority;
[0025] If so, then according to the arrival time of each separate memory access request in the first queue, each separate memory access request is retrieved from the first queue and executed in sequence;
[0026] Once all the separate memory access requests in the first queue have been completed, the separate memory access requests in the second queue, which has the lowest priority, are retrieved from the second queue in order of arrival time and executed.
[0027] In one possible implementation, the object hot set is determined according to the following steps:
[0028] If the number of data objects accessed by the application client meets the sampling condition, an object entry corresponding to the most recently accessed second data object is generated and inserted into the circular queue maintained by the application client; the object entry includes the key and timestamp of the second data object;
[0029] Using a preset thread, the object entry is retrieved from the circular queue, and the estimated frequency of the object entry is determined using the object entry and a target frequency counter set on a second memory node; the second memory node is the memory node where the second data object is located.
[0030] Based on the estimated frequency and set frequency of the object entry, determine whether to add the second data object to the object hot set, and based on the timestamp of the object entry, remove the object entry from the circular queue.
[0031] In one possible implementation, determining the estimated frequency of the object entry using the object entry and a target frequency counter set on the second memory node includes:
[0032] Obtain metadata of the target frequency counter from the second memory node; the metadata includes a first number of arrays, each array including a second number of slots;
[0033] Determine a first number of hash values corresponding to the key in the object entry, and determine the target slot corresponding to the key in each slot of the array;
[0034] Establish an index relationship between the first number of hash values and each of the target slots, and increment the value of each target slot.
[0035] Based on the index relationship, the current value of the object entry in each of the target slots on the target frequency counter is obtained, and the minimum value among the current values is used as the estimated frequency of the object entry.
[0036] In one possible implementation, determining whether to add the second data object to the object hot set based on the estimated frequency and a set frequency of the object entry includes:
[0037] If the estimated frequency of the object entry is greater than the set frequency, the second data object is inserted into a preset data structure maintained by the application client; each data object in the preset data structure constitutes the object hot set.
[0038] Secondly, embodiments of this disclosure also provide a memory access device for use in an application client running on a computing node, the device comprising:
[0039] The judgment module is used to determine whether the critical section is a contested critical section if the separate memory access request is located in a critical section after the separate memory access request is generated; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that is contested by multiple application clients.
[0040] The determination module is used to determine the target queue from two initial queues pre-designed for the first memory node based on the judgment result; the two initial queues have different queue priorities.
[0041] The submission module is used to submit the split memory access request to the target queue using the queue interface provided by the first memory node for the target queue; the first memory node is used to execute the split memory access request in the two initial queues according to the queue priority.
[0042] Thirdly, an optional implementation of this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the first aspect above, or any possible implementation of the first aspect.
[0043] Fourthly, an optional implementation of this disclosure also provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, performs the steps of the first aspect above, or any possible implementation of the first aspect.
[0044] The memory access method, apparatus, storage medium, and computer device provided in this disclosure, when an application client initiates a separate memory access request, adaptively submits the separate memory access request to a target queue with a corresponding queue priority in a first memory node based on whether the separate memory access request is located in a contention-critical section. This ensures that separate memory access requests located in a contention-critical section are placed in a high-priority queue, while separate memory access requests not located in a contention-critical section are placed in a low-priority queue. Furthermore, by having the first memory node execute the separate memory access requests in the initial queue according to the queue priority, it can be guaranteed that the first memory node can prioritize the execution of separate memory access requests located in the contention-critical section. This allows application clients to enter the contention-critical section more frequently, thereby improving application throughput, enhancing application performance in scenarios with an expanding number of clients, improving the scalability of the application's client count, and solving the problem of limited scalability of the number of clients faced in applications containing critical sections.
[0045] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0046] Figure 1 This is a flowchart illustrating a memory access method in an exemplary embodiment of this application;
[0047] Figure 2 This is a schematic diagram illustrating the overall process of a memory access method according to an exemplary embodiment of this application;
[0048] Figure 3 This is a schematic diagram illustrating the structure of a memory access device according to an exemplary embodiment of this application;
[0049] Figure 4 This is a schematic diagram of the structure of a computer device shown in an exemplary embodiment of this application. Detailed Implementation
[0050] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0051] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0052] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0053] Research has shown that Disaggregated Memory (DM) architecture holds promise as a foundational architecture for future data centers. In this architecture, the compute pool within a cluster consists of compute nodes equipped with high-performance central processing units (CPUs) and small amounts of memory; the memory pool consists of memory nodes equipped with large amounts of memory and CPUs with less powerful computing capabilities. The memory pool exposes memory semantics to the compute pool, enabling the compute pool to access data within the memory pool. This access includes operations such as reading, writing, comparing and swapping (CAS), and fetching and adding (FAA). In the DM architecture, the compute pool and memory pool can be interconnected via a high-speed network, typically using Remote Direct Memory Access (RDMA) technology. RDMA supports both one-way and two-way communication mechanisms. Bilateral communication requires cooperation between the sender and receiver, while unilateral communication does not. During communication, a compute node within the cluster can directly access memory data on a memory node, without the CPU of the accessed node needing to participate. For RDMA, the communication endpoints (i.e., when compute nodes and memory nodes communicate, they establish their own communication endpoints) are called Queue Pairs (QPs). The specific communication process can be initiated by submitting an RDMA request to the QP. In the DM architecture, application clients on compute nodes primarily initiate unilateral communication by submitting RDMA requests to the QP, thereby accessing the discrete memory on the memory node. Furthermore, RDMA networks provide Quality of Service (QoS) features to offer better service quality for specific network flows, such as dedicated bandwidth and controllable latency. A specific network flow can be understood as a pair of communication endpoints.
[0054] For applications using a DM architecture (hereinafter referred to as a split-memory application for ease of description), locking is a common technique for concurrent memory access control. It creates a critical section when accessing split memory, ensuring that only one application client can access shared data located in split memory (i.e., memory on memory nodes) at a time, preventing interference from other application clients. In DM, locks are typically implemented as distributed locks, with lock-related metadata stored in the split memory to allow different application clients within the cluster to share access and check whether a lock is currently held on data in the split memory. The distributed lock can be a lock for different data on different memory nodes.
[0055] In existing technologies, using locks to implement separate memory concurrency control is quite common, covering areas such as index structures, key-value stores, and transaction systems. For example, a typical technique is Sherman's separate memory B+ tree index, which not only uses locks to resolve write-write conflicts on tree nodes but also to protect tree structure modification operations, such as splitting and merging tree nodes. Another typical technique is FORD's separate memory transaction system, which employs the Optimistic Concurrency Control (OCC) protocol. During its transaction commit phase, to ensure transaction semantics, it needs to hold locks on all objects in the write set and / or read set before writing local updates back to separate memory. However, existing lock-based concurrency control methods all suffer from a significant performance degradation as the number of application clients increases. This is because the hardware resources (such as RDMA network cards, CPUs, and switches) available in the cluster for executing separate memory access requests are limited. For example, in a DM memory architecture based on RDMA, the RDMA network interface card (NIC) is the hardware device that executes discrete memory access requests, and its hardware processing units and hardware pipeline are limited. As the number of application clients increases, more discrete memory access requests will arise, thus prolonging the queuing latency experienced by these requests, and consequently extending the execution time of the critical sections containing these requests. If a critical section is a contention critical section (i.e., a critical section where multiple application clients compete for access), the extended execution time of this contention critical section will significantly reduce application throughput. Furthermore, since memory access latency and concurrency are higher in a DM architecture compared to local memory, the decrease in application throughput will be even more pronounced in a DM architecture. Therefore, how to solve the problem of limited client scalability faced by discrete memory applications containing critical sections has become a noteworthy technical issue.
[0056] Based on the above research, this disclosure provides a memory access method, apparatus, storage medium, and computer device. When an application client initiates a separate memory access request, the separate memory access request is adaptively submitted to a target queue with a corresponding queue priority in a first memory node, depending on whether the separate memory access request is located in a contention-critical section. This ensures that separate memory access requests located in the contention-critical section are placed in a high-priority queue, while separate memory access requests not located in the contention-critical section are placed in a low-priority queue. Furthermore, by having the first memory node execute the separate memory access requests in the initial queue according to the queue priority, it is guaranteed that the first memory node can prioritize the execution of separate memory access requests located in the contention-critical section. This allows application clients to enter the contention-critical section more frequently, thereby improving application throughput, enhancing application performance in scenarios with an expanding number of clients, improving the scalability of the application's client count, and solving the problem of limited scalability in applications containing critical sections.
[0057] The shortcomings of the above solutions are the result of the inventor's practical experience and careful research. Therefore, the discovery process of the above problems and the solutions proposed in this disclosure below are all contributions made by the inventor to this disclosure.
[0058] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0059] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0060] To facilitate understanding of this embodiment, a memory access method disclosed in this disclosure will first be described in detail. The execution subject of the memory access method provided in this disclosure is generally a terminal device or other processing device with certain computing capabilities. The terminal device may be a user equipment (UE), a mobile device, a terminal, a personal digital assistant (PDA), a handheld device, a computer device, an application client located on a computing node, etc. In some possible implementations, the memory access method can be implemented by the processor calling computer-readable instructions stored in the memory.
[0061] The memory access method provided in this disclosure embodiment will be described below using an application client running on a computing node as an example.
[0062] like Figure 1 The flowchart shown is a memory access method provided in an embodiment of this disclosure, which may include the following steps:
[0063] S101: After generating a separate memory access request, if the separate memory access request is located in a critical section, determine whether the critical section is a contested critical section; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that multiple application clients compete to access.
[0064] Here, a split memory access request is used to request access to a data object on a first memory node in the memory pool, which is defined as a target data object in this application. The first memory node includes multiple data objects, which are located at different memory locations within the first memory node. Accessing the target data object can be, for example, reading the target data object, writing the target data object, using CAS to access the target data object, or using FAA to access the target data object.
[0065] Split memory access requests are initiated by application clients of any split memory application running on a compute node. A compute node can run one or more application clients, and these application clients are running the same split memory application.
[0066] When an application client initiates a separate memory access request, it uses a lock-based concurrency control mechanism to lock or unlock the data object to be accessed, thus forming a critical section in the separate memory application. This critical section can be used to encrypt or deprotect the data object. Specifically, the application client executes a pre-defined piece of code during the locking or unlocking process; this code can be understood as the critical section. For the critical section, the separate memory application can actively call a pair of Application Programming Interfaces (APIs) named `begin_Critical Section` (`begin_CS`) and `end_Critical Section` (`end_CS`) to mark when the critical section begins and ends. For example, the separate memory application can call `begin_CS` when locking and `end_CS` when unlocking.
[0067] A critical section can contain one or more generated separate memory access requests. Specifically, the critical section in which any separate memory access request resides can be determined by the application client that initiated the request. A contentious critical section is a critical section where access is contested; that is, when multiple different application clients reside on the same computing node and need to compete for access to the same critical section, the critical section can be called a contentious critical section.
[0068] In practical implementation, for any application client running on any computing node, after the application client generates a split memory access request, the application client can first determine whether the split memory request is located in a critical section. For example, it can determine whether it is located in a critical section based on the lock metadata stored in the first memory node and the critical section's marking information. If not, the split memory access request can be processed directly using existing processing methods. If it is, it can further determine whether the critical section is a contested critical section. For example, it can determine whether the critical section is a contested critical section based on whether multiple application clients need to access the critical section.
[0069] Understandably, any existing method for determining a competing critical section can be used to determine whether a split memory access request is located in a critical section and therefore belongs to a competing critical section.
[0070] S102: Based on the judgment result, determine the target queue from the two initial queues pre-designed for the first memory node; the two initial queues have different queue priorities.
[0071] Here, this application designs a new access interface for discrete memory devices on the memory side to support the priority execution of discrete memory access requests located in contention-critical sections. Here, a discrete memory device is the device containing the memory space on a memory node, with one memory node corresponding to one discrete memory device. The aforementioned new access interface can be a dual-queue interface, which provides two pre-designed initial queues on the memory node to each compute node. Any application client running on the compute node can submit a discrete memory access request to either of the two initial queues through the dual-queue interface.
[0072] Each memory node has two pre-designed initial queues with different queue priorities. Each initial queue is used to store separate memory access requests of different priorities associated with the memory node. For example, the two initial queues may include a high-priority queue with a higher queue priority and a low-priority queue with a lower queue priority.
[0073] In practice, the queue priorities corresponding to different judgment results can be set in advance. After obtaining the judgment result, the priority of the target queue corresponding to the judgment result can be determined, and the initial queue with the target queue priority can be used as the target queue.
[0074] In one embodiment, S102 can be implemented according to the following steps:
[0075] If the judgment result is yes, the first queue with the higher priority among the two initial queues is designated as the target queue. Here, the queue with the higher priority among the two initial queues is defined as the first queue, and the queue with the lower priority is defined as the second queue. If the judgment result is yes, it means that the split memory access request is a request that accesses a contested critical section. Therefore, the first queue with the higher priority among the two initial queues can be designated as the target queue, which facilitates the first memory node to preferentially extract the split access request from the first queue and execute it first.
[0076] Alternatively, if the result is negative, the second queue with the lower priority among the two initial queues is designated as the target queue. Here, a negative result indicates that the split memory access request is not located in a contention-prone critical section, so the second queue with the lower priority among the two initial queues can be used as the target queue.
[0077] S103: Using the queue interface provided by the first memory node for the target queue, submit the separate memory access request to the target queue; the first memory node is used to execute the separate memory access requests in the two initial queues according to the queue priority.
[0078] In practice, the application client and the computing node where the application client resides can be used to submit split memory access requests to the target queue through the queue interface provided by the first memory node. Then, the first memory node can execute the split memory access requests in the two initial queues sequentially according to queue priority, thereby ultimately executing the split memory access requests submitted by each application client.
[0079] For example, the first memory node can first execute each separate memory access request in the high-priority queue, and after the high-priority queue is emptied, execute each separate memory access request in the low-priority queue.
[0080] In one embodiment, the queue interface for the two initial queues on any memory node can be provided through the following steps 1 to 4:
[0081] Step 1: Enable Quality of Service (QoS) for each network device in the cluster consisting of compute nodes and memory nodes.
[0082] Here, the DM architecture includes a memory pool and a compute pool. The memory pool contains multiple memory nodes, and the compute pool contains multiple compute nodes. The memory pool and compute pool together form a cluster. The network devices in the cluster are used to handle discrete memory access requests. Specifically, network devices may include CPUs, RDMA network cards, switches, etc. A compute node (or memory node) can have one CPU and one RDMA network card, and a single switch can connect multiple compute nodes (or memory nodes).
[0083] Quality of Service (QoS) is a function in a communication system used to ensure that communication services achieve the expected performance in terms of bandwidth, latency, jitter, and packet loss rate.
[0084] Specifically, the split memory device on the memory node can maintain two independent queue buffers for dual queues, introduce a multiplexer for queue selection, and add an extra bit to each split memory access request to encode the request priority. The dual queues can be viewed as two initial queues set up for the memory node, each initial queue corresponding to a queue buffer on the memory node. The extra encoding bit referenced in the split memory request can be determined by the application client after determining whether the split memory request is located in a contention-prone critical section. The memory node can determine whether the split memory access request should be prioritized based on the extra encoding bit referenced in the split memory request. The memory node can use the multiplexer to select a queue and obtain the split memory access request to be executed from the selected queue.
[0085] In this application, within an RDMA-based DM architecture, the dual-queue interface of memory nodes can be implemented by configuring the QoS characteristics of the RDMA network, thus eliminating the need for hardware modifications. Specifically, QoS functionality can be enabled for all network devices within the cluster, including RDMA network cards, switches, and CPUs, and configured to strict priority mode, thereby achieving separate processing of memory access requests strictly based on queue priority.
[0086] Step 2: By configuring a virtual channel arbitration table for the Quality of Service (QoS) function, start two virtual channels and establish a correspondence between the two virtual channels and the two initial queues based on the different queue priorities set for the two virtual channels.
[0087] In practical implementation, for each network device, the QoS function settings can be configured with a virtual lane (VL) arbitration table to enable two virtual lanes (VLs) and assign different queue priorities to each VL. For example, one VL can be assigned a low queue priority, and the other a high queue priority. Furthermore, a correspondence between VLs and initial queues can be established based on the queue priority of each VL and the queue priority of each initial queue. For example, a correspondence can be established between VLs with high queue priorities and initial queues, and between VLs with low queue priorities and initial queues. Understandably, since there is a one-to-one correspondence between queue buffers and initial queues, each VL also corresponds to a queue buffer. Specifically, a high-priority VL corresponds to the queue buffer of a high-priority queue, and a low-priority VL corresponds to the queue buffer of a low-priority queue.
[0088] Step 3: Assign two queue pairs to the application client as queue interfaces for two initial queues, and set different service levels for the two queue pairs respectively.
[0089] In practical implementation, for any application client running on a compute node, two queue pairs (QPs) can be allocated to that application client, and the two QPs serve as two queue interfaces provided by the memory node. After setting up the two QPs, they can be initialized. During initialization, a service level (SL) is set in the initialization parameters of each QP. The SL settings for the two QPs are different; for example, one QP is set to a high level, and the other to a low level.
[0090] Step 4: Establish a mapping relationship between queue pairs and virtual channels based on queue priority and service level.
[0091] In practical implementation, upon initializing two Queues (QPs), the mapping relationship between the two QPs and two Virtual Levels (VLs) can be determined based on their queue priorities and service levels. For example, a mapping relationship can be established between a QP with a high service level and a VL with a high queue priority, and vice versa. These mapping relationships are stored in a mapping table named SL-to-VL configured in the QoS settings of each QP. Thus, by using the SL-to-VL mapping table configured in the QoS settings, two SLs are mapped to two VLs. In this way, the two QPs can become a dual-queue interface provided by the memory node. Application clients can route separate memory access requests located within contention-critical sections to the high-queue-priority QP, and route other access requests to the low-queue-priority QP.
[0092] The method for setting up the queue interfaces of the two initial queues on the first memory node is similar to the process described above, and will not be repeated here.
[0093] Furthermore, the above-mentioned S103 can be implemented according to the following steps:
[0094] The separate memory access request is submitted to the target queue pair corresponding to the target queue. The network device then uses the mapping relationship between the target queue pair and the virtual channel, as well as the correspondence between the virtual channel and the initial queue, to submit the separate memory access request in the target queue pair to the target queue.
[0095] For example, when a split memory access request is located in a contention-critical section, and the target queue is a high-priority queue, the application client can submit the split memory access request to a high-priority QP with a high SL. Then, the network device on the compute node can, based on the QoS configuration, use the mapping relationship between SL and VL stored in the SL-to-VL mapping table to send the split memory access request to the high-priority VL corresponding to the high-priority QP. Then, based on the correspondence between VL and the initial queue, the high-priority VL is submitted to the queue buffer corresponding to the high-priority queue, thereby completing the submission of the split memory access request to the target queue.
[0096] Alternatively, if the split memory access request is not located in a contention-critical section and the target queue is a low-priority queue, the application client can submit the split memory access request to a low-priority QP with a low SL. Then, the network device on the compute node can, based on the QoS configuration, use the mapping relationship between SL and VL stored in the SL-to-VL mapping table to send the split memory access request to the low-priority VL corresponding to the low-priority QP. Based on the correspondence between VL and the initial queue, the low-priority VL is submitted to the queue buffer corresponding to the low-priority queue, thereby completing the submission of the split memory access request to the target queue.
[0097] In one embodiment, the first memory node can execute separate memory access requests in two initial queues according to queue priority, based on the following steps A to C:
[0098] Step A: Use a multiplexer to determine if there are any separate memory access requests in the first queue with the highest priority.
[0099] In practice, the first memory node can use a pre-defined multiplexer to select a queue. Specifically, it can first determine whether there is a separate memory access request in the first queue with high priority. If not, it means that there is no request currently in the contention critical section, so the multiplexer can be used to select the second queue with low priority, and obtain the separate memory access request from the second queue and execute it. If yes, then proceed to step B below.
[0100] Step B: If so, then retrieve each separate memory access request from the first queue in the order of its arrival time and execute it.
[0101] In practical implementation, if there is only one separate memory access request in the first queue, it can be directly retrieved and executed. If there are multiple separate memory access requests in the first queue, they can be retrieved and executed sequentially according to their arrival time. The arrival time is the time the separate memory access request was stored in the first queue. This allows for the sequential execution of each separate memory access request in the first queue in a first-come, first-served order.
[0102] Step C: After all the separate memory access requests in the first queue have been completed, retrieve the separate memory access requests from the second queue in the order of their arrival time, according to the queue priority of the lower queue, and execute them sequentially.
[0103] In practice, if all the separate memory access requests in the first queue have been completed, and there are separate memory access requests in the second queue with a lower priority, then the second queue is selected using a multiplexer. The separate memory access requests can be retrieved from the second queue in the order of their arrival time and executed sequentially.
[0104] In other words, the dual-queue (i.e., two initial queues) provided in this application adopts a strict-priority scheduling strategy. Split memory access requests in the high-priority queue are always executed before split memory access requests in the low-priority queue. Split memory access requests in the low-priority queue are only executed after the high-priority queue is completely emptied. Split memory access requests within the same priority queue are executed in a first-come-first-serve order. On the compute node side, the application client puts split memory access requests located in the contention critical section into the high-priority queue, thereby giving these access requests priority.
[0105] Understandably, the process of any memory node in the memory pool executing separate memory access requests in the two initial queues is similar to the execution process of the first memory node mentioned above, and will not be repeated here.
[0106] In one embodiment, this application can design a series of algorithms and data structures on the computing node side to identify competing critical sections in split-memory applications. The algorithms and data structures described above will be explained in detail below, in conjunction with the step of "determining whether the critical section belongs to a competing critical section" in S101:
[0107] Specifically, the step of "determining whether the critical section belongs to a competitive critical section" in S101 can be implemented according to the following steps S1 to S2:
[0108] S1: Get the hot set of objects currently stored; the hot set of objects includes the first data object whose access frequency from the application client is higher than the set frequency; the first data object is located in a memory node in the memory pool.
[0109] Here, each application client can maintain an object hot set, which stores data objects that are accessed by the application client more frequently than a set frequency.
[0110] The currently stored object hot set is the object hot set of the application client that generated the currently processed split memory access request. Each data object in this object hot set is defined as a first data object. For the application client that initiated the currently processed split memory access request, it may use different split memory access requests to access different data objects in different memory nodes. Therefore, the currently stored object hot set of the application client may include first data objects located in different memory nodes.
[0111] Access frequency is used to indicate how frequently an application client accesses the first data object. The frequency can be set empirically, and this application does not impose specific limitations.
[0112] In practice, for any application client, when determining whether a separate memory access request has accessed a contested critical section, the application client can obtain the hot set of objects it currently maintains.
[0113] In one embodiment, this application proposes a method to dynamically obtain a batch of objects most frequently accessed by a split-memory application through sampling, thus obtaining an object hot set, and treating the critical section accessing the object hot set as a contested critical section. Specifically, the object hot set of any application client can be determined and maintained through the following steps P1 to P3:
[0114] P1: If the number of data objects accessed by the application client meets the sampling condition, generate an object entry corresponding to the most recently accessed second data object and insert it into the circular queue maintained by the application client; the object entry includes the key and timestamp of the second data object.
[0115] Here, the sampling condition can be to perform sampling once for every preset number of data objects accessed. The preset number can be set based on experience, and this embodiment does not impose a specific limitation. For example, the preset number can be 16, 20, etc.
[0116] Application clients can initiate separate memory access requests to access different data objects in different memory nodes. The most recently accessed second data object is the data object accessed whenever the number of object accesses by the application client reaches a preset number. For example, if the preset number is 16, the 16th data object accessed by the application client can be used as the second data object, the 32nd data object accessed can be used as the second data object, the 48th data object accessed can be used as the second data object, and so on.
[0117] The preset number of data objects accessed by the application client each time can include data objects in different memory nodes, and / or different data objects in the same memory node.
[0118] Object entries can include a key and a timestamp for a second data object. Each application client can maintain a circular queue to store generated object entries. Whenever an object entry is identified, it is inserted at the tail of the circular queue maintained by the application client. When the number of object entries in the circular queue reaches its maximum capacity, new object entries can be inserted at the head of the queue if they are generated. The circular queue does not handle buffer wrap-around; instead, it filters out processed entries based on the timestamp.
[0119] In practice, the number of data objects accessed by the application client can be counted (or, in other words, the number of times the application client accesses a data object). Whenever this data volume meets the sampling condition, the most recently accessed data object is sampled and used as the second data object. For example, when the application client accesses a preset number of objects, a sample is taken, resulting in a second data object.
[0120] After obtaining the second data object, the hash value of its key can be calculated and used as the key. Simultaneously, a timestamp for the second data object can be generated using the `rdstc` command. Object entries for the second data object are then constructed using the key and timestamp, for example, in the form of:<key,timestamp> The hash value of the key of the second data object can be calculated using any hash algorithm in the prior art; for example, the hash value of the key of the second data object can be 64 bits.
[0121] Then, the object entry for the second data object can be inserted into the tail of the circular queue maintained by the application client.
[0122] P2: Using a preset thread, retrieve object entries from the circular queue, and use the object entries and the target frequency counter set on the second memory node to determine the estimated frequency of the object entries; the second memory node is the memory node where the second data object is located.
[0123] Here, the preset thread can be a background collection thread pre-established in the application client, which is used to collect object entries stored in the circular queue and determine the estimated frequency of the object entries.
[0124] A global frequency counter can be set on each memory node. This counter is used to estimate the frequency of occurrence of object entries inserted by a preset thread using the Count-Min Sketch algorithm.
[0125] Since there may be one or more object entries in the circular queue, the default thread will take out one object entry at a time for frequency estimation. Therefore, this application defines the memory node where the second data object corresponding to each object entry is located as the second memory node.
[0126] In practice, a background thread can be used to retrieve object entries from a circular queue. For each retrieved object entry, the second memory node containing the corresponding second data object can be determined, and the global frequency counter in that second memory node can be used as the target frequency counter. Then, the object entry is inserted into the target frequency counter, which records the estimated frequency corresponding to the object entry and returns it to the preset thread.
[0127] In one embodiment, P2 can be implemented according to the following steps:
[0128] P2-1: Obtain the metadata of the target frequency counter from the second memory node; the metadata includes a first number of arrays, each array including a second number of slots.
[0129] Here, both the first quantity and the second data can be preset quantities set based on experience. The global frequency counter is implemented as follows: the metadata of the global frequency counter is an array of d groups containing w slots, each slot encoding an unsigned integer, and the unsigned integer of each slot is initialized to zero; the metadata can be stored on memory nodes to allow concurrent and shared access by different application clients. Here, d is the first quantity, and w is the second quantity.
[0130] In practice, when inserting an object entry into the target frequency counter, the metadata of the counter can be obtained from the memory node where the target frequency counter is located.
[0131] P2-2: Determine the first number of hash values corresponding to the keys in the object entries, and in each array including the slots, determine the target slot corresponding to the key.
[0132] Here, the number of target slots corresponding to a key is the first quantity.
[0133] In practice, a first number of hash functions can be used to hash the keys in the object entries to be inserted, obtaining a first number of hash values. Simultaneously, a target slot can be selected from a second number of slots in each array to record the estimated frequency of the object entries to be inserted. The positions of the target slots selected in different arrays do not need to correspond one-to-one; for each array, the target slot can be selected randomly.
[0134] P2-3: Establish the index relationship between the first number of hash values and each target slot, and increment the value of each target slot.
[0135] In practice, an index relationship can be established between a first number of hash values and a first number of target slots. For example, one hash value indexes one target slot, and different hash values index different target slots. At the same time, an FFA operation can be performed on each target slot to increment the value of each slot by 1.
[0136] For example, the metadata for a target frequency counter is an array of d groups containing w slots. When an object entry is inserted, the key in the object entry is hashed using d different hash functions to calculate d hash values. These d hash values are then used as indices in the d arrays to map the key to one target slot in each of the d arrays; the FAA operation is then applied to the d target slots to increment their values by 1.
[0137] P2-4: Based on the index relationship, obtain the current value of the object entry in each target slot of the target frequency counter, and take the minimum value among the current values as the estimated frequency of the object entry.
[0138] In practice, when querying the estimated frequency of an object entry, the index relationship of the first number of hash values corresponding to the key of the object entry can be used to determine each target slot corresponding to the target frequency counter. Then, the current value of each target slot can be obtained, and the minimum value among these current values can be used as the estimated frequency of the object entry.
[0139] It should be noted that, since object entries corresponding to different second data objects within the same memory node may select the same target slot in some arrays, the frequency values of some slots may increase after an object entry is inserted into various target slots because these slots are used by other subsequently inserted object entries. Therefore, to improve the accuracy of the estimated frequency when querying the estimated frequency of an object entry, the minimum value among the current values in each target slot is used as the estimated frequency.
[0140] For example, when querying the estimated frequency of an object entry, the index relationship of the d hash values corresponding to the key of the object entry can be used to obtain d values from the target slots in the d arrays of the target frequency counter, and the minimum value among the d values can be used as the estimated frequency.
[0141] P3: Based on the estimated frequency and set frequency of the object entry, determine whether to add the second data object to the object hot set, and remove the object entry from the circular queue based on the timestamp of the object entry.
[0142] In practice, after obtaining the estimated frequency of an object entry using a preset thread, the application client can determine whether the estimated frequency of the object entry is greater than a set frequency. If so, it can be determined that the second data object corresponding to the object entry is a frequently accessed object, and therefore the second data object can be added to the object hot set. If the estimated frequency is not greater than the set frequency, it can be determined that the second data object corresponding to the object entry is not a frequently accessed object, and therefore it is not necessary to add the second data object to the object hot set.
[0143] In one embodiment, for the above-mentioned P3, the following steps can be taken:
[0144] If the estimated frequency of an object entry is greater than the set frequency, the second data object is inserted into the preset data structure maintained by the application client; the data objects in the preset data structure form an object hot set.
[0145] Here, the default data structure can be a heap structure. In the application client, the maintained heap structure can be used to achieve hot-set storage of objects.
[0146] In practice, if the estimated frequency of an object entry is greater than a set frequency, the second data object corresponding to the object entry can be inserted into the heap structure maintained by the application client. Furthermore, the data objects stored in the heap structure can be used as data objects in the object hot set. Whenever the object hot set needs to be retrieved, the application client can obtain it from the maintained heap structure.
[0147] S2: Determine whether the critical section is a contested critical section based on whether the critical section where the separate memory access request is located is related to any first data object.
[0148] Here, the critical section is related to the first data object, meaning it is a critical section specific to the first data object. Alternatively, it can be understood that a separate memory access request is used to access a target data object; if the target data object is identical to any of the first data objects in the object hot set, then the separate memory access request is related to that first data object.
[0149] In practice, it can be determined whether the target data object of the critical section where the split memory access request is located is the same as any of the first data objects. If so, it can be determined that the critical section where the split memory access request is located is a competing critical section (that is, a critical section that accesses the hot set of objects is considered a competing critical section). If not, it can be determined that the critical section where the split memory access request is located is not a competing critical section.
[0150] like Figure 2The diagram illustrates the overall process of a memory access method provided in this application embodiment. The application client can run a separate memory application. After the application client generates a separate memory access request, it can determine whether the data path of the separate memory access request is a critical path. The data path can be the data access operation (such as read operation, write operation, CAS operation, etc.) required by the separate memory access request and the target data object to be accessed. The critical path can be a path located in a critical section. Therefore, determining whether a data path is a critical path involves determining whether the separate memory access request is located in a critical section based on the data access operation required by the separate memory access request and the target data object to be accessed. If not, the separate memory access request can be directly submitted to the low-priority queue corresponding to the separate memory. The separate memory is located in a memory node. If so, it can be further determined whether the critical path accesses a hot set of objects, that is, whether the critical section where the separate memory access request is located is a contested critical section. If not, the separate memory access request can also be submitted to the low-priority queue. If an object hot set is accessed, it is determined that the split memory access request is located in a contention-critical section. Therefore, the split memory access request can be submitted to the high-priority queue corresponding to the split memory. The memory node where the split memory is located can allocate the application data stored in the split memory to the split memory access requests in each queue according to the queue priority, so as to realize the execution of the split memory access requests in each queue. At the same time, the object hot set can be maintained according to the control path, the circular queue maintained in the application client, the FAA auto-increment operation of the hot set identification layer, and the global frequency counter located in the memory node. Among them, the control path can be a begin_CS or end_CS operation, and the hot set identification layer is used to determine the data objects that need to be added to the object hot set based on the estimated frequency returned in the global frequency counter. For the specific steps of determining the object hot set, please refer to the description of the above embodiments, which will not be repeated here.
[0151] Understandably, the memory access method provided in this application can serve as a runtime environment. The separate memory application in the application client needs to use this runtime environment to run and implement the separate memory access process described in the above embodiments.
[0152] In summary, this application designs a dual-queue interface for discrete memory devices on the memory node side, enabling priority execution of discrete memory access requests located within contention-bound critical sections. By designing a series of algorithms and data structures on the compute node side, the contention-bound critical sections of the application can be accurately identified. Based on the memory access method provided in this application, the client scalability of discrete memory applications can be significantly improved.
[0153] Corresponding to the embodiments of the aforementioned memory access methods, this application also provides embodiments of memory access devices.
[0154] like Figure 3 The diagram shown is a structural schematic of a memory access device provided in an embodiment of this application, applied to an application client running on a computing node. The device includes:
[0155] The judgment module 301 is used to determine whether the critical section is a contested critical section if the separate memory access request is located in a critical section after the separate memory access request is generated; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that is contested by multiple application clients.
[0156] The determining module 302 is used to determine the target queue from two initial queues pre-designed for the first memory node based on the judgment result; the two initial queues have different queue priorities.
[0157] The submission module 303 is used to submit the split memory access request to the target queue using the queue interface provided by the first memory node for the target queue; the first memory node is used to execute the split memory access request in the two initial queues according to the queue priority.
[0158] In one possible implementation, the determining module 301, when determining whether the critical region belongs to a competitive critical region, is used to:
[0159] Obtain the currently stored hot set of objects; the hot set of objects includes a first data object whose access frequency from the application client is higher than a set frequency; the first data object is located in a memory node in the memory pool;
[0160] Based on whether the critical section where the split memory access request is located is related to any of the first data objects, it is determined whether the critical section is a contested critical section.
[0161] In one possible implementation, the determining module 302, when determining the target queue from two initial queues pre-designed for the first memory node based on the judgment result, is configured to:
[0162] If the determination result is yes, the first queue with the higher queue priority among the two initial queues is determined as the target queue;
[0163] If the determination result is negative, the second queue with the lower priority among the two initial queues is determined as the target queue.
[0164] In one possible implementation, the apparatus further includes an interface setting module 304 for providing queue interfaces for two initial queues of the first memory node through the following steps:
[0165] Enable Quality of Service (QoS) functions for each network device in the cluster consisting of each compute node and the memory node;
[0166] By configuring a virtual channel arbitration table for the quality of service function, two virtual channels are started, and a correspondence between the two virtual channels and the two initial queues is constructed according to the different queue priorities set for the two virtual channels.
[0167] Two queue pairs are allocated to the application client as queue interfaces for two initial queues, and different service levels are set for the two queue pairs respectively;
[0168] Based on the queue priority and the service level, establish a mapping relationship between the queue pairs and the virtual channels;
[0169] The submission module 303, when submitting the split memory access request to the target queue using the queue interface provided by the first memory node for the target queue, is used to:
[0170] The separate memory access request is submitted to the target queue pair corresponding to the target queue. The network device, based on the mapping relationship between the target queue pair and the virtual channel, and the correspondence between the virtual channel and the initial queue, submits the separate memory access request in the target queue pair to the target queue.
[0171] In one possible implementation, the device further includes an execution module 305;
[0172] The execution module 305 is configured to control the first memory node to execute the separate memory access requests in the two initial queues according to the queue priority, following the steps below:
[0173] Use a multiplexer to determine whether there are separate memory access requests in the first queue with high queue priority;
[0174] If so, then according to the arrival time of each separate memory access request in the first queue, each separate memory access request is retrieved from the first queue and executed in sequence;
[0175] Once all the separate memory access requests in the first queue have been completed, the separate memory access requests in the second queue, which has the lowest priority, are retrieved from the second queue in order of arrival time and executed.
[0176] In one possible implementation, the device further includes a hotspot setting module 306 for determining the object hotspot according to the following steps:
[0177] If the number of data objects accessed by the application client meets the sampling condition, an object entry corresponding to the most recently accessed second data object is generated and inserted into the circular queue maintained by the application client; the object entry includes the key and timestamp of the second data object;
[0178] Using a preset thread, the object entry is retrieved from the circular queue, and the estimated frequency of the object entry is determined using the object entry and a target frequency counter set on a second memory node; the second memory node is the memory node where the second data object is located.
[0179] Based on the estimated frequency and set frequency of the object entry, determine whether to add the second data object to the object hot set, and based on the timestamp of the object entry, remove the object entry from the circular queue.
[0180] In one possible implementation, the hot-set module 306, when determining the estimated frequency of the object entry using the object entry and the target frequency counter set on the second memory node, is configured to:
[0181] Obtain metadata of the target frequency counter from the second memory node; the metadata includes a first number of arrays, each array including a second number of slots;
[0182] Determine a first number of hash values corresponding to the key in the object entry, and determine the target slot corresponding to the key in each slot of the array;
[0183] Establish an index relationship between the first number of hash values and each of the target slots, and increment the value of each target slot.
[0184] Based on the index relationship, the current value of the object entry in each of the target slots on the target frequency counter is obtained, and the minimum value among the current values is used as the estimated frequency of the object entry.
[0185] In one possible implementation, the hot set setting module 306, when determining whether to add the second data object to the object hot set based on the estimated frequency and set frequency of the object entry, is configured to:
[0186] If the estimated frequency of the object entry is greater than the set frequency, the second data object is inserted into a preset data structure maintained by the application client; each data object in the preset data structure constitutes the object hot set.
[0187] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0188] Based on the same technical concept, embodiments of this application also provide a computer device. (Refer to...) Figure 4 The diagram shown is a structural schematic of a computer device provided in an embodiment of this application, comprising:
[0189] The system comprises a processor 401, a memory 402, and a bus 403. The memory 402 stores machine-readable instructions executable by the processor 401. The processor 401 executes these machine-readable instructions, and when executed, performs the following steps: S101: After generating a separate memory access request, if the separate memory access request is located in a critical section, determine whether the critical section is a contested critical section; the separate memory access request is used to access a target data object within a first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; a contested critical section is a critical section accessed by multiple application clients; S102: Based on the determination result, determine a target queue from two pre-designed initial queues for the first memory node; the two initial queues have different queue priorities; and S103: Using the queue interface provided by the first memory node for the target queue, submit the separate memory access request to the target queue; the first memory node executes the separate memory access requests in the two initial queues according to their queue priorities.
[0190] The aforementioned memory 402 includes a main memory 4021 and an external memory 4022. The main memory 4021, also known as internal memory, is used to temporarily store the computational data in the processor 401, as well as the data exchanged with external memory such as a hard disk 4022. The processor 401 exchanges data with the external memory 4022 through the main memory 4021. When the computer device is running, the processor 401 and the memory 402 communicate through the bus 403, so that the processor 401 executes the execution instructions mentioned in the above method embodiments.
[0191] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the memory access method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0192] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the software update method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0193] The computer program product can be implemented specifically through hardware, software, or a combination thereof. In one alternative embodiment, the computer program product is specifically embodied in a computer storage medium; in another alternative embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0194] The embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, the program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0195] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0196] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0197] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0198] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0199] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0200] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0201] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A memory access method, characterized in that, Application clients running on compute nodes include: After generating a separate memory access request, if the separate memory access request is located in a critical section, it is determined whether the critical section is a contested critical section; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that is contested by multiple application clients. Based on the judgment result, the target queue is determined from two initial queues pre-designed for the first memory node; the two initial queues have different queue priorities. Using the queue interface provided by the first memory node for the target queue, the separate memory access request is submitted to the target queue; the first memory node is used to execute the separate memory access request in the two initial queues according to the queue priority.
2. The method according to claim 1, characterized in that, The determination of whether the critical region belongs to a competitive critical region includes: Obtain the currently stored hot set of objects; the hot set of objects includes a first data object whose access frequency from the application client is higher than a set frequency; the first data object is located in a memory node in the memory pool; Based on whether the critical section where the split memory access request is located is related to any of the first data objects, it is determined whether the critical section is a contested critical section.
3. The method according to claim 1, characterized in that, Based on the judgment result, the target queue is determined from the two initial queues pre-designed for the first memory node, including: If the determination result is yes, the first queue with the higher queue priority among the two initial queues is determined as the target queue; If the determination result is negative, the second queue with the lower priority among the two initial queues is determined as the target queue.
4. The method according to claim 1, characterized in that, The queue interfaces for the two initial queues of the first memory node are provided through the following steps: Enable Quality of Service (QoS) functions for each network device in the cluster consisting of each compute node and the memory node; By configuring a virtual channel arbitration table for the quality of service function, two virtual channels are started, and a correspondence between the two virtual channels and the two initial queues is constructed according to the different queue priorities set for the two virtual channels. Two queue pairs are allocated to the application client as queue interfaces for two initial queues, and different service levels are set for the two queue pairs respectively; Based on the queue priority and the service level, establish a mapping relationship between the queue pairs and the virtual channels; The step of using the queue interface provided by the first memory node for the target queue to submit the separate memory access request to the target queue includes: The separate memory access request is submitted to the target queue pair corresponding to the target queue. The network device, based on the mapping relationship between the target queue pair and the virtual channel, and the correspondence between the virtual channel and the initial queue, submits the separate memory access request in the target queue pair to the target queue.
5. The method according to claim 1, characterized in that, The first memory node is used to execute the separate memory access requests in the two initial queues according to the queue priority, following the steps below: Use a multiplexer to determine whether there are separate memory access requests in the first queue with high queue priority; If so, then according to the arrival time of each separate memory access request in the first queue, each separate memory access request is retrieved from the first queue and executed in sequence; Once all the separate memory access requests in the first queue have been completed, the separate memory access requests in the second queue, which has the lowest priority, are retrieved from the second queue in order of arrival time and executed.
6. The method according to claim 2, characterized in that, The object hot set is determined according to the following steps: If the number of data objects accessed by the application client meets the sampling condition, an object entry corresponding to the most recently accessed second data object is generated and inserted into the circular queue maintained by the application client; the object entry includes the key and timestamp of the second data object; Using a preset thread, the object entry is retrieved from the circular queue, and the estimated frequency of the object entry is determined using the object entry and a target frequency counter set on a second memory node; the second memory node is the memory node where the second data object is located. Based on the estimated frequency and set frequency of the object entry, determine whether to add the second data object to the object hot set, and based on the timestamp of the object entry, remove the object entry from the circular queue.
7. The method according to claim 6, characterized in that, Determining the estimated frequency of the object entry using the object entry and the target frequency counter set on the second memory node includes: Obtain metadata of the target frequency counter from the second memory node; the metadata includes a first number of arrays, each array including a second number of slots; Determine a first number of hash values corresponding to the key in the object entry, and determine the target slot corresponding to the key in each slot of the array; Establish an index relationship between the first number of hash values and each of the target slots, and increment the value of each target slot. Based on the index relationship, the current value of the object entry in each of the target slots on the target frequency counter is obtained, and the minimum value among the current values is used as the estimated frequency of the object entry.
8. The method according to claim 6, characterized in that, Determining whether to add the second data object to the object hot set based on the estimated frequency and set frequency of the object entry includes: If the estimated frequency of the object entry is greater than the set frequency, the second data object is inserted into a preset data structure maintained by the application client; each data object in the preset data structure constitutes the object hot set.
9. A memory access device, characterized in that, The device is used in application clients running on compute nodes, and includes: The judgment module is used to determine whether the critical section is a contested critical section if the separate memory access request is located in a critical section after the separate memory access request is generated; the separate memory access request is used to access the target data object in the first memory node in the memory pool; the critical section is used to encrypt and protect the target data object; the contested critical section is a critical section that is contested by multiple application clients. The determination module is used to determine the target queue from two initial queues pre-designed for the first memory node based on the judgment result; the two initial queues have different queue priorities. The submission module is used to submit the split memory access request to the target queue using the queue interface provided by the first memory node for the target queue; the first memory node is used to execute the split memory access request in the two initial queues according to the queue priority.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 8.
11. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Hot spot data caching method, system and related device
CN111857597A
Cache replacement method, device and system
CN119536979A
Distributed metadata management method and storage system based on cloud computing
CN120335723A
Software and data processing system with priority queue dispatching
US20020083063A1
System and method for facilitating efficient management of data structures stored in remote memory
US20230019758A1