Memory key value storage system thread scheduling method and device and medium
By dividing the memory key-value storage system into a cache-resident layer and a memory-resident layer, and using inter-stage communication queues and finite state machines to manage request processing, the problems of low CPU cache efficiency and load balancing in existing technologies are solved. This achieves efficient load balancing and coordination of synchronization overhead, thereby improving system performance.
Patent Information
- Application Number
- CN202511006956.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
The existing thread scheduling model of memory key-value storage system suffers from low CPU cache efficiency due to the bundled execution of subtasks with different memory access modes, and it is difficult to balance load balancing and synchronization overhead under skewed workloads.
It adopts a two-layer architecture design, including a cache-resident layer (CR layer) and a memory-resident layer (MR layer). Request processing is managed through inter-stage communication queues and finite state machines, enabling real-time routing of high-frequency requests in the CR layer and processing of all data in the MR layer. Combined with non-blocking polling and dedicated cache resource management, it reduces cross-layer communication latency and synchronization overhead.
It significantly improves CPU pipeline efficiency, solves cache incompatibility issues, effectively mitigates load balancing and synchronization overhead, and enhances system scalability and responsiveness.
Smart Images

Figure CN120909719A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, and in particular, to a thread scheduling method for a memory key-value storage system, a device and a medium. BACKGROUND
[0002] The performance of a memory key-value storage (KVS) system depends largely on its thread model, which mainly reflects in two aspects: the scheduling manner of threads on CPU cores, and the allocation manner of client requests among threads.
[0003] In the era of slow hardware, the operating system usually introduces a preemptive multitasking mechanism to mask the high latency of devices through thread context switching, so as to improve the CPU utilization. However, with the network and storage hardware entering the era of nanosecond-level response, in order to fully exert the performance of hardware, user space libraries generally adopt a non-preemptive thread architecture to interact with devices through polling.
[0004] In the non-preemptive architecture, a model called Thread-Per-Queue (TPQ) (hereinafter referred to as NP-TPQ) is a common choice for building storage systems due to its high performance. In a typical NP-TPQ model, the system creates a certain number of worker threads, each of which reads and processes requests from its dedicated queue. To avoid synchronization overhead, these systems usually adopt a shared-nothing design, that is, each thread manages a part of data (i.e., data sharding), so as to realize lock-free data modification.
[0005] However, the NP-TPQ model in the prior art has the following defects:
[0006] Firstly, the cache-unfriendly nature leads to low CPU efficiency. The NP-TPQ model requires each worker thread to process a request in a run-to-completion manner, which means that a thread needs to execute a single function containing multiple tasks, for example, in a memory KVS, this will involve network polling, index lookup, data copying and sending response, etc. The memory access patterns of these sub-tasks are very different: the network polling phase can be designed to produce almost no cache misses, but the index lookup and data access phases need to access extensive memory space and are the main source of cache misses. Packing these phases with very different memory access patterns in the same function for execution can easily cause CPU cache thrashing, thus reducing the execution efficiency of the CPU pipeline.
[0007] Secondly, the conflict competition is difficult to effectively resolve. In handling skewed workloads, the shared-nothing (SN) design commonly adopted by NP-TPQ can cause uneven load distribution, resulting in some threads being overloaded while others are idle. Although the share-everything (SE) design can solve the load balancing problem, it introduces significant synchronization overhead, and the performance will quickly decline when the number of threads increases. In the NP-TPQ model, once a thread is blocked due to synchronization operations (such as lock contention), the entire request processing flow will be stalled, hindering other stages that could have been advanced. Therefore, the NP-TPQ model lacks a mechanism for effective scheduling in scenarios of uneven load and high contention.
[0008] Therefore, there is an urgent need for a new thread scheduling method for a memory key-value storage system to solve the problem of low cache efficiency caused by the overall execution of tasks in existing models, and to effectively resolve the contradiction between load balancing and synchronization overhead under different loads. SUMMARY
[0009] The present application proposes a thread scheduling scheme for a memory key-value storage system to solve the technical problems of low CPU cache efficiency caused by bundling the execution of sub-tasks of different memory access modes, and the difficulty in balancing load balancing and synchronization overhead under skewed workloads in the prior art.
[0010] According to one embodiment of the present application, a thread scheduling method for a memory key-value storage system is proposed, wherein the memory key-value storage system is organized into a two-layer architecture including a cache-resident layer (CR layer) containing first worker threads and a memory-resident layer (MR layer) containing second worker threads, and the method comprises:
[0011] The first worker thread obtains a request from a network buffer and performs index lookup on the key corresponding to the request in the CR layer;
[0012] When the key corresponding to the request is found in the index of the CR layer, the first worker thread directly processes the request;
[0013] When the key corresponding to the request is not found in the index of the CR layer, the first worker thread forwards the request to the MR layer through an inter-stage communication queue;
[0014] The second worker thread of the MR layer accesses the full key-value data to process the request forwarded from the CR layer, and after completing the request processing, updates the state identifier in the communication queue to implicitly notify the CR layer to send the response based on the request to the client.
[0015] In some embodiments, the method further comprises:
[0016] binding the first worker thread to a designated CPU core; and
[0017] allocating a dedicated CPU cache resource for the first worker thread to store key-value data satisfying a preset condition.
[0018] In some embodiments, the first worker thread is configured to manage the processing flow of the request in the CR layer using a finite state machine (FSM), including:
[0019] polling the state of the network receive buffer and the communication queue in a non-blocking manner;
[0020] performing an index lookup of the key corresponding to the request in the CR layer after the request is polled in the network receive buffer;
[0021] when the key corresponding to the request is found in the index of the CR layer, directly processing the request and returning immediately after completing the processing of the request, and continuing to poll the state of the network receive buffer and the communication queue in a non-blocking manner;
[0022] when the key corresponding to the request is not found in the index of the CR layer, forwarding the request to the MR layer and returning immediately without waiting for the processing result of the MR layer, and continuing to poll the state of the network receive buffer and the communication queue in a non-blocking manner;
[0023] after polling a state indicating that the MR layer completes the request, initiating an operation of sending the processing result stored in the network response buffer to the client, returning immediately, and continuing to poll the state of the network receive buffer and the communication queue in a non-blocking manner.
[0024] In some embodiments, the first worker thread forwards the request to the MR layer through an inter-stage communication queue, further including:
[0025] forwarding the request to the MR layer through a multi-producer, multi-consumer inter-stage communication queue, wherein the communication queue includes a plurality of dedicated, lock-free ring buffers, each first worker thread and second worker thread pair being assigned a dedicated ring buffer for transferring information.
[0026] In some embodiments, the request transferred through the ring buffer adopts a preset compact memory data structure, including a field for storing a hashed key, a field for indicating a request type, a field for representing a key-value pair data item size, and a field for pointing to a network receive buffer or a network response buffer.
[0027] In some embodiments, when a hash collision occurs in which hashed keys are the same but real keys are different, multiple data items with the same hashed key are organized into a linked list structure, and a search is performed in the linked list according to the real key included in the request.
[0028] In some embodiments, the first worker thread forwards the request to the MR layer through an inter-phase communication queue, further comprising:
[0029] The first worker thread pushes the request to the ring buffer corresponding to the second worker thread in a round-robin manner to achieve load balancing;
[0030] and / or packing multiple requests in a single slot of the ring buffer, sending multiple requests through a single push operation and correspondingly receiving multiple requests through a single pop operation to share communication overhead.
[0031] In some embodiments, the state identifier is a tail pointer of the ring buffer, and the second worker thread updates the state identifier in the communication queue after completing the request processing to implicitly inform the CR layer to send a response based on the request to the client, further comprising:
[0032] The second worker thread places the processing result in the network response buffer and updates the tail pointer of the ring buffer used to receive the request from the first worker thread after completing the request processing;
[0033] The first worker thread determines that the request has been processed based on the update of the tail pointer, and initiates an operation of sending the processing result stored in the network response buffer to the client.
[0034] In some embodiments, the second worker thread processes the request, further comprising:
[0035] Processing the request in one or more of batch processing, data prefetching, and coroutines to share cache miss overhead when accessing full key-value data.
[0036] In some embodiments, processing the request in one or more of batch processing, data prefetching, and coroutines, further comprising:
[0037] Obtaining a batch of requests from the communication queue;
[0038] For each request in the batch of requests obtained from the communication queue, creating an index coroutine for performing an index search;
[0039] Before reading data from full key-value data in the index coroutine, inserting a memory prefetching instruction, and the index coroutine yielding execution right immediately after inserting the memory prefetching instruction;
[0040] The second worker thread is used as a scheduler, and after the index coroutine yields the execution right, the execution right is immediately switched to another index coroutine that is ready, so as to hide the memory access delay.
[0041] In some embodiments, the MR layer processes the request forwarded from the CR layer, and further comprises performing zero-copy data transmission by using the second worker thread, the zero-copy data transmission comprising:
[0042] Data is directly copied between the network receiving buffer or the network response buffer and the key-value storage, wherein for a read request, data is copied from the key-value storage to the network response buffer, and for a write request, data is copied from the network receiving buffer to the key-value storage.
[0043] In some embodiments, the CR layer and the MR layer both adopt a fully shared design, and dynamic switching between the two types of worker threads is achieved by configuring the number of first worker threads and second worker threads, wherein control over concurrent access is achieved by:
[0044] An index structure supporting multithreaded concurrent access is adopted;
[0045] Locks and version numbers are embedded in each data item, and read and write operations on the data item are concurrently controlled by combining atomic instructions.
[0046] In some embodiments, concurrent control over read and write operations on the data item further comprises:
[0047] For a write operation, an atomic instruction is used to directly update, or lock the data item by using an atomic compare-and-swap (CAS) operation and then update, according to the size of the data item;
[0048] For a read operation, the version number is read before and after accessing the data item in a lock-free manner, and the atomicity of the read is ensured by comparing the version numbers read before and after.
[0049] According to an embodiment of the present application, an electronic device is provided, the device comprising a memory for storing computer instructions executable on a processor, and a processor for implementing the method according to any one of the above embodiments when executing the computer instructions.
[0050] According to an embodiment of the present application, a computer readable storage medium is provided, which stores a computer program executable by a processor, the program being executed by the processor to implement the method according to any one of the above embodiments.
[0051] The memory key-value storage system (KVS) thread scheduling scheme proposed in the application organizes the memory key-value storage system into a two-layer architecture including a cache resident layer (CR layer) and a memory resident layer (MR layer), and effectively overcomes the low cache efficiency problem caused by the overall execution of tasks in the existing model through a complete cross-layer request processing and notification process. Among them, by letting the CR layer as the unified entrance of all network requests and performing shunting, it ensures that the task of processing high-frequency hotspot data always resides in the CR layer, and effectively isolates it from the task of accessing full data, avoiding the pollution of a large number of cache misses when accessing full data to the CPU cache processing high-frequency requests, thereby solving the cache-unfriendly problem of the prior art and significantly improving the CPU pipeline efficiency; by updating the state identifier in the communication queue to realize implicit notification, the explicit message reply overhead between stages is avoided, solving the problem of high communication cost of traditional phased architecture, so that the two-layer architecture model proposed in the application is feasible and efficient in the modern high-speed hardware environment.
[0052] In some embodiments of the application, by binding the working threads of the CR layer to dedicated CPU cores and allocating dedicated cache resources, a physically isolated and interference-free path is created for the processing of high-frequency requests, further improving the low-latency performance; by using a finite state machine and a non-blocking polling mechanism in the CR layer, high responsiveness and high throughput of the front-end processing logic are realized; by specific design of the inter-stage communication queue, such as using a dedicated lock-free ring buffer, polling distribution to achieve load balancing, batch packaging requests to amortize overhead, etc., the delay and resource consumption of cross-layer communication are further reduced; by using compact memory data structures and designing a solution to hash conflicts, the communication efficiency is improved while the robustness of the system is guaranteed; by deeply integrating batch processing, coroutine and data prefetching technology in the MR layer, the inevitable high latency when accessing the main memory is actively hidden, and the back-end processing capability is significantly improved; by using zero-copy data transmission, unnecessary memory copying is reduced, and the CPU and memory bus load is reduced; by using a full-sharing design combined with fine-grained concurrency control means (such as lock-free reading and atomic writing), while achieving better load balancing than the non-sharing model, the high synchronization overhead of the traditional full-sharing model is effectively avoided, and the scalability of the system is improved. BRIEF DESCRIPTION OF DRAWINGS
[0053] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the application and serve to explain the principles of the application.
[0054] Figure 1 A flowchart of a memory key-value storage system thread scheduling method according to an embodiment of the application is shown.
[0055] Figure 2 A system architecture diagram of a non-preemptive thread architecture stage teardown method is shown according to one example embodiment of the present application.
[0056] Figure 3 A state transition diagram of a finite state machine (FSM) employed by a cache resident layer (CR layer) worker thread is shown according to one example embodiment of the present application.
[0057] Figure 4 A diagram of an inter-stage efficient communication queue is shown according to one example embodiment of the present application.
[0058] Figure 5 is a structural diagram of an electronic device according to at least one embodiment of the present application. DETAILED DESCRIPTION
[0059] The example embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, unless otherwise indicated. The following description of example embodiments is not representative of all embodiments consistent with the present application. Rather, it is merely an example of apparatus and methods consistent with some aspects of the present application as detailed in the appended claims.
[0060] In the scheme of the present application, the CR (Cache-Resident Layer) layer and the MR layer (Memory-Resident Layer) are two specialized logical layers that divide the stages in the KVS request processing flow into two different processing modes. The CR layer is a logical layer designed specifically for handling high-frequency access (hotspot) data and network communication; the MR layer is a backend logical layer that works in conjunction with the CR layer, responsible for managing the full amount of data and index of the system, and processing complex requests that are not hit in the CPU cache and forwarded by the CR layer. The "first worker thread" refers to the worker thread organized in the CR layer; the "second worker thread" refers to the worker thread organized in the MR layer.
[0061] Figure 1A flow chart of a method for thread scheduling of a memory key-value store system is shown. According to the concept of the present application, the memory key-value store system is organized as a two-tier architecture comprising a cache-resident tier (CR tier) containing first worker threads and a memory-resident tier (MR tier) containing second worker threads. The memory key-value store system is a data structure for storing, retrieving and managing associative arrays, and is often used in scenarios requiring high-speed read-write capability, such as distributed caching, session storage, etc. In the embodiments of the present application, to solve the problems of cache-unfriendliness and resource contention existing in the prior art, the system is logically divided into two specialized tiers, the CR tier and the MR tier, which work cooperatively. The CR tier can be used to process frequently accessed key-value data and network communication requests, while the memory-resident tier (MR tier) is responsible for managing all index and key-value data items in the system. Through the two-tier architecture design, tasks with different memory access characteristics (e.g., different access frequencies) can be decoupled, so that, for example, performance-sensitive hot data access can be quickly processed in an interference-free, high cache hit rate environment, while access to the full amount of data is isolated to the backend.
[0062] As shown in Figure 1 the method for thread scheduling of the memory key-value store system comprises the following steps S101-S104.
[0063] In step S101, the first worker thread obtains a request from a network receive buffer and performs index lookup on a key corresponding to the request in the CR tier.
[0064] The first worker thread can use a finite state machine (FSM) to manage the processing flow of the request in the CR tier. In some embodiments, the first worker thread can poll the network receive buffer in a non-blocking manner to actively discover and obtain newly arrived client requests. Using a non-blocking manner, a single scan operation on the queue will immediately return without causing the thread to be blocked in waiting when the queue is empty, to ensure that the CR tier can respond in a timely manner. When all the queues to be polled are empty, the first worker thread can continuously and repeatedly poll to ensure that new requests can be discovered with the lowest delay.
[0065] In some embodiments, the first worker thread can be bound to a specified CPU core to avoid the operating system arbitrarily scheduling the thread among multiple cores, eliminating the switching overhead and CPU cache invalidation caused by thread migration; and the first worker thread can be allocated dedicated CPU cache resources, for example, a specific LLC cache way can be allocated to the bound CPU core for storing key-value data satisfying a preset condition, to create a dedicated cache space for the CR layer that will not be disturbed by other tasks in the system (including MR layer tasks), ensuring that hot data can be stably resident in the cache. The amount of data managed by the CR layer can be kept small to fit the cache size, for example, only the hottest key-value pairs and network buffers can be managed.
[0066] According to the processing flow of the finite state machine, after obtaining the request, the first worker thread can search for the key corresponding to the request in the index structure managed by the CR layer itself, which is small in size and has been loaded into the dedicated CPU cache. If the search hits, that is, when the key corresponding to the request is found in the index of the CR layer, the processing can be completed directly in the CR layer (step S102), and if the search misses, that is, when the key corresponding to the request is not found in the index of the CR layer, the request can be forwarded to the MR layer (step S103).
[0067] Step S102, when the key corresponding to the request is found in the index of the CR layer, the first worker thread directly processes the request.
[0068] For example, when the key-value data operated by the request is identified as hot data with high frequency access, the index and data of which are managed by the CR layer, for example, in some embodiments, they are resident in dedicated CPU cache resources, the key corresponding to the request can be found in the index of the CR layer, that is, the index search result is a hit.
[0069] In the case of a hit path, the first worker thread directly processes the request, which can mainly include data access and sending a response. The first worker thread can first read or write the data item already in the cache. To ensure the consistency of data in a multi-threaded environment, the read-write operation can be performed through a unified, thread-safe concurrent control mechanism.
[0070] After the data access operation is completed, the first worker thread can immediately perform the operation of sending a response. It can construct a response message according to the operation result, for example, the data returned by the read request or the success status of the write request, and start the operation of sending the response to the client. The start operation can be asynchronous and non-blocking, and the first worker thread submits a sending task to the network subsystem, and then the finite state machine to which it belongs can immediately return to the initial polling state, ready to process the next event, without waiting for the actual completion of the network sending.
[0071] Since the index lookup, data access and start-up response are all completed in the CPU cache, and the non-blocking asynchronous I / O mode is adopted, the request for the hot data can be completed with extremely low delay, and the system efficiency is significantly improved.
[0072] In step S103, when the key corresponding to the request is not found in the index of the CR layer, the first worker thread forwards the request to the MR layer through an inter-stage communication queue.
[0073] When the index lookup result of step S101 is a miss, according to the present embodiment, the first worker thread forwards the request to the MR layer, which is processed by the MR layer with full data access capability. The forwarding process can be completed through a specially designed inter-stage communication queue (also referred to as CR-MR queue in the present application). In some embodiments, the communication queue can be designed as a multi-producer, multi-consumer queue, and can include multiple dedicated, lock-free ring buffers. Each first worker thread and second worker thread is allocated a dedicated ring buffer for transmitting information. For example, if there are n first worker threads and m second worker threads, a communication matrix of n x m can be formed. The use of dedicated, lock-free ring buffers can effectively avoid lock contention between threads and ensure that information can be safely and efficiently transmitted between first worker threads and second worker threads.
[0074] In some embodiments, when forwarding the request, the first worker thread can push the request into the ring buffer corresponding to each of the multiple different second worker threads in a polling manner, so as to avoid all missed requests from flowing to the same second worker thread, thereby facilitating load balancing.
[0075] In some embodiments, each slot of the ring buffer can be designed to accommodate multiple requests to share the fixed overhead of a single communication. The first worker thread can package multiple missed requests in a single slot and complete the forwarding of multiple requests through a single push operation. Similarly, the second worker thread can correspondingly receive a batch of multiple requests through a single pop operation.
[0076] In some embodiments, the request transmitted through the queue is designed as a compact memory data structure to minimize the memory overhead of communication. The data structure can include the following fields:
[0077] A key field for storing the key after hash operation, for example, the length can be 8 bytes (i.e. 64 bits), and for any original key with a length greater than the length, a hash function can be used to calculate the hash value of the length.
[0078] a type field for indicating the type of the request, e.g. 8 bits in length, to distinguish between a get operation or a put operation;
[0079] a size field for indicating the size of the key-value pair data item, e.g. 24 bits in length, to indicate the size of the value portion of the key-value pair associated with the request;
[0080] a buf field for pointing to the original network buffer, e.g. 32 bits in length, as a pointer or index to the slot position of the request in the original network buffer, which points to the network receive buffer for a put operation and to the network response buffer for a get operation. For example, the second worker thread of the MR layer can find the original key or the complete value data without hashing in case of handling hash collision or performing a put operation, etc. according to the buf field.
[0081] In some embodiments, when a hash collision occurs in which the hashed keys are the same but the real keys are different, multiple data items with the same hashed key can be organized in a linked list structure, and the real key (e.g. the buf field) included in the request is used to search in the linked list. By comparing the real keys one by one, the only correct data item can be located.
[0082] According to some embodiments, after the first worker thread successfully pushes the request to the communication queue, it does not wait for any processing result from the MR layer, and its finite state machine (FSM) can immediately return to the polling state to continue processing other newly arrived events (e.g. a new request in the network receive buffer or a new updated state identifier in the inter-stage communication queue). Through this non-blocking forwarding mechanism, it is beneficial to ensure that the CR layer worker thread is not slowed down by the slow MR layer operation, thereby maximizing the throughput of the system.
[0083] In step S104, the second worker thread of the MR layer accesses the full key-value data to process the request forwarded from the CR layer, and after completing the request processing, it updates the state identifier in the communication queue to implicitly notify the CR layer to send the response based on the request to the client.
[0084] According to this step, the second worker thread of the memory resident layer (MR layer) accesses the full data, which is usually large and cannot be stored in the CPU cache, to process the miss request, and after completing the request processing, it notifies the CR layer to start the operation of sending the response to the client.
[0085] According to embodiments of the present application, the second worker thread does not send any form of explicit message to inform the CR layer that the request has been processed, but implicitly informs the CR layer by updating a specific state identifier of the inter-phase communication queue. In one implementation, the second worker thread uses the tail pointer of the ring buffer where it receives the request as the state identifier. After the second worker thread finishes processing the request, it places the processing result in the network response buffer and updates the tail pointer of the ring buffer where it receives the request from the first worker thread; the first worker thread learns of the update of the tail pointer, for example, by polling, determines that the request has been processed, and initiates the operation of sending the processing result stored in the network response buffer to the client.
[0086] The implicit notification mechanism used in the embodiments avoids the need to create and pass a special reply message between the two processing layers by having the second worker thread update a state identifier (e.g., the tail pointer) of the communication queue, significantly reduces the CPU and memory overhead of inter-thread communication, shortens the delay of completing a state synchronization to the time of a single atomic operation, ensures that inter-phase communication does not become a performance bottleneck of the system in a high-speed hardware environment, and significantly simplifies the communication logic of the system, improving the scalability and robustness of the entire layered architecture.
[0087] In some implementations, the second worker thread processing the request further includes processing the request using one or more of batch processing, data prefetching, and coroutines to amortize the cache miss overhead when accessing the full amount of key-value data to amortize and hide the latency caused by cache misses.
[0088] In some examples, the request can be processed in combination with batch processing, data prefetching, and coroutines to further hide memory access latency. The second worker thread first obtains a batch of requests to be processed from the inter-phase communication queue. Then, instead of directly executing these requests, the second worker thread creates an independent index coroutine for each request. The second worker thread, as a scheduler, manages the concurrent execution of these coroutines. Inside one index coroutine, before it needs to read data from the main memory (e.g., dereference a pointer to access a child node of a B-Tree), the index coroutine can execute a memory prefetch instruction and immediately yield its execution right, for example, by co_yield. At this time, the second worker thread, as a scheduler, can immediately capture this event (i.e., the index coroutine yielding its execution right) and switch the CPU execution right to another index coroutine that is ready (e.g., its data can have been prefetched into the cache). By performing the above quick, memory access latency-based switching between multiple index coroutines, the CPU can perform the computing tasks of another index coroutine while waiting for the data loading of one index coroutine, thereby effectively hiding the memory access latency and greatly improving the processing throughput of the MR layer.
[0089] After locating the specific data item through the index query, the second worker thread can perform the actual data read / write operation. In some embodiments, zero-copy data transfer can be performed by the second worker thread to avoid the overhead of memory copy. Zero-copy data transfer refers to copying data directly between the network receive buffer or the network response buffer and the key-value storage, where for a read request, data is copied from the key-value storage to the network response buffer, and for a write request, data is copied from the network receive buffer to the key-value storage. According to the present embodiment, data is copied directly between the network buffer and the final key-value storage, completely bypassing the intermediate transfer between the CR layer and the MR layer. This design can also avoid introducing additional cache misses in the CR layer. For a write request, the MR layer only reads the network receive buffer without modifying it, so it does not invalidate it in the CR layer cache; for a read request, after data is written to the network response buffer, the contents of the network response buffer are ultimately sent out by the network hardware (such as RNIC), and the CPU core of the CR layer does not directly access the contents of the network response buffer, so no cache miss is generated.
[0090] In some embodiments, both the CR layer and the MR layer adopt a fully shared design. According to this design, all data in the system, including both the hot data managed by the CR layer and the full volume data managed by the MR layer, are stored in a unified memory address space and are not physically partitioned or fragmented by threads. Therefore, when a request is forwarded from the CR layer to the MR layer, the threads of the two layers operate on the same unique instance of the data object in memory, thereby avoiding expensive data serialization and copying between different threads or processing units and ensuring the efficiency of cross-layer collaboration.
[0091] Meanwhile, the number of first worker threads allocated to the CR thread pool and the number of second worker threads allocated to the MR thread pool are configurable. By adjusting the size of the two thread pools, dynamic conversion between the two types of worker threads can be achieved. A first worker thread originally responsible for the CR layer task can change its role and task allocation according to the scheduling decision of the system and take on the task of the MR layer to become a second worker thread; the opposite is also true to achieve dynamic and flexible resource allocation.
[0092] For example, under a load with a low cache hit rate, the bottleneck of the system can be in the processing capacity of the MR layer. At this time, more threads can be configured as second worker threads to perform MR layer tasks. Conversely, under a load with highly concentrated hotspots, the task of the CR layer can be heavier, and at this time, the number of first worker threads can be increased. Through flexible resource allocation, the present application can better adapt to diversified and dynamically changing workloads.
[0093] In the embodiment, concurrent access can be controlled to ensure data consistency and thread safety in a full sharing design by the following method:
[0094] An index structure supporting multi-thread concurrent access is adopted;
[0095] Lock and version number are embedded in each data item, and atomic instruction is combined to perform concurrent control on read and write operations of the data item.
[0096] For example, according to the embodiment, an index structure such as B+ tree, hash table, etc. can be selected. For example, for key-value pair data itself, according to the embodiment, metadata such as lock and version number can be embedded in each data item, and atomic instruction is combined to perform fine-grained concurrent control. In a specific implementation, different strategies can be adopted for read and write operations.
[0097] In some embodiments, for write operation, atomic instruction can be adopted to directly update or lock the data item through atomic compare-and-swap (CAS) operation and then update according to the size of the data item;
[0098] For read operation, version number is read before and after accessing the data item in a lock-free manner, and the atomicity of reading is ensured by comparing the version numbers read before and after.
[0099] In summary, referring to the flow of Figure 1 , the embodiment fully demonstrates the whole process from client request entering the system to the final response. The flow starts from the CR layer as the unified entrance of the system, and all requests are instantaneously shunted through efficient index lookup by the first working thread. For high-frequency requests that hit the cache, the flow is processed in a non-blocking manner in the CR layer, and data access and response initiation are directly completed; for requests that do not hit, they are efficiently and non-blockingly forwarded to the MR layer through a low-overhead inter-stage communication queue. The MR layer processes these requests and notifies the CR layer to send the response based on the request to the client through an efficient implicit notification mechanism, which effectively isolates memory access tasks with different characteristics, significantly improves CPU cache utilization, and solves the contradiction between load balancing and synchronization overhead in a high-concurrency scenario.
[0100] Figure 2 is a system architecture diagram of a non-preemptive thread architecture stage disassembly method in an exemplary embodiment of the present application. The exemplary embodiment is used to enable those skilled in the art to more intuitively understand the system architecture and data flow of the present application.
[0101] As shown in the figure, according to the present application, the memory key-value storage system is divided into two parts: the cache resident layer (CR layer) on the left and the memory resident layer (MR layer) on the right.
[0102] The global scheduler is responsible for the macro management of the thread resources of the whole system. For example, it can dynamically adjust the number of threads allocated to the CR layer and MR layer thread pools according to the current workload characteristics of the system.
[0103] As Figure 2 The CR layer, shown in the left orange box, is the front end and fast path of the whole system. It is used to handle all external incoming client requests and quickly respond to the hotspot data requests with the highest access frequency. The CR layer mainly includes an RPC network buffer, a thread pool, a hotspot index, and hotspot data. The RPC network buffer, including a network receiving buffer and a network response buffer, is the unified entrance for all external requests and the exit for the final response. The thread pool contains multiple first worker threads. The first worker threads are bound to the CPU cores dedicated to the CR layer to ensure that they run in an interference-free environment. The hotspot index and hotspot data represent the data managed by the CR layer, which is a subset of the full index and full data managed by the MR. In this example, the hotspot index and hotspot data contain the most active and frequently accessed key-value pairs in the system, and the total amount of data is controlled to be within the size that can be completely loaded into the dedicated CPU cache.
[0104] As Figure 2 The MR layer, shown in the right green box, is the back end and full data center of the whole system, used to handle all data requests that the CR layer cannot handle. The MR layer mainly includes a thread pool, a full index, and full data. The thread pool contains multiple second worker threads. The full index and full data represent all the key-value data stored by the system. The threads of the MR layer have the ability to access the entire data set.
[0105] Figure 2 Different color dashed arrows are used to intuitively show the two request processing paths:
[0106] 1. Hit path (orange dashed line): After a request enters, the first worker thread of the CR layer polls the request. The first worker thread can first query its local hotspot index. If the lookup hits (i.e., "hit" marked in the figure), the thread can directly access the hotspot data area to complete the read-write operation, and then send the reply to the client through the RPC network buffer. This path is completely closed in the CR layer, and because it operates on data already in the CPU cache, the delay is extremely low.
[0107] 2. Miss path (orange to green dashed line): If the first worker thread finds a miss in the hot index (i.e. "miss" as indicated in the figure), the request can be encapsulated and pushed into the inter-stage communication queue. The second worker thread of the MR layer can get this request from the queue and complete the processing by querying the full index and accessing the full data. After the processing is completed, the MR layer implicitly notifies the CR layer of the completion status through the communication queue, and finally the CR layer initiates the reply to the client.
[0108] Figure 2 The process of decoupling the system into the CR layer focusing on the hot data and the MR layer responsible for the full data, and cooperating with the help of the efficient communication queue to solve the cache competition and performance bottleneck problem caused by the task bundled execution in the prior art is shown.
[0109] In some embodiments, the first worker thread is configured to manage the processing flow of the request in the CR layer by using a finite state machine (FSM), including:
[0110] Polling the state identifiers of the network receiving buffer and the communication queue in a non-blocking manner;
[0111] After polling the request in the network receiving buffer, performing an index lookup on the key corresponding to the request in the CR layer;
[0112] When the key corresponding to the request is found in the index of the CR layer, directly processing the request, and returning immediately after completing the processing of the request, and continuing to poll the state identifiers of the network receiving buffer and the communication queue in a non-blocking manner;
[0113] When the key corresponding to the request is not found in the index of the CR layer, forwarding the request to the MR layer, and returning immediately without waiting for the processing result of the MR layer, and continuing to poll the state identifiers of the network receiving buffer and the communication queue in a non-blocking manner;
[0114] After polling the state identifier indicating that the MR layer completes the request, initiating the operation of sending the processing result stored in the network response buffer to the client, returning immediately, and continuing to poll the state identifiers of the network receiving buffer and the communication queue in a non-blocking manner.
[0115] Figure 3 The state transition diagram of the finite state machine (FSM) for managing the internal logic of the first worker thread according to an exemplary embodiment of the present application is shown. The flow is a non-blocking event loop.
[0116] As shown in the figure, the complete flow of the FSM is as follows:
[0117] The first worker thread first polls the network receive buffer and the communication queue in a non-blocking manner, corresponding to Figure 3 The polling operation in the middle takes the receiving request and the scanning queue state as the core, and the FSM continuously and alternately checks the two event sources.
[0118] When the request is polled in the network receive buffer, the FSM enters the processing path of a new request, corresponding to Figure 3 In the "Y" (Yes) path starting from the receiving request state, the FSM is transferred to the indexing state to perform an indexing lookup in the CR layer according to the key of the request. According to the lookup result, the flow is divided into two paths:
[0119] Hit path: as shown in Figure 3 , the flow is transferred to the data state along the hit arrow to access data and finally reaches the reply state. After completing the processing (i.e., starting the response), the FSM returns immediately to continue polling new events to ensure that the event loop of the thread is not interrupted.
[0120] Miss path: as shown in Figure 3 , the flow is transferred to the insertion queue state along the miss arrow to push the request into the CR-MR queue (i.e., the inter-stage communication queue). The first worker thread returns immediately to continue polling new events to ensure that the execution flow of the CR layer thread is not blocked by the MR layer operation.
[0121] When the first worker thread polls and finds, through the scanning queue state, that the state identifier of the communication queue indicates that the MR layer has completed the processing of the request, corresponding to Figure 3 In the "Y" (Yes) path starting from the scanning queue, the FSM enters the asynchronous reply processing path.
[0122] At this time, the FSM is transferred to the reply state to start the operation of sending the processing result stored in the network response buffer to the client. Like other paths, after this operation is started, the first worker thread also returns immediately to continue polling, preparing to process the next event.
[0123] Through the finite state machine as shown in Figure 3 , a complete and high-performance event processing model is realized. According to the embodiment, all links such as polling, hit processing, non-blocking forwarding, and asynchronous reply processing are integrated into a non-blocking FSM loop, further ensuring that the first worker thread of the CR layer can maximize the use of CPU resources to achieve extremely high processing throughput and system response capability.
[0124] Figure 4The structure and working principle of the inter-stage efficient communication queue according to an example embodiment of the present application is shown. As shown, the communication mechanism contains various optimization designs to improve throughput and reduce latency.
[0125] Figure 4 The left side shows the structure of the entire communication queue, which is a multi-producer, multi-consumer communication matrix. The matrix establishes multiple dedicated point-to-point communication channels between n first worker threads (producers, performing push operations) and m second worker threads (consumers, performing pop operations).
[0126] Each independent communication channel is implemented through a high-performance lock-free ring buffer. The data structure uses a head pointer (head) and a tail pointer (tail) to manage the queue state.
[0127] Each transferred slot in the ring buffer can accommodate a batch containing multiple requests. That is, the first worker thread of the CR layer can first accumulate a certain number of missed requests locally, and then push the entire packaged batch of requests into the queue at once through a push operation. Similarly, the second worker thread of the MR layer can also obtain a batch of tasks for processing at once through a pop operation. Through batch processing, the fixed CPU overhead of a single queue operation can be shared, significantly improving communication efficiency.
[0128] Figure 4 The dashed box on the far right shows the memory data structure of an independent request in the batch in detail. In this example, the total size of the data structure is only 16 bytes, which is divided into multiple fields: 8 bytes of key, 8 bits of type, 24 bits of size, and 32 bits of buf, to further reduce memory occupancy and data copy overhead.
[0129] Figure 4 The three optimizations of dedicated lock-free queue matrix, batch processing, and compact data structure are shown to build a high-performance inter-stage communication mechanism.
[0130] Figure 5 The electronic device provided for at least one embodiment of the present application includes a memory and a processor. The memory is used to store computer instructions executable on the processor. The processor is used to implement the memory key-value storage system thread scheduling method described in any embodiment or implementation of the present application when executing the computer instructions.
[0131] At least one embodiment of the present application also provides a computer readable storage medium having a computer program stored thereon. The program is executed by a processor to implement the memory key-value storage system thread scheduling method described in any embodiment or implementation of the present application.
[0132] Those skilled in the art will appreciate that one or more embodiments of the present specification can be provided as methods, systems or computer program products. Accordingly, one or more embodiments of the present specification can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, one or more embodiments of the present specification can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer readable program code.
[0133] Each of the embodiments described in the present specification is described in a progressive manner, and the same or similar parts between the embodiments can be mutually referred to, and each embodiment focuses on the differences from other embodiments. In particular, for the data processing device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0134] Although the present specification contains many specific embodiments, these should not be construed as limiting the scope or the required protection of any of the inventions, but merely as describing features of specific embodiments of the particular inventions. Some of the features described in the various embodiments within the present specification can also be implemented in a single embodiment. On the other hand, various features described in a single embodiment can also be implemented in multiple embodiments or in any suitable subcombination. Furthermore, although features can function as described in certain combinations and even initially claimed as such, in some cases one or more features from a claimed combination can be removed therefrom and the claimed combination can be directed to a subcombination or variation of a subcombination.
[0135] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring or implying that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0136] The above descriptions are only preferred embodiments of one or more embodiments of the present specification, and are not intended to limit one or more embodiments of the present specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of one or more embodiments of the present specification shall be included in the scope of protection of one or more embodiments of the present specification.
Claims
1. A method for thread scheduling in a memory key-value storage system, the method comprising: The memory key-value storage system is organized as a two-layer architecture including a cache-resident CR layer containing first worker threads and a memory-resident MR layer containing second worker threads, and the method comprises: The first worker thread receives a buffer obtaining request from the network and performs an index lookup on a key corresponding to the request in the CR layer; When the key corresponding to the request is found in the index of the CR layer, the first worker thread directly processes the request; When the key corresponding to the request is not found in the index of the CR layer, the first worker thread forwards the request to the MR layer through an inter-stage communication queue; The second worker thread of the MR layer accesses the full key-value data to process the request forwarded from the CR layer, and after completing the request processing, updates the state identifier in the communication queue to implicitly notify the CR layer to send the response based on the request to the client.
2. The method of claim 1, wherein, The method further comprises: binding the first worker thread to a specified CPU core; and allocating dedicated CPU cache resources for the first worker thread to store key-value data meeting preset conditions.
3. The method of claim 1, wherein, The first worker thread is configured to manage the processing flow of the request in the CR layer using a finite state machine (FSM), including: polling the state identifiers of the network receiving buffer and the communication queue in a non-blocking manner; after polling the request in the network receiving buffer, performing an index lookup on the key corresponding to the request in the CR layer; when the key corresponding to the request is found in the index of the CR layer, directly processing the request, and immediately returning after completing the processing of the request to continue polling the network receiving buffer and the state identifiers of the communication queue in a non-blocking manner; when the key corresponding to the request is not found in the index of the CR layer, forwarding the request to the MR layer, and without waiting for the processing result of the MR layer, immediately returning to continue polling the network receiving buffer and the state identifiers of the communication queue in a non-blocking manner; after polling the state identifier indicating that the MR layer completes the request, starting the operation of sending the processing result stored in the network response buffer to the client, immediately returning to continue polling the network receiving buffer and the state identifiers of the communication queue in a non-blocking manner.
4. The method of claim 1, wherein, The first worker thread forwards the request to the MR layer through the inter-stage communication queue, further comprising: forwarding the request to the MR layer through a multi-producer, multi-consumer inter-stage communication queue, wherein the communication queue includes a plurality of dedicated, lock-free ring buffers, and each first worker thread and second worker thread is assigned a dedicated ring buffer for information transfer.
5. The method of claim 4, wherein, The request transferred through the ring buffer adopts a preset compact memory data structure, and the data structure includes: a field for storing a hashed key, a field for indicating a request type, a field for representing a key-value pair data item size, and a field for pointing to a network receiving buffer or a network response buffer.
6. The method of claim 5, wherein, When a hash collision occurs, i.e., the hashed keys are the same but the real keys are different, multiple data items with the same hashed key are organized into a linked list structure, and a search is performed in the linked list according to the real key contained in the request.
7. The method of claim 4, wherein, The first worker thread forwards the request to the MR layer through an inter-stage communication queue, further comprising: The first worker thread pushes the request to the ring buffer corresponding to the second worker thread in a polling manner to achieve load balancing; And / or multiple requests are packaged in a single slot of the ring buffer, and multiple requests are sent through a single push operation and correspondingly received through a single pop operation to share the communication overhead.
8. The method of claim 4, wherein, The state identifier is the tail pointer of the ring buffer, and the second worker thread updates the state identifier in the communication queue after completing the request processing to implicitly inform the CR layer to send the response based on the request to the client, further comprising: The second worker thread places the processing result in the network response buffer and updates the tail pointer of the ring buffer used to receive the request from the first worker thread after completing the request processing; The first worker thread determines that the request has been processed based on the update of the tail pointer, and initiates an operation of sending the processing result stored in the network response buffer to the client.
9. The method of claim 1, wherein, The second worker thread processes the request, further comprising: One or more of batch processing, data prefetching, and coroutines are used to process the request to share the cache miss overhead when accessing the full key-value data.
10. The method of claim 9, wherein, One or more of batch processing, data prefetching, and coroutines are used to process the request, further comprising: A batch of requests are obtained from the communication queue; For each request in the batch of requests obtained from the communication queue, an index coroutine for performing index lookup is created; Before reading data from the full key-value data in the index coroutine, a memory prefetching instruction is inserted, and the index coroutine yields execution right immediately after the memory prefetching instruction is inserted; The second worker thread is used as a scheduler, and the execution right is switched to another index coroutine that is ready after the index coroutine yields execution right to hide memory access delay.
11. The method of claim 1, wherein, The request forwarded from the CR layer is processed in the MR layer, further comprising performing zero-copy data transmission using the second worker thread, the zero-copy data transmission comprising: Data is directly copied between the network receiving buffer or the network response buffer and the key-value storage area, wherein for a read request, data is copied from the key-value storage area to the network response buffer, and for a write request, data is copied from the network receiving buffer to the key-value storage area.
12. The method of claim 1, wherein, Both the CR layer and the MR layer adopt a fully shared design, and dynamic conversion between the first worker thread and the second worker thread is achieved by configuring the number of the first worker thread and the second worker thread, wherein control of concurrent access is achieved by: An index structure supporting multi-thread concurrent access is adopted; A lock and a version number are embedded in each data item, and atomic instructions are combined to perform concurrent control on read and write operations of the data item.
13. The method of claim 12, wherein, Concurrent control on read and write operations of the data item, further comprising: For a write operation, according to the size of the data item, an atomic instruction is used to directly update, or a data item is locked through an atomic compare-and-swap (CAS) operation and then updated; For a read operation, a version number is read before and after accessing the data item in a lock-free manner, and the atomicity of reading is ensured by comparing the version numbers read before and after.
14. An electronic device, comprising: The device comprises a memory for storing computer instructions executable on a processor, and a processor for implementing the method of any one of claims 1 to 13 when executing the computer instructions.
15. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1 to 13.
Citation Information
Cited By
Cache performance detection method and device and storage medium
CN121434037A