Reconfigurable remote procedure call system construction method, equipment and medium

By introducing a shared receive buffer and role switching mechanism into the RPC system, the problems of memory overhead and dynamic adjustment of worker threads in the RPC system are solved, achieving low memory overhead and flexible refactoring, and improving the system's resource utilization and stability.

CN120909813APending Publication Date: 2025-11-07TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511006739.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-21
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing RDMA-based RPC systems suffer from problems such as huge memory overhead, difficulty in dynamically adjusting worker threads, and complex information synchronization, which affect the system's memory utilization efficiency and stability.

Method used

It adopts a single-queue receive buffer shared by multiple clients, dynamically adjusts the thread ratio through a role switching mechanism between network threads and worker threads, and combines shared receive queue and multi-packet receive queue technology to achieve low memory overhead and flexible reconfiguration.

Benefits of technology

It significantly improves memory resource utilization, reduces system memory usage, enhances processing efficiency and system stability, and enables flexible allocation and smooth adjustment of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909813A_ABST
    Figure CN120909813A_ABST
Patent Text Reader

Abstract

The invention discloses a reconfigurable remote procedure call system construction method, equipment and a medium. The method is applied to a storage system and comprises the following steps: setting a single queue receiving buffer shared by a plurality of clients at a server node; obtaining requests from the single queue by n network threads in the server according to a preset numerical value association relationship including the serial number i, the total number n and the request position m of the network threads; forwarding the obtained request to one or more working threads for processing; and performing role switching between the network thread and the working thread in response to the adjustment instruction so as to dynamically adjust the quantity ratio of the network thread and the working thread. According to the method and the device, flexible and reliable task assignment is realized through a request acquisition mode based on the preset numerical value association relationship, and through sharing the memory pool among the threads and transparent internal role switching for the client, the memory overhead is remarkably reduced, flexible and dynamic adjustment of resources is realized, and the flexibility of the system and the resource utilization rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and particularly relates to a reconfigurable remote procedure call system construction method, device and medium. BACKGROUND

[0002] Remote Procedure Call (RPC) is a key component in distributed computing. In order to improve the interaction efficiency between nodes, the existing high-performance RPC system usually utilizes Remote Direct Memory Access (RDMA) technology.

[0003] RDMA is a technology that allows a computer to directly access the memory of a remote computer without CPU intervention. In traditional network communication, data needs to be copied multiple times between the user space and the kernel space of the sending and receiving parties, and involves kernel intervention, which consumes a lot of CPU time and memory bandwidth, significantly increases communication delay and reduces system performance. RDMA technology bypasses the above bottlenecks of the traditional network stack through zero-copy and kernel bypass. The main operations provided by RDMA include bilateral Send / Receive and unilateral Read and Write.

[0004] In the prior art, there are various RPC frameworks constructed based on RDMA. These frameworks attempt to utilize the unilateral or bilateral operation primitives of RDMA to improve the throughput and delay performance of RPC.

[0005] However, the existing RDMA-based RPC system still has the following defects in implementation:

[0006] 1. Huge memory pool management overhead: The existing RDMA RPC system has deficiencies in memory pool management. For example, some systems need to allocate a receiving buffer for each client on each worker thread. With the increase in the number of clients, this approach will result in huge memory overhead, far exceeding the CPU cache size, and significantly reducing the memory utilization efficiency of the system. Some general RPC libraries also need to allocate a buffer of up to tens of MB for each worker thread, resulting in large memory consumption.

[0007] 2. Difficulty in dynamically adjusting the number of worker threads: The existing RPC system has difficulty in supporting dynamic adjustment of the number of worker threads. When the system needs to increase or decrease the number of worker threads according to real-time load changes, these systems cannot adapt well, and simple reconstruction not only cannot effectively solve the problem, but also introduces additional performance overhead, limiting the ability of the system to flexibly adjust resources on demand.

[0008] 3. Complex information synchronization and error-prone: In some RPC systems that allow clients to specify server worker threads, when the number of server-side worker threads changes, all clients need to synchronize this information. This greatly increases the complexity and synchronization cost of the system, and is prone to request processing errors due to inconsistent information, affecting the overall performance and stability of the system.

[0009] Therefore, there is an urgent need for a low-memory-overhead, flexible reconfigurable remote procedure call system construction scheme to solve the problems of high memory overhead, difficulty in dynamically adjusting worker threads, and complex information synchronization in the prior art. SUMMARY

[0010] The present application proposes a low-memory-overhead reconfigurable remote procedure call (RPC) system construction method scheme to overcome the problems of high memory overhead and difficulty in dynamically adjusting worker threads according to load in the prior art based on RDMA-based RPC systems.

[0011] According to an embodiment of the present application, a reconfigurable remote procedure call system construction method is proposed, which is applied to a storage system and includes:

[0012] A single queue receiving buffer shared by connections of multiple clients is set in a server node, wherein a network interface controller of the server node sequentially appends requests from the multiple clients to the end of the receiving buffer;

[0013] n network threads in the server node acquire requests from the receiving buffer in a cyclic manner, wherein the ith network thread acquires a request at position m associated with the network thread number i and the network thread number n satisfying a preset numerical relationship;

[0014] The network thread forwards the acquired request to one or more worker threads, and the worker thread processes the request;

[0015] In response to an adjustment instruction, the network thread and the worker thread switch roles to dynamically adjust the number ratio of the network thread and the worker thread.

[0016] In some embodiments, the preset numerical relationship is set as:

[0017] When m mod n=i is satisfied, the ith network thread acquires a request at position m in the receiving buffer.

[0018] In some embodiments, the receiving buffer is configured as a shared receiving queue (SRQ), wherein the server node associates the connection queue pairs of the multiple clients to the shared receiving queue.

[0019] In some embodiments, the method further comprises:

[0020] The management thread puts the slot of the receive buffer into the shared receive queue in address-incrementing order via a receive (recv) primitive;

[0021] The server accepts requests sent by the clients in the form of send primitives;

[0022] The network interface controller of the server node stores the received request data into the head slot of the shared receive queue via direct memory access (DMA); and

[0023] After the network interface controller finishes processing the data of the current slot, the management thread puts the slot back into the shared receive queue via a receive primitive.

[0024] In some embodiments, the slots of the receive buffer employ a multi-packet receive queue (MP-RQ) technique, such that a single receive buffer slot associated with a receive primitive accommodates multiple requests from the clients.

[0025] In some embodiments, the method further comprises:

[0026] A dedicated response buffer is allocated for each network thread, which is reused in processing different batches of requests.

[0027] In some embodiments, the sum of the number of network threads and the number of worker threads remains constant.

[0028] In some embodiments, the role switching between network threads and worker threads in response to the adjustment instruction comprises switching one or more network threads into worker threads by:

[0029] The management thread updates a global variable representing the number of network threads n and a global variable representing the number of worker threads;

[0030] The management thread notifies all current network threads to switch when processing to the first synchronization slot in the single-queue receive buffer;

[0031] The network threads notified to switch invoke the entry function of the worker thread when processing to the first synchronization slot, become worker threads, and the remaining network threads update their saved local copy of the variable representing the number of network threads n when processing to the first synchronization slot to be consistent with the corresponding global variable.

[0032] In some embodiments, the method further comprises:

[0033] The original worker thread waits for all network threads that have been notified to move to become worker threads, and when it confirms that its own task queue is empty, it updates its local variable copy of the global variable representing the number of worker threads to synchronize with the corresponding global variable.

[0034] In some implementations, switching roles between network threads and worker threads in response to adjustment instructions includes switching one or more worker threads to network threads in the following ways:

[0035] The management thread updates the global variable representing the number of network threads, n, and the global variable representing the number of worker threads to notify the role switch.

[0036] The management thread notifies all original network threads to update the configuration when using the second synchronization slot.

[0037] The worker thread that is notified to switch continues to process its current task until its task queue is empty and all the original network threads have reached the second synchronization slot. Then, it calls the entry function of the network thread, becomes a network thread, and obtains requests based on the updated number of network threads n.

[0038] According to one embodiment of this application, an electronic device is provided, the device including a memory and a processor, the memory being used to store computer instructions executable on the processor, and the processor being used to implement the method as described in any of the preceding claims when executing the computer instructions.

[0039] According to one embodiment of this application, a computer-readable storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the method as described in any of the preceding claims.

[0040] The proposed low-memory-overhead reconfigurable remote procedure call system construction scheme significantly improves memory resource utilization and reduces system memory usage by employing a single-queue receive buffer shared by all network threads and a request retrieval mechanism based on numerical correlations. This avoids reserving large independent memory buffers for each client or thread. Furthermore, the lock-free request retrieval method based on pre-and-drop numerical correlations avoids task assignment contention between threads, improving processing efficiency.

[0041] Meanwhile, the dynamic adjustment mechanism proposed in this application allows server nodes to smoothly adjust the resource allocation between front-end network processing and back-end business processing internally by switching roles between network threads and worker threads based on real-time load, significantly enhancing the system's flexibility and stability. The entire dynamic adjustment process is transparent to the client, requiring no complex coordination and synchronization, thus achieving a smooth and stable adjustment process.

[0042] Further, the various specific embodiments of the present application also have at least the following beneficial effects: by adopting the shared receiving queue (SRQ) and multi-packet receiving queue (MP-RQ) technologies, the data receiving efficiency is further improved and the overhead of communication primitives is reduced; by setting a predefined synchronization slot as a switching point, the state consistency of all related threads in the adjustment process is ensured, and data processing errors are avoided; by reallocating and switching between network threads and worker threads, the system allows the CPU computing resources to be dynamically balanced between the network I / O processing and business logic processing, effectively alleviating CPU cache competition, and achieving fine management of system resources and performance optimization. In particular, the switching mechanism from the worker thread to the network thread ensures data integrity and system stability during complex role conversion by waiting for the source task queue to be emptied and the target thread pool to be stable. BRIEF DESCRIPTION OF DRAWINGS

[0043] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present description and serve to explain the principles of the present description.

[0044] Figure 1 A flowchart of a low-memory-overhead reconfigurable remote procedure call system construction method according to an embodiment of the present application is shown.

[0045] Figure 2 A structural diagram of a single-queue message pool management technology according to an exemplary embodiment of the present application is shown.

[0046] Figure 3 is a structural diagram of an electronic device according to at least one embodiment of the present application. DETAILED DESCRIPTION

[0047] The exemplary embodiments will be described in detail herein with reference to the attached drawings. When the description refers to the drawings, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments are not representative of all embodiments consistent with the present application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the present application, as detailed in the appended claims.

[0048] In the scheme of the present application, the storage system is organized as a two-layer architecture to decouple the request processing flow. The first layer is a front-end logic layer responsible for processing network communication, which contains network threads for obtaining requests from network interface controllers and forwarding them to the second layer; the second layer is a back-end logic layer responsible for processing business logic, which contains worker threads for processing the forwarded requests.

[0049] Figure 1A flowchart illustrating a method for constructing a reconfigurable remote procedure call system with low memory overhead according to an embodiment of this application is shown. This method is applied to a storage system. Figure 1 As shown, the method may include steps S101 to S103.

[0050] Step S101: Set up a single-queue receive buffer on the server node that is shared by multiple client connections, wherein the network interface controller of the server node sequentially appends requests from multiple clients to the end of the receive buffer.

[0051] According to this embodiment, a high-efficiency message receiving pool is first constructed on the server node. Unlike the prior art practice of allocating an independent buffer for each client or worker thread, this embodiment sets up a single-queue receive buffer on the server node that is shared by all client connections. This receive buffer can be a pre-allocated physical memory area on the server.

[0052] When multiple clients (such as Figure 2 C1, C2...C n When requests are sent concurrently to the server, the server node's Network Interface Controller (NIC, such as RNIC 3 which supports RDMA) appends these requests from different clients sequentially to the end of the single-queue receive buffer in the order of arrival using Direct Memory Access (DMA) technology, forming a logical request queue.

[0053] In some implementations, this sharing mechanism is achieved using a shared receive queue (SRQ) technology. For example, RNICs such as NVIDIA RNICs support SRQ functionality. In implementation, the system can associate all client connections (i.e., queue pairs) with this SRQ.

[0054] In some implementations, the management thread within the server node can divide the receive buffer into multiple slots and pre-place these slots into the SRQ in ascending address order using the RDMA receive primitive (recv primitive), preparing them for data reception. When a client sends a request using the send primitive, the server's RNIC hardware automatically writes the received request data into the first slot of the SRQ queue, a process requiring no CPU intervention. After the data in a slot has been processed, the management thread can return that slot to the SRQ using the recv primitive, achieving cyclical utilization.

[0055] In addition, in order to further reduce the submission overhead of the recv primitive, especially in the case of processing a large number of small requests, the slot of each receiving buffer can also use the multi-packet receiving queue (MP-RQ) technology, so that a single slot can accommodate multiple requests, thereby improving the overall efficiency of data transmission and processing.

[0056] In step S102, the n network threads in the server node acquire requests from the receiving buffer in a cyclic manner, wherein the ith network thread acquires the request at position m associated with the preset numerical relationship between the network thread number i, the network thread number n.

[0057] As shown in Figure 2 When the requests are queued in the single-queue receiving buffer, the total number of network threads (i.e., the first layer of worker threads) responsible for request acquisition in the current server node is n, and the n network threads acquire the requests in an efficient and decentralized manner.

[0058] All network threads acquire requests in the queue in a cyclic polling manner. The ith network thread is responsible for acquiring and processing the request at position m in the queue, wherein the number of threads i, the number of network threads n, and the position of the request m satisfy a preset numerical relationship. In some embodiments, the numerical relationship is a modulo operation, i.e., m mod n = i. For example, the 0th worker thread processes requests at positions 0, n, 2n,..., the 1st worker thread processes requests at positions 1, n+1, 2n+1,..., and so on.

[0059] The allocation method based on the deterministic mathematical relationship proposed in this embodiment allows each worker thread to independently calculate the task it needs to process, avoiding lock contention and communication overhead between threads, and eliminating the need for a central dispatcher.

[0060] In some embodiments, when the request is processed and the response is returned to the client, in order to balance performance and memory overhead, a dedicated response buffer with a small size (e.g., 64KB) can be allocated for each network thread. The response buffer can be reused in different batches of requests. For example, after completing the processing, the worker thread writes the response result into the pre-allocated, fixed-size response buffer, and then the network thread starts the operation of sending the processing result in the corresponding response buffer to the client. After sending this batch of response data to the corresponding client, the network thread does not release the memory of this buffer, but empties it and uses it to fill the response data of the next batch of requests, thereby achieving the reuse of the buffer. This reusable design avoids the performance overhead and management complexity caused by independent dynamic memory allocation and release for each response, thereby efficiently completing the response processing while controlling memory usage.

[0061] Step S103, the network thread forwards the acquired request to one or more worker threads, and the request is processed by the worker threads.

[0062] In this embodiment, the network thread is mainly used for processing network I / O tasks. After acquiring a request, the network thread can forward the request to a second-layer worker thread. The forwarding can be achieved through a thread-safe task queue shared by both. The worker thread pool continuously takes requests from the task queue and executes specific business logic, and can write the results to a response buffer.

[0063] In this way, the reception and parsing of network packets (handled by the network thread) and the business logic processing (handled by the worker thread) are effectively decoupled, enabling the system to dynamically adjust the number ratio of the two types of threads according to different types of load, achieving optimal resource utilization and performance.

[0064] Step S104, in response to the adjustment instruction, role switching is performed between the network thread and the worker thread to dynamically adjust the number ratio of the network thread and the worker thread.

[0065] This embodiment can dynamically and smoothly adjust the number of network threads and worker threads according to system load, to achieve flexible allocation of computing resources. This process is completely transparent to the client, without the need for complex coordination and synchronization, thereby enhancing the stability and resource utilization of the system.

[0066] In some embodiments, the sum of the number of network threads and the number of worker threads remains unchanged.

[0067] According to this embodiment, the total thread pool size used by the system at runtime is fixed. When the system needs to dynamically adjust between network threads and worker threads according to real-time load, it does not request the operating system to create new threads or destroy existing threads, but adjusts the resource ratio through a role switching mechanism. For example, when one or more network threads need to be added, the system can convert some of the worker threads in the current worker thread pool to the same number of network threads; conversely, when one or more worker threads need to be added, the system can convert some of the network threads in the current network thread pool to the same number of worker threads. This embodiment, on the one hand, avoids the high system overhead caused by thread creation and destruction, making the adjustment of resource ratio rapid and with minimal overhead; on the other hand, it ensures that the total resource occupation of the system remains stable, easy to deploy and manage, and enhances the predictability and stability of the system.

[0068] The underlying design of each network thread in this embodiment independently and lock-free acquires requests according to a predetermined numerical association, making the system dynamically reconfigurable. The reconfiguration process can be coordinated by a management thread and can include switching in both directions.

[0069] When the system monitors that the pressure of backend business processing increases, or the pressure of front-end network requests decreases, a part of the ability of processing network I / O can be converted into the ability of processing business logic. The process is as follows.

[0070] The management thread first updates the global variable representing the number of network threads n and the global variable representing the number of worker threads, for example, reduces n from 5 to 4, and increases the number of worker threads by 1 accordingly, and the subsequent steps are described in this example.

[0071] The management thread notifies all current network threads (i.e. network threads 0, 1, 2, 3, 4) that the specified synchronization slot in the receive buffer (for distinction from the synchronization slot of the reverse switch, it can be called the first synchronization slot) will be used as the unified switching point.

[0072] Network threads 0, 1, 2, 3, and 4 continue to work according to the old n value until their respective processing progress reaches the first synchronization slot. At this time, as the notified network thread 4, when processing reaches the first synchronization slot, it calls the entry function of the worker thread, smoothly changes its role from a network thread to a worker thread, and starts executing the task of the worker thread. The remaining network threads 0-3 can read the updated global variable n=4 when processing reaches the first synchronization slot, update their locally saved copy of n, and acquire requests according to the new n=4 rule, thereby seamlessly redistributing the load of the entire request processing. In some embodiments, each network thread can determine whether it is a notified thread or a remaining thread according to the updated number of network threads n. As shown in this example, the updated n value is 4, so the network thread numbered 4 can determine that it is a notified thread, and the network threads 0-3 can determine that they are all remaining threads.

[0073] The original worker thread will wait until all notified network threads have successfully changed to new worker thread members and confirm that its task queue is empty before updating its locally saved copy of the variable about the total number of worker threads, to ensure consistency within the worker thread pool, so that the entire worker thread pool can work collaboratively based on consistent configurations after new members join.

[0074] When the system monitors that the network request pressure increases and more threads are needed to process network I / O, a reverse switch can be performed, i.e. a part of the worker threads are switched to network threads. The process is as follows.

[0075] Similarly, the management thread first updates the global variables representing the number of network threads (i.e. the n value) and the number of worker threads.

[0076] The management thread notifies all original network threads to configure update at the specified second synchronization slot, so that the network thread pool is ready to receive new members.

[0077] The working thread notified to switch does not immediately stop the current work processing, but waits to enter a safe switching state, the conditions of which include:

[0078] Condition one: the working thread has finished processing all existing tasks in its task queue, ensuring the integrity of its own work and not losing any processing request;

[0079] Condition two: all original network threads have been processed to the second synchronization slot, that is, the network thread pool is ready to receive new members.

[0080] After entering the safe switching state, the working thread notified to switch can call the network thread entry function to become a network thread. Subsequently, the newly added network thread will start to obtain requests from the shared receiving buffer according to the updated total number of network threads n.

[0081] Through dynamic role switching between network threads and working threads, the system can dynamically tilt and rebalance the CPU computing resources between network I / O processing and business logic processing according to real-time load changes. This specialized thread division and flexible resource rebalancing can effectively alleviate CPU cache competition caused by mixed execution of different types of tasks, for example, network processing tasks sensitive to delay and business computing tasks consuming CPU can be isolated, thereby significantly improving the overall performance and resource utilization of the system.

[0082] Please refer to Figure 2 , Figure 2 The main results of the single queue message pool management technology in the exemplary embodiments of the present application are shown. A plurality of clients C1, C2,..., C n Concurrently send requests, which are sequentially appended to a single queue receiving buffer (i.e., receiving queue) shared by all network threads. At the first level, the i-th network thread obtains the request at position m in the receiving buffer from the receiving queue according to the preset m mod n = i rule, where n is the number of network threads.

[0083] The following is an exemplary description of the working process when the number of network threads n is dynamically adjusted through specific examples.

[0084] Example one: the number of network threads n is reduced from 5 to 4.

[0085] Initial state: The system initially has 5 network threads (thread 0, 1, 2, 3, 4), i.e. n = 5. According to the allocation rule of m mod n = i, the workload is allocated as follows at this time:

[0086] Thread 0 gets the requests at positions 0, 5, 10, 15,...

[0087] Thread 1 gets the requests at positions 1, 6, 11, 16,...

[0088] Thread 2 gets the requests at positions 2, 7, 12, 17,...

[0089] Thread 3 gets the requests at positions 3, 8, 13, 18,...

[0090] Thread 4 gets the requests at positions 4, 9, 14, 19,...

[0091] Trigger adjustment: Suppose the management thread monitors that the system load is reduced and decides to reduce the number of network threads from 5 to 4. The management thread will update the value of the global variable n to 4 and set a predefined synchronization slot, for example, set position m = 20 as the synchronization slot.

[0092] Switching process:

[0093] Threads 0, 1, 2, 3 (remaining network threads): These threads continue to work according to the rule of n = 5. When the processing progress of each thread reaches or passes the synchronization slot m = 20, the thread reads the global variable and updates its local copy of n to 4. Thereafter, it will acquire requests according to the new rule m mod 4 = i.

[0094] Thread 4 (network thread to be switched): This thread also acquires and forwards the last request it is responsible for (i.e. the request at position 19) according to the rule of n = 5. When its processing progress reaches the synchronization slot m = 20, the thread finds that the new n value is 4, and its own number "4" is already beyond the new valid range [0, 3]. At this time, the thread can call the entry function of the worker thread and switch to the worker thread.

[0095] Final state: After the synchronization slot m = 20, the system stabilizes to 4 network threads. The allocation of subsequent requests becomes:

[0096] Request at position 20: 20 mod 4 = 0, acquired by thread 0.

[0097] Request at position 21: 21 mod 4 = 1, acquired by thread 1.

[0098] Request at position 22: 22 mod 4 = 2, acquired by thread 2.

[0099] Request at position 23: 23 mod 4 = 3, taken by thread 3.

[0100] Request at position 24: 24 mod 4 = 0, taken again by thread 0.

[0101] Workload is smoothly rebalanced across the 4 threads.

[0102] Example Two: The number of network threads n increases from 3 to 4.

[0103] Initial state: The system has 3 network threads (threads 0, 1, 2), n = 3.

[0104] Triggered adjustment: The management thread monitors the system load increase and decides to increase n to 4. The management thread can update the value of the global variable n to 4, set the position m = 30 as the synchronization slot, and notify (for example, implicit notification) one worker thread to prepare to switch to the first layer of network threads (i.e., the new thread 3).

[0105] Switching process:

[0106] Threads 0, 1, 2 (original network threads): These network threads continue to work according to the rule n = 3. When their respective processing progress reaches the synchronization slot m = 30, these threads will synchronize and update their local copy of n to 4, and continue to work according to the new rule m mod 4 = i.

[0107] Newly added thread (original worker thread): After learning that it will be moved, the thread will not switch immediately. To ensure data consistency, the thread will continue to process its second-layer tasks until its task queue is empty, and it will also wait for the original network threads 0, 1, 2 to reach the synchronization slot m = 30. After both conditions are met, the thread calls the network thread's entry function, officially becoming the new network thread, i.e., thread 3.

[0108] Final state: After the synchronization slot m = 30, the system stabilizes to 4 network threads. The workload is redistributed:

[0109] For example, the request at position 31: according to the old rule n = 3, it should be handled by network thread 1 (31 mod 3 = 1). But since 31 is after the synchronization slot, it needs to be calculated according to the new rule n = 4, 31 mod 4 = 3, the responsibility of this request is dynamically and seamlessly transferred to the newly added network thread 3.

[0110] The embodiment of the application builds a low memory overhead reconfigurable remote procedure call system through a unique single queue message pool management technology and a working thread dynamic configuration mechanism. The system not only solves the problems of difficult flexible deployment of working threads and complex information synchronization through an internal dynamic adjustment mechanism transparent to the client, but also effectively solves the problem of huge memory overhead in the prior art through a shared memory pool, and significantly improves the resource utilization, flexibility and stability of the system.

[0111] Figure 3 The electronic device provided by at least one embodiment of the application includes a memory and a processor, the memory is used to store computer instructions executable on the processor, and the processor is used to implement the low memory overhead reconfigurable remote procedure call system construction method described in any embodiment or implementation manner of the application.

[0112] At least one embodiment of the application further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the low memory overhead reconfigurable remote procedure call system construction method described in any embodiment or implementation manner of the application.

[0113] Those skilled in the art should understand that one or more embodiments of the present specification can be provided as a method, a system or a computer program product. Therefore, one or more embodiments of the present specification can adopt a completely hardware embodiment, a completely software embodiment or an embodiment combining software and hardware aspects. Moreover, one or more embodiments of the present specification can adopt a computer program product implemented on one or more computer usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer usable program codes.

[0114] Each embodiment in the present specification is described in a progressive manner, and the same or similar parts between each embodiment can be referred to each other. Each embodiment mainly describes the difference from other embodiments. Especially, the data processing device embodiment is basically similar to the method embodiment, so the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0115] While this specification contains many specifics, these should not be construed as limitations on the scope of any invention, but rather as descriptions of particular implementations of particular embodiments thereof. Certain features that are, for clarity, described above in the context of separate embodiments, can also be provided in combination in a single embodiment. Conversely, various features that are, for brevity, described above in the context of a single embodiment, can also be provided separately or in any suitable subcombination. In addition, while features can be described above as being implemented in specific combinations, one or more features from a combination can in some cases be implemented in a different combination, or in an different embodiment, and such modifications and permutations are also within the scope of the disclosure. In the claims, means-plus-function clauses are used where for clarity it is more

[0116] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring such order, nor that all illustrated operations be performed, to implement such a methodology. In certain circumstances, multitasking and parallel processing can be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated in a single software product or packaged into multiple software products.

[0117] The above description is that of the best mode contemplated of carrying out one or more embodiments of the present specification, and of the principles thereof, and is subject to many modifications. Therefore, many options and implementations are possible that are within the scope of the present specification, and such options and implementations along with modifications, if any, are intended to be included in the scope of patent protection sought. Thus, the scope of the present specification is to be interpreted only by document statutes and regulations in effect at the time of the filing of this specification.

Claims

1. A reconfigurable remote procedure call system construction method, characterized by, The method is applied to a storage system, comprising: setting a single-queue receiving buffer shared by connections from a plurality of clients at a server node, wherein a network interface controller of the server node sequentially appends requests from the plurality of clients to the end of the receiving buffer; n network threads within the server node acquire requests from the receiving buffer in a cyclic manner, wherein an i-th network thread acquires a request at a position m satisfying a preset numerical relationship between the network thread number i, the network thread number n, and the position m; the network thread forwards the acquired request to one or more worker threads, and the worker thread processes the request; in response to an adjustment instruction, switching roles between the network thread and the worker thread to dynamically adjust the number ratio of the network thread and the worker thread.

2. The method of claim 1, wherein, the preset numerical relationship is set as: when m mod n = i is satisfied, the i-th network thread acquires the request at the position m in the receiving buffer.

3. The method according to claim 1 or 2, characterized in that, the receiving buffer is configured as a shared receiving queue (SRQ), wherein the server node associates the connection queues of the plurality of clients to the shared receiving queue.

4. The method of claim 3, wherein, The method further comprises: a management thread puts slots of the receiving buffer into the shared receiving queue in address-increasing order through a receive (recv) primitive; the server accepts requests sent by the client in the form of a send primitive; the network interface controller of the server node stores the received request data into the head slot of the shared receiving queue through direct memory access (DMA); and after the network interface controller processes the data of the current slot, the management thread re-puts the slot into the shared receiving queue through the receive primitive.

5. The method of claim 4, wherein, The slots of the receiving buffer use a multi-packet receiving queue (MP-RQ) technology, so that a single receiving buffer slot associated with the receive primitive can accommodate multiple requests from the client.

6. The method of claim 1, wherein, The method further comprises: allocating a dedicated response buffer for each network thread, which is reused in different batches of request processing.

7. The method of claim 1, wherein, The sum of the number of network threads and the number of worker threads remains unchanged.

8. The method of claim 1, wherein, in response to an adjustment instruction, switching roles between the network thread and the worker thread, including switching one or more network threads to worker threads by: the management thread updates a global variable representing the number n of network threads and a global variable representing the number of worker threads; the management thread notifies all current network threads to switch when processing to the first synchronization slot in the single-queue receiving buffer; the notified network thread calls the entry function of the worker thread when processing to the first synchronization slot, becomes a worker thread, and the remaining network thread updates the local variable copy of the global variable representing the number n of network threads to be consistent with the corresponding global variable when processing to the first synchronization slot.

9. The method of claim 8, wherein, The method further comprises: the original worker thread updates the local variable copy of the global variable representing the number of worker threads saved by itself to be synchronized with the corresponding global variable when confirming that the task queue of itself is empty after all notified network threads are converted into worker threads.

10. The method of claim 1, wherein, In response to the adjustment instruction, the role switching is performed between the network threads and the worker threads, including switching one or more worker threads into network threads by the following ways: The management thread updates a global variable representing the number n of network threads and a global variable representing the number of worker threads to inform the role switching; The management thread informs all original network threads to perform configuration update at the second synchronization slot; The worker thread which is informed to switch continues to process its current task until its task queue is empty and all original network threads reach the second synchronization slot, then calls the entry function of the network thread, turns into a network thread, and acquires requests according to the updated number n of network threads.

11. An electronic device, comprising: The device comprises a memory and a processor, the memory is used to store computer instructions executable on the processor, and the processor is used to implement the method of any one of claims 1 to 10 when executing the computer instructions.

12. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the method of any one of claims 1 to 10.