Method for Separating Performance of Transmission Queue of RDMA Network Card and RDMA Network Card

The method addresses performance interference between different types of applications on an RDMA network card by separating the transmission queue performance based on tenant types and using specific scheduling algorithms, resulting in improved performance isolation and expression.

JP7690641B2Active Publication Date: 2025-06-10YUSUR TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024070295
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2023-06-27
Filing Date
2024-04-24
Publication Date
2025-06-10
Estimated Expiration
2044-04-24

Smart Images

  • Figure 0007690641000001
    Figure 0007690641000001
  • Figure 0007690641000002
    Figure 0007690641000002
  • Figure 0007690641000003
    Figure 0007690641000003
Patent Text Reader

Abstract

To provide a remote direct memory access (RDMA) network card.SOLUTION: A sending RDMA network card: caches the work queue element WQEs of identified latency-sensitive tenants and bandwidth-sensitive tenants into a latency-sensitive group and a bandwidth-sensitive group; determines, by using an inter-group scheduler, whether to schedule either the WQEs in the latency-sensitive group or the WQEs in the bandwidth-sensitive group; requests, by using a first scheduler and a scheduling algorithm, to schedule the latency-sensitive WQEs in the latency-sensitive group; schedules, by using a second scheduler and a scheduling algorithm, the bandwidth-sensitive WQEs; and transmits, by the inter-group scheduler, the WQE scheduling result to the RDMA network card on a transmitter side and outputs all corresponding data packets corresponding to the latency-sensitive WQEs.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] (Cross-reference) This disclosure claims the priority of the Chinese patent application No. 202310762123.8, titled "Method for Performance Isolation of Transmission Queues of RDMA Network Cards and RDMA", filed on June 27, 2023, and the entire content of the application is incorporated herein by reference.

[0002] (Technical Field) The present invention relates to the technical field of improving the performance of the transmission side of an RDMA network card, and particularly to a method for performance isolation of the transmission queue of an RDMA network card and an RDMA network card.

Background Art

[0003] Remote Direct Memory Access (RDMA) technology is a high-performance network technology. The RDMA network card (RDMA NIC) designed based on RDMA technology offloads the network protocol transport layer through hardware, realizes zero-copy and operating system kernel bypass network transport services, and effectively reduces the occupancy rate of the server CPU during network data transmission. RDMA technology uses queues as the interface for software-hardware interaction. The designed queues include a send queue (SQ), a receive queue (RQ), and a completion queue (CQ). Here, SQ and RQ form a queue pair (QP) in a paired manner. The basic composition of SQ and RQ is the work queue element (WQE). Each WQE in SQ represents one RDMA data transmission request. WQE can be considered as the task that the software wants the hardware to execute and the "task description" containing detailed information about the task. The data transmission request is sent from the software to the hardware. Figure 2 is a schematic diagram of the application mode of the RDMA network card under the multi-application scenario of cloud computing in the prior art. The RDMA network cards on the startup side and the response side of cloud computing can both execute multiple applications APP1, …, APPN. The structures of the RDMA network card on the startup side and the RDMA network card on the response side correspond. The SQ is stored in the cache block (Memory) of the RDMA network card on the startup side. SQ consists of WQEs. Different applications store the WQEs that need to send content in SQ. Correspondingly, the RQ is stored in the cache block (Memory) of the RDMA network card on the response side. RQ also consists of WQEs.The sending side (sending module) of the RDMA network card schedules and processes WQEs in multiple ways to complete the corresponding sending tasks from the upper-layer application. Under the scenario of cloud computing, different applications may be of different types. Specifically, the types of applications are mainly divided into two types: delay-sensitive and bandwidth-sensitive. Different types of applications can be executed simultaneously on the same RDMA network card. The message volume of delay-sensitive applications is usually relatively small, focusing on the completion delay of each request of the application, and the message volume of bandwidth-sensitive applications is usually relatively large, focusing on the available bandwidth of the application.

[0004] In the process of realizing the present invention, the inventor found that there is an obvious performance interference phenomenon when different types of applications, that is, bandwidth-sensitive applications and delay-sensitive applications, are executed on the same RDMA network card, and there is also an obvious performance interference between bandwidth-sensitive applications and bandwidth-sensitive applications.

[0005] The application executed on the RDMA network card is called a tenant of the RDMA network card. Under the scenario of cloud computing, how to improve the RDMA network card to solve the obvious performance interference between bandwidth-sensitive applications and bandwidth-sensitive applications and improve the performance expression when different types of tenants are executed on the same RDMA network card simultaneously is a problem to be solved.

Summary of the Invention

Problems to be Solved by the Invention

[0006] In view of this, embodiments of the present invention provide a method for separating the performance of an RDMA network card transmission queue and an RDMA network card to eliminate or improve one or more defects existing in the prior art.

Means for Solving the Problem

[0007] One aspect of the present invention provides a method for separating the performance of an RDMA network card transmission queue, the method comprising: The transmitting-side RDMA network card recognizes the tenant type based on the tenant type identifier included in the queue pair identification, caches the working queue element WQE of the recognized delay-sensitive tenant's transmission queue as a delay-sensitive WQE in the delay-sensitive group of the cache module, and caches the WQE of the recognized bandwidth-sensitive tenant's transmission queue as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module, wherein at least a part of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in the waiting station, the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE; Determining whether to schedule either the WQEs in the delay-sensitive group or the WQEs in the bandwidth-sensitive group using an inter-group scheduler, and sending a scheduling request to a first scheduler when it is determined to schedule the WQEs in the delay-sensitive group, and sending a scheduling request to a second scheduler when it is determined to schedule the WQEs in the bandwidth-sensitive group; The first scheduler schedules the delay-sensitive WQEs in the delay-sensitive group using a first scheduling algorithm based on the scheduling request of the inter-group scheduler, and returns the WQE scheduling result to the inter-group scheduler; Using a token bucket, the second scheduler utilizes a second scheduling algorithm to schedule bandwidth-sensitive WQEs at the waiting station, and returns the WQE scheduling results to the inter-group scheduler. The group transmits the WQE scheduling results from the first scheduler and the second scheduler to the WQE processing module on the transmitting-side RDMA network card by the inter-group scheduler. The WQE processing module processes delay-sensitive WQEs based on the WQE scheduling results, outputs all data packets corresponding to the delay-sensitive WQEs, processes bandwidth-sensitive WQEs based on the WQE scheduling results, outputs one data packet corresponding to the bandwidth-sensitive WQE based on the path maximum transmission unit, and if there are data packets that have not been output for the WQEs generated by the bandwidth-sensitive tenant, rewrites the WQEs corresponding to the unoutput data packets to the waiting station.

[0008] In some embodiments of the present invention, the drive interface on the RDMA network card receives a tenant identification request, provides tenant identification including tenant type identification to the tenant based on the identification request, creates a queue pair having the queue pair identification including the tenant type identification for each tenant to which the tenant identification is assigned, and allocates the queue pair resources for storing WQEs for the queue pair.

[0009] In some embodiments of the present invention, the tenant type identification is a binary number with a preset length, and the tenant type identification is the last bit of the tenant identification.

[0010] In some embodiments of the present invention, the first scheduling algorithm is a fair polling scheduling algorithm. The step of using the first scheduling algorithm to schedule the delay-sensitive WQEs within the delay-sensitive group based on the scheduling requirements of the inter-group scheduler by the aforementioned first scheduler and returning the WQE scheduling result to the inter-group scheduler is as follows: The first scheduler receives the scheduling requirements of the inter-group scheduler, traverses the queues of the delay-sensitive WQEs of different tenants in the active link list maintained by the first scheduler based on the fair polling scheduling algorithm, and polls the queues of the delay-sensitive WQEs of different tenants within the delay-sensitive group. Here, the queue numbers of the delay-sensitive WQEs to be processed by different tenants within the delay-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one delay-sensitive WQE. The WQE of the queue at the head of the active link list is returned to the inter-group scheduler as the WQE scheduling result. When W is cached in the delay-sensitive group, the corresponding node is deleted from the active link list.

[0011] In some embodiments of the present invention, the second scheduling algorithm is a token bucket scheduling algorithm. The step of scheduling the bandwidth-sensitive WQEs in the bandwidth-sensitive group at the standby station using the second scheduling algorithm by the second scheduler as described above and returning the WQE scheduling result to the inter-group scheduler includes: receiving, by the second scheduler, the scheduling request of the inter-group scheduler, traversing the token counters of the queues of the bandwidth-sensitive WQEs of different tenants in the active link list maintained by the second scheduler based on the token bucket scheduling algorithm, and refreshing the tokens according to a predetermined refresh period and the number of tokens refreshed each time to update the token counter count. Here, the queue numbers of the bandwidth-sensitive WQEs to be processed by different tenants in the bandwidth-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one bandwidth-sensitive WQE.

[0012] In some embodiments of the present invention, the method further includes maintaining, by the standby WQE state module, the states of all WQEs at the standby station for the WQEs in the bandwidth-sensitive group at the standby station, and outputting all standby WQE states including two states of idle and occupied outside, and performing programming control of the WQE state based on the set and reset interfaces provided by the standby WQE state module.

[0013] In some embodiments of the present invention, the step of determining to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group using the inter-group scheduler described above includes: the inter-group scheduler determines to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on each scheduling result.

[0014] In some embodiments of the present invention, updating the dynamic weights based on each scheduling result as described above includes reducing the weight of the delay-sensitive group or the bandwidth-sensitive group corresponding to the scheduling result based on each scheduling result.

[0015] Another aspect of the present invention provides an RDMA network card, the RDMA network card recognizing a tenant type based on a tenant type identifier included in a queue pair identifier, caching, as delay-sensitive WQEs, working queue elements WQEs of a transmission queue of a recognized delay-sensitive tenant in a delay-sensitive group of a cache module, and caching, as bandwidth-sensitive WQEs, WQEs of a transmission queue of a recognized bandwidth-sensitive tenant in a bandwidth-sensitive group of the cache module, wherein at least some of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in a waiting station, WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, each node of the queue being one WQE, a control module; the waiting station for recording at least some of the WQEs to be processed in the bandwidth-sensitive group; a between-group scheduler for determining to schedule either a WQE in the delay-sensitive group or a WQE in the bandwidth-sensitive group, transmitting a scheduling request to a first scheduler when determining to schedule a WQE in the delay-sensitive group, transmitting a scheduling request to a second scheduler when determining to schedule a WQE in the bandwidth-sensitive group, and transmitting WQE scheduling results from the first scheduler and the second scheduler to a WQE processing module; a first scheduler for scheduling delay-sensitive WQEs in the delay-sensitive group using a first scheduling algorithm based on the scheduling request of the between-group scheduler and returning a WQE scheduling result to the between-group scheduler; a second scheduler for scheduling bandwidth-sensitive WQEs in the bandwidth-sensitive group in the waiting station using a second scheduling algorithm and returning a WQE scheduling result to the between-group scheduler; processing delay-sensitive WQEs based on the WQE scheduling result and outputting all data packets corresponding to the delay-sensitive WQEs,And process bandwidth-sensitive WQEs based on the WQE scheduling results, output one data packet corresponding to the bandwidth-sensitive WQE based on the path maximum transmission unit, and if there are data packets that have not been output, write back the WQEs corresponding to the unoutput data packets to the waiting station. The WQE processing module for this purpose is included.

[0016] In some embodiments of the present invention, the waiting station includes a first waiting WQE read / write channel for transmitting bandwidth-sensitive WQEs within the bandwidth-sensitive group in the cache module to the waiting station, and a second waiting WQE read / write channel for writing back the WQEs corresponding to the unoutput data packets to the waiting station. It includes a true dual-port SRAM, a waiting WQE state module for maintaining the states of all WQEs in the waiting station and outputting all waiting WQE states outside, and a set and reset interface for performing programming control of the WQE state. Here, the waiting WQE state includes a set and reset interface including two states: idle and occupied.

[0017] In some embodiments of the present invention, the first scheduling algorithm is a fair polling scheduling algorithm. The second scheduling algorithm is a token bucket scheduling algorithm. The inter-group scheduler determines which of the WQEs in the delay-sensitive group or the WQEs in the bandwidth-sensitive group to schedule based on the weighted polling scheduling algorithm based on the dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on the scheduling results of each time.

Advantages of the Invention

[0018] The performance separation method of the RDMA network card transmission queue proposed by the present invention stores, schedules, and processes WQEs from delay-sensitive tenants and bandwidth-sensitive tenants respectively, thereby realizing performance separation between delay-sensitive applications and bandwidth-sensitive applications on the transmission side of the RDMA network card. At the same time, for the WQEs of bandwidth-sensitive tenants, a method is designed in which they are written back to the waiting station instead of being completed in one transmission and rescheduled, effectively avoiding the long-term occupation of the transmission-side bandwidth resources by one WQE and realizing performance separation between bandwidth-sensitive applications. The RDMA network card proposed by the present invention solves the obvious performance interference problem between bandwidth-sensitive applications and bandwidth-sensitive applications through the improvement of performance separation, and can improve the performance expression of different types of tenants running simultaneously on the same RDMA network card.

[0019] Additional advantages, objects, and features of the present invention will be partially described in the following description, and will become partially apparent to those skilled in the art after consideration, or can be known from the implementation of the present invention. The objects and other advantages of the present invention can be realized and obtained by the structure specifically pointed out in the specification and the accompanying drawings.

[0020] Those skilled in the art will understand that the objects and advantages that can be realized by the present invention are not limited to those specifically described above, and the above and other objects that can be realized by the present invention will be more clearly understood from the following detailed description.

Brief Description of the Drawings

[0021] The accompanying drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation to the present invention. The accompanying drawings are described.

[0022]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Embodiments for Carrying Out the Invention

[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be described in more detail below in combination with embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and the description thereof are for the purpose of explaining the present invention and do not constitute a limitation to the present invention.

[0024] Here, it should be further explained that, in order to avoid the present invention being obscured by unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the accompanying drawings, and other details not related to the present invention are omitted.

[0025] It should be emphasized that the term "comprising / including", as used herein, means the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.

[0026] When different types of applications (e.g., bandwidth-sensitive applications and latency-sensitive applications) run on the same RDMA network card, there is an obvious performance interference phenomenon. There is also an obvious performance interference between bandwidth-sensitive applications. To address this problem, the inventor selected an existing RDMA network card with a bandwidth of 100 Gbps for experimental evaluation and ran a latency-sensitive application and a bandwidth-sensitive simulation application simultaneously on the experimental RDMA network card. As a result of the experiment, when running simultaneously with a bandwidth-sensitive application, the 50th percentile latency and 99th percentile latency of the latency-sensitive application increased by a factor of 5 compared to when running independently. At the same time, the inventor ran two bandwidth-sensitive simulation applications simultaneously on the experimental RDMA network card. As a result of the experiment, as the bandwidth occupied by one of the bandwidth-sensitive applications increased, the bandwidth of the other bandwidth-sensitive application decreased, that is, there was an obvious performance interference between the two bandwidth-sensitive applications, and the transmission-side bandwidth resources of the RDMA network card could not be fairly shared.

[0027] Therefore, in order to overcome the above performance interference and ensure the performance of RDMA with low latency and high bandwidth, the present invention provides a method for separating the performance of the transmission queue of an RDMA network card so as to achieve performance separation between latency-sensitive applications and bandwidth-sensitive applications and between bandwidth-sensitive applications in cloud computing. In an embodiment of the present invention, performance separation between applications is achieved by optimizing the transmission-side structure of the RDMA network card.

[0028] The improvement of the present invention for the transmission-side structure of the RDMA network card is as follows: (1) A mechanism for classifying queue pairs is proposed to achieve identifying and classifying transmission queues. (2) A WQE cache mechanism for distinguishing and handling different applications (or different types of applications) was designed. (3) For the characteristics of different types of application WQEs, different scheduling mechanisms were respectively proposed. As examples, scheduling based on fair polling and scheduling based on token bucket (two in-group scheduling methods), and deployable weighted polling inter-group scheduling for realizing different types of application differentiation services (inter-group scheduling method) were proposed. (4) Optionally, some embodiments of the present invention further proposed a method for slicing WQEs so as to avoid the situation that bandwidth-sensitive application WQEs occupy WQE processing module resources for a long time. (5) Optionally, some embodiments of the present invention also designed to realize differentiated services for delay-sensitive applications and bandwidth-sensitive applications based on RDMA by constructing dynamic weights.

[0029] FIG. 1 is a flowchart of a method for separating the performance of an RDMA network card transmission queue according to the present invention, and the method includes the following steps:

[0030] Step S110, the transmitting-side RDMA network card recognizes the tenant type based on the tenant type identifier included in the queue pair identifier, caches the working queue element WQE of the transmission queue of the recognized delay-sensitive tenant as a delay-sensitive WQE in the delay-sensitive group of the cache module, and caches the WQE of the transmission queue of the recognized bandwidth-sensitive tenant as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module, where at least some of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in the waiting station, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE.

[0031] Here, in the present invention, an application running on an RDMA network card (or an application executed on an RDMA network card) is regarded as a tenant of the RDMA network card. For each application to be executed on the RDMA network card, first, a queue pair (QP) needs to be created by the driver, and thereby an application needs to apply for queue pair (transmission queue and reception queue) resources. In an embodiment of the present invention, each tenant that first reaches the RDMA network card first applies for a tenant identifier (Tenant ID, or called tenant identification) through a drive interface on the RDMA network card (for example, the request_tenant drive interface) before creating a queue pair resource for storing WQEs on the RDMA network card, and the RDMA network card registered in the computer operating system assigns a tenant identifier to the tenant that first reaches the RDMA network card. Here, the tenant identification includes a tenant type identifier, and a queue pair resource for storing WQEs is created for each tenant to which the tenant identification is assigned, and the tenant is classified into a delay-sensitive tenant and a bandwidth-sensitive tenant based on the tenant type identifier. The process of assigning the tenant type identifier is that the first drive interface located on the RDMA network card receives the request for the tenant's application tenant type identifier, and the control module on the RDMA network card assigns the tenant type identifier to the tenant and stores all the tenants and the tenant type identifiers corresponding to the tenants.

[0032] Step S120: Determine whether to schedule either the WQE within the delay-sensitive group or the WQE within the bandwidth-sensitive group using an inter-group scheduler, and when it is determined to schedule the WQE within the delay-sensitive group, send a scheduling request to the first scheduler, and when it is determined to schedule the WQE within the bandwidth-sensitive group, send a scheduling request to the second scheduler. Optionally, the inter-group scheduler employs an inter-group weighted polling scheduling algorithm or an inter-group fair polling scheduling algorithm.

[0033] Step S130: The first scheduler schedules the delay-sensitive WQE within the delay-sensitive group using the first scheduling algorithm based on the scheduling request of the inter-group scheduler, and returns the WQE scheduling result to the inter-group scheduler. In an embodiment of the present invention, the first scheduling algorithm is a fair polling scheduling algorithm.

[0034] Step S140: The second scheduler schedules the bandwidth-sensitive WQE at the standby station using the second scheduling algorithm, and returns the WQE scheduling result to the inter-group scheduler. In an embodiment of the present invention, the second scheduling algorithm is a token bucket scheduling algorithm.

[0035] Step S150: The group scheduler transmits the WQE scheduling results from the first scheduler and the second scheduler to the WQE processing module on the sending-side RDMA network card. Here, the preset number of data packets may be one or multiple. For the transmission tasks of bandwidth-sensitive tenants, after transmitting data of one length, the transmission is interrupted, aiming to avoid a single task from occupying the bandwidth resources on the sending side of the RDMA network card for a long time. The Maximum Transmission Unit (MTU) refers to the maximum data packet size that the network can transmit. It is a parameter determined by the network device and the communication environment, with the unit of byte.

[0036] Step S160: The WQE processing module processes the delay-sensitive WQE based on the WQE scheduling results and outputs all data packets corresponding to the delay-sensitive WQE. Also, it processes the bandwidth-sensitive WQE based on the WQE scheduling results, outputs one data packet corresponding to the bandwidth-sensitive WQE based on the path maximum transmission unit, and if there are unoutput data packets, writes back the WQE corresponding to the unoutput data packets to the waiting station.

[0037] Before step S110, the method further includes receiving, by a first drive interface on the RDMA network card, an identification request of a tenant that first reaches the RDMA network card, and providing a tenant identification including a tenant type identifier to the tenant based on the identification request assignment; and for each tenant to which the tenant identification is assigned, creating a queue pair having the queue pair identification including the tenant type identification, and allocating the queue pair resource for storing a WQE to the queue pair. The present invention classifies and identifies queue pairs of an RDMA network card to distinguish queue pairs used by delay-sensitive applications and bandwidth-sensitive applications, and further distinguishes and schedules transmission queues of different service types. The present invention takes an application running on the RDMA network card as a tenant of the RDMA network card. In an embodiment of the present invention, an application on the RDMA network card can actively apply for a tenant type identifier (Tenant ID) by scheduling a request_tenant drive interface according to the present invention before creating a queue pair resource. The Tenant ID is a numerical value of M bits, that is, the numerical range of the Tenant is 0 to 2 m and the type of the tenant can be identified by the last bit of the Tenant ID. Specifically, when the last bit of the Tenant ID is 0, a delay-sensitive tenant is identified, and when the last bit is 1, a bandwidth-sensitive tenant is identified. In a specific embodiment of the present invention, in a callback function that provides a create_qp create QP, the Tenant ID provided by the application is shown in FIG. 3 as part of a QP identifier (QP ID). Optionally, the tenant type identifier (Tenant ID) is packaged into the i-th to j-th bits of the queue pair identifier (QP ID), and the numerical relationship is j - i = M - 1.

[0038] FIG. 3 is a schematic diagram of the packaging relationship between a tenant type identifier and a queue pair identifier in an embodiment of the present invention. The tenant type identifier (Tenant ID) is packaged as part of the queue pair identifier (QP ID). Optionally, in the present invention, it is possible to recognize whether the tenant type is bandwidth-sensitive or delay-sensitive based on the tenant type identifier packaged in the queue pair identifier. Further, the tenant (i.e., application) number is also packaged in the queue pair identifier, and based on this, it is possible to recognize specifically which application the WQE corresponding to the queue pair identifier is derived from, and to know whether the source of the WQE is a bandwidth-sensitive application or a delay-sensitive application.

[0039] In step S120, the step of determining to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group using the inter-group scheduler described above includes: the inter-group scheduler determines to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on each scheduling result. The above-mentioned updating the dynamic weights based on each scheduling result includes reducing the weight of the delay-sensitive group or the bandwidth-sensitive group corresponding to the scheduling result based on each scheduling result. Here, for the dynamic weights respectively configured for the WQEs from the delay-sensitive group and the bandwidth-sensitive group, the dynamic weight of the WQE from the delay-sensitive group has a positive correlation with the number of WQEs to be processed from the delay-sensitive tenant in the WQE cache module, and the dynamic weight of the WQE from the bandwidth-sensitive group has a positive correlation with the number of WQEs to be processed from the bandwidth-sensitive tenant in the WQE cache module. The dynamic weights of the WQEs from the delay-sensitive group and the bandwidth-sensitive group are initialized to equal values each time they are updated to 0, updated the dynamic weight of the WQE from the delay-sensitive group each time it is scheduled by the fair polling scheduler, and updated the dynamic weight of the WQE from the bandwidth-sensitive group each time it is scheduled by the token bucket scheduler.

[0040] In step S130, the step of scheduling the delay-sensitive WQEs within the delay-sensitive group using the first scheduling algorithm based on the scheduling request of the inter-group scheduler by the aforementioned first scheduler and returning the WQE scheduling result to the inter-group scheduler is that the first scheduler receives the scheduling request of the inter-group scheduler, traverses the queues of the delay-sensitive WQEs of different tenants in the active link list maintained by the first scheduler based on the fair polling scheduling algorithm, and polls the queues of the delay-sensitive WQEs of different tenants within the delay-sensitive group. Here, the queue numbers of the delay-sensitive WQEs to be processed by different tenants within the delay-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one delay-sensitive WQE. Return the WQE of the queue at the active link list head to the inter-group scheduler as the WQE scheduling result. When W is cached in the delay-sensitive group, delete the corresponding node from the active link list. Optionally, the correspondence between the nodes and queues of the active link list may be realized by storing the queue number, or may be realized by storing a pointer to the head of the queue.

[0041] In step S140, the second scheduling algorithm is the token bucket scheduling algorithm. The step of the second scheduler using the second scheduling algorithm to schedule the bandwidth-sensitive WQEs within the bandwidth-sensitive group in the waiting station and returning the WQE scheduling result to the inter-group scheduler, as described above, includes the second scheduler receiving the scheduling request of the inter-group scheduler, traversing the token counters of the queues of the bandwidth-sensitive WQEs of different tenants in the active link list maintained by the second scheduler based on the token bucket scheduling algorithm, and refreshing the tokens according to the predetermined refresh period and the number of tokens refreshed each time to update the token counter count. Here, the queue numbers of the bandwidth-sensitive WQEs to be processed by different tenants within the bandwidth-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one bandwidth-sensitive WQE. Here, for the WQE corresponding to the head of the traversed queue, if the count of the token counter is greater than or equal to the size of the path maximum transmission unit, the WQE is scheduled to the WQE processing module; if the count of the token counter is less than the size of the path maximum transmission unit, the corresponding WQE node is moved to the end of the queue. The count update rule of the token bucket counter is to adjust the token counter of the corresponding bandwidth-sensitive tenant to the smaller value of the added value of the preset maximum value of the token counter and the original token counter and the number of refreshed tokens after each traversal in the process where the clock interrupt trigger sequentially traverses the token counters of the active link list.

[0042] In one embodiment of the present invention, for a standby station, the method includes maintaining, by a standby WQE state module, the states of all WQEs in the standby station for WQEs within a bandwidth-sensitive group in the standby station, and outputting all standby WQE states including two states of idle and occupied outside, and further includes performing programming control of the WQE state based on a set and reset interface provided by the standby WQE state module. Here, set is a method of mapping an input to an output by forcibly changing the input from the outside, and reset is changing the value input by the program to the initial state at power-on. A PLC, that is, a programmable logic controller, stores a program therein, executes user-oriented instructions such as logic operations, sequence control, timing, counting, and arithmetic operations, and employs a programmable memory for controlling various types of machines or production processes by digital or analog input / output. The purpose of such a design is to realize a PLC program logic with clearer logic and clearer steps through the instructions of two operations of set and reset.

[0043] The following introduces specific embodiments for implementing the method proposed by the present invention. First, the overall technical solution will be described with reference to FIG. 4. FIG. 4 shows the transmission-side structure of an RDMA network card that supports the realization of performance separation for delay-sensitive applications and bandwidth-sensitive application transmission queues WQE. The WQE cache module is used to mask the read delay of RDMA based on the PCIe interface by caching WQEs of different tenants. Based on tenant identification on the RDMA network card, the present invention groups and caches the WQEs of delay-sensitive tenants and bandwidth-sensitive tenants, and divides them into a delay-sensitive group and a bandwidth-sensitive group. Within the delay-sensitive group and the bandwidth-sensitive group, the WQEs of different tenants are similarly stored separately. For the WQEs of different tenants in the delay-sensitive group, a fair polling scheduler is adopted. The standby station is used to store the WQEs of different tenants in the bandwidth-sensitive group. For the WQEs within the standby station, a token bucket scheduler is adopted. Between the delay-sensitive group and the bandwidth-sensitive group, an inter-group weighted polling scheduler is adopted. The WQE processing model realizes the processing of WQEs. The processing process adopts slice scheduling, and the unprocessed WQEs are written into the standby station for subsequent scheduling.

[0044] FIG. 4 is a structural schematic diagram of the transmission side of an RDMA network card that supports delay-sensitive application and bandwidth-sensitive application performance separation in an embodiment of the present invention. The cache module 01 stores the WQEs from the bandwidth-sensitive tenant by the bandwidth-sensitive group 012, stores the WQEs from the delay-sensitive tenant by the delay-sensitive group 011, and categorizes and stores the WQEs from different tenants using queues inside the bandwidth-sensitive group 012 and the delay-sensitive group 011. The WQEs in the delay-sensitive group 011 are scheduled by the first scheduler 02, the WQEs of the bandwidth-sensitive group are first scheduled to the waiting station 03, and then scheduled by the second scheduler 04, and the group-inter scheduler 05 schedules the WQEs from the delay-sensitive group 011 and the bandwidth-sensitive group 012 to the WQE processing module 06. The WQE processing module 06 completes the processing of the scheduled WQEs of the delay-sensitive group 011 at one time. However, if it cannot finish processing the WQEs of the bandwidth-sensitive group 012 within the limited length, it writes the WQEs back to the waiting station 03 and waits for rescheduling and processing again.

[0045] FIG. 5 is a structural schematic diagram of the transmission side of an RDMA network card that supports delay-sensitive application and bandwidth-sensitive application performance separation in a specific embodiment of the present invention. The cache module 01 stores the WQEs from the bandwidth-sensitive tenants by the bandwidth-sensitive group 012, stores the WQEs from the delay-sensitive tenants by the delay-sensitive group 011, and categorizes and stores the WQEs from different tenants using queues inside the bandwidth-sensitive group 012 and the delay-sensitive group 011. The WQEs within the delay-sensitive group 011 are scheduled by the fair polling scheduler 021 (based on the fair polling algorithm), the WQEs of the bandwidth-sensitive group are first scheduled into the waiting station 03, and then scheduled by the token bucket scheduler 041 (based on the token bucket scheduling algorithm), and the weighted polling scheduler 051 between groups schedules the WQEs from the delay-sensitive group 011 and the bandwidth-sensitive group 012 to the WQE processing module 06. The WQE processing module 06 has already processed the scheduled WQEs of the delay-sensitive group 011, but for the WQEs of the bandwidth-sensitive group 012, if they are not processed within the specified length, the WQEs are written back to the waiting station 03 to wait for being rescheduled and processed again.

[0046] FIG. 6 is a flowchart of fair polling scheduling within a delay-sensitive group in one embodiment of the present invention. In one embodiment of the present invention, the step of polling all queues in the delay-sensitive group using a fair polling scheduler includes polling the head of the queue in the delay-sensitive group by periodically traversing an active link list maintained by the fair polling scheduler, where each node of the active link list corresponds to a queue in one delay-sensitive group, and maintaining the active link list by inserting a node into the active link list for a queue having a WQE to be processed and deleting a node corresponding to a queue having no WQE to be processed. Specifically, in an embodiment of the present invention, by using a fair polling scheduling algorithm to schedule the WQEs of the delay-sensitive group, it is ensured that the WQEs of each delay-sensitive tenant are treated equally by the WQE processing model. The present invention needs to maintain one active link list in the implementation of the fair polling scheduling algorithm. The elements in the active link list (which may be a single link list, a dual link list, or a circular link list) are the numbers of different tenant WQE queues in the delay-sensitive group, and the active link list only includes the active queue numbers where there are elements to be processed in the queue. In particular, when a WQE is written into buffer queue Q_i, if Q_i is empty before writing, the Q_i queue number is inserted into the active link list. FIG. 6 illustrates the execution flow of the fair polling scheduling algorithm used by the present invention, which includes the following:

[0047] Step S610, the inter-group scheduler requests scheduling.

[0048] Step S620: Remove the queue Queue_head of the active link list head. Queue_head is the scheduling result of the current fair polling.

[0049] Step S630: Output the scheduling result Queue_head to the WQE processing module.

[0050] Step S640: After outputting the scheduling result, determine whether Queue_head is empty. If not, return to Step S610.

[0051] Step S650: If so, move the head of the active link list to the tail and return to Step S610.

[0052] The present invention proposes to adopt a processing method for bandwidth-sensitive WQE slices (and write-back) in the design of the WQE processing module, thereby avoiding a bandwidth-sensitive WQE from occupying the WQE processing module for a long time. The WQE slice processing method proposed by the present invention relates to two crucial elements for storing the waiting station of the executable WQE and the WQE processing model that supports the WQE slice.

[0053] FIG. 7 is a schematic diagram of a standby station structure according to an embodiment of the present invention. The standby station adopts a true dual-port SRAM (including a first standby WQE read / write channel 031 and a second standby WQE read / write channel 032) to realize the storage of standby WQEs. The two read / write channels (the first standby WQE read / write channel 031 and the second standby WQE read / write channel 032) realize the read / write of WQEs in the bandwidth-sensitive group at different positions of the standby station 03. Here, the first standby WQE read / write channel 031 is used to store WQEs from the cache module to the standby station, and the second standby WQE read / write channel 032 is used to write back the WQEs that have not been all processed by the WQE processing module to the standby station. The standby WQE status module 033 maintains the status of all standby WQEs through the set interface 034 and the reset interface 035. The standby WQEs include two states: idle and occupied. The standby WQE status module further provides a standby WQE status output interface 036 for outputting all standby WQE statuses outside.

[0054] FIG. 8 is a flowchart of the WQE processing module realizing WQE slicing according to an embodiment of the present invention. Optionally, the present invention can also slice only the tasks corresponding to the WQEs of the bandwidth-sensitive tenant. The flowchart includes the following steps:

[0055] Step S810, the WQE processing module receives the WQE to be processed.

[0056] Step S820, the WQE processing module outputs one data packet corresponding to the WQE based on the path maximum transmission unit (PMTU).

[0057] Step S830, after the data packet is output, check whether the remaining message length of the WQE is greater than 0. If not, return to step S810 and continue to prepare for receiving the next WQE to be processed.

[0058] Step S840, when it is checked that the remaining message length of the WQE is greater than 0, continue to check whether the WQE is from the bandwidth-sensitive group. If not, return to Step S820.

[0059] Step S850, when it is checked that the WQE is from the bandwidth-sensitive group, rewrite the unprocessed WQE to the waiting station and return to Step S810.

[0060] Figure 9 is a flowchart of token counter maintenance in an embodiment of the present invention, specifically showing the update rules of the token counter, which includes the following steps:

[0061] Step S910, the clock interrupts the trigger (sequentially traverses the token counters of the active link list).

[0062] Step S920, initialize the link list access pointer rd_ptr to 1.

[0063] Step S930, read the token_bucket of the rd_ptr-th element of the active link list.

[0064] Step S940, update the token counter in the way of token_bucket = min(B, token_bucket + F), that is, update the token_bucket to the smaller value between B and token_bucket + F.

[0065] Step S950, increment rd_ptr.

[0066] Step S960, check whether the active link list has been traversed. If so, return to Step S910. If not, return to Step S930.

[0067] FIG. 10 is a bandwidth-sensitive WQE scheduling execution flow based on the token bucket algorithm in an embodiment of the present invention, including the following steps:

[0068] Step S1010, the inter-group scheduler requests scheduling.

[0069] Step S1020, remove the queue Queue_head of the active link list head.

[0070] Step S1030, determine whether the token counter token_bucket corresponding to Queue_head is larger than the path maximum transmission unit (PMTU).

[0071] Step S1040, if the token counter token_bucket corresponding to Queue_head is larger than the path maximum transmission unit (PMTU), schedule the WQE corresponding to Queue_head.

[0072] Step S1050, if the token counter token_bucket corresponding to Queue_head is less than or equal to the path maximum transmission unit (PMTU), insert Queue_head at the end of the link list and return to step S1020.

[0073] Step S1060, subtract the PMTU from the token_bucket.

[0074] Step S1070, insert Queue_head at the end of the link list and return to step S1010.

[0075] For step S140, in the embodiments of the present invention, by using a token bucket-based scheduling algorithm to schedule bandwidth-sensitive WQEs in the standby station, a performance separation effect of fairly sharing the available bandwidth by active bandwidth-sensitive tenants is achieved. Embodiments of the present invention maintain one token counter token_bucket for each bandwidth-sensitive tenant, the maximum value of the counter is B, and at the same time maintain one timer, and the refresh period of the timer is T. In some embodiments, when the number of active bandwidth-sensitive group tenants is NA, all bandwidth-sensitive tenants can allocate the bandwidth to BW, and each active bandwidth-sensitive tenant can allocate the bandwidth to BW / NA. In the embodiments of the present invention, the number of tokens refreshed each time in a given token bucket algorithm is F. Based on the principle of token bucket scheduling, BW / NA should be approximated to F / T. The present invention needs to maintain one active link list in the realization of the token bucket scheduling algorithm. The elements in the active link list are the queue numbers of different tenant standby WQEs in the bandwidth-sensitive group, and the active link list only contains the active queue numbers in which the WQEs are in the standby processing state. In particular, when a WQE with a bandwidth-sensitive queue number Q_j is written to the standby station, Q_j is inserted into the active link list. After the WQE with the bandwidth-sensitive queue number Q_k is processed, Q_k is deleted from the active link list.

[0076] Figure 11 shows a scheduling process based on weighted polling in an embodiment of the present invention. The present invention schedules the WQEs of the delay-sensitive group and the bandwidth-sensitive group in a weighted polling manner, and can configure the weights of the two groups, thereby improving the flexibility of differentiated services. The dynamic weights of the two groups, i.e., the weight w_ls for the delay-sensitive group and the weight w_bs for the bandwidth-sensitive group, are maintained, and by default, w_ls = w_bs = 1. At the same time, the present invention needs to maintain the dynamic weights of the two groups, i.e., ls_cnt for the delay-sensitive group and bs_cnt for the bandwidth-sensitive group. Regarding the dynamic weights, ls_cnt and bs_cnt are initialized and configured with the same parameter, and each time a WQE is scheduled, the weight value ls_cnt or bs_cnt of the corresponding category is decremented. The process of weighted polling scheduling includes the following steps as shown in Figure 11:

[0077] Step S1110, the WQE processing module requests scheduling.

[0078] Step S1120, determine whether ls_cnt is greater than or equal to bs_cnt. If so, sequentially execute steps S1130 and S1140; otherwise, sequentially execute steps S1150 and S1160.

[0079] Step S1130, schedule the delay-sensitive group.

[0080] Step S1140, decrement ls_cnt and execute step S1170.

[0081] Step S1150, schedule the bandwidth-sensitive group.

[0082] Step S1160, decrement bs_cnt and execute step S1170.

[0083] Step S1170, determine whether ls_cnt and bs_cnt are both equal to 0 at the same time. If not, return to step S1110. If so, execute step S1180.

[0084] Step S1180, re-initialize ls_cnt and bs_cnt to w_ls and w_bs.

[0085] Furthermore, in yet another embodiment of the present invention, in order for two RDMA network cards to achieve efficient dual redundancy backup, it is necessary to ensure that these two network cards have the same physical address and IP address. In the case of the upper-layer application system, the feature of "single network card" is presented in the system. Conversely, when one network card in the system switches to another block network card and operates, if the IP address changes, the system cannot send and receive data normally. When the IP address remains unchanged and the physical address changes, it causes a change in the ARP binding table in the protocol stack. When re-corresponding the relationship between the IP address and the network card physical address in the ARP binding table, the switching time between the two network cards can be extended. However, the physical address of each network card is unique worldwide and is stored in the PROM of the network card. In order for the two network cards to have the same physical address, when initializing the network card, read the physical address of one of the network cards from the PROM, and write the content of the physical address to the physical address register and the data structure variable of another block network card. In this case, these two network cards have exactly the same physical address. Based on this design, the present invention can realize the switching of dual network cards.

[0086] Another aspect of the present invention provides an RDMA network card. FIG. 12 is a structural schematic diagram of an RDMA network card in an embodiment of the present invention. The RDMA network card includes a cache module 01, a first scheduler 02, a waiting station 03, a second scheduler 04, an inter-group scheduler 05, a WQE processing module 06, a control module 07, a transmission module 08, and a reception module 09. Here, The control module 07 is used to recognize the tenant type based on the tenant type identifier included in the queue pair identifier, and cache the working queue element WQE of the transmission queue of the recognized delay-sensitive tenant as a delay-sensitive WQE in the delay-sensitive group of the cache module 01, and cache the WQE of the transmission queue of the recognized bandwidth-sensitive tenant as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module 01. Here, at least a part of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module 01 are stored in the waiting station 03, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE. The waiting station 03 is used to store at least a part of the WQEs to be processed in the bandwidth-sensitive group. Here, the structure of the waiting station 03 is shown in detail in FIG. 7 and includes a true dual-port SRAM, a waiting WQE state module 033, a set interface 034, and a reset interface 035. The true dual-port SRAM includes a first waiting WQE read / write channel 031 for transmitting the bandwidth-sensitive WQEs in the bandwidth-sensitive group in the cache module to the waiting station, and a second waiting WQE read / write channel 032 for writing back the WQEs corresponding to the unoutput data packets to the waiting station. The waiting WQE state module 033 is used to maintain the states of all the WQEs in the waiting station and output all the waiting WQE states to the outside.

[0087] The set interface 034 and the reset interface 035 are used to perform programming control of the WQE state. Here, the waiting WQE state includes two states: idle and occupied.

[0088] The inter-group scheduler 05 determines whether to schedule either a WQE in the delay-sensitive group or a WQE in the bandwidth-sensitive group. When scheduling a WQE in the delay-sensitive group, it sends a scheduling request to the first scheduler. When it decides to schedule a WQE in the bandwidth-sensitive group, it sends a scheduling request to the second scheduler. It is also used to transmit the WQE scheduling results from the first scheduler 02 and the second scheduler 04 to the WQE processing module. The first scheduler 05 uses the first scheduling algorithm to schedule the delay-sensitive WQEs in the delay-sensitive group based on the scheduling request from the inter-group scheduler, and is used to return the WQE scheduling result to the inter-group scheduler 05. The second scheduler 04 uses the second scheduling algorithm to schedule the bandwidth-sensitive WQEs in the bandwidth-sensitive group at the waiting station 03, and is used to return the WQE scheduling result to the inter-group scheduler 05. The WQE processing module 06 processes the delay-sensitive WQEs based on the WQE scheduling results and outputs all the data packets corresponding to the delay-sensitive WQEs. It also processes the bandwidth-sensitive WQEs based on the WQE scheduling results and outputs one data packet corresponding to the bandwidth-sensitive WQE based on the path maximum transmission unit. If there are data packets that have not been output, it is used to write back the WQEs corresponding to the unoutput data packets to the waiting station.

[0089] The control module 07 is used to recognize the tenant type based on the tenant type identification included in the queue pair identification, cache the working queue element WQE of the transmission queue of the recognized delay-sensitive tenant as a delay-sensitive WQE in the delay-sensitive group of the cache module, and cache the WQE of the transmission queue of the recognized bandwidth-sensitive tenant as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module. Here, at least some of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in the waiting station, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE.

[0090] The transmission module 08 has a conventional configuration of an RDMA network card for transmitting data corresponding to the WQE.

[0091] The receiving module 09 has a conventional configuration of an RDMA network card for receiving data from other network cards.

[0092] Corresponding to the above method, the present invention further provides a performance separation system for an RDMA network card transmission queue, the system includes a computer device including a memory for storing computer instructions and a processor for executing the computer instructions stored by the memory, and when the computer instructions are executed by the processor, the system realizes the steps of the method.

[0093] The embodiments of the present invention further provide a computer-readable storage medium storing a computer program that realizes the steps of the above method when executed by a processor. The computer-readable storage medium may be a tangible storage medium such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the technical field.

[0094] Compared with the prior art, the RDMA network card transmission queue performance separation method, system and storage medium proposed by the present invention have the following advantages: (1) The WQEs from delay-sensitive tenants and bandwidth-sensitive tenants are respectively stored, scheduled and processed so as to realize performance separation between delay-sensitive applications and bandwidth-sensitive applications on the transmission side of the RDMA network card. (2) For the WQEs of bandwidth-sensitive tenants, a method is designed in which they are written back to the waiting station without being completed in one transmission and rescheduled, so as to ensure that bandwidth-sensitive applications based on RDMA can fairly obtain the bandwidth resources on the transmission side, effectively avoid a WQE of a bandwidth-sensitive application from occupying the bandwidth resources on the transmission side for a long time, effectively avoid a bandwidth-sensitive application from occupying the WQE processing module resources for a long time, and realize performance separation between bandwidth-sensitive applications. (3) A fair polling scheduler is used on the transmission side of the RDMA network card to schedule the corresponding WQEs, ensuring that delay-sensitive applications can fairly obtain the opportunity to be processed by the WQE processing module.

[0095] Those skilled in the art should understand that each exemplary component, system, and method described in connection with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether it is implemented in a hardware manner or a software manner depends on the specific application of the technical solution and the design constraints. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementations should not be considered as exceeding the scope of the present invention. When implemented in a hardware manner, it may be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in a software manner, the elements of the present invention are used to execute the programs or code segments of the required tasks. The programs or code segments may be stored in a machine-readable medium or transmitted on a transmission medium or communication link-up by a data signal carried by a carrier.

[0096] It is necessary to clarify that the present invention is not limited to the specific configurations and processes described above and shown in the figures. The detailed description of known methods is omitted here for the sake of brevity. In the above-described embodiments, some specific steps are described and shown as examples. However, the method steps of the present invention are not limited to the specific steps described and illustrated. Those skilled in the art can make various changes, modifications, and additions or change the order between steps after understanding the spirit of the present invention.

[0097] In the present invention, the features described and / or exemplified with respect to one embodiment can be used in the same or similar way in one or more other embodiments, and / or combined with the features of other embodiments, or replace the features of other embodiments.

[0098] The above description is only a preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, various modifications and changes are possible to the embodiments of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should all be included within the protection scope of the present invention.

[0099] (Appendix) (Appendix 1) A method for separating the performance of an RDMA network card transmission queue, comprising: The transmitting-side RDMA network card recognizes the tenant type based on the tenant type identifier included in the queue pair identifier, and caches the working queue element WQE of the transmission queue of the recognized delay-sensitive tenant as a delay-sensitive WQE in the delay-sensitive group of the cache module, and caches the WQE of the transmission queue of the recognized bandwidth-sensitive tenant as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module, wherein at least a part of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in the waiting station, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE, and Using an inter-group scheduler to determine whether to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group, and sending a scheduling request to the first scheduler when it is determined to schedule the WQE in the delay-sensitive group, and sending a scheduling request to the second scheduler when it is determined to schedule the WQE in the bandwidth-sensitive group, and The first scheduler schedules the delay-sensitive WQEs in the delay-sensitive group using the first scheduling algorithm based on the scheduling request of the inter-group scheduler, and returns the WQE scheduling result to the inter-group scheduler, and The second scheduler schedules the bandwidth-sensitive WQEs in the waiting stations using a second scheduling algorithm, and returns the WQE scheduling results to the inter-group scheduler. The inter-group scheduler transmits the WQE scheduling results from the first scheduler and the second scheduler to the WQE processing module on the transmitting-side RDMA network card. The WQE processing module processes the latency-sensitive WQEs based on the WQE scheduling results, outputs all data packets corresponding to the latency-sensitive WQEs, processes the bandwidth-sensitive WQEs based on the WQE scheduling results, outputs one data packet corresponding to the bandwidth-sensitive WQEs based on the path maximum transmission unit, and if there are data packets that have not been output, rewrites the WQEs corresponding to the data packets that have not been output to the waiting stations. A method for performance separation of an RDMA network card transmission queue.

[0100] (Appendix 2) The method includes: receiving a tenant identification request by a drive interface on the RDMA network card, and providing tenant identification including tenant type identification to the tenant based on the identification request; For each tenant to which tenant identification is assigned, creating a queue pair having the queue pair identification including the tenant type identification, and allocating queue pair resources for storing WQEs to the queue pair. The method according to Appendix 1.

[0101] (Appendix 3) The tenant type identification is a binary number with a preset length, and the tenant type identification is the last bit of the tenant identification. The method according to Appendix 2.

[0102] (Appendix 4) The first scheduling algorithm is a fair polling scheduling algorithm. The step of using the first scheduling algorithm to schedule delay-sensitive WQEs within a delay-sensitive group based on the scheduling requirements of the inter-group scheduler by the first scheduler as described above, and returning the WQE scheduling result to the inter-group scheduler is as follows: Receiving, by the first scheduler, the scheduling requirements of the inter-group scheduler, traversing the queues of delay-sensitive WQEs of different tenants in the active link list maintained by the first scheduler based on the fair polling scheduling algorithm, and polling the queues of delay-sensitive WQEs of different tenants within the delay-sensitive group. Here, the queue numbers of the delay-sensitive WQEs to be processed by different tenants within the delay-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one delay-sensitive WQE. Returning the WQE of the polled active link list head queue to the inter-group scheduler as the WQE scheduling result, and deleting the corresponding node from the active link list. The method according to Appendix 1 includes the above steps.

[0103] (Appendix 5) The second scheduling algorithm is a token bucket scheduling algorithm. The step of using the second scheduling algorithm to schedule bandwidth-sensitive WQEs within a bandwidth-sensitive group at a standby station by the second scheduler as described above, and returning the WQE scheduling result to the inter-group scheduler is as follows: Receiving, by a second scheduler, scheduling requests of an inter-group scheduler, traversing token counters of queues of bandwidth-sensitive WQEs of different tenants in an active link list maintained by the second scheduler based on a token bucket scheduling algorithm, and refreshing tokens according to a predetermined refresh period and the number of tokens refreshed each time to update the token counter count, where queue numbers of bandwidth-sensitive WQEs to be processed by different tenants in a bandwidth-sensitive group are recorded in the active link list, and one node of the active link list corresponds to a non-empty queue of one bandwidth-sensitive WQE, the method described in Appendix 1.

[0104] (Appendix 6) The method further includes For WQEs in a bandwidth-sensitive group at a standby station, maintaining, by a standby WQE state module, the states of all WQEs at the standby station, and outputting all standby WQE states including two states of idle and occupied outside, and further performing programming control of the WQE state based on a set and reset interface provided by the standby WQE state module, the method described in Appendix 1.

[0105] (Appendix 7) The step of determining to schedule either WQEs in a delay-sensitive group or WQEs in a bandwidth-sensitive group using the inter-group scheduler as described above includes The inter-group scheduler determines to schedule either WQEs in a delay-sensitive group or WQEs in a bandwidth-sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on each scheduling result, the method described in Appendix 1.

[0106] (Appendix 8) The method described in Appendix 7, which includes updating dynamic weights based on the scheduling results of each time, reducing the weights of delay-sensitive groups or bandwidth-sensitive groups corresponding to the scheduling results based on the scheduling results of each time.

[0107] (Appendix 9) An RDMA network card, A control module for recognizing tenant types based on tenant type identifications included in queue pair identifications, caching the working queue element WQEs of the transmission queues of the recognized delay-sensitive tenants as delay-sensitive WQEs in the delay-sensitive group of the cache module, and caching the WQEs of the transmission queues of the recognized bandwidth-sensitive tenants as bandwidth-sensitive WQEs in the bandwidth-sensitive group of the cache module. Here, at least some of the WQEs to be processed cached in the bandwidth-sensitive group of the cache module are stored in the waiting station, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE. The waiting station for recording at least some of the WQEs to be processed in the bandwidth-sensitive group, A cross-group scheduler for determining whether to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group, sending a scheduling request to the first scheduler when determining to schedule the WQE in the delay-sensitive group, sending a scheduling request to the second scheduler when determining to schedule the WQE in the bandwidth-sensitive group, and transmitting the WQE scheduling results from the first scheduler and the second scheduler to the WQE processing module. A first scheduler for scheduling the delay-sensitive WQEs in the delay-sensitive group using a first scheduling algorithm based on the scheduling request of the cross-group scheduler and returning the WQE scheduling result to the cross-group scheduler. Using a second scheduling algorithm, schedule the bandwidth-sensitive WQEs within the bandwidth-sensitive group at the standby station, and a second scheduler for returning the WQE scheduling result to the inter-group scheduler; A WQE processing module that processes delay-sensitive WQEs based on the WQE scheduling result to output all data packets corresponding to the delay-sensitive WQEs, processes bandwidth-sensitive WQEs based on the WQE scheduling result, outputs one data packet corresponding to the bandwidth-sensitive WQE based on the path maximum transmission unit, and, if there are data packets that have not been output, writes back the WQEs corresponding to the unoutput data packets to the standby station. The RDMA network card includes these components.

[0108] (Appendix 10) The standby station includes: A true dual-port SRAM including a first standby WQE read / write channel for transmitting the bandwidth-sensitive WQEs within the bandwidth-sensitive group in the cache module to the standby station, and a second standby WQE read / write channel for writing back the WQEs corresponding to the unoutput data packets to the standby station; A standby WQE state module for maintaining the states of all WQEs at the standby station and outputting all standby WQE states externally; A set and reset interface for performing programming control of the WQE state. Here, the standby WQE state includes a set and reset interface including two states: idle and occupied. The RDMA network card according to Appendix 9 includes these components.

[0109] (Appendix 11) The first scheduling algorithm is a fair polling scheduling algorithm; The second scheduling algorithm is a token bucket scheduling algorithm; The inter-group scheduler determines to schedule either the WQE in the delay-sensitive group or the WQE in the bandwidth-sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on each scheduling result, the RDMA network card described in Appendix 9.

Claims

1. A method for performance isolation of an RDMA network card transmit queue, comprising: The sending RDMA network card recognizes a tenant type according to the tenant type identification included in the queue pair identification, caches the working queue element WQE of the transmission queue of the recognized delay-sensitive tenant as a delay-sensitive WQE in the delay-sensitive group of the cache module, and caches the WQE of the transmission queue of the recognized bandwidth-sensitive tenant as a bandwidth-sensitive WQE in the bandwidth-sensitive group of the cache module, where at least a part of the WQE to be processed cached in the bandwidth-sensitive group of the cache module is stored in a waiting station, and the WQEs of different tenants in the delay-sensitive group and the bandwidth-sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE; determining to schedule either a WQE in the delay sensitive group or a WQE in the bandwidth sensitive group using an inter-group scheduler, and sending a scheduling request to a first scheduler when determining to schedule a WQE in the delay sensitive group and sending a scheduling request to a second scheduler when determining to schedule a WQE in the bandwidth sensitive group; Schedule the delay-sensitive WQEs in the delay-sensitive group by the first scheduler according to the scheduling request of the inter-group scheduler using a first scheduling algorithm, and return the WQE scheduling result to the inter-group scheduler; scheduling the bandwidth-sensitive WQEs at the waiting stations by a second scheduler using a second scheduling algorithm, and returning the WQE scheduling results to the inter-group scheduler; transmitting the WQE scheduling results from the first scheduler and the second scheduler to a WQE processing module on the sending RDMA network card by an inter-group scheduler; A method for performance isolation of an RDMA network card transmission queue, comprising: processing a delay-sensitive WQE by a WQE processing module according to a WQE scheduling result, and outputting all data packets corresponding to the delay-sensitive WQE; processing a bandwidth-sensitive WQE according to the WQE scheduling result, and outputting one data packet corresponding to the bandwidth-sensitive WQE according to a path maximum transmission unit; and, if there is a data packet that has not been output, writing back the WQE corresponding to the data packet that has not been output to the waiting station.

2. The method comprises: receiving a request for an identification of a tenant by a drive interface on the RDMA network card and providing a tenant identification, including a tenant type identification, to the tenant based on the identification request; 2. The method of claim 1, further comprising: for each tenant assigned a tenant identification, creating a queue pair having the queue pair identification including a tenant type identification, and allocating queue pair resources for storing WQEs to the queue pair.

3. The method of claim 2 , wherein the tenant type identification is a binary number of a preset length, and the tenant type identification is the last bit of the tenant identification.

4. The first scheduling algorithm is a fair polling scheduling algorithm, and the above-mentioned step of scheduling the delay-sensitive WQE in the delay-sensitive group by the first scheduler according to the scheduling request of the inter-group scheduler using the first scheduling algorithm and returning the WQE scheduling result to the inter-group scheduler includes: receiving a scheduling request of an inter-group scheduler by a first scheduler, traversing queues of delay-sensitive WQEs of different tenants in an active linked list maintained by the first scheduler based on a fair polling scheduling algorithm, and polling queues of delay-sensitive WQEs of different tenants in a delay-sensitive group, wherein queue numbers of delay-sensitive WQEs to be processed by different tenants in a delay-sensitive group are recorded in the active linked list, and one node of the active linked list corresponds to a non-empty queue of one delay-sensitive WQE; 2. The method of claim 1, further comprising: returning a WQE of a queue at the head of the polled active linked list to the inter-group scheduler as a WQE scheduling result, and removing a corresponding node from the active linked list.

5. The second scheduling algorithm is a token bucket scheduling algorithm, and the step of scheduling the bandwidth-sensitive WQEs in the bandwidth-sensitive group at the waiting station by the second scheduler using the second scheduling algorithm and returning the WQE scheduling result to the inter-group scheduler includes:

2. The method of claim 1, further comprising: receiving a scheduling request of an inter-group scheduler by a second scheduler; traversing token counters of queues of bandwidth-sensitive WQEs of different tenants in an active linked list maintained by the second scheduler based on a token bucket scheduling algorithm; refreshing tokens and updating token counter counts according to a predetermined refresh period and a token refresh number each time; wherein the active linked list records queue numbers of bandwidth-sensitive WQEs to be processed by different tenants in a bandwidth-sensitive group, and one node of the active linked list corresponds to a non-empty queue of one bandwidth-sensitive WQE.

6. The method comprises: The method of claim 1, further comprising: for WQEs in a bandwidth sensitive group in a waiting station, maintaining the states of all WQEs in the waiting station by a waiting WQE state module, and outputting all waiting WQE states including two states, idle and occupied, and performing programming control of the WQE states based on a set and reset interface provided by the waiting WQE state module.

7. The step of determining whether to schedule a WQE in a delay sensitive group or a WQE in a bandwidth sensitive group using an inter-group scheduler, as described above, comprises:

2. The method of claim 1, further comprising: determining to schedule either a WQE in the delay sensitive group or a WQE in the bandwidth sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay sensitive group and the bandwidth sensitive group; and updating the dynamic weights based on each scheduling result.

8. The method according to claim 7, wherein the updating of the dynamic weight based on each scheduling result includes reducing a weight of the delay-sensitive group or the bandwidth-sensitive group corresponding to the scheduling result based on each scheduling result.

9. 1. An RDMA network card, comprising: a control module for recognizing a tenant type based on a tenant type identification included in the queue pair identification, caching the working queue element WQEs of the transmit queue of the recognized delay sensitive tenant as delay sensitive WQEs in a delay sensitive group of the cache module, and caching the WQEs of the transmit queue of the recognized bandwidth sensitive tenant as bandwidth sensitive WQEs in a bandwidth sensitive group of the cache module, wherein at least a part of the WQEs to be processed that are cached in the bandwidth sensitive group of the cache module are stored in a waiting station, and the WQEs of different tenants in the delay sensitive group and the bandwidth sensitive group included in the cache module are cached according to different queues, and each node of the queue is one WQE; said waiting station for recording the WQEs to be processed of at least a portion of said bandwidth sensitive group; an inter-group scheduler for determining to schedule either a WQE in the delay sensitive group or a WQE in the bandwidth sensitive group, sending a scheduling request to the first scheduler when determining to schedule a WQE in the delay sensitive group, sending a scheduling request to the second scheduler when determining to schedule a WQE in the bandwidth sensitive group, and transmitting WQE scheduling results from the first scheduler and the second scheduler to a WQE processing module; a first scheduler for scheduling delay-sensitive WQEs in the delay-sensitive group using a first scheduling algorithm according to a scheduling request of the inter-group scheduler, and returning the WQE scheduling result to the inter-group scheduler; a second scheduler for scheduling bandwidth-sensitive WQEs in the bandwidth-sensitive group at the waiting station using a second scheduling algorithm and returning the WQE scheduling results to the inter-group scheduler; and a WQE processing module for processing a delay-sensitive WQE based on a WQE scheduling result to output all data packets corresponding to the delay-sensitive WQE, and for processing a bandwidth-sensitive WQE based on a WQE scheduling result to output one data packet corresponding to the bandwidth-sensitive WQE based on a path maximum transmission unit, and, if there is a data packet that has not been output, writing back the WQE corresponding to the data packet that has not been output to the waiting station.

10. The waiting station includes: a true dual port SRAM including a first standby WQE read / write channel for transmitting bandwidth sensitive WQEs in a bandwidth sensitive group in the cache module to a standby station, and a second standby WQE read / write channel for writing back to the standby station WQEs corresponding to data packets that have not been output; a waiting WQE status module for maintaining the status of all WQEs in the waiting station and outputting all waiting WQE statuses to the outside; 10. The RDMA network card of claim 9, further comprising a set and reset interface for programming control of a WQE state, wherein the wait WQE state includes two states: idle and occupied.

11. the first scheduling algorithm is a fair polling scheduling algorithm; the second scheduling algorithm is a token bucket scheduling algorithm; The RDMA network card of claim 9, wherein the inter-group scheduler determines to schedule either a WQE in the delay-sensitive group or a WQE in the bandwidth-sensitive group based on a weighted polling scheduling algorithm based on dynamic weights configured for the delay-sensitive group and the bandwidth-sensitive group, and updates the dynamic weights based on each scheduling result.

Citation Information

Patent Citations

  • Data center multi-application QoS guarantee system and method based on intelligent network card unloading

    CN114666281A