Data transmission method, device and product
By locking the channel group in RDMA technology and releasing the channel after data transmission is completed, the problem of time-consuming locking operations is solved, the flexibility and efficiency of data transmission are improved, and the synchronous transmission process is optimized.
Patent Information
- Application Number
- CN202410516849.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-10-28
Smart Images

Figure CN120849077A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computers, and more specifically to data transmission methods, apparatus and computer program products. Background Art
[0002] Typically, network communication involves numerous system context switches and data copying between user space and kernel space during data transmission, leading to high latency and CPU load. In contrast, Remote Direct Memory Access (RDMA) technology significantly improves data transmission efficiency and reduces latency by allowing user programs to bypass the operating system kernel and interact directly with the network card for network communication.
[0003] In RDMA technology, queue pairs (QPs) play a crucial role. Each QP consists of a send queue and a receive queue, which together handle data transmission and reception. The use of QPs enables RDMA communication to perform data transfer directly between memory without operating system intervention. Summary of the Invention
[0004] Embodiments of this disclosure provide a data transmission method, apparatus, and computer program product. In a first aspect of an embodiment of this disclosure, a data transmission method is provided. The method includes obtaining a first allocation request for a channel from a first thread. The method further includes locking a first set of channels for the first thread, wherein multiple channels in the first set of channels correspond to multiple remote direct memory access (RDMA) devices, the first thread submits first data to a first RDMA device corresponding to a first channel in the first set of channels, and a first completion channel corresponds to the first channel. The method further includes detecting whether the first data exists in the first completion channel. The method further includes releasing the first set of channels in response to the presence of first data in the first completion channel.
[0005] In a second aspect of embodiments of this disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage device for storing one or more programs that, when executed by the one or more processors, cause the one or more processors to perform actions, including obtaining a first allocation request for a channel from a first thread. These actions also include locking a first set of channels for the first thread, wherein a plurality of channels in the first set of channels correspond to a plurality of remote direct memory access (RDMA) devices, the first thread submitting first data to a first RDMA device corresponding to a first channel in the first set of channels, and a first completion channel corresponding to the first channel. These actions also include detecting whether first data exists in the first completion channel. These actions further include releasing the first set of channels in response to the presence of first data in the first completion channel.
[0006] In a third aspect of embodiments of this disclosure, a computer program product is provided, tangibly stored on a non-volatile computer-readable medium and including machine-executable instructions that, when executed, cause a machine to perform actions, including obtaining a first allocation request for a channel from a first thread. These actions also include locking a first set of channels for the first thread, wherein a plurality of channels in the first set of channels correspond to a plurality of remote direct memory access (RDMA) devices, the first thread submitting first data to a first RDMA device corresponding to a first channel in the first set of channels, and a first completion channel corresponding to the first channel. These actions also include detecting whether the first data exists in the first completion channel. These actions further include releasing the first set of channels in response to the presence of first data in the first completion channel.
[0007] It should be understood that the description in the Summary of the Invention section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0008] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 This is a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0010] Figure 2 This is a flowchart of a data transmission method according to some embodiments of the present disclosure;
[0011] Figure 3 This is a schematic diagram of a switching queue pair according to an embodiment of the present disclosure;
[0012] Figure 4 This is a schematic diagram of data transmission according to an embodiment of the present disclosure;
[0013] Figure 5 This is a schematic block diagram of an example device that can be used to implement embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
[0016] In related technologies, when a thread used for synchronous data transmission begins submitting data to an RDMA device, it locks a queue pair corresponding to that RDMA device. The thread must ensure that not only is the data submitted to the RDMA device, but the RDMA device also successfully sends the submitted data before leaving the RDMA device and ending the current round of data transmission. This locking mechanism ensures the consistency and security of data transmission. In RDMA-based communication, both the sending and receiving devices maintain their own queue pairs and transmit data through these queue pairs. This locking mechanism ensures the atomicity of data transmission, meaning that data transmission will not be interrupted by other operations. However, this locking operation is very time-consuming and computationally resource-intensive; improving data transmission efficiency is a pressing issue.
[0017] To address this, this disclosure proposes a data transmission method. An embodiment of the method includes obtaining a first allocation request for a channel from a first thread. The method further includes locking a first group of channels for the first thread, wherein multiple channels in the first group correspond to multiple Remote Direct Memory Access (RDMA) devices, the first thread submits first data to a first RDMA device corresponding to a first channel in the first group, and a first completion channel corresponds to the first channel. The method further includes detecting whether the first data exists in the first completion channel. The method further includes releasing the first group of channels in response to the presence of first data in the first completion channel. Therefore, in a scenario of synchronous data transmission, locking the first group of channels for each thread that needs to transmit data allows the thread used for synchronous data transmission to switch to another channel to transmit data without unlocking and locking operations. Thus, using the method of this disclosure can improve the flexibility and efficiency of data transmission.
[0018] Figure 1 This is a schematic diagram of an example environment 100 that can be implemented according to embodiments of the present disclosure. Figure 1 As shown, environment 100 may include thread 101, network 102, manager 103, channel set 104, first group of channels 105, and data receiving device 106. Channel set 104 is communicatively coupled to data receiving device 106 via network 102. Network 102 may be, for example, a wide area network (WAN), local area network (LAN), wireless network, public telephone network, intranet, and any other type of network well known to those skilled in the art.
[0019] In this embodiment, the data transmission method is primarily executed by the manager 103. The manager 103 may be, for example, a control over the channel set 104, used for resource management of the channel set 104. The manager 103 may be applied in the mirroring process of storing data, for example. In this embodiment, the method executed by the manager 103 includes the following steps: The manager 103 obtains a first allocation request for a channel from a first thread (hereinafter referred to as "thread") 101. RDMA technology provides the ability to directly access remote memory, allowing the sending thread to directly write data to the receiving thread's memory when it needs to send data. This operation is performed directly at the hardware level, bypassing the intervention of the operating system kernel and CPU. To utilize the channel set 104, thread 101 can first communicate with the manager 103, which is used for control, to request channel allocation.
[0020] Manager 103 locks the first group of channels 105 for thread 101. Multiple channels in the first group of channels 105 correspond to multiple Remote Direct Memory Access (RDMA) devices. Thread 101 submits first data to the first RDMA device corresponding to the first channel in the first group of channels. The first completion channel corresponds to the first channel. Each RDMA device corresponding to the multiple RDMA devices in channel set 104 can be configured with several channels for data transmission. The first group of channels 105 can correspond to multiple RDMA devices. In some embodiments, the first group of channels 105 corresponds to four RDMA devices, comprising four channels, each corresponding to one of the four RDMA devices. This allows thread 101 to flexibly switch to another RDMA device (i.e., the RDMA device corresponding to another channel in the first group of channels 105) to transmit data without repeatedly unlocking and locking, improving data transmission efficiency and flexibility.
[0021] The manager 103 checks whether the first data exists in the first completion channel. If the first completion channel is a first completion queue, a corresponding completion queue entity exists for each successfully transmitted data. The first completion queue reads data according to a first-in, first-out (FIFO) principle. The manager 103 can check whether the first data was successfully transmitted by reading the completion queue entity in the first completion queue. The process of reading the completion queue entity is asynchronous, meaning the RDMA device does not need to wait for the detection to complete before continuing to execute other tasks. In this operation, the manager 103 can access the first channel in the first group of channels 105 and then poll the first completion channel corresponding to that first channel.
[0022] In response to the presence of first data on the first completion channel, manager 103 releases the first group of channels 105. If the first data exists on the first completion channel, it means that the first data has been transmitted and the data transmission task of thread 101 has been completed. There is no need to continue locking the first group of channels 105, so the first group of channels 105 can be released.
[0023] like Figure 1As shown, in environment 100, network 102 can be used to transmit data between channel set 104 and data receiving device 106. Network 102 has a theoretical bandwidth, which refers to the maximum transmission speed supported by network 102. It represents the maximum amount of data that network 102 can transmit under ideal conditions, usually measured in bits per second (bps). For example, if the theoretical bandwidth of network 102 is 100 Mbps, it means that under ideal conditions it can transmit one hundred megabits of data per second. However, in reality, due to other factors that may exist in the network (e.g., signal interference, bandwidth sharing, transmission delay, etc.), the actual transmission speed may not reach 100 Mbps.
[0024] As understood by those skilled in the art, the manager 103 can be integrated into a set of RDMA devices corresponding to the channel set 104, using the processor of the RDMA device without using the processor of other devices, thereby enabling data transfer without consuming the computing resources of other devices, such as enabling mirrored storage of data.
[0025] Figure 2 This is a flowchart of a data transmission method according to some embodiments of this disclosure. For example... Figure 2 As shown, flowchart 200 includes blocks 202-208. In block 202, a first allocation request for a channel is obtained from the thread. In this operation, the thread does not need to specify information such as the identifier of the RDMA device to be used in the first allocation request as required in related technologies; it only needs to inform, for example, the manager, that the thread needs a channel to transmit data.
[0026] In box 204, a first group of channels is locked for the first thread. Multiple channels in this first group correspond to multiple Remote Direct Memory Access (RDMA) devices. The first thread submits first data to the first RDMA device corresponding to the first channel in the first group. The first completion channel corresponds to the first channel. This group of channels includes multiple channels corresponding to different RDMA devices. The locking operation ensures that the thread exclusively uses these channels during data transmission, avoiding contention for channels by other threads, thus guaranteeing the stability and reliability of data transmission. After channel locking, the thread selects the first channel as the data transmission channel and submits the first data to the first RDMA device corresponding to the first channel in the first group, instead of submitting the first data on all or some of the locked channels. The first data here can be any type of data packet, such as a file, video stream, database record, etc. Through the RDMA device, data can be directly transferred from the sender's memory to the receiver's memory without operating system processing, thus greatly improving data transmission efficiency.
[0027] In box 206, the presence of first data in the first completion channel is detected. In some embodiments, the first completion channel is a first completion queue, which has a first-in-first-out data structure. The success of the first data transmission can be detected by reading the completion queue entity in the first completion queue. The process of reading the completion queue entity is asynchronous, meaning that the RDMA device can continue to execute other tasks without waiting for the detection to complete.
[0028] In box 208, in response to the presence of first data in the first completion channel, the first group of channels is released. Once completion information related to the first data is detected in the first completion channel, it is known that the first data has been successfully transmitted. At this point, a release operation can be triggered to release the previously locked first group of channels back to the channel resource pool so that other threads can re-request and use these channels.
[0029] Therefore, in the scenario of synchronous data transmission, the first set of channels is locked for each thread that needs to transmit data. In this way, the thread used for synchronous data transmission can switch to another channel to transmit data without unlocking and locking operations. Therefore, the method disclosed in this paper can improve the flexibility and efficiency of data transmission.
[0030] This disclosure is based on RDMA technology. The first group of channels can be a first group of queue pairs (QPs), which may include multiple queue pairs. In some embodiments, the first group of queue pairs is provided with an iterator. In this embodiment, assuming the current iterator points to the first queue pair, when the first queue pair fails, the iterator of the first group of queues is modified to point to a second queue pair in the first group, wherein the second queue pair is different from the first queue pair. Further, a thread submits first data to a second RDMA device corresponding to the second queue pair. When the first group of queue pairs includes four queue pairs, the iterator can increment by 1 with each iteration, and return to 1 when it reaches 4 and needs to iterate again.
[0031] The implementation provides a solution for handling situations where a queue pair fails. By setting an iterator, queue pairs can be automatically switched within the locked first queue pair. Threads do not need to repeatedly request locks on queue pairs outside the first queue pair, nor do they need to make any allocation requests within the first queue pair, thereby improving the efficiency and flexibility of data transmission.
[0032] In some embodiments, for any assigned set of queue pairs, the RDMA device corresponding to each queue pair is different from each other. This is very helpful for threads. Sometimes, the failure of a queue pair is due to a failure of the corresponding RDMA device, so all queue pairs set by that RDMA device are considered faulty. However, the threads and the manager cannot recognize this situation without further assistance, resulting in threads being unable to transmit data even if randomly assigned to the same faulty RDMA device. In related technologies, neither locking the channel group nor limiting each channel in the channel group to a different RDMA device can lead to repeated locking and unlocking for many threads when an RDMA device fails, resulting in a significant waste of computational resources. In this embodiment, restricting each queue pair to different RDMA devices allows threads to directly switch to another RDMA device when automatically switching queue pairs within the first set using an iterator, without needing to repeatedly switch queue pairs.
[0033] Figure 3 This is a schematic diagram of switching queue pairs according to an embodiment of the present disclosure. It shows an iterator 350. The first group of queue pairs 302 includes four queue pairs: a first queue pair 3022, a second queue pair 3024, a third queue pair 3026, and a fourth queue pair 3028. Queue pairs 3022, 3024, 3026, and 3028 correspond to the first RDMA device 310; queue pairs 3042, 3044, 3046, and 3048 correspond to the second RDMA device 320; queue pairs 3062, 3064, 3066, and 3068 correspond to the third RDMA device 330; and queue pairs 3082, 3084, 3086, and 3088 correspond to the fourth RDMA device 340. Similarly, the second group of queue pairs 304 includes four queue pairs: a first queue pair 3042, a second queue pair 3044, a third queue pair 3046, and a fourth queue pair 3048. The rest will not be elaborated further.
[0034] Thread 360 can currently submit data (i.e., the first data) to the first queue pair 3022 in the locked first queue pair 302, and the value of iterator 350 is 1. During the data transmission process, if the first RDMA device 310 fails, the iterator 350 will automatically increment by 1, becoming 2, that is, move to position 3024 according to the dashed arrow. Thread 360 then submits data to the second queue pair 3024.
[0035] In some embodiments, the first completion channel is a first completion queue corresponding to a first queue pair, and the first thread has a first list indicating the transmission status of each piece of data in the first data. After detecting whether the first data exists in the first completion channel, this embodiment further includes updating the first list to indicate that the first part of the data was successfully transmitted if the first completion queue contains a first portion of the first data, wherein the first portion of the data includes at least one piece of data. Further, according to the first list, the thread submits data other than the first portion of the first data to the second RDMA device. In this embodiment, the success or failure of the transmission of each piece of data is determined by setting a list to monitor the transmission status. In this way, if the queue pair is switched, it is not necessary to retransmit all the data, but only a portion of the data is transmitted, which can improve data transmission efficiency.
[0036] Figure 4 This is a schematic diagram of data transmission according to an embodiment of the present disclosure. In this embodiment, the iterator currently indicates a first queue pair. At block 402, a second allocation request for the queue pair is obtained from the first thread. That is, the manager receives another allocation request from the thread. At block 404, the first queue pair is locked for the first thread. Typically, the manager can lock any queue pair for the first thread; this embodiment relates to the processing when the locked queue pair is again the first queue pair.
[0037] In box 406, the iterator of the first queue group is modified to indicate the third queue pair, which is different from the first queue pair. The first thread submits the second data to the third RDMA device corresponding to the third queue pair. The third completion queue corresponds to the third queue pair. That is, when the thread requests channel allocation again, it can lock the same set of channels, but the specific channel allocated needs to be changed. In box 408, it checks whether the second data exists in the third completion queue. Because this thread is a synchronous thread, it needs to ensure that the second data is completely transmitted before the locking ends, so it is still necessary to check whether the third data exists in the completion queue. In box 410, in response to the existence of the second data in the third completion queue, the first queue pair is released. At this time, the thread's task has been completed, and the queue pair can be allocated to other threads, so the first queue pair can be released.
[0038] In this embodiment, each time a thread requests the allocation of a queue pair, if the allocation is to the same queue pair group as before, the queue pair ultimately used for data transmission will be automatically switched. This can distribute the workload among the queue pairs in the group, balance the load of each RDMA device, and improve the working efficiency of the RDMA device group.
[0039] Similarly, in some embodiments, allocation requests for queue pairs are obtained from another thread (i.e., a "second thread"). This other thread can also lock the first queue pair. However, the iterator for the first queue pair needs to be modified to indicate a fourth queue pair, which may be different from the first queue pair, but the same as or different from the second and third queue pairs. This embodiment also includes the other thread submitting third data to a fourth RDMA device corresponding to the fourth queue pair, with a fourth completion queue corresponding to the fourth queue pair. The presence of third data in the fourth completion queue is detected. This embodiment also includes releasing the first queue pair in response to the fourth RDMA device transmitting third data. In this embodiment, when different threads request the same queue pair, the workload of each queue pair in the group is distributed through iterator control, which can improve the efficiency of the RDMA device group.
[0040] This disclosure also provides embodiments for detecting whether an RDMA device has transmitted data. In some embodiments, the first channel is a first queue pair, the first queue pair includes a transmit queue, and the first completion channel is a first completion queue corresponding to the first queue pair. This embodiment first detects whether first data exists in the transmit queue. If first data exists in the transmit queue, it is determined that the RDMA device corresponding to the first queue pair has received the first data, that is, the thread has submitted the first data to the RDMA device, but the RDMA device has not yet transmitted it. If first data does not exist in the transmit queue, it is detected whether a completion queue entity for the first data exists in the first completion queue. If a completion queue entity for the first data exists in the first completion queue, it is determined that the first RDMA device has transmitted the first data. This embodiment can provide a specific scheme for detecting the data transmission status, which can improve the accuracy of data monitoring.
[0041] To ensure the validity of transmitted data, in some embodiments, the transmission status of the data being transmitted in the sending queue is first checked. If the data transmission in the sending queue fails, the data is retransmitted. Due to network failures or other reasons, data transmission may sometimes be interfered with; retransmission can ensure that the receiver receives complete and ordered data. If the data transmission is successful, a completion queue entity for the transmitted data is established in the first completion queue. In this embodiment, the establishment of each completion queue entity signifies the successful transmission of one piece of data, which can improve transmission efficiency.
[0042] Accordingly, this disclosure also provides specific schemes for detecting the data transmission status. In some embodiments, to detect the data transmission status, a completion queue can be accessed, and a completion queue entity for each piece of data corresponding to the first data can be retrieved from the completion queue. If a completion queue entity for each piece of data is retrieved, it is determined that the first data has been transmitted. Further, the completion queue entity for each piece of data is removed from the first completion queue. In this embodiment, each completion queue entity can be read once, eliminating the need for repeated readings, which has high processing efficiency. In some embodiments, if a completion queue entity for each piece of data is not retrieved, the first completion queue is retrieved repeatedly. This allows monitoring of whether the transmission of all data was successful.
[0043] Figure 5 This is a schematic block diagram of an example device 500 that can be used to implement embodiments of the present disclosure. As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 502 or loaded from storage unit 508 into random access memory (RAM) 503. The RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. Input / output (I / O) interface 505 is also connected to bus 504.
[0044] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0045] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of method 200 described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform method 200 by any other suitable means (e.g., by means of firmware).
[0046] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload programmable logic devices (CPLDs), and so on.
[0047] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0048] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing. Furthermore, although operations are depicted in a specific order, this should be understood as requiring that such operations be performed in the specific order shown or in sequential order, or requiring that all illustrated operations be performed to achieve the desired result. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.
[0049] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to computer-readable storage media within the respective computing / processing device.
[0050] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.
[0051] Various aspects of this disclosure have been described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0052] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0053] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operations or steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0054] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0055] Various embodiments of the present disclosure have been described above. The description is exemplary and exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or improvement of the technology in the market of the various embodiments, or to enable others skilled in the art to understand the various embodiments disclosed herein.
Claims
1. A data transmission method, comprising: Obtain the first allocation request for the channel from the first thread; For the first thread to lock the first group of channels, wherein multiple channels in the first group of channels correspond to multiple remote direct memory access (RDMA) devices, the first thread submits first data to the first RDMA device corresponding to the first channel in the first group of channels, and the first completion channel corresponds to the first channel; Detect whether the first data exists in the first completed channel; as well as In response to the presence of first data in the first completed channel, the first group of channels is released.
2. The method according to claim 1, wherein the first group of channels is a first group of queue pairs, the first group of queue pairs is provided with an iterator, the iterator indicating the first queue pair, the method further comprising: In response to a failure of the first queue pair, the iterator of the first group of queues is modified to indicate a second queue pair in the first group of queues, wherein the second queue pair is different from the first queue pair; The first thread submits the first data to the second RDMA device corresponding to the second queue pair.
3. The method according to claim 2, wherein the first completion channel is a first completion queue corresponding to the first queue pair, the first thread is provided with a first list, the first list indicates the transmission status of each piece of data in the first data, and after detecting whether the first data exists in the first completion channel, it further includes: In response to the presence of a first portion of the first data in the first completion queue, the first list is updated to indicate that the first portion of the data has been successfully transmitted, wherein the first portion of the data includes at least one piece of data. And the submission of the first data by the first thread to the second RDMA device corresponding to the second queue pair includes: According to the first list, the first thread submits data other than the first portion of the first data to the second RDMA device.
4. The method of claim 2, wherein the iterator indicates the first queue pair, the method further comprising: Obtain a second allocation request for the queue pair from the first thread; The first thread locks the first group of queues; The iterator of the first queue group is modified to indicate the third queue pair, wherein the third queue pair is different from the first queue pair, the first thread submits the second data to the third RDMA device corresponding to the third queue pair, and the third completion queue corresponds to the third queue pair; Detect whether the second data exists in the third completion queue; as well as In response to the presence of the second data in the third completion queue, the first queue pair is released.
5. The method according to claim 4, further comprising: Obtain a third allocation request for the queue pair from the second thread; The second thread locks the first group of queues; The iterator of the first queue group is modified to indicate the fourth queue pair, wherein the fourth queue pair is different from the first queue pair, the second thread submits third data to the fourth RDMA device corresponding to the fourth queue pair, and the fourth completion queue corresponds to the fourth queue pair; Detect whether the third data exists in the fourth completion queue; as well as In response to the fourth RDMA device having transmitted the third data, the first queue pair is released.
6. The method of claim 1, wherein in the first set of queue pairs, the RDMA devices corresponding to each queue pair are different from each other.
7. The method according to claim 1, wherein the first channel is a first queue pair, the first queue pair includes a sending queue, the first completion channel is a first completion queue corresponding to the first queue pair, and the method further includes: Detect whether the first data exists in the sending queue; In response to the presence of the first data in the transmission queue, it is determined that the RDMA device corresponding to the first queue pair has received the first data; In response to the absence of the first data in the sending queue, it is detected whether there is a completion queue entity for the first data in the first completion queue; In response to the existence of a completion queue entity for the first data in the first completion queue, it is determined that the first RDMA device has transmitted the first data.
8. The method according to claim 7, further comprising: Detect the transmission status of the data to be sent in the sending queue; In response to a transmission failure in the transmission queue, the transmission data is retransmitted; as well as In response to the successful transmission of data, a completion queue entity for the transmitted data is established in the first completion queue.
9. The method according to claim 8, further comprising: Access the first completed queue; Retrieve the completion queue entity for each piece of data in the first completion queue for the first data; In response to the retrieval of the completion queue entity for each piece of data, it is determined that the first data has been transmitted; and Remove the completion queue entity for each data item from the first completion queue.
10. The method of claim 9, further comprising: In response to the failure to retrieve the completion queue entity for each piece of data, the first completion queue is retrieved again.
11. An electronic device, comprising: At least one processor; as well as Coupled to the at least one processor and having instructions stored thereon, the instructions, when executed by the at least one processor, cause the electronic device to perform actions, the actions including: Obtain the first allocation request for the channel from the first thread; For the first thread to lock the first group of channels, wherein multiple channels in the first group of channels correspond to multiple remote direct memory access (RDMA) devices, the first thread submits first data to the first RDMA device corresponding to the first channel in the first group of channels, and the first completion channel corresponds to the first channel; Detect whether the first data exists in the first completed channel; and In response to the presence of first data in the first completed channel, the first group of channels is released.
12. The electronic device of claim 11, wherein the first group of channels is a first group of queues, the first group of queues is provided with an iterator, the iterator indicating the first group of queues, and the action further includes: In response to a failure of the first queue pair, the iterator of the first group of queues is modified to indicate a second queue pair in the first group of queues, wherein the second queue pair is different from the first queue pair; The first thread submits the first data to the second RDMA device corresponding to the second queue pair.
13. The electronic device of claim 12, wherein the first completion channel is a first completion queue corresponding to the first queue pair, the first thread is provided with a first list indicating the transmission status of each piece of data in the first data, and the action after detecting whether the first data exists in the first completion channel further includes: In response to the presence of a first portion of the first data in the first completion queue, the first list is updated to indicate that the first portion of the data has been successfully transmitted, wherein the first portion of the data includes at least one piece of data. And the submission of the first data by the first thread to the second RDMA device corresponding to the second queue pair includes: According to the first list, the first thread submits data other than the first portion of the first data to the second RDMA device.
14. The electronic device of claim 12, wherein the iterator indicates the first queue pair, and the action further includes: Obtain a second allocation request for the queue pair from the first thread; The first thread locks the first group of queues; The iterator of the first queue group is modified to indicate the third queue pair, wherein the third queue pair is different from the first queue pair, the first thread submits the second data to the third RDMA device corresponding to the third queue pair, and the third completion queue corresponds to the third queue pair; Detect whether the second data exists in the third completion queue; as well as In response to the presence of the second data in the third completion queue, the first queue pair is released.
15. The electronic device according to claim 14, further comprising: Obtain a third allocation request for the queue pair from the second thread; The second thread locks the first group of queues; The iterator of the first queue group is modified to indicate the fourth queue pair, wherein the fourth queue pair is different from the first queue pair, the second thread submits third data to the fourth RDMA device corresponding to the fourth queue pair, and the fourth completion queue corresponds to the fourth queue pair; Detect whether the third data exists in the fourth completion queue; as well as In response to the fourth RDMA device having transmitted the third data, the first queue pair is released.
16. The electronic device of claim 11, wherein in the first set of queue pairs, the RDMA devices corresponding to each queue pair are different from each other.
17. The electronic device of claim 11, wherein the first channel is a first queue pair, the first queue pair includes a transmission queue, the first completion channel is a first completion queue corresponding to the first queue pair, and the action further includes: Detect whether the first data exists in the sending queue; In response to the presence of the first data in the transmission queue, it is determined that the RDMA device corresponding to the first queue pair has received the first data; In response to the absence of the first data in the sending queue, it is detected whether there is a completion queue entity for the first data in the first completion queue; In response to the existence of a completion queue entity for the first data in the first completion queue, it is determined that the first RDMA device has transmitted the first data.
18. The electronic device according to claim 17, wherein the action further comprises: Detect the transmission status of the data to be sent in the sending queue; In response to a transmission failure in the transmission queue, the transmission data is retransmitted; as well as In response to the successful transmission of data, a completion queue entity for the transmitted data is established in the first completion queue.
19. The electronic device according to claim 18, wherein the action further comprises: Access the first completed queue; Retrieve the completion queue entity for each piece of data in the first completion queue for the first data; In response to the retrieval of the completion queue entity for each piece of data, it is determined that the first data has been transmitted; and Remove the completion queue entity for each data item from the first completion queue.
20. A computer program product tangibly stored on a non-volatile computer-readable medium and comprising machine-executable instructions that, when executed, cause a machine to perform actions, the actions comprising: Obtain the first allocation request for the channel from the first thread; For the first thread to lock the first group of channels, wherein multiple channels in the first group of channels correspond to multiple remote direct memory access (RDMA) devices, the first thread submits first data to the first RDMA device corresponding to the first channel in the first group of channels, and the first completion channel corresponds to the first channel; Detect whether the first data exists in the first completed channel; as well as In response to the presence of first data in the first completed channel, the first group of channels is released.