Method and system for constructing high-performance data exchange based on RoCE
By establishing a channel link and cache area between the sending end and the receiving end, combining the RoCE protocol and the doorbell mechanism, the problem of low data transmission efficiency in the heterogeneous acquisition and memory computing architecture is solved, high bandwidth and low latency data exchange is realized, and the system's computing and storage capabilities are improved.
Patent Information
- Application Number
- CN202510433590.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, the heterogeneous acquisition and memory computing architecture based on SRIO cannot effectively exert the flexible allocation and data exchange efficiency of heterogeneous computing resources such as FPGAs, GPUs/NPUs, and the RDMA network card cannot directly send the cached data of FPGAs, resulting in poor data transmission performance and efficiency.
RoCE is used to build a high-performance data exchange method. By establishing a preset number of channel links between the sending end and the receiving end, and configuring multiple data cache areas and command cache areas, the task instructions and doorbell mechanism are used to achieve efficient packaging and transmission of data packets, avoid task instructions interaction, and utilize the low-latency characteristics of the RoCE protocol.
It realizes high bandwidth and low latency data transmission between hardware units such as FPGA, CPU, GPU, etc., breaks through the performance bottleneck of single-channel data exchange, improves data transmission efficiency and system computing and storage capabilities for multiple synchronous acquisition.
Smart Images

Figure CN120342982A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data acquisition and computing, and particularly relates to a method and system for constructing high-performance data exchange based on RoCE. Background Art
[0002] With the continuous development of electronic information fields such as communication and radar, the demand trends for multi-channel parallel acquisition of high-speed signals, large bandwidth, and low-latency transmission are becoming increasingly obvious, posing higher requirements for aspects such as high-speed data real-time processing, parallel computing capabilities, and high-performance storage of signal acquisition, processing, and analysis systems.
[0003] The heterogeneous acquisition, storage, and computing architecture based on SRIO interconnection is a traditional technical solution. This solution uses SRIO as a high-speed interconnection interface, supports various topological interconnection structures such as point-to-point and ring-shaped, and has advantages such as high bandwidth and low latency. With the development of optical communication technology and the emergence of RoCEv2 technology, compared with RDMA interfaces on 40GE / 100G Ethernet and other networks, SRIO has the disadvantages of a lower bandwidth upper limit and greater difficulty in rate upgrading. At the same time, the transmission distance of SRIO is limited and requires specific software and hardware support, which restricts its generality between different devices, and its scalability and compatibility in software and hardware are inferior to those of networks.
[0004] In addition, in the above system, FPGA is generally used for preprocessing the acquired signals. However, the data bandwidth after preprocessing is still very large, and the data often needs to be split into multiple SRIO links for transmission. With the development of information technology, hardware such as FPGA, GPU, and NPU has gradually become high-performance computing power resources in the field of signal processing. The combined application and dynamic expansion of these heterogeneous computing power resources have also become key points in system design. In traditional technical solutions, due to the point-to-point and hardware dependence of SRIO, the flexible allocation of data and the dynamic combination of computing power resources cannot be exerted in a complex system composed of heterogeneous computing powers such as FPGA, GPU / NPU, thus affecting data exchange efficiency; at the same time, in the existing technology, it is impossible to simply send the cached data directly connected to the FPGA using an RDMA network card, and during the data transmission process, the CPU side needs to continuously issue task instructions to implement corresponding data transmission, resulting in poor performance and efficiency. Summary of the Invention
[0005] The purpose of the present invention is to propose a method and system for constructing high-performance data exchange based on RoCE to solve the problems raised in the background art.
[0006] To achieve the above purpose, the technical solutions adopted by the present invention are as follows:
[0007] A method for constructing high-performance data exchange based on RoCE, which is applied between a sending end and a receiving end. The method for constructing high-performance data transmission based on RoCE includes:
[0008] Establish a preset number of first channel links between the sending end and the receiving end, and configure multiple first data buffer areas located at the sending end, multiple second data buffer areas located at the receiving end, and multiple command buffer areas located at the sending end for each first channel link. Among them, the first data buffer areas, the second data buffer areas, and the command buffer areas in the same first channel link correspond to each other one by one;
[0009] For each first channel link:
[0010] Configure task instructions according to the corresponding first data buffer area and second data buffer area, and each task instruction is cached in the command buffer area one by one. Number each task instruction and each first data buffer area, and the corresponding first data buffer area and task instruction have the same number;
[0011] The sending end stores the sequentially received data into the first data buffer area in the order of the numbers;
[0012] Monitor the data in each first data buffer area in the order of the numbers according to the content of the task instruction, and judge whether the preset conditions are met;
[0013] When the preset conditions are met, trigger the doorbell, obtain the task instruction with the corresponding number, obtain the data from the corresponding first data buffer area according to the content of the obtained task instruction, perform a packing operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. After the current numbered task instruction and the corresponding data are packed, obtain the next numbered task instruction.
[0014] Preferably, the content of each task instruction includes the data length, the address of the first data buffer area, the operation code, the address of the second data buffer area, and the number of the task instruction. Among them, the data length is the preset condition, and the data lengths in the task instructions in the same first channel link are the same.
[0015] Preferably, each time the data in the first data buffer area is monitored, check whether the data in the first data buffer area meets the data length of the task instruction;
[0016] If it is met, trigger the doorbell, obtain the task instruction with the corresponding number, obtain the corresponding data according to the address of the first data buffer area in the obtained task instruction, perform a packing operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end, and the data packet is a RoCE protocol packet.
[0017] Preferably, for each first-channel link, after receiving a data packet, the receiving end parses the data packet to obtain data and a task instruction, and writes the parsed data into the second data buffer at the corresponding address for storage according to the operation code and the address of the second data buffer in the parsed task instruction, and then the receiving end returns an information packet indicating that the data packet has been received to the sending end.
[0018] Preferably, for each first-channel link of the sending end, the number of task instructions is n, the number of currently triggered doorbells is m, and the number of currently returned information packets is h. When the sending end receives a returned information packet each time, the following process is executed:
[0019] Calculate the difference k between the number m of currently triggered doorbells and the number h of currently returned information packets, and compare the difference k with the number n of task instructions:
[0020] When k < n, the data receiving port of the corresponding first-channel link of the sending end is in an open state;
[0021] When k ≥ n, the data receiving port of the corresponding first-channel link of the sending end is in a closed state.
[0022] Preferably, the method for constructing a high-performance data switch based on RoCE further includes setting priorities for all first-channel links. When the data in the first data buffers corresponding to at least two first-channel links all meet the data length, doorbells are triggered in sequence according to the priority order.
[0023] A system for constructing a high-performance data switch based on RoCE includes an acquisition module and a first computing module, and the acquisition module and the first computing module are interconnected through a switching network. Among them, the acquisition module and the first computing module perform data exchange using the method for constructing a high-performance data switch based on RoCE;
[0024] When the acquisition module serves as the sending end, the acquisition module includes a driving unit, a data receiving unit located in the first FPGA, a doorbell triggering unit, and a first packet processing unit, where:
[0025] The driving unit is used to establish a preset number of first-channel links between the acquisition module and the first computing module according to service requirements, and configure multiple first data buffers located in the acquisition module, multiple command buffers located in the acquisition module, and multiple second data buffers located in the first computing module for each first-channel link, where the first data buffers, second data buffers, and command buffers in the same first-channel link correspond to each other one by one;
[0026] The driving unit is further configured to, for each first channel link, configure task instructions according to the corresponding first data buffer and second data buffer, cache each task instruction in the command buffer one by one, and send one of the task instructions to the doorbell trigger unit. The task instructions and each first data buffer are numbered respectively, and the corresponding first data buffer and the task instruction number are the same;
[0027] The data receiving unit is configured to, for each first channel link, sequentially receive the data to be sent and store the data in the first data buffer in the order of the numbers;
[0028] The doorbell trigger unit is configured to, for each first channel link, monitor the data in each first data buffer according to the content of the task instruction and in the order of the numbers, determine whether the preset condition is met, and when the preset condition is met, trigger the doorbell to notify the first packet processing unit;
[0029] The first packet processing unit is configured to, for each first channel link, after receiving the doorbell notification, obtain the task instruction with the corresponding number, obtain the data from the corresponding first data buffer according to the content of the obtained task instruction, perform a packet operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. After the task instruction of the current number and the corresponding data are packet-operated, then obtain the next numbered task instruction.
[0030] Preferably, when the first computing module serves as the receiving end, the first computing module includes a second packet processing unit located in the second FPGA, where:
[0031] The second packet processing unit is configured to obtain the data packet corresponding to the first channel link from the switching network, parse the data packet to obtain the data and the task instruction, and write the parsed data into the second data buffer at the corresponding address for storage according to the operation code in the parsed task instruction and the address of the second data buffer, and return an information packet of the received data packet to the sending end.
[0032] Preferably, the first computing module further includes a first computing processing unit. When the first computing module serves as the receiving end, for each first channel link, the first computing processing unit monitors each second data buffer, and when it monitors that there is data stored, it obtains the data for calculation.
[0033] Preferably, the system for constructing a high-performance data exchange based on RoCE further includes a second computing module and a storage module connected to the switching network. The second computing module includes a first configuration unit, a first network card, and a second computing processing unit, where:
[0034] The first configuration unit is used to establish a preset number of second channel links between the second computing module and the storage module according to service requirements;
[0035] The first network card is used to receive data packets of the corresponding channel link from the switching network for parsing, and send the parsed data to the second computing processing unit for calculation;
[0036] The first network card is also used to send the data calculated by the second computing processing unit to the switching network through the second channel link;
[0037] The storage module includes a second network card and a data disk, where:
[0038] The second network card is used to receive data packets of the corresponding channel link from the switching network for parsing, and send the parsed data to the data disk for storage;
[0039] The first computing module further includes a second configuration unit, which is used to establish a preset number of third channel links between the first computing module and the second computing module and / or the storage module according to service requirements. After the first computing processing unit calculates, the first computing processing unit sends the calculated data to the switching network through the third channel link.
[0040] Compared with the prior art, the beneficial effects of the present invention are:
[0041] The method and system for building high-performance data switching based on RoCE pre-configure task instructions, and realize notification by triggering doorbells during the data transmission process, so that task instruction interaction is no longer required during the data transmission process, thereby improving performance and realizing efficient data transmission. In addition, data is transmitted through multiple channel links, and the low latency of the RoCE protocol is utilized to further improve the data transmission efficiency;
[0042] The system for building high-performance data switching based on RoCE of the present invention realizes high-bandwidth and low-latency data transmission between hardware units such as FPGA, CPU, and GPU among different modules in a distributed computing processing system for high-speed signals, breaking through the performance bottleneck of single-channel high-speed data transmission and switching. The highest single-channel data exchange can reach 10GB / s, and it is relatively easy to improve the system's ability to perform real-time calculation and storage on high-speed large-bandwidth data collected synchronously in multiple channels. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a schematic flowchart of the method for building high-performance data switching based on RoCE of the present invention;
[0044] Figure 2 It is a module diagram of the system for building high-performance data switching based on RoCE of the present invention;
[0045] Figure 3 This is the module diagram between the sender and receiver of the present invention. Specific Embodiments
[0046] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0047] In one embodiment, as Figure 1 and Figure 3 shown, a method for constructing high-performance data switching based on RoCE is provided, which is applied between the FPGA of the sender and the receiver. The specific method includes:
[0048] Step 1: Establish a preset number of first channel links between the sender and the receiver, and configure multiple first data buffers located at the sender, multiple second data buffers located at the receiver, and multiple command buffers located at the sender for each first channel link, where the first data buffers, second data buffers, and command buffers in the same first channel link correspond to each other one by one;
[0049] It should be noted that the number of first channel links can be set according to actual needs, but the maximum does not exceed 2047, and the numbers of the first data buffers, second data buffers, and command buffers are also set according to actual needs, and the numbers of the three are the same and correspond to each other one by one.
[0050] For each first channel link (the data exchange between the first channel links does not affect each other, where steps 4-step 6 are the data exchange process implemented by one of the first channel links, that is, the data exchange of each first channel link is like steps 4-step 6):
[0051] Step 2: Configure task instructions according to the corresponding first data buffer and second data buffer, and each task instruction is cached in the command buffer one by one. Each task instruction and each first data buffer are numbered, and the corresponding first data buffer and task instruction have the same number (for example, if the numbers are 1, 2...a in sequence, then the numbers of the first data buffer and the task instruction are both a);
[0052] Among them, the content of each task instruction includes the data length (the data length in the task instruction is set according to the user's needs, and in the same first channel link, the data lengths of each task instruction are kept the same, and at the same time the data length does not exceed the total length of the data to be processed by each first channel link), the address of the first data buffer, the operation code, the address of the second data buffer, and the number of the task instruction.
[0053] Step 3: The sender stores the successively received data (where which first-channel link is specifically used to send the data is specified by the user, and the user directly inputs the data to be sent as the port of the specified first-channel link sender) into the first data buffer in the order of the numbers.
[0054] It should be noted that each time the sender receives a data, it is stored in the corresponding first data buffer in the order of the numbers. For the first received data, it is stored in the first data buffer numbered 1. For the second received data, it is stored in the first data buffer numbered 2.
[0055] Step 4: Monitor the data in each first data buffer in the order of the numbers according to the content of the task instruction, and determine whether the preset condition is met (the preset condition is the data length in the task instruction content).
[0056] It should be noted that each time the data in the first data buffer is monitored, check whether the data in the first data buffer meets the data length of the task instruction. In this embodiment, as long as data is received, the data is successively stored in each first data buffer in the order of the numbers. As long as new data is stored, it is successively polled and monitored in the order of the numbers to check whether the data length is met. Each first data buffer, second data buffer, command buffer, and task instruction can be time-division multiplexed, improving the resource reuse rate.
[0057] Step 5: When the preset condition is met, trigger the doorbell (for the same first-channel link, regardless of whether the previous doorbell notification has been processed, as long as it is monitored that the data length is met each time, a doorbell notification is triggered. For each doorbell notification, the task instruction is obtained in the order of the numbers, and wait until the current task instruction and the corresponding data are packaged together before obtaining the task instruction of the next number. And the doorbell is triggered by the FPGA of the sender), and obtain the task instruction of the corresponding number. According to the content of the obtained task instruction, obtain the data from the corresponding first data buffer, and perform a packaging operation on the obtained data and the task instruction to obtain a data packet (and the data packet is a packet of the RoCE protocol, specifically the data packet of the RoCEv2 protocol in this embodiment), and send the data packet to the receiver. After the current numbered task instruction and the corresponding data are packaged, obtain the task instruction of the next number (in the order of the numbers); that is, obtain the task instruction of the corresponding number, obtain the corresponding data according to the address of the first data buffer in the obtained task instruction, and perform a packaging operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiver.
[0058] Step 6: After the receiving end receives the data packet, it parses the data packet to obtain the data and the task instruction, and writes the parsed data into the second data buffer at the corresponding address for storage according to the operation code (the operation code is a write operation, a read operation, etc.) in the parsed task instruction and the address of the second data buffer. Then, the receiving end returns an information packet indicating that the data packet has been received to the sending end (where the information packet is to inform the sending end that the receiving end has received the sent data, and the information packet is a packet of the RoCE protocol). According to the service requirements, when subsequent applications need to be performed on the data in the second data buffer (such as calculating the data), the data is taken out from each second data buffer for application, and each second data buffer is in an empty area state. When all the a first data buffers of the sending end have been stored in sequence for one round, for the data received subsequently, it continues to be stored in sequence from the first data buffer numbered 1 to the first data buffer numbered a, and then steps 4 to 6 are repeated.
[0059] It should be noted that for each first channel link of the sending end, the number of task instructions is n, the number of currently triggered doorbells is m, and the number of currently returned information packets is h. When the sending end receives each returned information packet, the following process is executed:
[0060] Calculate the difference k between the number m of currently triggered doorbells and the number h of currently returned information packets, and compare the difference k with the number n of task instructions:
[0061] When k < n, the data receiving port of the corresponding first channel link of the sending end is in an open state (that is, the input port of the corresponding first channel link can continue to receive new data);
[0062] When k ≥ n, the data receiving port of the corresponding first channel link of the sending end is in a closed state (that is, the input port of the corresponding first channel link cannot continue to receive new data).
[0063] It should be noted that by controlling the data receiving port of the sending end, it is possible to prevent the problem of confusion and damage caused by continuing to receive data when the data has not been successfully sent.
[0064] In this embodiment, the method for constructing high-performance data exchange based on RoCE further includes setting priorities for all first channel links. When the data in the first data buffers corresponding to at least two first channel links all meet the data length, the doorbells are triggered in sequence according to the priority order (that is, the doorbell triggered first performs the packaging operation first). When the buffer capacity of the path with a lower priority is not enough to cache, a discard policy will be adopted.
[0065] In another embodiment, such as Figure 2As shown in the figure, based on a method for constructing high-performance data exchange based on RoCE, a system for constructing high-performance data exchange based on RoCE (specifically, a storage-computation system) is provided, including:
[0066] The system includes an acquisition module and a first computing module (which can be a first computing module mainly based on FPGA for computing power), and the acquisition module and the first computing module are interconnected through a switching network (a lossless Ethernet network). Among them, the acquisition module and the first computing module use the method of constructing high-performance data exchange based on RoCE for data exchange;
[0067] When the acquisition module is used as the sending end, the acquisition module includes a driving unit, a data receiving unit located in the first FPGA, a doorbell triggering unit, and a first packet processing unit, where:
[0068] The driving unit (located in the CPU) is used to establish a preset number of first channel links between the acquisition module and the first computing module according to service requirements (based on service requirements, using MAC, IP, and QPN numbers as the unique port attributes to construct a point-to-point data exchange channel end between different services), and configure multiple first data buffer areas located in the acquisition module, multiple command buffer areas located in the acquisition module, and multiple second data buffer areas located in the first computing module for each first channel link. Among them, the first data buffer areas, second data buffer areas, and command buffer areas in the same first channel link correspond to each other one by one;
[0069] The driving unit is also used to configure task instructions for each first channel link according to the corresponding first data buffer area and second data buffer area, cache each task instruction in the command buffer area one by one, and send one of the task instructions to the doorbell triggering unit (since the data lengths in each task instruction in each first channel link are the same, only one task instruction needs to be sent to the doorbell triggering unit to facilitate the subsequent judgment of the data lengths in each first data buffer area by the doorbell triggering unit), number each task instruction and each first data buffer area, and the corresponding first data buffer area and task instruction have the same number;
[0070] The data receiving unit is used to sequentially receive the data to be sent for each first channel link and store them in the first data buffer area in the order of the numbers;
[0071] The above process is the process of the data receiving unit processing each first channel link respectively, and the processes of the data receiving unit processing each first channel link are independent of each other and do not interfere with each other. That is, it is equivalent to that the data receiving unit includes sub-data receiving units corresponding to each first channel link one by one. Each sub-data receiving unit respectively receives data for its own first channel link, and the receiving process is the same as the above process and will not be described repeatedly.
[0072] The doorbell trigger unit, for each first channel link, is configured to monitor the data in each first data buffer in sequence according to the content of the task instruction and in the order of the numbers, determine whether the preset condition is met, and when the preset condition is met, trigger a doorbell to notify the first packet processing unit;
[0073] The above process is the process of the doorbell trigger unit processing each first channel link respectively, and the processes of the doorbell trigger unit processing each first channel link are independent of each other and do not interfere with each other. That is, it is equivalent that the doorbell trigger unit includes sub-doorbell trigger units corresponding one by one to each first channel link, and each sub-doorbell trigger unit processes its respective first channel link: each sub-doorbell trigger unit is configured to monitor the data in each first data buffer in the corresponding first channel link in sequence according to the content of the task instruction in the corresponding first channel link and in the order of the numbers, determine whether the preset condition is met, and when the preset condition is met, trigger a doorbell to notify the corresponding sub-processing unit;
[0074] The first packet processing unit, for each first channel link, is configured to, after receiving the doorbell notification, obtain the task instruction with the corresponding number, obtain data from the corresponding first data buffer according to the content of the obtained task instruction, perform a packing operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. After the task instruction with the current number and the corresponding data are packed, obtain the next numbered task instruction;
[0075] The above process is the process of the first packet processing unit processing each first channel link respectively, and the processes of the first packet processing unit processing each first channel link are independent of each other and do not interfere with each other. That is, it is equivalent that the first packet processing unit includes first sub-processing units corresponding one by one to each first channel link, and each sub-processing unit processes its respective first channel link, and the processing process is the same as the above process and will not be described repeatedly.
[0076] In this embodiment, when the first computing module serves as the receiving end, the first computing module includes a second packet processing unit located in the second FPGA, where (the processes of the first computing module processing each first channel link are independent of each other and do not interfere with each other):
[0077] The second packet processing unit is configured to obtain the data packet corresponding to the first channel link from the switching network, parse the data packet to obtain data and a task instruction, and write the parsed data into the second data buffer at the corresponding address for storage according to the operation code in the parsed task instruction and the address of the second data buffer, and return an information packet indicating that the data packet has been received to the sending end;
[0078] The above process is the process in which the second packet processing unit processes each first channel link respectively, and the process in which the second packet processing unit processes each first channel link is independent of each other and does not interfere with each other, that is, it is equivalent to the second packet processing unit including a second sub-processing unit corresponding to each first channel link one by one, each second sub-processing unit processes its own first channel link respectively, and the processing process is consistent with the above process, and will not be described again.
[0079] In this embodiment, the first calculation module further includes a first calculation processing unit. When the first calculation module acts as a receiving end, for each first channel link, the first calculation processing unit monitors each second data buffer area, and when monitoring that there is data storage, the first calculation processing unit obtains the data for calculation;
[0080] The above process is the process in which the first computing processing unit processes each first channel link respectively, and the processes in which the first computing processing unit processes each first channel link are independent of each other and do not interfere with each other, that is, it is equivalent to the first computing processing unit including first sub-computing processing units corresponding to each first channel link one by one, and each first sub-computing processing unit processes its own first channel link respectively, and the processing process is consistent with the above process, and will not be described again.
[0081] In this embodiment, the first computing module also includes a second configuration unit, which is used to establish a preset number of third channel links between the first computing module and the second computing module and / or storage module according to business needs. After the first computing processing unit calculates, the first computing processing unit sends the calculated data to the switching network through the third channel link.
[0082] In this embodiment, the system for building high-performance data exchange based on RoCE also includes a second computing module (which may be a second computing module based on computing power such as GPU) and a storage module connected to the switching network. The second computing module includes a first configuration unit, a first network card (wherein the first network card is an RDMA network card, which may be a 100GE network card of the Kunpeng processor or a dedicated chip network card such as MLNX) and a second computing processing unit, wherein:
[0083] A first configuration unit is used to establish a preset number of second channel links between the second computing module and the storage module according to business requirements (i.e., according to customer requirements, when data processed by the second computing module needs to be sent to the storage module for storage, a corresponding channel is established);
[0084] The first network card is used to receive a data packet of a corresponding channel link (i.e., a third channel link) from the switching network for parsing, and send the parsed data to the second computing processing unit for calculation;
[0085] The first network card is also used to send the data calculated by the second computing and processing unit to the switching network through the second channel link;
[0086] The storage module includes a second network card (where the second network card is an RDMA network card) and a data disk, where:
[0087] The second network card is used to receive the data packets of the corresponding channel link (i.e., the second channel link or the third channel link) from the switching network for parsing, and send the parsed data to the data disk for storage.
[0088] Among them, the data processes of the second computing module and the storage module for each channel link are independent of each other and do not interfere with each other.
[0089] The specific limitations on a method for constructing high-performance data switching based on RoCE also apply to the limitations on a system for constructing high-performance data switching based on RoCE, and will not be repeated here.
[0090] The method and system for constructing high-performance data switching based on RoCE pre-configure task instructions, and use doorbells to trigger notifications during data transmission, so that task instruction interactions are no longer required during data transmission, thereby improving performance and achieving efficient data transmission. In addition, data is transmitted through multiple channel links, and the low latency of the RoCE protocol is utilized to further improve the efficiency of data transmission;
[0091] The system for constructing high-performance data switching based on RoCE realizes high-bandwidth and low-latency data transmission between hardware units such as FPGA, CPU, and GPU among different modules in a distributed computing and processing system for high-speed signals, breaking through the performance bottleneck of single-channel high-speed data transmission and switching. The highest single-channel data exchange can reach 10 GB / s, and it is relatively easy to improve the system's ability to perform real-time calculation and storage on high-speed large-bandwidth data collected synchronously in multiple channels.
[0092] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0093] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for constructing high-performance data switching based on RoCE, characterized in that: The method for constructing high-performance data transmission based on RoCE is applied between a sending end and a receiving end. The method for constructing high-performance data transmission based on RoCE includes: Establish a preset number of first channel links between the sending end and the receiving end, and configure multiple first data buffers located at the sending end, multiple second data buffers located at the receiving end, and multiple command buffers located at the sending end for each first channel link. Among them, the first data buffers, second data buffers, and command buffers in the same first channel link correspond to each other one by one; For each first channel link: Configure task instructions according to the corresponding first data buffer and second data buffer, and each task instruction is cached in the command buffer one by one. Number each task instruction and each first data buffer, and the numbers of the corresponding first data buffer and task instruction are the same; The sending end stores the sequentially received data into the first data buffer in the order of the numbers; Monitor the data in each first data buffer in the order of the numbers according to the content of the task instruction, and judge whether the preset conditions are met; When the preset conditions are met, trigger the doorbell, obtain the task instruction with the corresponding number, obtain the data from the corresponding first data buffer according to the content of the obtained task instruction, perform a packaging operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. After the current numbered task instruction and the corresponding data are packaged, obtain the next numbered task instruction.
2. The method for constructing high-performance data switching based on RoCE according to claim 1, wherein: The content of each task instruction includes the data length, the address of the first data buffer, the operation code, the address of the second data buffer, and the number of the task instruction. Among them, the data length is the preset condition, and the data lengths in each task instruction in the same first channel link are the same.
3. The method for constructing high-performance data switching based on RoCE according to claim 2, wherein: Each time the data in the first data buffer is monitored, check whether the data in the first data buffer meets the data length of the task instruction; If it is satisfied, trigger the doorbell, obtain the task instruction with the corresponding number, obtain the corresponding data according to the address of the first data buffer in the obtained task instruction, perform a packaging operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. And the data packet is a RoCE protocol packet.
4. The method for constructing high-performance data switching based on RoCE according to claim 2, characterized in that: For each first channel link, after receiving the data packet, the receiving end parses the data packet to obtain the data and the task instruction, and writes the parsed data into the second data buffer at the corresponding address for storage according to the operation code and the address of the second data buffer in the parsed task instruction. Then the receiving end returns an information packet of the received data packet to the sending end.
5. The method for constructing high-performance data switching based on RoCE according to claim 1, wherein: For each first channel link of the sending end, the number of task instructions is n, the number of currently triggered doorbells is m, and the number of currently returned information packets is h. When the sending end receives the returned information packet each time, the following process is executed: Calculate the difference k between the number m of currently triggered doorbells and the number h of currently returned information packets, and compare the difference k with the number n of task instructions: When k < n, the port for the sending end to receive data corresponding to the first channel link is in an open state; When k≥n, the port of the sending end corresponding to the first channel link for receiving data is in a closed state.
6. The method for constructing high-performance data switching based on RoCE according to claim 2, wherein: The method for constructing high-performance data switching based on RoCE further includes setting priorities for all first channel links. When the data in the first data buffer areas corresponding to at least two first channel links all meet the data length, doorbells are triggered in sequence according to the priority order.
7. A system for building high-performance data switching based on RoCE, characterized in that: The system includes an acquisition module and a first computing module, and the acquisition module and the first computing module are interconnected through a switching network. Among them, the acquisition module and the first computing module use the method for constructing high-performance data switching based on RoCE as described in any one of claims 1 to 6 for data exchange; When the acquisition module is used as the sending end, the acquisition module includes a driving unit, a data receiving unit located in the first FPGA, a doorbell triggering unit, and a first packet processing unit, where: The driving unit is used to establish a preset number of first channel links between the acquisition module and the first computing module according to service requirements, and configure multiple first data buffer areas located in the acquisition module, multiple command buffer areas located in the acquisition module, and multiple second data buffer areas located in the first computing module for each first channel link. Among them, the first data buffer areas, second data buffer areas, and command buffer areas in the same first channel link correspond to each other one by one; The driving unit is further used to configure task instructions for each first channel link according to the corresponding first data buffer area and second data buffer area, cache each task instruction in the command buffer area one by one, and send one of the task instructions to the doorbell triggering unit. Number each task instruction and each first data buffer area, and the corresponding first data buffer area and task instruction have the same number; For each first channel link, the data receiving unit is used to sequentially receive the data to be sent and store it in the first data buffer area in the order of the numbers; For each first channel link, the doorbell triggering unit is used to monitor the data in each first data buffer area according to the content of the task instruction and in the order of the numbers, determine whether the preset conditions are met, and when the preset conditions are met, trigger the doorbell to notify the first packet processing unit; For each first channel link, the first packet processing unit is used to, after receiving the doorbell notification, obtain the task instruction with the corresponding number, obtain the data from the corresponding first data buffer area according to the content of the obtained task instruction, perform a packing operation on the obtained data and the task instruction to obtain a data packet, and send the data packet to the receiving end. After the current numbered task instruction and the corresponding data are packed, obtain the next numbered task instruction.
8. The system for constructing high-performance data switching based on RoCE according to claim 7, wherein: When the first computing module is used as the receiving end, the first computing module includes a second packet processing unit located in the second FPGA, where: The second packet processing unit is configured to obtain packets corresponding to the first channel link from the switching network, parse the packets to obtain data and task instructions, write the parsed data into the second data buffer at the corresponding address for storage according to the operation code and the address of the second data buffer in the parsed task instructions, and return an information packet indicating that the packets have been received to the sending end.
9. The system for constructing high-performance data switching based on RoCE according to claim 7, wherein: The first computing module further includes a first computing processing unit. When the first computing module serves as the receiving end, for each first channel link, the first computing processing unit monitors each second data buffer. When it monitors that data is stored, it obtains the data for computing.
10. The system for constructing high-performance data switching based on RoCE according to claim 7, wherein: The system for constructing a high-performance data switch based on RoCE further includes a second computing module and a storage module connected to the switching network. The second computing module includes a first configuration unit, a first network card, and a second computing processing unit, where: The first configuration unit is configured to establish a preset number of second channel links between the second computing module and the storage module according to service requirements. The first network card is configured to receive packets corresponding to the channel link from the switching network for parsing, and send the parsed data to the second computing processing unit for computing. The first network card is further configured to send the data computed by the second computing processing unit to the switching network through the second channel link. The storage module includes a second network card and a data disk, where: The second network card is configured to receive packets corresponding to the channel link from the switching network for parsing, and send the parsed data to the data disk for storage. The first computing module further includes a second configuration unit. The second configuration unit is configured to establish a preset number of third channel links between the first computing module and the second computing module and / or the storage module according to service requirements. After the first computing processing unit finishes computing, it sends the computed data to the switching network through the third channel link.