Network card multi-node interconnection communication method and device, computer equipment and storage medium
By initializing two sets of data structure instances in the network card, the communication type is automatically identified and the path is matched, which solves the problem that existing network cards cannot be compatible with local multi-process and remote RDMA communication, reduces hardware costs and deployment difficulty, and is suitable for distributed storage and high-performance computing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- RAMAXEL TECH SHENZHEN
- Filing Date
- 2026-01-22
- Publication Date
- 2026-04-21
AI Technical Summary
Existing network cards cannot simultaneously support local multi-process RDMA communication within the same computer and remote RDMA communication between different computers, resulting in high hardware costs, difficulty in system deployment and upgrades, and poor compatibility.
By initializing two sets of data structure instances, the local and remote communication types are automatically identified and the corresponding communication paths are matched. The compatibility of local and remote RDMA communication is achieved using SSS network cards and Soft-RoCE technology, avoiding the need to replace hardware.
It achieves compatibility between local multi-process RDMA communication and remote RDMA communication under the same network card port, reducing the cost and difficulty of system deployment and upgrade, and adapting to the needs of scenarios such as distributed storage and high-performance computing.
Smart Images

Figure CN121907754A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of RDMA interconnection communication technology, and in particular to a method, apparatus, computer equipment, and storage medium for multi-node interconnection communication of network cards. Background Technology
[0002] In inter-process data exchange scenarios within computer systems, Remote Direct Memory Access (RDMA) technology, by bypassing the CPU to directly transfer data between memory, significantly reduces communication latency and improves data transfer efficiency, making it one of the core communication technologies in fields such as distributed storage and high-performance computing. Inter-process RDMA communication is mainly divided into two scenarios: local communication, i.e., communication between different processes within the same computer, and remote communication, i.e., communication between processes on different computers. Both local and remote RDMA communication rely on a Network Interface Card (NIC) that supports RDMA functionality, and the computer system needs to deploy RDMA-related library functions and drivers to implement core functions such as RDMA communication protocol parsing and data transfer control.
[0003] However, existing network interface cards (NICs) typically only support remote RDMA communication, meaning they can only handle inter-process RDMA data transfer between different computers and cannot support local multi-process RDMA communication on a single port of the NIC within the same computer. To address this issue, the common solution is to replace the NIC with a dedicated one that supports single-port multi-process RDMA communication. However, dedicated NICs are significantly more expensive than standard NICs, drastically increasing system deployment or upgrade costs. Furthermore, for existing systems with numerous NICs that do not support this feature, replacing the NIC requires extensive hardware modifications, which are not only time-consuming and difficult but also waste existing hardware resources and result in poor compatibility.
[0004] Therefore, there is an urgent need for a communication solution that can be compatible with local multi-process RDMA communication on a single port of the same computer's network card and remote RDMA communication between different computers without replacing the existing hardware network card, and can reduce the cost and difficulty of system deployment and upgrade. Summary of the Invention
[0005] This invention provides a method, apparatus, computer equipment, and storage medium for multi-node interconnection communication of network interface cards (NICs) to solve the technical problems of high system deployment and upgrade costs and difficulties in existing multi-node interconnection communication of NICs.
[0006] Firstly, a method for interconnecting and communicating multiple network interface cards (NICs) is provided, including: Perform communication library initialization operations, and configure a first data structure instance for remote communication and a second data structure instance for local communication; When a communication request is received, check whether the current communication type is remote communication or local communication; Based on the first data structure instance and the second data structure instance, a preset first communication path is selected to process remote data, or a second communication path is selected to process local data.
[0007] Secondly, a network card multi-node interconnection communication device is provided, including: The configuration module is used to perform communication library initialization operations, configuring a first data structure instance for remote communication and a second data structure instance for local communication. The inspection module is used to check whether the current communication type is remote communication or local communication when a communication request is received. The allocation module is used to select a preset first communication path to process remote data or select a second communication path to process local data based on the first data structure instance and the second data structure instance.
[0008] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described network interface card (NIC) multi-node interconnection communication method.
[0009] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described network interface card (NIC) multi-node interconnection communication method.
[0010] The beneficial effects of this invention compared with the prior art are as follows: This invention automatically identifies local and remote communication types and matches corresponding communication paths by initializing two sets of data structure instances, perfectly compatible with two types of RDMA communication scenarios, without the need to replace network card hardware, thus avoiding the high cost of dedicated network cards and the waste of existing hardware resources, saving large-scale system transformation procedures, reducing deployment and upgrade difficulty, and adapting to the core scenario requirements such as distributed storage and high-performance computing.
[0011] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of the present invention more obvious and understandable, preferred embodiments are described in detail below. Attached Figure Description
[0012] Figure 1 This is a flowchart illustrating a method for multi-node interconnection communication of network cards in one embodiment of the present invention; Figure 2 yes Figure 1 A schematic diagram of a specific implementation method for step S10; Figure 3 yes Figure 2 A schematic diagram of a specific implementation method for step S12; Figure 4 yes Figure 1 A schematic diagram of a specific implementation method for step S20; Figure 5 yes Figure 4 A schematic diagram of a specific implementation method for step S30; Figure 6 This is a schematic diagram of the system framework structure of a network card multi-node interconnection communication method in one embodiment of the present invention; Figure 7 This is a schematic diagram of the interface function call flow of a network card multi-node interconnection communication method in one embodiment of the present invention; Figure 8 yes Figure 7 A flowchart illustrating the specific implementation of the application process initialization interface function call; Figure 9 yes Figure 7 A flowchart illustrating the specific implementation of runtime interface function calls in application processes. Figure 10 This is a schematic diagram of a network card multi-node interconnection communication device according to an embodiment of the present invention; Figure 11 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 12 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. The technical solutions in the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0014] It should be understood that, when used in this specification and the appended claims, the terms “comprising” and “including” indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0015] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0016] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0017] Please see Figure 1 As shown, Figure 1 This is a schematic flowchart illustrating a network interface card (NIC) multi-node interconnection communication method provided in an embodiment of the present invention. The NIC multi-node interconnection communication method includes the following steps: S10: Perform communication library initialization operations, and configure a first data structure instance for remote communication and a second data structure instance for local communication.
[0018] For step S10, by pre-configuring two independent data structure instances, which are respectively bound to remote communication and local communication scenarios, the delay caused by temporary creation of instances during subsequent communication is avoided. At the same time, it provides directly available resource support for path switching after communication type identification, ensuring the compatible operation of the two types of communication scenarios. From the underlying architecture, it solves the defect that the existing network card cannot support local and remote RDMA communication at the same time, and provides structural guarantee for the smooth switching of local single-port multi-process communication and inter-computer communication.
[0019] The communication library is an intermediate layer library located between the RDMA common initialization library functions and the network card driver. Its core function is to achieve intelligent path selection and data distribution based on communication type by simultaneously initializing and managing two sets of independent data structure instances. In this embodiment, the communication library is sssroce.so, a dynamic link library designed for compatibility with single-port multi-process RDMA communication, used to encapsulate the underlying communication logic. The first data structure instance, namely the sss_* instance, is bound to the remote communication scenario to support RDMA data transmission between different computers; the second data structure instance, namely the rxe_* instance, is bound to the local communication scenario, implementing RDMA communication within the same computer based on the open-source Soft-RoCE (rxe). Together, these three constitute the core data foundation for compatibility with both types of communication, ensuring dedicated resource support for different scenarios.
[0020] Understandably, the first data structure instance, namely the sss_* instance, is the native driver and supporting components of the SSS network card. It is the RDMA function carrier that depends on the SSS network card hardware and is responsible for realizing hardware-level remote RDMA communication between different computers. The second data structure instance, namely the rxe_* instance, uses rxe, short for Soft-RoCE, a software-based RDMA technology. Rxe does not directly depend on the SSS network card hardware, but in this solution, rxe leverages the single-port resources of the SSS network card to achieve RDMA communication between local processes under the same SSS network card. Both SSS and rxe are implementations of RDMA communication. In this embodiment, the two are integrated through the sssroce.so library. That is, when the communication type is remote, the hardware RDMA capability of SSS is invoked; when the communication type is local, the software RDMA capability of rxe is invoked, thus achieving a single network card capable of supporting both types of communication scenarios.
[0021] It is also understood that, in this embodiment, the network interface card (NIC) is an SSS NIC, and communication between computers is directly conducted through the drivers of each computer's SSS NIC. See [link to relevant documentation]. Figure 6 As shown, process a1 of computer A transmits data to computer A's SSS network card, computer A's SSS network card sends the data to computer B's SSS network card, and computer B's SSS network card transmits the data to the memory of process b3 of computer B. The receiving process of computer A is the reverse. It can also be understood that inter-process communication within the same computer using RDMA relies on the Rxe driver. For example, if process b1 of computer B wants to transmit data to process b2, process b1 first submits the data to the SSS network card, transferring the data from process b1's memory to the virtual communication link, so that the data is forwarded to process b2's memory; the receiving process of process b2 transmitting data to process b1 is the reverse.
[0022] In some embodiments of the present invention, such as Figure 2 As shown, a specific configuration scheme is provided. In S10, the communication library initialization operation is performed, configuring a first data structure instance for remote communication and a second data structure instance for local communication, specifically including the following steps S11-S12.
[0023] S11: Execute the application process initialization procedure and call the RDMA common initialization library function.
[0024] For step S11, the application process actively calls the standardized RDMA common initialization library function to ensure that the initialization process conforms to the general specifications of RDMA communication. This also provides a unified calling interface for subsequent integration with the sssroce.so communication library, avoiding initialization failures due to application layer differences and ensuring the universality and compatibility of the process. In the application scenario of the sss network card, the application process calls the RDMA common initialization library function to trigger the subsequent sssroce.so instance initialization process. This provides the necessary prerequisites for generating two sets of instances—the first data structure instance and the second data structure instance—ensuring the normal startup of the subsequent communication process.
[0025] The initialization process of the application process refers to the preparatory process for establishing RDMA communication after the application starts, including calling basic library functions and requesting resources. The RDMA common initialization library functions are a standardized set of basic library functions in the field of RDMA communication, providing common interfaces for device opening, resource allocation, etc., and providing a unified calling standard for network cards and communication libraries from different manufacturers. RDMA (Remote Direct Memory Access) is a technology that enables direct transfer of data between different computer memory without the involvement of the CPU.
[0026] S12: Call the communication library through the RDMA common initialization library function, execute the communication library's pre-set instance initialization process, and generate the first data structure instance and the second data structure instance.
[0027] For step S12, the call to the RDMA public library is forwarded to sssroce.so through the callback function mechanism, which realizes the connection between the standardized interface and the patented innovative function. It can generate two sets of data structure instances at the same time, breaking the limitation that the traditional network card can only initialize a single instance, and providing key resource support for distinguishing local / remote communication in the future.
[0028] See Figure 7 As shown, during application process initialization, i.e., at the start of the instance initialization process, the RDMA common initialization library function is called first. This RDMA common initialization library function calls the callback in sssroce.so via a callback function mechanism; therefore, instance initialization is implemented internally by sssroce.so. It can be understood that the callback function mechanism, a function call method, automatically triggers the pre-registered interface function with the same name in sssroce.so when the RDMA common initialization library function is called, achieving call forwarding. The communication library maps to interface functions with the same names as the RDMA common initialization library function. The instance initialization process, predefined in the communication library, is used to create and initialize the complete set of data structures corresponding to SSS and RXE.
[0029] In some embodiments of the present invention, such as Figure 3 As shown, a specific scheme for generating a first data structure instance and a second data structure instance is provided. In S12, the communication library is called through the RDMA common initialization library function to execute the communication library's pre-set instance initialization process to generate the first data structure instance and the second data structure instance. Specifically, it includes the following steps S121-S126.
[0030] S121: The device opens the interface by calling the RDMA common initialization library function, and forwards it to the interface function mapped in the communication library through the callback function mechanism to generate the device context of the first data structure instance and the device context of the second data structure instance.
[0031] In step S121, by simultaneously generating device contexts corresponding to SSS and RXE, connection channels are established between the two instances and the hardware (SSS network card) and the virtual driver (RXE). This ensures that the hardware resources of the SSS network card can be accessed during remote communication, and the virtual resources of the RXE can be accessed during local communication, providing underlying support for subsequent resource allocation and communication operations.
[0032] See Figure 8 As shown, in this embodiment, the device open interface refers to the ibv_open_device interface in the RDMA common initialization library function, used to open the specified network interface card (NIC) device and create a device context. The device context is a structure that records core data such as connection information and resource configuration of the NIC device or virtual driver. It is the top-level resource management unit in the RDMA environment, similar to the process control block in an operating system. It provides a unique identifier and operation entry point for the creation and lifecycle management of all subsequent RDMA resources (such as protection domains, memory regions, queue pairs, etc.). Without a device context, the application process will not be able to identify and operate the specific NIC device. The function in the communication library mapped to the device open interface, namely the callback function corresponding to ibv_open_device in sssroce.so, namely sss_alloc_context() / rxe_alloc_context(), is responsible for generating the device context sss_context of the first data structure instance and the device context rxe_context of the second data structure instance.
[0033] S122: Call the protection domain allocation interface of the RDMA common initialization library function, and forward it to the interface function mapped in the communication library through the callback function mechanism to generate the protection domain of the first data structure instance and the protection domain of the first data structure instance.
[0034] Understandably, protection domains can isolate resources from different communication processes, avoiding resource conflicts between local and remote communication. Step S122 simultaneously allocates protection domains to both instances, ensuring that the remote communication resources corresponding to sss and the local communication resources corresponding to rxe are independent of each other, improving communication stability and security, and avoiding communication anomalies caused by resource reuse.
[0035] See Figure 8 As shown in this embodiment, the protection domain allocation interface refers to the ibv_alloc_pd interface in the RDMA common initialization library function, which is used to allocate independent resource isolation domains for RDMA communication. A protection domain (PD) is a resource isolation unit in RDMA communication; all subsequent memory registration, queue creation, and other operations must be associated with the protection domain to ensure that resources in different communication processes do not interfere with each other. The protection domain of the first data structure instance, i.e., sss_pd, is associated with the sss device context and is used for remote communication; the protection domain of the second data structure instance, i.e., rxe_pd, is associated with the rxe device context and is used for local communication. The function in the communication library corresponding to the protection domain allocation interface, i.e., the callback function in sssroce.so corresponding to ibv_alloc_pd, i.e., sss_alloc_pd() / rxe_alloc_pd(), is responsible for generating the protection domains of the first data structure instance / second data structure instance.
[0036] S123: Call the memory registration interface of the RDMA common initialization library function, and forward it to the interface function mapped in the communication library through the callback function mechanism to generate the memory region of the first data structure instance and the memory region of the first data structure instance.
[0037] Understandably, by registering the application's memory region with the RDMA system, the SSS network card or RXE driver can directly access that memory without CPU copying. Step S123 registers memory regions for both instances to ensure that both remote and local communication have dedicated, directly accessible memory, guaranteeing efficient data transmission while avoiding memory access conflicts.
[0038] See Figure 8As shown, in this embodiment, the memory registration interface is the ibv_reg_mr interface in the RDMA public initialization library function, used to register the application's memory address space with the RDMA system and obtain the identifier of the memory region that can be directly accessed. The function in the communication library that maps to the memory registration interface, namely the callback function corresponding to ibv_reg_mr in sssroce.so, namely sss_reg_mr() / rxe_reg_mr(), is responsible for generating the memory regions of the first data structure instance / second data structure instance. The memory region (MR) is the RDMA-accessible memory structure generated after registration, containing information such as memory address, length, and access permissions, and is the core carrier of RDMA data transmission; the memory region of the first data structure instance, i.e., sss_mr, is used for data storage for remote communication and can be directly accessed by the sss network card; the memory region of the second data structure instance, i.e., rxe_mr, is used for data storage for local communication and can be directly accessed by the rxe driver.
[0039] S124: Call the completion queue creation interface of the RDMA common initialization library function, and forward it to the interface function mapped in the communication library through the callback function mechanism to generate the completion queue of the first data structure instance and the completion queue of the first data structure instance.
[0040] For step S124, complete queues are created for the two sets of instances respectively to ensure that the operation results of local communication and remote communication (such as successful sending and successful receiving) can be monitored and processed independently, avoid confusion in the feedback of the results of the two types of communication, improve the accuracy of communication status judgment, and ensure the reliability of data transmission.
[0041] See Figure 8 As shown, in this embodiment, the completion queue creation interface refers to the ibv_create_cq interface in the RDMA common initialization library functions, which is used to create the completion queue for RDMA operations. The function in the communication library that maps to the completion queue creation interface, namely the callback function corresponding to ibv_create_cq in sssroce.so, i.e., sss_create_cq() / rxe_create_cq(), is responsible for generating the completion queues for the first data structure instance / second data structure instance. The completion queue (CQ) is a queue that records the operation completion status in RDMA communication. Each time a send or receive operation is completed, a completion event is inserted into the corresponding completion queue for the application to query. The completion queue of the first data structure instance, i.e., sss_cq, is used to monitor the operation status of remote communication; the completion queue of the second data structure instance, i.e., rxe_cq, is used to monitor the operation status of local communication.
[0042] S125: Call the queue pair creation interface of the RDMA common initialization library function, and forward it to the interface function mapped in the communication library through the callback function mechanism to generate the queue pair of the first data structure instance and the queue pair of the first data structure instance.
[0043] The queue pair (QP) is the core unit responsible for data transmission and reception in RDMA communication. It consists of a transmit queue (SQ) and a receive queue (RQ). Each QP corresponds to an independent communication link, and all data transmission and reception requests are submitted through the queue pair. Step S125 creates queue pairs for the two instances respectively, so that local communication and remote communication have independent operation channels, avoiding request submission conflicts, and providing operation objects for subsequent determination of communication type through the modify_qp interface.
[0044] See Figure 8 As shown, in this embodiment, the queue pair creation interface refers to the ibv_create_qp interface in the RDMA common initialization library functions, which is used to create queue pairs for RDMA communication. The function in the communication library that maps to the queue pair creation interface, namely the callback function in sssroce.so corresponding to ibv_create_qp, i.e., sss_create_qp() / rxe_create_qp(), is responsible for generating queue pairs of the first data structure instance / second data structure instance. The queue pair of the first data structure instance, i.e., sss_qp, is used for remote communication, associating the SSS device context and protection domain; the queue pair of the second data structure instance, i.e., rxe_qp, is used for local communication, associating the RXE device context and protection domain.
[0045] S126: Based on the context, protection domain, memory region, completion queue and queue pair, obtain the first data structure instance and the second data structure instance.
[0046] For step S126, the previously generated results (context, PD, MR, CQ, QP) are integrated into complete first and second data structure instances. Through integration, two sets of fully functional and resource-independent communication support instances are formed, providing a complete resource package that can be used immediately for path selection after subsequent communication type identification, and avoiding communication failures caused by missing components or incorrect associations during subsequent communication processes.
[0047] Understandably, the first data structure instance, composed of sss_context, sss_pd, sss_mr, sss_cq, and sss_qp, is a complete resource set supporting remote communication; the second data structure instance, composed of rxe_context, rxe_pd, rxe_mr, rxe_cq, and rxe_qp, is a complete resource set supporting local communication.
[0048] S20: When a communication request is received, check whether the current communication type is remote communication or local communication.
[0049] Step S20 addresses the technical pain point that existing network cards cannot automatically distinguish between local and remote communication. By dynamically checking the communication type, it achieves automatic adaptation of one network card to two communication scenarios without manual configuration or switching, improving the system's flexibility and compatibility. It also provides a basis for decision-making in selecting the corresponding communication path.
[0050] It is understandable that a communication request is an RDMA data transfer request initiated by an application process, which may come from other processes on the same computer (local communication request) or processes on other computers (remote communication request); remote communication is RDMA communication between different computers, following the original hardware and software process of the SSS network card and relying on the SSS instance; local communication is RDMA communication between different processes under a single port of the network card on the same computer, following the RXE open source driver process and relying on the RXE instance.
[0051] In some embodiments of the present invention, such as Figure 4 As shown, a specific inspection scheme is provided. In S20, when a communication request is received, the current communication type is checked to see if it is remote communication or local communication. Specifically, this includes the following steps S21-S23.
[0052] S21: Call the queue state modification interface in the RDMA common initialization library function, and forward it to the communication library through the callback function mechanism.
[0053] See Figure 7 As shown, when the application process is in the running state, it connects the application layer with the core logic of sssroce.so by calling the RDMA standardized queue state modification interface and the callback of sssroce.so. The design idea is to use the return result of the existing RDMA interface as the judgment basis, without adding a new custom interface, ensuring compatibility with existing RDMA applications and reducing the system transformation cost.
[0054] It is understandable that the queue pair state modification interface refers to the modify_qp interface in the RDMA public library, which is used to modify the working state of the queue pair (such as ready to receive, ready to send, etc.); the callback function mechanism is used to automatically forward the call to the modify_qp interface to the corresponding mapped callback function in sssroce.so, ensuring that the core judgment logic is executed in sssroce.so; the queue pair state is the working stage identifier of the queue pair, and different states correspond to different communication preparation stages. In this embodiment, the IBV_QPS_RTR (ready to receive) state is used.
[0055] S22: Based on the communication library, call the interface function that maps the queue state modification interface in the first data structure instance and the RDMA common initialization library function.
[0056] For step S22, the design idea of prioritizing remote communication is adopted. The modify_qp interface corresponding to the first data structure instance is called first. Since the existing network card originally supports remote communication, prioritizing remote communication can preserve the original communication performance to the greatest extent. Only when remote communication cannot be realized (local communication scenario) will the process be switched to the rxe process, which takes into account both compatibility and performance requirements.
[0057] Understandably, sssroce.so pre-binds the modify_qp interface of the RDMA public library with the sss_modify_qp function of the sss instance to ensure the accuracy of call forwarding. See also Figure 8 As shown, the queue state modification interface in the RDMA common initialization library function is ibv_modify_qp(IBV_QPS_RTR), and the mapping callback function in sssroce.so is sss_modify_qp(sss_qp). If sss_modify_qp(sss_qp) returns failure, then rex_modify_qp(rex_qp) is called.
[0058] S23: Obtain the execution result of the queue state modification interface, and determine whether the execution result is successful; wherein, if the execution result is successful, the current communication type is remote communication; if the execution result is unsuccessful, the current communication type is local communication.
[0059] For step S23, by parsing the return result of the modify_qp interface, the automatic distinction between local and remote communication is realized. Taking advantage of the characteristic that the sss network card does not support the modification of the status of local communication (returns failure), the interface return result is directly associated with the communication type. No additional address judgment or configuration is required, thus realizing imperceptible type recognition.
[0060] The execution result is the return value of the modify_qp interface. For example, in the RDMA standard, a successful return will set the sss_sign flag. If the error code EPERM is returned, it indicates that it is a single-port communication on the same computer's network card. It will initiate rxe_modify_qp(rxe_qp) and set the rxe_sign flag. When a queue pair on a single port of the same computer attempts to modify to the IBV_QPS_RTR state, the sss network card does not support this operation and returns this error code.
[0061] In a further embodiment, the queue state modification interface is a modify_qp interface, and the parameter of the modify_qp interface is set to the IBV_QPS_RTR state; wherein, if modify_qp(IBV_QPS_RTR) returns a success, the execution result is successful, and if it returns an error code, the execution result is failed.
[0062] Understandably, the IBV_QPS_RTR state (Ready to Receive) is a fundamental step in establishing a connection for RDMA communication. The difference in network card support for local / remote communication is most evident when this state is modified, and it can accurately distinguish the communication type, so it is used as a criterion for judgment.
[0063] The modify_qp interface is a standard interface in the RDMA common initialization library functions. Its function prototype is intibv_modify_qp (struct ibv_qp *qp, struct ibv_qp_attr *attr, int attr_mask), which is used to modify the attributes (including the state) of queue pairs.
[0064] S30: Based on the first data structure instance and the second data structure instance, select a preset first communication path to process remote data, or select a second communication path to process local data.
[0065] For step S30, based on the previously determined communication type, the data transmission request is routed to the corresponding communication path, realizing the dynamic switching of one set of network card ports and two sets of communication paths. This solves the defect that the existing network card can only support a single path, while ensuring that remote communication continues to use the original efficient process, and local communication enables the RXE compatible process, thus balancing performance and compatibility.
[0066] The first communication path corresponds to the remote communication processing flow, consisting of the sss_qp, sss_post_send, etc. of the sss instance and the hardware and software resources of the sss network card. It is an efficient communication path supported by the original network card. The second communication path corresponds to the local communication processing flow, consisting of the rxe_qp, rxe_post_send, etc. of the rxe instance and the rxe open-source driver. It is a compatible communication path based on Soft-RoCE. Data processing refers to the entire process of sending, receiving, and confirming the status of RDMA data.
[0067] In some embodiments of the present invention, such as Figure 5 As shown, a specific inspection scheme is provided. In S30, based on the first data structure instance and the second data structure instance, a preset first communication path is selected to process remote data, or a second communication path is selected to process local data. Specifically, it includes the following steps S31-S32.
[0068] S31: Based on the selected first communication path, select the sending interface, the completion queue polling interface, the receiving interface, and the completion queue polling interface under the first data structure instance; or, based on the selected second communication path, select the sending interface, the completion queue polling interface, the receiving interface, and the completion queue polling interface under the second data structure instance.
[0069] Step S31 involves the specific adaptation of the communication path, selecting the corresponding functional interface for different paths to ensure consistency between the interface and the instance. Each communication path has a dedicated interface (SSS interface corresponds to an SSS instance, and RXE interface corresponds to an RXE instance), avoiding communication failures caused by mixing interfaces. This also standardizes the communication process and improves code maintainability and communication stability.
[0070] See Figure 9As shown in this embodiment, the sending interface refers to the `post_send` series of functions, used to send data submission requests to the queue. The first communication path corresponds to `sss_post_send`, and the second communication path corresponds to `rxe_post_send`. The receiving interface refers to the `post_recv` series of functions, used to submit data reception requests to the queue. The first communication path corresponds to `sss_post_recv`, and the second communication path corresponds to `rxe_post_recv`. The completion queue polling interface refers to the `poll_cq` series of functions, used to query events in the completion queue to confirm whether the sending or receiving operation is complete. The first communication path corresponds to `sss_poll_cq`, and the second communication path corresponds to `rxe_poll_cq`. It can be understood that the calls to the sending interface, completion queue polling interface, receiving interface, and completion queue polling interface all involve first calling the `sending`, `completion queue polling`, `receiving`, and `completion queue polling` interfaces in the RDMA common initialization library functions, and then calling back the logic of the `sending`, `completion queue polling`, `receiving`, and `completion queue polling` interfaces in the communication library.
[0071] S32: In sequence, call the sending interface to submit data, call the completion queue polling interface to check the sending status, call the receiving interface to prepare to receive, and call the completion queue polling interface again to check the receiving status.
[0072] Step S32 follows the standard RDMA communication process, namely send-acknowledge-receive-acknowledge, ensuring the integrity and reliability of data transmission. By calling the interface step by step and checking the status, transmission anomalies, such as sending failures and receiving timeouts, can be detected in a timely manner. At the same time, it unifies the execution flow of local and remote communication, allowing applications to process communication according to the same logic without distinguishing between communication types, thus improving the ease of use and compatibility of the method.
[0073] Understandably, submitting data refers to writing the description information (such as memory address and length) of the data to be transmitted into the send queue of the queue pair through the send interface, triggering the RDMA driver or network card to execute data transmission; checking the sending status means querying the completion queue through the polling interface to confirm whether the data has been successfully sent to the other end; preparing to receive means writing the description information of the receive buffer into the receive queue of the queue pair through the receive interface to inform the driver or network card of the storage location of the received data; checking the receive status means confirming whether the data has been successfully received into the local buffer through the polling interface.
[0074] See Figure 6As shown, taking local communication as an example, the service program (process b1) of computer B first calls rxe_post_recv to prepare to receive data, and the client program (process b2) calls rxe_post_send to submit data. Then, both parties call rxe_poll_cq to check the status. During remote communication, the service program (process a1) of computer A calls sss_post_recv, and the client program (process b3) of computer B calls sss_post_send. Similarly, the status is checked through sss_poll_cq. The standardized process ensures the stability and repeatability of the two types of communication, and ultimately all test cases pass.
[0075] In a further embodiment, the sending interface of the first data structure instance is the sss_post_send function, the receiving interface is the sss_post_recv function, and the queue polling interface is the sss_poll_cq function; the sending interface of the second data structure instance is the rxe_post_send function, the receiving interface is the rxe_post_recv function, and the queue polling interface is the rxe_poll_cq function.
[0076] The interface functions are the native interfaces of the sss network card driver and the rxe open-source driver, respectively. Direct calls can maximize their respective performance advantages, while avoiding compatibility issues caused by custom interfaces and ensuring compatibility with the existing RDMA ecosystem.
[0077] Understandably, `sss_post_send`, `sss_post_recv`, and `sss_poll_cq` are native interfaces provided by the SSS network card driver, specifically designed for RDMA operations of SSS instances, adapted for remote communication scenarios, and possess hardware acceleration capabilities. Instances of `sss_post_send`, `sss_post_recv`, and `sss_poll_cq` are `sss_qp`, `sss_qp`, and `sss_cq`, respectively. Meanwhile, `rxe_post_send`, `rxe_post_recv`, and `rxe_poll_cq` are open-source software... The RoCE (rxe) driver provides native interfaces specifically for RDMA operations on rxe instances, adapting to local communication scenarios. It enables RDMA interconnect communication within a single computer without hardware support. Instances of rxe_post_send / rxe_post_recv / rxe_poll_cq are rxe_qp, rxe_qp, and rxe_cq, respectively. It can also be understood that the first data structure instance (sss instance) is bound to the sss series interfaces, and the second data structure instance (rxe instance) is bound to the rxe series interfaces, ensuring consistency between operations and instances.
[0078] As can be seen, in the above solution, by initializing two sets of data structure instances, the local and remote communication types are automatically identified and the corresponding communication paths are matched, which is perfectly compatible with two types of RDMA communication scenarios. There is no need to replace the network card hardware. This avoids the high cost of dedicated network cards and the waste of existing hardware resources, saves the process of large-scale system transformation, reduces the difficulty of deployment and upgrade, and adapts to the core scenario requirements such as distributed storage and high-performance computing.
[0079] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0080] In one embodiment, the present invention provides a network interface card (NIC) multi-node interconnection communication device 100, which corresponds one-to-one with the NIC multi-node interconnection communication method described in the above embodiments. For example... Figure 10 As shown, the network card multi-node interconnection communication device 100 includes a configuration module 101, a checking module 102, and an allocation module 103. Detailed descriptions of each functional module are as follows: Configuration module 101 is used to perform communication library initialization operations, and to configure a first data structure instance for remote communication and a second data structure instance for local communication.
[0081] The inspection module 102 is used to check whether the current communication type is remote communication or local communication when a communication request is received.
[0082] The allocation module 103 is used to select a preset first communication path to process remote data or select a second communication path to process local data based on the first data structure instance and the second data structure instance.
[0083] In one embodiment, the configuration module 101 is specifically used for: The application process initialization procedure is executed, and the RDMA common initialization library functions are called; The communication library is called through the RDMA common initialization library function to execute the pre-set instance initialization process of the communication library and generate the first data structure instance and the second data structure instance.
[0084] The step of calling the communication library through the RDMA common initialization library function, executing the pre-set instance initialization process of the communication library, and generating the first data structure instance and the second data structure instance includes: The device opens an interface by calling the RDMA common initialization library function, and forwards the call to the interface function mapped in the communication library through a callback function mechanism to generate the device context of the first data structure instance and the device context of the second data structure instance. The protection domain allocation interface of the RDMA common initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the protection domain of the first data structure instance and the protection domain of the first data structure instance. The memory registration interface of the RDMA public initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the memory region of the first data structure instance and the memory region of the first data structure instance. The completion queue creation interface of the RDMA common initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the completion queue of the first data structure instance and the completion queue of the first data structure instance. The queue pair creation interface calls the RDMA common initialization library function, and forwards it to the interface function mapped in the communication library through the callback function mechanism to generate the queue pair of the first data structure instance and the queue pair of the first data structure instance. Based on the context, protection domain, memory region, completion queue, and queue pair, the first data structure instance and the second data structure instance are obtained.
[0085] In one embodiment, the inspection module 102 is specifically used for: The queue state modification interface in the RDMA common initialization library function is called, and forwarded to the communication library through a callback function mechanism. Based on the communication library, call the interface function that maps the queue state modification interface in the first data structure instance and the RDMA common initialization library function; Obtain the execution result of the queue state modification interface and determine whether the execution result is successful; if the execution result is successful, the current communication type is remote communication; if the execution result is unsuccessful, the current communication type is local communication.
[0086] The queue state modification interface is the modify_qp interface, and the parameter of the modify_qp interface is set to the IBV_QPS_RTR state. If modify_qp(IBV_QPS_RTR) returns a success, the execution result is successful; if it returns an error code, the execution result is unsuccessful.
[0087] In one embodiment, the configuration module 103 is specifically used for: Based on the selected first communication path, select the sending interface, the queue polling interface, the receiving interface, and the queue polling interface under the first data structure instance; or, based on the selected second communication path, select the sending interface, the queue polling interface, the receiving interface, and the queue polling interface under the second data structure instance. The process involves sequentially calling the sending interface to submit data, calling the completion queue polling interface to check the sending status, calling the receiving interface to prepare for receiving, and calling the completion queue polling interface again to check the receiving status.
[0088] In this context, the sending interface of the first data structure instance is the sss_post_send function, the receiving interface is the sss_post_recv function, and the queue polling interface is the sss_poll_cq function; the sending interface of the second data structure instance is the rxe_post_send function, the receiving interface is the rxe_post_recv function, and the queue polling interface is the rxe_poll_cq function.
[0089] Specific limitations regarding the network interface card (NIC) multi-node interconnection communication device 100 can be found in the limitations of the NIC multi-node interconnection communication method described above, and will not be repeated here. Each module in the aforementioned NIC multi-node interconnection communication device 100 can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0090] In one embodiment, a computer device 200 is provided, which may be a server, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device 200 includes a processor 220, memory, and a network interface 250 connected via a system bus 210. The processor 220 provides computing and control capabilities. The memory of the computer device 200 includes non-volatile and / or volatile storage media and internal memory 240. The non-volatile storage media 230 stores an operating system 231, computer programs 232, and a database 233. The internal memory 240 provides an environment for the operation of the operating system and computer programs in the non-volatile storage media 230. The network interface 250 of the computer device 200 is used for communication with external clients via a network connection. When the computer program is executed by the processor 220, it implements the functions or steps of a network interface card (NIC) multi-node interconnection communication method server. That is, when the processor 220 executes the computer program, it performs the following steps: Perform communication library initialization operations, and configure a first data structure instance for remote communication and a second data structure instance for local communication; When a communication request is received, check whether the current communication type is remote communication or local communication; Based on the first data structure instance and the second data structure instance, a preset first communication path is selected to process remote data, or a second communication path is selected to process local data.
[0091] In one embodiment, a computer device 300 is provided, which may be a client, and its internal structure diagram may be as follows: Figure 12 As shown. The computer device includes a processor 320, memory, network interface 350, display screen 370, and input device 360 connected via a system bus 310. The processor 320 provides computing and control capabilities. The memory includes a non-volatile storage medium 330 and internal memory 340. The non-volatile storage medium 330 stores an operating system 331 and a computer program 332. The internal memory provides an environment for the operation of the operating system 331 and the computer program 332 in the non-volatile storage medium 330. The network interface 350 of the computer device 300 is used for communication with an external server via a network connection. When the computer program is executed by the processor 320, it implements the functions or steps of a network interface card (NIC) multi-node interconnection communication method on the client side. That is, when the processor 320 executes the computer program 332, it implements the following steps: Perform communication library initialization operations, and configure a first data structure instance for remote communication and a second data structure instance for local communication; When a communication request is received, check whether the current communication type is remote communication or local communication; Based on the first data structure instance and the second data structure instance, a preset first communication path is selected to process remote data, or a second communication path is selected to process local data.
[0092] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0093] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0094] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0095] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for multi-node interconnection communication of network interface cards, characterized in that, include: Perform communication library initialization operations, and configure a first data structure instance for remote communication and a second data structure instance for local communication; When a communication request is received, check whether the current communication type is remote communication or local communication; Based on the first data structure instance and the second data structure instance, a preset first communication path is selected to process remote data, or a second communication path is selected to process local data.
2. The network card multi-node interconnection communication method according to claim 1, characterized in that, The execution of the communication library initialization operation, configuring a first data structure instance for remote communication and a second data structure instance for local communication, includes: The application process initialization procedure is executed, and the RDMA common initialization library functions are called; The communication library is called through the RDMA common initialization library function to execute the pre-set instance initialization process of the communication library and generate the first data structure instance and the second data structure instance.
3. The network card multi-node interconnection communication method according to claim 2, characterized in that, The step of calling the communication library through the RDMA common initialization library function, executing the pre-set instance initialization process of the communication library, and generating the first data structure instance and the second data structure instance includes: The device opens an interface by calling the RDMA common initialization library function, and forwards the call to the interface function mapped in the communication library through a callback function mechanism to generate the device context of the first data structure instance and the device context of the second data structure instance. The protection domain allocation interface of the RDMA common initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the protection domain of the first data structure instance and the protection domain of the first data structure instance. The memory registration interface of the RDMA public initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the memory region of the first data structure instance and the memory region of the first data structure instance. The completion queue creation interface of the RDMA common initialization library function is called, and forwarded to the interface function mapped in the communication library through the callback function mechanism to generate the completion queue of the first data structure instance and the completion queue of the first data structure instance. The queue pair creation interface calls the RDMA common initialization library function, and forwards it to the interface function mapped in the communication library through the callback function mechanism to generate the queue pair of the first data structure instance and the queue pair of the first data structure instance. Based on the context, protection domain, memory region, completion queue, and queue pair, the first data structure instance and the second data structure instance are obtained.
4. The network card multi-node interconnection communication method according to claim 3, characterized in that, The step of checking whether the current communication type is remote communication or local communication when a communication request is received includes: The queue state modification interface in the RDMA common initialization library function is called, and forwarded to the communication library through a callback function mechanism. Based on the communication library, call the interface function that maps the queue state modification interface in the first data structure instance and the RDMA common initialization library function; Obtain the execution result of the queue state modification interface and determine whether the execution result is successful; if the execution result is successful, the current communication type is remote communication; if the execution result is unsuccessful, the current communication type is local communication.
5. The network interface card (NIC) multi-node interconnection communication method according to claim 4, characterized in that, The queue state modification interface is the modify_qp interface, and the parameter of the modify_qp interface is set to the IBV_QPS_RTR state; wherein, if modify_qp(IBV_QPS_RTR) returns a success, the execution result is successful, and if it returns an error code, the execution result is failed.
6. The network card multi-node interconnection communication method according to claim 5, characterized in that, The step of selecting a preset first communication path to process remote data or selecting a second communication path to process local data based on the first data structure instance and the second data structure instance includes: Based on the selected first communication path, select the sending interface, the queue polling interface, the receiving interface, and the queue polling interface under the first data structure instance; or, based on the selected second communication path, select the sending interface, the queue polling interface, the receiving interface, and the queue polling interface under the second data structure instance. The process involves sequentially calling the sending interface to submit data, calling the completion queue polling interface to check the sending status, calling the receiving interface to prepare for receiving, and calling the completion queue polling interface again to check the receiving status.
7. The network interface card (NIC) multi-node interconnection communication method according to claim 6, characterized in that, The first data structure instance uses the sss_post_send function as its sending interface, the sss_post_recv function as its receiving interface, and the sss_poll_cq function as its queue polling interface. The second data structure instance uses the rxe_post_send function as its sending interface, the rxe_post_recv function as its receiving interface, and the rxe_poll_cq function as its queue polling interface.
8. A network card multi-node interconnection communication device, characterized in that, include: The configuration module is used to perform communication library initialization operations, configuring a first data structure instance for remote communication and a second data structure instance for local communication. The inspection module is used to check whether the current communication type is remote communication or local communication when a communication request is received. The allocation module is used to select a preset first communication path to process remote data or select a second communication path to process local data based on the first data structure instance and the second data structure instance.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the network interface card multi-node interconnection communication method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the network card multi-node interconnection communication method as described in any one of claims 1 to 7.