An end-to-end large-scale RDMA network interconnection method

By establishing multiple virtual networks in the RDMA network and adopting isolation mechanisms and encryption strategies, dynamically adjusting traffic resources and quickly switching backup links, the problems of resource competition and performance degradation in the RDMA network are solved, and efficient and secure data transmission is achieved.

CN119629059BActive Publication Date: 2025-10-17GUANGXI POWER GRID CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411721489.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-17
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

In the prior art, when multiple virtual links share one RDMA link resource, resource contention and performance degradation may occur, especially when the virtual link performance is degraded under high load conditions.

Method used

Virtualization technology is used to establish multiple virtual networks in the RDMA areas of the sending and receiving ends. Each virtual network is marked with a unique identifier, and an isolation mechanism is used to allocate independent resources. Data is encrypted according to task requirements, and access control policies are established. Traffic resource allocation is dynamically adjusted through the SDN controller, and rapid switching to the backup virtual link occurs in the event of a failure.

Benefits of technology

It effectively solves the problem of resource competition, improves the security and performance of data transmission, ensures the stability and reliability of virtual links under high load conditions, and quickly recovers from network failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119629059B_ABST
    Figure CN119629059B_ABST
Patent Text Reader

Abstract

The application discloses an end-to-end large-scale RDMA network interconnection method, which comprises the following steps: a plurality of virtual networks are established in the RDMA area of a sending end and a receiving end by using a virtualization technology; each virtual network is marked; each virtual network corresponds to a group of virtual links; different virtual networks are identified by a virtual network identifier; an isolation mechanism is used between the virtual networks; independent resources, including bandwidth and a buffer area, are allocated to each virtual network; data on the RDMA link is encrypted according to the task demand of different virtual networks; an access control strategy is established to limit access authority; an end-to-end flow control mechanism is established between the sending end and the receiving end; the routing is programmed by an SDN controller; the resource allocation in the virtual network is dynamically adjusted; a fault recovery mechanism is established; when a network fault occurs, the system can quickly switch to a standby virtual link.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the technical field of communication networks, and in particular to an end-to-end large-scale RDMA network interconnection method. BACKGROUND

[0002] With the rapid development of technologies such as cloud computing, big data and artificial intelligence, the requirements for the data transmission performance of computer networks are becoming higher and higher. As a high-performance computer network communication technology, RDMA can realize direct memory access of data, reduce the involvement of CPU, and significantly improve the throughput and reduce the delay of data transmission.

[0003] In Chinese Patent Publication No. CN115623057A, a method, device and equipment based on RDMA are disclosed. In the core technical solution of the application, for the scenario that multiple communication links need to be established between the upper-layer applications on the client and the server, a virtual link is abstracted on the RDMA link, multiple virtual connections are allowed to be created on one RDMA link, and the resources of one RDMA link are shared by multiple virtual links, so that the two upper-layer applications realize multi-channel data transmission through the multiple virtual links carried on the same RDMA link. In this way, the problem of needing multiple communication links between the two upper-layer applications can be solved, and the number of RDMA links can be reduced because the same RDMA link carries multiple virtual links, which to some extent avoids the problem of slow RDMA connection, especially in the case of a large number of required links, the overall connection speed can be significantly improved. In the prior art, multiple virtual links share the resources of one RDMA link, and when multiple virtual links request resources at the same time, resource competition problems occur in the transmission process. In the case of high load, the performance of the virtual link will decrease.

[0004] The application proposes a solution to the above-mentioned shortcomings: a virtualization technology is used to establish multiple virtual networks in the RDMA area of the sending end and the receiving end, and each virtual network is marked with a unique identifier tag. Each virtual network corresponds to a group of virtual links. An isolation mechanism is used between the virtual networks, independent resources are allocated to each virtual network, data on the RDMA link is encrypted according to the task requirements of different virtual networks, an access control policy is established to limit access permissions, an end-to-end flow control mechanism is established between the sending end and the receiving end, the routing is programmed by an SDN controller, the flow resource allocation in the virtual network is dynamically adjusted, and a fault recovery mechanism is established to quickly switch to a backup virtual link when a network fault occurs, thereby solving the problem of resource competition and performance decrease of the virtual link caused by multiple virtual links sharing the resources of one RDMA link. SUMMARY

[0005] The present application provides an end-to-end large-scale RDMA network interconnection method to solve the above technical problems.

[0006] The technical scheme of the present application is as follows: an end-to-end large-scale RDMA network interconnection method, comprising:

[0007] S1, a plurality of virtual networks are established in the RDMA area of the sending end and the receiving end by using virtualization technology, each virtual network is marked, and each virtual network corresponds to a group of virtual links;

[0008] S2, different virtual networks are identified by a virtual network identifier, an isolation mechanism is used between the virtual networks, and each virtual network is allocated independent resources, including bandwidth and buffer;

[0009] S3, according to the task requirements of different virtual networks, data on the RDMA link is encrypted, an access control strategy is established, and access permission is limited;

[0010] S4, an end-to-end flow control mechanism is established between the sending end and the receiving end, the routing is programmed by an SDN controller, and the resource allocation in the virtual network is dynamically adjusted;

[0011] S5, a fault recovery mechanism is established, and when a network fault occurs, the backup virtual link is quickly switched to.

[0012] Further, in the S1 step, the plurality of virtual networks are established in the RDMA area of the sending end and the receiving end by using virtualization technology, a connection request is sent from the sending end to the receiving end, a virtual network is established on the RDMA link, the RDMA context of the virtual network is configured, including configuring independent queue QP and independent completion queue CP, each independent queue QP is connected to the corresponding queue QP of the remote node in the RDMA area of the sending end and the receiving end by using the RDMA protocol RoCE, when the virtual network is established on the RDMA link, a unique identifier VN I tag is allocated to mark each virtual network, each virtual network corresponds to a group of virtual links, the group of virtual links includes the configured main virtual link and the backup virtual link, VTEP is configured for the virtual network, VXLAN tunnel is established, and the data packet is transmitted in the VTEP and the VXLAN tunnel through the unique identifier;

[0013] Further, the group of virtual links corresponding to each virtual network is set as follows: a virtual network is created, a bridge mode is selected for configuration, a virtual machine is directly connected to a network interface card of a physical machine, a unique name and IP address are allocated to each virtual network, and the connectivity with an external network is checked on each virtual machine by using a ping command or other network tools.

[0014] Further, the step S2, the different virtual networks are identified by the virtual network identifier, the isolation mechanism is used between the virtual networks, the different subnets are identified according to the unique identifier VN I tag of each virtual network, the bridge isolation mechanism is selected as the isolation mechanism, a virtual bridge is created on the physical machine, the network interface of the virtual machine is configured to be connected to the created virtual bridge when the virtual machine is created, the independent IP address is configured for each virtual machine in the same subnet, a plurality of independent virtual bridges are created, the different virtual machines are connected to the different virtual bridges, and the independent bandwidth and buffer resources are allocated according to the bandwidth requirement and the buffer usage of each virtual network.

[0015] Further, the step S3, the data on the RDMA link is encrypted according to the task requirement of the different virtual networks, the access control strategy is established, the symmetric encryption algorithm is implemented according to the task requirement of the different virtual networks, the data is converted into ciphertext by using the encryption algorithm and the encryption key, and the encrypted data is transmitted in the VTEP and the VXLAN tunnel through the unique identifier.

[0016] The sending end uses the AES symmetric encryption algorithm and the key to encrypt the data, the key is subjected to permutation, substitution and XOR operation to encrypt the data, and the encrypted data is transmitted in the VTEP and the VXLAN tunnel to the receiving end through the unique identifier of the RDMA link.

[0017] The receiving end uses the same AES symmetric encryption algorithm and the key to decrypt the data, and the decrypted data restores the plaintext data.

[0018] The access control strategy is established, the role-based access control (RBAC) model is established, the user, the role and the permission in the model are defined, the user is allocated to the corresponding role according to the permission requirement of the user, the mapping between the user and the role in the model is established, the permission is allocated to the corresponding role according to the permission range, the mapping between the role and the permission is established, the access control strategy is formulated according to the relationship mapping of the RBAC model, the different access permissions are limited according to the access control strategy, and the data security is improved.

[0019] Further, the step S4, the end-to-end flow control mechanism is established between the sending end and the receiving end, the route is programmed by the SDN controller, the flow resources are allocated by using the SDN controller according to the data transmission rate, the delay requirement and the reliability requirement of the end-to-end flow control, and the specific steps are as follows.

[0020] S401, the SDN controller is configured, and the network state information is collected, the SND controller interacts with the network device router (SND switch) through the southbound interface, and the load information on the link is integrated.

[0021] S402, flow recognition, the SDN recognizes and classifies the flow in the virtual network in a fine-grained manner, the flow type includes interactive flow and bulk transmission flow, the interactive flow is small data packet and frequent transmission, the bulk transmission flow is large data packet and long transmission interval, and a flow table is deployed on the network according to the programming ability of the SDN controller to identify and mark different types of flow;

[0022] S403, flow control, according to the identification of the flow, the type and priority of the flow are obtained for the classification and scheduling of the flow, the classification of the flow adopts the identification port number, the scheduling of the flow is through big data analysis, the priority of the flow is divided into standard priority and high priority, and the SDN performs flow control strategy through the centralized controller and the programming ability;

[0023] S404, flow optimization, the SDN controller optimizes the network structure and the flow path, and dynamically adjusts the network topology according to the load and delay of the flow.

[0024] Further, the S5 step of establishing the fault recovery mechanism, when the network fails, quickly switches to the standby virtual link, establishes the BFD fast detection technology, and monitors the state of the network virtual link in real time. When the main virtual link of the network is detected to have a fault, a switching mechanism is triggered to quickly switch to the standby virtual link to timely respond to the occurrence of network failure.

[0025] Further, the establishment of the fault recovery mechanism specifically comprises: a static BFD session is created by manually adding opposite neighbor information through static configuration, when the opposite interface also starts the BFD and correctly responds to the BFD packet of the opposite end, the static BFD configuration is completed, the BFD control packet is periodically sent to each other to judge the session state, when a plurality of packets are not received, the virtual link has a fault, a switching mechanism is triggered, through the configured network device, the main virtual link is automatically switched to the standby virtual link when the main virtual link fails, if the automatic switching cannot meet the situation, the administrator manually triggers the switching mechanism through an easy-to-operate switching interface or tool, and the main virtual link is quickly switched to the standby virtual link when the main virtual link fails.

[0026] Beneficial effects

[0027] The present application aims at the shortcomings that multiple virtual links share a resource of an RDMA link in the prior art, but when multiple virtual links request resources at the same time, resource competition occurs in the transmission process, by adopting virtualization technology to establish multiple virtual networks in the RDMA area of the sending end and the receiving end, and using a unique identifier to mark each virtual network, each virtual network corresponds to a group of virtual links, the problem of resource competition in the prior art is solved, the virtual networks are isolated through the isolation mechanism, and the data is encrypted by using the encryption algorithm, so as to ensure the security of data transmission, an access control strategy is established to limit the access permission of the user, the security of data is improved, an end-to-end flow control mechanism is established between the sending end and the receiving end, an SDN controller is used to dynamically adjust the resource allocation in the virtual network, and a fault recovery mechanism is established, when a fault occurs in the network, a switching mechanism is triggered to quickly switch to a standby virtual link, so that the problem of performance decline of the virtual link under high load is solved, and the present application solves the problems of multiple virtual links sharing a resource of an RDMA link and resource competition and performance decline of the virtual link by that each virtual network corresponds to a group of virtual links, and a standby virtual link is established in each group of virtual links, the routing is programmed by the SDN controller to dynamically adjust the resource allocation in the virtual network. BRIEF DESCRIPTION OF DRAWINGS

[0028] Figure 1 A structural block diagram of an end-to-end large-scale RDMA network interconnection method in an embodiment of the present application is shown in the figure.

[0029] Figure 2 A step flow chart of an end-to-end large-scale RDMA network interconnection method in an embodiment of the present application is shown in the figure.

[0030] Figure 3 A step flow chart of a symmetric encryption algorithm of an end-to-end large-scale RDMA network interconnection method in an embodiment of the present application is shown in the figure.

[0031] Figure 4 A step flow chart of an SDN controller distributing flow resources of an end-to-end large-scale RDMA network interconnection method in an embodiment of the present application is shown in the figure. DETAILED DESCRIPTION

[0032] In order to make the purpose and advantages of the present application more clear and explicit, the present application is further described below in combination with embodiments; it should be understood that the specific embodiments described herein are only used to explain the present application, and do not limit the present application.

[0033] The preferred implementation methods of the present application are described below with reference to the accompanying drawings. Those skilled in the art should understand that these implementation methods are only used to explain the technical principles of the present application, and do not limit the protection scope of the present application.

[0034] It can be understood that the terms "first", "second" and the like used in the present application can be used herein to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from another element.

[0035] As used herein, the singular forms "a", "an" and "the" can include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "comprise / comprising" or "have / having" specifies the presence of stated features, integers, steps, operations, components, parts, or combinations thereof, but does not exclude the presence or addition of one or more other features, integers, steps, operations, components, parts, or combinations thereof. At the same time, the term "and / or" used in the specification includes any and all combinations of the related listed items.

[0036] Referring to Figures 1-4 As shown in the figure, an end-to-end large-scale RDMA network interconnection method comprises:

[0037] S1, a plurality of virtual networks are established in the RDMA area of the sending end and the receiving end by using virtualization technology, each virtual network is marked, and each virtual network corresponds to a group of virtual links;

[0038] S2, different virtual networks are identified by virtual network identifiers, an isolation mechanism is used between the virtual networks, and each virtual network is allocated independent resources including bandwidth and buffer;

[0039] S3, according to the task requirements of different virtual networks, data on the RDMA link is encrypted, an access control strategy is established, and access permission is limited;

[0040] S4, an end-to-end flow control mechanism is established between the sending end and the receiving end, the routing is programmed by an SDN controller, and the resource allocation in the virtual network is dynamically adjusted;

[0041] S5, a fault recovery mechanism is established, and when a network fault occurs, the standby virtual link is quickly switched to.

[0042] As Figures 1-2 shown, in the S1 step, a plurality of virtual networks are established in the RDMA area of the sending end and the receiving end by using virtualization technology, each virtual network is marked, and each virtual network corresponds to a group of virtual links;

[0043] Specifically, in step S1, virtualization technology is used to establish multiple virtual networks in the RDMA areas of the sending and receiving ends. The sending end sends a connection request to the receiving end, establishes a virtual network on the RDMA link, and configures the RDMA context of the virtual network, including configuring an independent queue QP and an independent completion queue CP. The RDMA protocol RoCE is used in the RDMA areas of the sending and receiving ends to connect each independent queue QP to the queue QP corresponding to the remote node. When establishing a virtual network on the RDMA link, a unique identifier VNI label is assigned to mark each virtual network. Each virtual network corresponds to a group of virtual links, which include a configured primary virtual link and a backup virtual link. VTEP is configured for the virtual network, and a VXLAN tunnel is established. Data packets are transmitted in the VTEP and VXLAN tunnel using the unique identifier.

[0044] Each virtual network corresponds to a set of virtual links: create a virtual network and select bridge mode for configuration, connect the virtual machine directly to the network interface card of the physical machine, assign a unique name and IP address to each virtual network, and on each virtual machine, use the ping command or other network tools to check connectivity with the external network.

[0045] like Figures 1-2 As shown, in step S2, different virtual networks are identified by virtual network identifiers, an isolation mechanism is adopted between virtual networks, and each virtual network is allocated independent resources, including bandwidth and buffer;

[0046] Specifically, in step S2, different virtual networks are identified by virtual network identifiers, an isolation mechanism is adopted between virtual networks, different subnets are identified according to the unique identifier VNI tag of each virtual network, the isolation mechanism selects a bridge isolation mechanism, a virtual bridge is created on the physical machine, and when creating a virtual machine, the network interface of the virtual machine is configured to connect to the previously created virtual bridge, an independent IP address is configured for each virtual machine in the same subnet, multiple independent virtual bridges are created, different virtual machines are connected to different virtual bridges, and independent broadband and buffer resources are allocated according to the broadband requirements and buffer usage of each virtual network.

[0047] like Figures 1-3 As shown, in the S3 step, data on the RDMA link is encrypted according to the task requirements of different virtual networks, and an access control policy is established to restrict access rights;

[0048] Specifically, in the S3 step, the data on the RDMA link is encrypted according to the task requirements of different virtual networks, an access control policy is established, and a symmetric encryption algorithm is implemented according to the task requirements of different virtual networks, so that the data is converted into ciphertext by using the encryption algorithm and the encryption key, and the encrypted data is transmitted through the unique identifier in the VTEP and the VXLAN tunnel.

[0049] The sending end uses the AES symmetric encryption algorithm and the key to encrypt the data, the key is encrypted through the processes of permutation, substitution and XOR operation, and the encrypted data is transmitted through the unique identifier in the VTEP and the VXLAN tunnel to the receiving end, wherein the symmetric encryption algorithm AES uses the ECB mode, and the encryption steps are as follows:

[0050] S301, key and plaintext preparation, initial key and plaintext are prepared, and the plaintext is filled with PKCS#7 if it is insufficient in bytes;

[0051] S302, key expansion, the initial key is expanded into multiple sub-keys;

[0052] S303, initial round, performing a round key addition operation, XORing the plaintext with the first expanded sub-key to obtain a state matrix;

[0053] S304, main round, performing 10, 12 or 14 rounds of iteration according to the key length, each round including 4 steps, and performing byte substitution, row displacement, column confusion and round key addition;

[0054] S305, final round, performing 3 steps in the last round, namely byte substitution, row displacement and round key addition;

[0055] The receiving end uses the same AES symmetric encryption algorithm and the key to decrypt the data, and the decrypted data restores the plaintext data, wherein the decryption process is the inverse process of the encryption process, and specifically as follows:

[0056] Initial round: performing a round key addition operation;

[0057] Main round: performing inverse round key addition, inverse column confusion, inverse row displacement and inverse byte substitution operations;

[0058] Final round: performing inverse row displacement, inverse byte substitution and inverse round key addition operations;

[0059] An access control strategy is established, a role-based access control (RBAC) model is established, users, roles and permissions in the model are defined, users are assigned to corresponding roles according to the permission requirements of the users, a mapping between the users and the roles in the model is established, permissions are assigned to corresponding roles according to the permission range, a mapping between the roles and the permissions is established, an access control strategy is formulated according to the relationship mapping of the RBAC model, when a user attempts to access a system resource, the system judges whether the user has access permission according to the roles and the permissions possessed by the user, if the user has the corresponding permissions, the access request is allowed, otherwise, the access request is denied, by limiting different access permissions, the security of data is improved, wherein all the user information of the users, the roles and the permissions involved in the specification are legally authorized.

[0060] As shown in Figures 1-4 The S4 step establishes an end-to-end flow control mechanism between the sending end and the receiving end, the routing is programmed by the SDN controller, and the resource allocation in the virtual network is dynamically adjusted;

[0061] Specifically, in the S4 step, the end-to-end flow control mechanism between the sending end and the receiving end is established, the routing is programmed by the SDN controller, the flow resources are allocated by the SDN controller according to the data transmission rate, delay requirement and reliability requirement of the end-to-end flow control, and the specific steps are as follows:

[0062] S401, configure the SDN controller, collect network state information, and the SND controller interacts with the network device router (SND switch) through the southbound interface to integrate the load information on the link;

[0063] S402, flow identification, the SDN identifies and classifies the flow in the virtual network in a fine-grained manner, the flow types include interactive flow and bulk transmission flow, the interactive flow is small data packet and frequent transmission, the bulk transmission flow is large data packet and long transmission interval, and a flow table is deployed on the network according to the programming capability of the SDN controller to identify and mark different types of flows;

[0064] S403, flow control, according to the identification of the flow, the type and priority of the flow are obtained for flow classification and scheduling, the flow classification adopts identification of port number, the flow scheduling is through big data analysis, the priority of the flow is divided into standard priority and high priority, and the SDN performs flow control strategy through the centralized controller and the programming capability;

[0065] S404, flow optimization, the SDN controller optimizes the network structure and the flow path, and dynamically adjusts the network topology according to the load and delay of the flow;

[0066] In step S404, the network topology is dynamically adjusted by using the shortest path Dijkstra algorithm:

[0067] Two sets P and Q are introduced. The data of sets P and Q include traffic load and delay. Set P records the vertices for which the shortest path has been found, and set Q records the vertices for which the shortest path has not yet been found.

[0068] Initialization: The distance of the starting point is set to 0, and the distances of other vertices are set to infinity (indicating that the point has not been found). During initialization, the set P only contains the starting point p, and the set Q contains all vertices except the starting point p;

[0069] Execution process: Select the vertex k with the shortest distance to the starting point p from the set Q and add the vertex k to the set P. At the same time, remove the vertex k from the set Q, update the distance from each vertex in the set Q to the starting point p, update the vertices in the set Q and the paths corresponding to the vertices, and then find the vertex with the shortest path from the set Q.

[0070] The steps of initialization and execution process are repeated until all vertices have been visited, and traffic resources are transmitted through the shortest vertex found in set Q until they can be transmitted throughout the entire network topology.

[0071] like Figures 1-4 As shown, in the step S5, a fault recovery mechanism is established to quickly switch to a backup virtual link when a network failure occurs;

[0072] Specifically, in step S5, a fault recovery mechanism is established to quickly switch to a backup virtual link when a network fault occurs. By establishing BFD fast detection technology, the status of the network virtual link is monitored in real time. When a fault is detected in the primary virtual link of the network, the switching mechanism is triggered to quickly switch to the backup virtual link, thereby promptly responding to the occurrence of network faults.

[0073] In step S5, the fault recovery mechanism is specifically established as follows: manually adding the peer neighbor information through static configuration to create a static BFD session. When the peer interface also enables BFD and correctly responds to the BFD message of the local end, the static BFD configuration is completed. BFD control messages are periodically sent to each other to determine the session status. If multiple consecutive messages are not received, the virtual link has failed, triggering the switching mechanism. Through the configured network equipment, it automatically switches to the backup virtual link when the primary virtual link fails. If the automatic switching cannot meet the requirements, the administrator manually triggers the switching mechanism through an easy-to-use switching interface or tool to quickly switch to the backup virtual link when the primary virtual link fails.

[0074] In the S5 step, the configuration of the standby virtual link is specifically: the standby virtual link selects a shorter path and avoids network bottlenecks, and reserves sufficient network bandwidth resources to ensure that the standby virtual link can be started in time when the main virtual link fails, and the availability of the standby virtual link is tested regularly by sending test data packets and simulating network main virtual link failure, and an alarm mechanism is set to send an alarm notification to notify the administrator in time when the standby virtual link fails, and the alarm is sent through the mode including email and short message.

[0075] As shown in Figures 1-4 The present application solves the problem that multiple virtual links share one RDMA link resource in the prior art, which causes resource competition in the transmission process and performance degradation of the virtual link when the resource is requested at the same time, and works as follows:

[0076] A plurality of virtual networks are established in the RDMA area of the sending end and the receiving end by using virtualization technology, each virtual network is marked, and each virtual network corresponds to a group of virtual links;

[0077] Different virtual networks are identified by a virtual network identifier, and an isolation mechanism is used between the virtual networks, and each virtual network is allocated independent resources including bandwidth and buffer;

[0078] According to the task requirements of different virtual networks, the data on the RDMA link is encrypted, an access control policy is established, and access permission is limited;

[0079] An end-to-end flow control mechanism is established between the sending end and the receiving end, the routing is programmed by an SDN controller, and the resource allocation in the virtual network is dynamically adjusted;

[0080] A fault recovery mechanism is established, and the standby virtual link is quickly switched when the network fails.

[0081] Thus, the technical solution of the present application has been described in combination with the preferred embodiments shown in the drawings, but it is easy for those skilled in the art to understand that the protection scope of the present application is obviously not limited to these specific embodiments. Those skilled in the art can make equivalent changes or replacements to the related technical features without departing from the principles of the present application, and the technical solutions after the changes or replacements will fall within the protection scope of the present application.

[0082] The above description is only the preferred embodiments of the present application and is not used to limit the present application; for those skilled in the art, the present application can have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and rules of the present application shall be included in the protection scope of the present application.

Claims

1. A method for end-to-end large-scale RDMA network interconnection, characterized by: S1. Use virtualization technology to establish multiple virtual networks in the RDMA area of ​​the sending and receiving ends, mark each virtual network, and each virtual network corresponds to a group of virtual links; S2. Different virtual networks are identified by virtual network identifiers. An isolation mechanism is used between virtual networks, and each virtual network is allocated independent resources, including bandwidth and buffer. S3. Encrypt data on the RDMA link and establish access control policies to restrict access rights based on the task requirements of different virtual networks. According to the task requirements of different virtual networks, data on the RDMA link is encrypted, access control policies are established, and symmetric encryption algorithms are implemented according to the task requirements of different virtual networks. The data is converted into ciphertext using the encryption algorithm and encryption key. The encrypted data is transmitted in the VTEP and VXLAN tunnels using a unique identifier; The sender uses the AES symmetric encryption algorithm and key to encrypt the data. The key is permuted, replaced, and XORed. The encrypted data is then transmitted to the receiver through the VTEP and VXLAN tunnel using the unique identifier of the RDMA link. The receiving end uses the same AES symmetric encryption algorithm and key to decrypt the data, and the decrypted data is restored to the plaintext data; S4. Establish an end-to-end flow control mechanism between the sender and the receiver, program the routing through the SDN controller, and dynamically adjust the resource allocation in the virtual network; S5. Establish a fault recovery mechanism to quickly switch to the backup virtual link when a network failure occurs.

2. The end-to-end large-scale RDMA network interconnection method according to claim 1, characterized in that: The method uses virtualization technology to establish multiple virtual networks in the RDMA areas of the sending end and the receiving end, sends a connection request from the sending end to the receiving end, establishes a virtual network on the RDMA link, configures the RDMA context of the virtual network, including configuring an independent queue QP and an independent completion queue CP, uses the RDMA protocol RoCE in the RDMA areas of the sending end and the receiving end to connect each independent queue QP to the queue QP corresponding to the remote node, and when establishing the virtual network on the RDMA link, assigns a unique identifier VNI tag to mark each virtual network, each virtual network corresponds to a group of virtual links, and the group of virtual links includes a configured primary virtual link and a backup virtual link, and configures VTEP for the virtual network and establishes a VXLAN tunnel. Data packets are transmitted in the VTEP and the VXLAN tunnel using the unique identifier.

3. The end-to-end large-scale RDMA network interconnection method according to claim 1, characterized in that: Different virtual networks are identified by virtual network identifiers, and an isolation mechanism is adopted between virtual networks. Different subnets are identified according to the unique identifier VNI tag of each virtual network. The isolation mechanism selects a bridge isolation mechanism, and a virtual bridge is created on the physical machine. When creating a virtual machine, the network interface of the virtual machine is configured to connect to the previously created virtual bridge, and an independent IP address is configured for each virtual machine in the same subnet. Multiple independent virtual bridges are created, and different virtual machines are connected to different virtual bridges. Independent broadband and buffer resources are allocated according to the broadband requirements and buffer usage of each virtual network.

4. The end-to-end large-scale RDMA network interconnection method according to claim 1, characterized in that: The end-to-end flow control mechanism is established between the sending end and the receiving end, the routing is programmed through the SDN controller, and the flow resources are allocated by the SDN controller according to the data transmission rate, delay requirements, and reliability requirements of the end-to-end flow control.

5. The end-to-end large-scale RDMA network interconnection method according to claim 4, characterized in that: The specific steps of allocating traffic resources by using the SDN controller include: S401: Configure the SDN controller to collect network status information. The SDN controller interacts with network device routers through the southbound interface to integrate load information on the link. S402. Traffic identification: SDN performs fine-grained identification and classification of traffic in the virtual network. Traffic types include interactive traffic and bulk traffic. Interactive traffic is small data packets that are frequently transmitted, while bulk traffic is large data packets that are transmitted at long intervals. Based on the programming capabilities of the SDN controller, flow tables are deployed on the network to identify and mark different types of traffic. S403, traffic control, based on traffic identification, obtains traffic type and priority for traffic classification and scheduling. Traffic classification uses port identification, traffic scheduling is based on big data analysis, and traffic priority is divided into standard priority and high priority. SDN uses a centralized controller and programming capabilities to implement traffic control strategies; S404, traffic optimization, the SDN controller dynamically adjusts the network topology according to the traffic load and delay by optimizing the network structure and traffic path.

6. The end-to-end large-scale RDMA network interconnection method according to claim 1, characterized in that: The fault recovery mechanism is established to quickly switch to the backup virtual link when a network failure occurs. By establishing BFD fast detection technology, the status of the network virtual link is monitored in real time. When the main virtual link of the network is detected to fail, the switching mechanism is triggered to quickly switch to the backup virtual link.

7. The end-to-end large-scale RDMA network interconnection method according to claim 6, characterized in that: The fault recovery mechanism is established by manually adding the neighbor information of the other end through static configuration to create a static BFD session. When the interface of the other end also enables BFD and correctly responds to the BFD message of the local end, the static BFD configuration is completed. BFD control messages are periodically sent to each other to determine the session status. If multiple consecutive messages are not received, the virtual link has failed, triggering the switching mechanism. Through the configured network equipment, it automatically switches to the backup virtual link when the primary virtual link fails. When the automatic switching cannot meet the requirements, the administrator manually triggers the switching mechanism through an easy-to-operate switching interface or tool to quickly switch to the backup virtual link when the primary virtual link fails.

Citation Information

Patent Citations

  • RDMA-based connection establishment method and device, equipment and storage medium

    CN115623057A

  • Network system, communication method, network node, and storage medium

    CN116566763A

  • Communication method, computer device, storage medium and program product

    CN118714106A