Port-binding disaster recovery method and apparatus for storage system, and device and non-volatile readable storage medium
By binding ports in RoCE-SAN and RoCE multi-controlled storage systems to form communication paths and group management, the problem of service imbalance caused by network link independence is solved, smooth path switching and load balancing are achieved, and the stability and performance of the system are improved.
Patent Information
- Application Number
- PCT/CN2024/122463
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-09-29
- Publication Date
- 2025-07-03
AI Technical Summary
In RoCE-SAN and RoCE multi-controlled storage systems, the independent network links lead to insufficient mutual backup redundancy, unable to achieve business balance, and prone to single-channel congestion, affecting service performance and stability, and limiting disaster recovery capabilities.
By binding the port of the first device to the port of the second device to form several communication paths, the performance of each path is calculated, and grouped into available groups and spare groups, and the packet is adjusted in response to path exceptions, so as to achieve smooth path switching and load balancing.
It improves the stability and disaster recovery capabilities of communication services, switches to use alternate paths when link congestion, ensures the smoothness and stability of customer services, reduces the risk of path failures, and improves storage performance.
Smart Images

Figure CN2024122463_03072025_PF_FP_ABST
Abstract
Description
A method, device, equipment and non-volatile readable storage medium for storage system port binding disaster recovery
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on December 28, 2023, with application number 202311829151.3, entitled “A method, apparatus, device and medium for port binding disaster recovery of a storage system”, the entire contents of which are incorporated by reference into this application. Technical Field
[0003] The present application relates to the field of computers, and more specifically to a method, apparatus, device, and non-volatile readable storage medium for port binding disaster recovery in a storage system. Background Art
[0004] SAN (storage area network) networks based on RoCE (RDMA over Converged Ethernet) technology, which can transfer data from one server to another or from storage to a server with minimal CPU usage, have become a trend in storage systems. RoCE-SAN networks must ensure interoperability between storage-side network ports and server-side network ports. A link failure inevitably leads to failure of the network link between storage and server. Because RoCE-SAN links are independent of each other, they lack the redundancy needed to improve customer service stability and balance services. This can lead to uneven network resource allocation and single-path congestion, impacting service performance and the customer experience. Link failures directly threaten customer service stability.
[0005] RoCE multi-controller (interconnected) storage is a comprehensive storage system that uses RoCE networking to achieve network connectivity between polymorphic storage and to build storage clusters. RoCE multi-controller storage also has independent paths between storage clusters, which can lead to service imbalances, single-path congestion, and lease expiration. This significantly limits the disaster recovery capabilities of multi-controller storage and further restricts the stability of the storage system.
[0006] Summary of the Invention
[0007] In view of this, the purpose of the embodiments of the present application is to propose a method, device, equipment and non-volatile readable storage medium for storage system port binding disaster recovery. By using the technical solution of the present application, the stability and disaster recovery capability of communication services can be improved in scenarios where multiple paths fail, and in scenarios where link congestion and performance pressure are high, the backup path can be switched to share the business pressure. The storage performance can be improved at the port level to ensure the smoothness and stability of customer services, and the smooth switching of communication paths can be achieved. Risk warning reminders can be issued to reduce the risk of path failure or loss.
[0008] Based on the above objectives, one aspect of an embodiment of the present application provides a method for port binding disaster recovery in a storage system, comprising the following steps:
[0009] Binding the port of the first device to the port of the second device to form a plurality of communication paths, and calculating the performance of each communication path;
[0010] grouping the communication paths based on performance of the communication paths and performing data transmission based on the grouping;
[0011] In response to an abnormality occurring in a communication path, grouping of the communication paths is adjusted based on a cause of the abnormality and the number of communication paths in each group.
[0012] According to one embodiment of the present application, the step of grouping the communication paths based on the performance of the communication paths and performing data transmission based on the grouping includes:
[0013] Sort the communication paths by performance from high to low;
[0014] selecting a threshold number of communication paths ranked top in performance as an available group for data transmission;
[0015] Use the remaining communication paths as a backup group.
[0016] According to one embodiment of the present application, selecting a threshold number of communication paths ranked higher in performance as an available group for data transmission includes:
[0017] determining the communication paths ranked in top 50% of performance as the available group;
[0018] The step of using the remaining communication paths as a backup group includes:
[0019] The communication paths ranked in the bottom 50% of performance are determined as the backup group.
[0020] According to one embodiment of the present application, determining the communication paths ranked in the top 50% of performance as the available group includes:
[0021] Setting the communication paths ranked in the top 50% of performance to an active state, wherein the communication paths in the active state are used for data transmission;
[0022] The step of determining the communication paths ranked in the bottom 50% of performance as the backup group includes:
[0023] The communication paths ranked in the bottom 50% of the performance ranking are set to an enabled state, wherein the communication paths in the enabled state are used to take over the abnormal communication paths in the available group for data transmission when an abnormality occurs to the communication paths in the available group.
[0024] According to one embodiment of the present application, in response to an abnormality occurring in a communication path, the step of adjusting the grouping of the communication paths based on the abnormality cause and the number of communication paths in each group includes:
[0025] In response to an abnormality occurring in the communication path, determining a cause of the abnormality occurring in the communication path;
[0026] In response to the abnormality occurring in the communication path being caused by path disconnection, executing a first preset strategy;
[0027] In response to the abnormality occurring in the communication path being caused by path congestion, a second preset strategy is executed.
[0028] According to one embodiment of the present application, in response to the abnormality of the communication path being caused by path disconnection, the step of executing the first preset strategy includes:
[0029] In response to the abnormality of the communication path being caused by path disconnection, marking the disconnected communication path as an unavailable path, and determining a group of the disconnected communication path;
[0030] In response to the disconnected communication paths being grouped into a standby group, acquiring the number of communication paths in the standby group, and performing a first preset operation based on the acquired number;
[0031] In response to the disconnected communication paths being grouped into an available group, the number of communication paths of the standby group is acquired, and a second preset operation is performed based on the acquired number.
[0032] According to one embodiment of the present application, in response to the disconnected communication paths being grouped into a standby group, obtaining the number of communication paths in the standby group, and performing the first preset operation based on the obtained number includes:
[0033] In response to the disconnected communication paths being grouped into a standby group, obtaining the number of communication paths of the standby group;
[0034] In response to the number of communication paths of the backup group being less than one, sending a warning of no redundant path;
[0035] In response to the number of communication paths of the standby group being equal to 1, the warning of no redundant paths is cleared, and a warning of insufficient redundant paths is sent.
[0036] According to one embodiment of the present application, in response to the disconnected communication paths being grouped into an available group, obtaining the number of communication paths in the standby group, and performing the second preset operation based on the obtained number includes:
[0037] In response to the disconnected communication paths being grouped into an available group, obtaining the number of communication paths of the standby group;
[0038] In response to the number of communication paths of the standby group being equal to 1, switching the communication path of the standby group to the available group, setting the communication path to an available state, and sending a warning of no redundant path;
[0039] In response to the number of communication paths in the standby group being greater than 1, calculating performance of all communication paths in the standby group;
[0040] Switch the communication path with the highest performance to the available group and set it to the available state;
[0041] In response to the number of communication paths of the backup group being less than one, a warning of no redundant paths is sent.
[0042] According to one embodiment of the present application, in response to the abnormality of the communication path being caused by path congestion, the step of executing the second preset strategy includes:
[0043] In response to the abnormality occurring in the communication path being caused by path congestion, calculating the performance of the congested communication path and all communication paths in the backup group;
[0044] In response to the communication path with the highest performance being the communication path where congestion occurs, no processing is performed;
[0045] In response to the communication path with the highest performance being a communication path in the standby group, the communication path with the highest performance in the standby group is switched to the available group and set to an available state.
[0046] According to one embodiment of the present application, the step of binding a port of the first device to a port of the second device to form a communication path includes:
[0047] Binding the RoCE port of the first device to the RoCE port of the second device to form a communication path;
[0048] Create a virtual network card for the RoCE port of the first device and a virtual network card for the RoCE port of the second device respectively;
[0049] A determination is made as to whether the first device is able to communicate with the second device.
[0050] According to one embodiment of the present application, the step of binding the RoCE port of the first device to the RoCE port of the second device to form a communication path includes:
[0051] A one-to-one connection is established between each RoCE port of the first device and each RoCE port of the second device to form a plurality of communication paths.
[0052] According to one embodiment of the present application, the step of determining whether the first device can communicate with the second device includes:
[0053] Configure IP addresses for the first virtual network card corresponding to the RoCE port of the first device and the second virtual network card corresponding to the RoCE port of the second device respectively;
[0054] Determining whether the network of the first device is connectable to the network of the second device;
[0055] In response to the network of the first device being able to communicate with the network of the second device, information of the first device is sent to the second device, so that the second device is configured according to the information of the first device.
[0056] According to one embodiment of the present application, the first device is a host, the second device is a storage node, a RoCE port of the host and a RoCE port of the storage node are connected via an optical fiber, and before configuring an IP address for a first virtual network card corresponding to the RoCE port of the first device and a second virtual network card corresponding to the RoCE port of the second device, the method further includes:
[0057] Configuring BOND port binding through the RoCE port of the host to generate a first virtual network card of the host;
[0058] The BOND port binding is configured through the RoCE port of the storage node to generate a second virtual network card of the storage node.
[0059] According to one embodiment of the present application, the first device is a host, and the second device is a storage node. The step of sending information of the first device to the second device so that the second device is configured according to the information of the first device includes:
[0060] The storage node uses the host's information to configure host management;
[0061] The host uses the first command to discover the storage node, and uses the second command to connect to the storage node to complete the configuration of the host and the storage node.
[0062] According to one embodiment of the present application, the storage node uses the host information to configure host management, including:
[0063] The storage node obtains the unique identifier of the host and configures host management using the unique identifier of the host.
[0064] According to one embodiment of the present application, the host uses a first command to discover a storage node, and uses a second command to connect to the storage node to complete the configuration of the host and the storage node, including:
[0065] The host uses an nvme discover command to discover the storage node and uses an nvme connect command to connect to the storage node, wherein the first command includes the nvme discover command and the second command includes the nvme connect command.
[0066] According to one embodiment of the present application, the first device is a first storage node, the second device is a second storage node, and the step of determining whether the first device can communicate with the second device includes:
[0067] Configure IP addresses for the first virtual network card corresponding to the RoCE port of the first storage node and the second virtual network card corresponding to the RoCE port of the second storage node respectively;
[0068] Determining whether the network of the first storage node is connectable to the network of the second storage node;
[0069] In response to the network of the first storage node being able to communicate with the network of the second storage node, a cluster is created based on the first storage node and the second storage node.
[0070] Another aspect of the embodiments of the present application further provides a port binding disaster recovery device, the device comprising:
[0071] a binding module configured to bind the port of the first device to the port of the second device to form a plurality of communication paths, and calculate the performance of each communication path;
[0072] a grouping module configured to group the communication paths based on performance of the communication paths and perform data transmission based on the groups;
[0073] The adjustment module is configured to adjust the grouping of the communication paths based on a cause of the abnormality and the number of communication paths in each group in response to an abnormality occurring in the communication paths.
[0074] Another aspect of the embodiments of the present application further provides a computer device, the computer device comprising:
[0075] at least one processor; and
[0076] The memory stores computer instructions that can be run on the processor, and when the instructions are executed by the processor, the steps of any of the above methods are implemented.
[0077] Another aspect of the embodiments of the present application further provides a computer non-volatile readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of any of the above methods are implemented.
[0078] The present application has the following beneficial technical effects: the method for storage system port binding disaster recovery provided by the embodiment of the present application, by respectively binding the port of the first device with the port of the second device to form several communication paths, and calculating the performance of each communication path; grouping the communication paths based on the performance of the communication paths and transmitting data based on the groups; in response to an abnormality in the communication path, adjusting the grouping of the communication paths based on the cause of the abnormality and the number of communication paths in each group. The technical solution can improve the stability and disaster recovery capability of the communication business in the scenario of multiple path failures, can switch to the use of backup paths to share the business pressure in the scenario of link congestion and high performance pressure, can improve the performance of storage at the port level, ensure the fluency and stability of customer business, can realize smooth switching of communication paths, and provide risk warning reminders to reduce the risk of path failure or loss. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can be obtained based on these drawings without paying any creative work.
[0080] FIG1 is a schematic flow chart of a method for port binding disaster recovery according to an embodiment of the present application;
[0081] FIG2 is a schematic diagram of binding a host to a storage node port according to an embodiment of the present application;
[0082] FIG3 is a schematic diagram of a RoCE-SAN scenario deployment according to an embodiment of the present application;
[0083] FIG4 is a schematic diagram of binding a storage node to a storage node port according to an embodiment of the present application;
[0084] FIG5 is a schematic diagram of a RoCE multi-controller interconnection scenario deployment according to an embodiment of the present application;
[0085] FIG6 is a schematic diagram of communication path grouping according to one embodiment of the present application;
[0086] FIG7 is a schematic diagram of monitoring a communication path of a standby group according to one embodiment of the present application;
[0087] FIG8 is a schematic diagram of monitoring available group communication paths according to one embodiment of the present application;
[0088] FIG9 is a schematic diagram of a port binding disaster recovery device according to an embodiment of the present application;
[0089] FIG10 is a schematic diagram of a computer device according to an embodiment of the present application;
[0090] FIG11 is a schematic diagram of a computer non-volatile readable storage medium according to an embodiment of the present application. DETAILED DESCRIPTION
[0091] In order to make the objectives, technical solutions and advantages of the present application more clear, the embodiments of the present application are further described in detail below in combination with the embodiments and with reference to the accompanying drawings.
[0092] Based on the above objectives, the first aspect of the embodiments of the present application provides an embodiment of a method for port binding disaster recovery in a storage system. FIG1 shows a schematic flow chart of the method.
[0093] As shown in FIG1 , the method may include the following steps:
[0094] S1 binds the port of the first device and the port of the second device respectively to form a plurality of communication paths, and calculates the performance of each communication path. In one embodiment, in a RoCE-SAN scenario, as shown in FIG2 , the first device may be a host, and the second device may be a storage node. The RoCE ports of the storage node and the host may be connected via optical fiber lines. For example, the dual-control storage node 1 has four RoCE ports and the host has four RoCE ports. The four RoCE ports of the two devices are connected respectively to form four communication paths. The four RoCE ports of the storage node are then selected to configure BOND port binding to generate the virtual network card Seth0 of the storage node. The four RoCE ports of the host are selected to configure BOND port binding to generate the virtual network card Seth0 of the host. As shown in FIG3 , in several application scenarios, an IP is then configured for the virtual network card Seth0 of the storage node and the virtual network card Seth0 of the host. The IPs of the storage node and the host need to be in the same network segment. The host pings the IP of the storage node. If the ping fails, check whether the physical link or the configured IP is correct. If the ping succeeds, obtain the unique identifier of the host. The storage node uses the unique identifier of the host to configure host management. Finally, the host discovers the storage node using the nvme discover command. After the connect command is used to connect the storage node, the deployment is completed. In another embodiment, in the RoCE multi-controller interconnection scenario, as shown in FIG4 , the first device can be a storage node, and the second device can be another storage node. For example, the two storage nodes have a total of four nodes, namely Node1-Node4, and each node has 4 RoCE ports. The RoCE ports of the four nodes are connected to the same RoCE switch through optical fiber lines. The four RoCE ports of each node are configured with BOND port binding to generate the node's virtual network card Seth0. The virtual network card Seth0 of each node is configured with an IP, and the IP needs to be in the same network segment, as shown in FIG5 . Then, each node needs to ping each other's IP. If the IP cannot be pinged, the IP configuration of the virtual network card or the physical link is checked to see if it is correct. If the IP can be pinged, a cluster is created and then managed by other nodes. In the above two scenarios, several communication links will be formed, and the performance of each communication link needs to be calculated.
[0095] S2 groups communication paths based on their performance and performs data transmission based on the groups. In several application scenarios, the performance of communication paths can be ranked from high to low, and a threshold number of communication paths ranked at the top of the performance ranking are selected as the available group for data transmission, and the remaining communication paths are used as the standby group. As shown in Figure 6, for example, the top 50% of paths in terms of performance are set as the active group (available group), assigned the active state, responsible for IO transmission, and the bottom 50% of paths in terms of performance are set as the enabled group (standby group), assigned the enabled state, as standby paths. When an abnormality occurs in a communication path in the active group, the communication path in the enabled group can take over the IO transmission work of the abnormal path. This can improve the stability and disaster recovery capabilities of communication services under multiple path failures, and can switch to using standby paths to share business pressure in scenarios with high link congestion and performance pressure.
[0096] S3 adjusts the grouping of the communication paths based on the cause of the abnormality and the number of communication paths in each group in response to an abnormality occurring in the communication paths. In several application scenarios, if the cause of the abnormality occurring in the communication paths is path disconnection, the grouping of the disconnected communication paths is determined. If the grouping of the disconnected communication paths is a standby group, the number of communication paths of the standby group is obtained. If the number of communication paths of the standby group is less than 1, a warning of no redundant paths is sent. If the number of communication paths of the standby group is equal to 1, the warning of no redundant paths is cleared, a warning of insufficient redundant paths is sent, and the disconnected communication paths are marked as unavailable paths. If the grouping of the disconnected communication paths is an available group, the number of communication paths of the standby group is obtained. If the number of communication paths of the standby group is equal to 1, the communication paths of the standby group are switched to the available group and set to an available state, and a warning of no redundant paths is sent. If the number of communication paths of the standby group is greater than 1, the performance of all communication paths in the standby group is calculated, the communication path with the highest performance is switched to the available group and set to an available state. If the number of communication paths of the standby group is less than 1, a warning of no redundant paths is sent. If the abnormality in the communication path is caused by path congestion, the performance of the congested communication path and all communication paths in the standby group are calculated. If the communication path with the highest performance is the congested communication path, no processing is performed. If the communication path with the highest performance is a communication path in the standby group, the communication path with the highest performance in the standby group is switched to the available group and set to the available state.
[0097] By using the technical solution of the present application, the stability and disaster recovery capability of communication services can be improved under multiple path failures. In scenarios where link congestion and performance pressure are high, the backup path can be switched to share the service pressure. The storage performance can be improved at the port level to ensure the smoothness and stability of customer services. The smooth switching of communication paths can be achieved, and risk warning reminders can be issued to reduce the risk of path failure or loss.
[0098] In an optional embodiment of the present application, the step of grouping the communication paths based on the performance of the communication paths and performing data transmission based on the groups includes:
[0099] Sort the communication paths by performance from high to low;
[0100] selecting a threshold number of communication paths ranked top in performance as an available group for data transmission;
[0101] The remaining communication paths are used as backup groups. After sorting by performance, a certain number of communication paths can be selected as available groups for data transmission. In some embodiments, the top 50% of paths in terms of performance are set as the active group (available group), assigned the active state, and are responsible for IO transmission. The bottom 50% of paths in terms of performance are set as the enabled group (backup group), assigned the enabled state, and serve as backup paths. When an abnormality occurs in a communication path in the active group, the communication path in the enabled group can take over the IO transmission work of the abnormal path. This can improve the stability and disaster recovery capabilities of communication services under multiple path failures, and can switch to using backup paths to share business pressure in scenarios with high link congestion and performance pressure.
[0102] In an optional embodiment of the present application, in response to an abnormality occurring in a communication path, the step of adjusting the grouping of the communication paths based on the abnormality cause and the number of communication paths in each group includes:
[0103] In response to an abnormality occurring in the communication path, determining a cause of the abnormality occurring in the communication path;
[0104] In response to the abnormality occurring in the communication path being caused by path disconnection, executing a first preset strategy;
[0105] In response to a communication path abnormality caused by path congestion, a second preset strategy is executed. Communication path abnormalities are typically caused by path disconnection and path congestion. A path disconnection occurs when the communication path is unable to transmit data due to some reason, while a path congestion occurs when the communication path can transmit data, but at a slower speed than normal. During use, the status of each communication link can be detected at regular intervals. This application sets different strategies for these two abnormality causes.
[0106] In an optional embodiment of the present application, in response to the abnormality of the communication path being caused by path disconnection, the step of executing the first preset strategy includes:
[0107] In response to the abnormality occurring in the communication path being caused by path disconnection, determining a group of the disconnected communication path;
[0108] In response to the disconnected communication paths being grouped into a standby group, acquiring the number of communication paths in the standby group, and performing a first preset operation based on the acquired number;
[0109] In response to the disconnected communication paths being grouped into an available group, the number of communication paths of the standby group is acquired, and a second preset operation is performed based on the acquired number.
[0110] In an optional embodiment of the present application, as shown in FIG7 , in response to the disconnected communication paths being grouped into a standby group, obtaining the number of communication paths in the standby group, and performing the first preset operation based on the obtained number includes:
[0111] In response to the disconnected communication paths being grouped into a standby group, obtaining the number of communication paths of the standby group;
[0112] In response to the number of communication paths in the standby group being less than 1, a warning indicating a lack of redundant paths is issued. If the communication path abnormality is caused by a path disconnection, and the disconnected communication path is in the standby group, the number of communication paths in the current standby group is obtained. If the number is less than 1, indicating that there are no communication paths available for switching, a warning indicating a lack of redundant paths is issued.
[0113] In an optional embodiment of the present application, it further includes:
[0114] In response to the number of communication paths in the standby group being equal to one, the warning regarding the lack of redundant paths is cleared and a warning regarding insufficient redundant paths is issued. If the number of communication paths in the standby group is equal to one, this indicates that one communication path is available for switching when needed. If a warning regarding the lack of redundant paths was previously issued, the warning is cleared and a warning regarding insufficient redundant paths is issued to notify the administrator to add redundant paths.
[0115] In an optional embodiment of the present application, it further includes:
[0116] Mark the disconnected communication path as an unavailable path. Regardless of whether the disconnected communication path is in the available group or the standby group, it needs to be marked as an unavailable path and a corresponding warning will be issued to notify the administrator to check the status of the path.
[0117] In an optional embodiment of the present application, as shown in FIG8 , in response to the disconnected communication paths being grouped into an available group, obtaining the number of communication paths in the standby group, and performing the second preset operation based on the obtained number includes:
[0118] In response to the disconnected communication paths being grouped into an available group, obtaining the number of communication paths of the standby group;
[0119] In response to the number of communication paths of the standby group being equal to 1, switching the communication path of the standby group to the available group and setting it to an available state;
[0120] A warning about a lack of redundant paths is issued. If the disconnected path is in the available group, the system first obtains the number of communication paths in the standby group. If the number of communication paths in the standby group is one, the communication path is switched to the available group and set to available status to replace the disconnected communication path. A warning about a lack of redundant paths is also issued, notifying the administrator to add a redundant path.
[0121] In an optional embodiment of the present application, it further includes:
[0122] In response to the number of communication paths in the standby group being greater than 1, calculating performance of all communication paths in the standby group;
[0123] The communication path with the highest performance is switched to the available group and set to the available state. If the number of communication paths in the standby group is greater than one, the performance of all communication paths in the standby group is recalculated, and the communication path with the highest performance is switched to the available group and set to the available state to replace the disconnected communication path.
[0124] In an optional embodiment of the present application, it further includes:
[0125] In response to the number of communication paths in the backup group being less than one, a warning indicating a lack of redundant paths is issued. If the number of communication paths in the backup group is less than one, this indicates that there are no alternative communication paths available for switching, and a warning indicating a lack of redundant paths is issued. The above steps implement multipath management for bound ports, monitor path anomalies, and report abnormality alerts, facilitating smooth path switching, providing risk warnings, and reducing the risk of path failure or loss.
[0126] In an optional embodiment of the present application, in response to the abnormality of the communication path being caused by path congestion, the step of executing the second preset strategy includes:
[0127] In response to the abnormality occurring in the communication path being caused by path congestion, calculating the performance of the congested communication path and all communication paths in the backup group;
[0128] In response to the communication path with the highest performance being the communication path where congestion occurs, no processing is performed;
[0129] In response to the communication path with the highest performance being a communication path in the standby group, the communication path with the highest performance in the standby group is switched to the available group and set to an available state. If the abnormality in the communication path is caused by path congestion, the performance of the congested communication path and all communication paths in the standby group is calculated. If the communication path with the highest performance is the congested communication path, no action is taken. If the communication path with the highest performance is a communication path in the standby group, the communication path with the highest performance in the standby group is switched to the available group and set to an available state to replace the disconnected communication path, and the communication path that is experiencing congestion is simultaneously added to the standby group.
[0130] In an optional embodiment of the present application, the step of binding a port of the first device to a port of the second device to form a communication path includes:
[0131] Binding the RoCE port of the first device to the RoCE port of the second device to form a communication path;
[0132] Create a virtual network card for the RoCE port of the first device and a virtual network card for the RoCE port of the second device respectively;
[0133] A determination is made as to whether the first device is able to communicate with the second device.
[0134] In an optional embodiment of the present application, the step of binding a RoCE port of the first device to a RoCE port of the second device to form a communication path includes:
[0135] Establish a one-to-one connection between each RoCE port of the first device and each RoCE port of the second device to form multiple communication paths. For example, the first port of the first device is connected to the first port of the second device, the second port of the first device is connected to the second port of the second device, and so on.
[0136] In an optional embodiment of the present application, the step of determining whether the first device can communicate with the second device includes:
[0137] Configure IP addresses for the first virtual network card corresponding to the RoCE port of the first device and the second virtual network card corresponding to the RoCE port of the second device respectively;
[0138] Ping the IP address of the second device via the first device;
[0139] In response to the first device being able to ping the second device's IP address, the first device's information is sent to the second device, causing the second device to configure itself based on the first device's information. The ID configured for the virtual network card must be on the same network segment. If the ping fails, check whether the physical link or configured IP address is correct.
[0140] In an optional embodiment of the present application, the first device includes a host, and the second device includes a storage node.
[0141] In an optional embodiment of the present application, the step of sending information of the first device to the second device so that the second device is configured according to the information of the first device includes:
[0142] The storage node uses the host's information to configure host management;
[0143] The host uses the first command to discover the storage node, and uses the second command to connect to the storage node to complete the configuration of the host and the storage node. In one embodiment, in a RoCE-SAN scenario, the first device may be a host, and the second device may be a storage node. The RoCE ports of the storage node and the host may be connected via optical fiber. For example, if the dual-controller storage node 1 has four RoCE ports and the host has four RoCE ports, the four RoCE ports of the two devices are connected respectively to form four communication paths. Then, the four RoCE ports of the storage node are selected to configure BOND port binding to generate the virtual network card Seth0 of the storage node. The four RoCE ports of the host are selected to configure BOND port binding to generate the virtual network card Seth1 of the host. Then, an IP is configured for the virtual network card Seth0 of the storage node, and an IP is configured for the virtual network card Seth1 of the host. The IPs of the storage node and the host need to be in the same network segment. The host pings the IP of the storage node. If the ping fails, check whether the physical link or the configured IP is correct. If the ping succeeds, obtain the unique identifier of the host. The storage node uses the unique identifier of the host to configure host management. Finally, the host discovers the storage node using the nvme discover command, and connects to the storage node using the nvme connect command to complete the deployment.
[0144] In an optional embodiment of the present application, the first device is a first storage node, and the second device is a second storage node.
[0145] In an optional embodiment of the present application, the step of determining whether the first device can communicate with the second device includes:
[0146] Configure IP addresses for the first virtual network card corresponding to the RoCE port of the first storage node and the second virtual network card corresponding to the RoCE port of the second storage node respectively;
[0147] Ping the IP address of the second storage node via the first storage node;
[0148] In response to the first storage node being able to ping the IP of the second storage node, a cluster is created based on the first storage node and the second storage node. In another embodiment, in a RoCE multi-controller interconnection scenario, the first device can be a storage node, and the second device can be another storage node. For example, the two storage nodes have a total of four nodes, namely Node1-Node4, and each node has 4 RoCE ports. The RoCE ports of the four nodes are connected to the same RoCE switch through optical fiber lines. The four RoCE ports of each node are configured with BOND port binding to generate the node's virtual network card Seth0. The virtual network card Seth0 of each node is configured with an IP, and the IP needs to be in the same network segment. Then, each node needs to ping each other's IP. If the IP cannot be pinged, the IP configuration of the virtual network card or the physical link is checked to see if it is correct. If the IP can be pinged, a cluster is created and then managed by other nodes. In the above two scenarios, several communication links will be formed, and the performance of each communication link needs to be calculated.
[0149] In other embodiments, the first device may be multiple hosts, the second device may be multiple storage nodes, the first device may be multiple storage nodes, and the second device may also be multiple storage nodes. That is, there is no limit on the number of devices in the first device and the second device.
[0150] This application implements a port binding mode based on the RoCE network port, combining four or more ports into one virtual network port, designing multi-path management, and providing high-performance and stable services to the outside world. It has the following effects:
[0151] Effect 1: For RoCE-SAN services based on port binding, the active and enabled group paths implement mutual backup. In the event of multiple path failures, the backup path takes over, improving the stability and disaster recovery capabilities of the RoCE-SAN service.
[0152] Effect 2: For RoCE-SAN services based on port binding, active and enabled group paths achieve load balancing. In scenarios where link congestion and performance pressure are high, enabled paths are switched to share the service pressure. This improves storage performance at the port level by at least two times, ensuring the smoothness and stability of customer services.
[0153] Effect three: It realizes multi-path management of bound ports, monitors path anomalies and reports abnormal alarms, assists in smooth path switching, provides risk warnings, and reduces the risk of path failure or loss.
[0154] Effect 4: RoCE cluster interconnection based on port binding can virtualize multiple physical ports into a single logical port. The path specifications of cluster interconnection are no longer dependent on software specifications, and cluster paths can be expanded to four times, eight times, or even more than the original, with no theoretical upper limit.
[0155] It should be noted that those skilled in the art will understand that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program. The above-mentioned program can be stored in a non-volatile computer readable storage medium. When the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The non-volatile readable storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM). The embodiment of the above-mentioned computer program can achieve the same or similar effects as any of the corresponding aforementioned method embodiments.
[0156] In addition, the method disclosed in the embodiment of the present application can also be implemented as a computer program executed by a CPU, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed by the CPU, the above-mentioned functions defined in the method disclosed in the embodiment of the present application are performed.
[0157] Based on the above objectives, a second aspect of an embodiment of the present application provides a port binding disaster recovery device. As shown in FIG9 , the device 200 includes:
[0158] a binding module configured to bind the port of the first device to the port of the second device to form a plurality of communication paths, and calculate the performance of each communication path;
[0159] a grouping module configured to group the communication paths based on performance of the communication paths and perform data transmission based on the groups;
[0160] The adjustment module is configured to adjust the grouping of the communication paths based on a cause of the abnormality and the number of communication paths in each group in response to an abnormality occurring in the communication paths.
[0161] Based on the above objectives, the third aspect of the embodiments of the present application provides a computer device. Figure 10 shows a schematic diagram of an embodiment of the computer device provided by the present application. As shown in Figure 10, the embodiment of the present application includes the following apparatus: at least one processor 21; and a memory 22, the memory 22 storing computer instructions 23 executable on the processor, which, when executed by the processor, implement any of the above methods.
[0162] Based on the above objectives, a fourth aspect of the embodiments of the present application provides a non-volatile computer-readable storage medium. FIG11 is a schematic diagram of an embodiment of the non-volatile computer-readable storage medium provided by the present application. As shown in FIG11 , the non-volatile computer-readable storage medium 31 stores a computer program 32 that, when executed by a processor, performs any of the above methods.
[0163] In addition, the method disclosed in the embodiments of the present application may also be implemented as a computer program executed by a processor, which may be stored in a non-volatile computer-readable storage medium. When the computer program is executed by the processor, the above-mentioned functions defined in the method disclosed in the embodiments of the present application are performed.
[0164] In addition, the above method steps and system units can also be implemented using a controller and a computer non-volatile readable storage medium for storing a computer program that enables the controller to implement the above steps or unit functions.
[0165] Those skilled in the art will also appreciate that the various exemplary logic blocks, modules, circuits and algorithmic steps described in conjunction with the disclosure herein can be implemented as electronic hardware, computer software or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, a general description has been given of the functions of various schematic components, blocks, modules, circuits and steps. Whether this function is implemented as software or as hardware depends on the application and the design constraints imposed on the entire system. Those skilled in the art can implement the function in a variety of ways for every application, but this implementation decision should not be interpreted as causing a departure from the disclosed scope of the present application's embodiments.
[0166] In one or more exemplary designs, the function can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the function can be stored as one or more instructions or codes on a computer non-volatile readable storage medium or transmitted via a computer non-volatile readable storage medium. Computer non-volatile readable storage media include computer storage media and communication media, and the communication media include any non-volatile readable storage medium that helps to transfer a computer program from one location to another. The storage medium can be any available non-volatile readable storage medium that can be accessed by a general or special-purpose computer. As an example and not limitation, the computer non-volatile readable storage medium can include RAM, ROM, EEPROM (Electrically Erasable Programmable Read-Only Memory), CD-ROM (Compact Disc-Read Only Memory) or other optical disk storage devices, magnetic disk storage devices or other magnetic storage devices, or any other non-volatile readable storage medium that can be configured to carry or store the required program code in the form of instructions or data structures and can be accessed by a general or special-purpose computer or a general or special-purpose processor. In addition, any connection may be appropriately referred to as a computer non-volatile readable storage medium. For example, if a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves are used to transmit software from a website, server, or other remote source, then the above coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are all included in the definition of non-volatile readable storage medium. As used herein, disks and optical disks include compact disks (CDs), laser disks, optical disks, digital versatile disks (DVDs), floppy disks, and Blu-ray disks, where disks typically reproduce data magnetically, while optical disks reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer non-volatile readable storage media.
[0167] The above are exemplary embodiments disclosed in the present application, but it should be noted that various changes and modifications may be made without departing from the scope of the present application as defined in the claims. The functions, steps and / or actions of the method claims according to the disclosed embodiments described herein do not need to be performed in any particular order. In addition, although the elements disclosed in the embodiments of the present application may be described or required in individual form, they may also be understood as multiple unless expressly limited to the singular.
[0168] It should be understood that, as used herein, the singular forms "a" and "an" are intended to include the plural forms as well, unless the context clearly supports an exception. It should also be understood that, as used herein, "and / or" is intended to include any and all possible combinations of one or more of the associated listed items.
[0169] The serial numbers of the embodiments disclosed in the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0170] Those skilled in the art will understand that all or part of the steps for implementing the above embodiments may be accomplished by hardware, or may be accomplished by programs instructing related hardware, and the programs may be stored in a computer non-volatile readable storage medium, and the above-mentioned non-volatile readable storage medium may be a read-only memory, a disk, or an optical disk, etc.
[0171] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the disclosure of the embodiments of the present application (including the claims) is limited to these examples; based on the ideas of the embodiments of the present application, the technical features of the above embodiments or different embodiments can also be combined, and there are many other variations of different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of the embodiments of the present application.
Claims
1. A method for disaster tolerance of storage system port binding, characterized in that It includes the following steps: Bind the ports of the first device and the second device respectively to form a number of communication paths, and calculate the performance of each communication path; Group the communication paths based on the performance of the communication paths and perform data transmission based on the grouping; In response to an exception occurring in the communication path, adjust the grouping of the communication paths based on the cause of the exception and the number of communication paths in each group.
2. The method according to claim 1, wherein The step of grouping the communication paths based on the performance of the communication paths and performing data transmission based on the grouping includes: Sort the performance of the communication paths from high to low; Select the threshold number of communication paths with the top performance rankings as the available group for data transmission; Use the remaining communication paths as the standby group.
3. The method according to claim 2, wherein The step of selecting the threshold number of communication paths with the top performance rankings as the available group for data transmission includes: Determine the communication paths with the top 50% performance rankings as the available group; The step of using the remaining communication paths as the standby group includes: Determine the communication paths with the bottom 50% performance rankings as the standby group.
4. The method according to claim 3, wherein The step of determining the communication paths with the top 50% performance rankings as the available group includes: Set the communication paths with the top 50% performance rankings to the active state, wherein the communication paths in the active state are used for data transmission; The step of determining the communication paths with the bottom 50% performance rankings as the standby group includes: Set the communication paths with the bottom 50% performance rankings to the enabled state, wherein the communication paths in the enabled state are used to take over the abnormal communication paths in the available group for data transmission when an abnormality occurs in the communication paths in the available group.
5. The method according to claim 2, wherein The step of, in response to an exception occurring in the communication path, adjusting the grouping of the communication paths based on the cause of the exception and the number of communication paths in each group includes: In response to an exception occurring in the communication path, determine the cause of the exception that occurred in the communication path; In response to the cause of the exception that occurred in the communication path being a path disconnection, execute a first preset policy; In response to the cause of the exception that occurred in the communication path being a path congestion, execute a second preset policy.
6. The method according to claim 5, characterized in that, The step of, in response to the cause of the exception that occurred in the communication path being a path disconnection, executing a first preset policy includes: In response to the cause of the exception that occurred in the communication path being a path disconnection, mark the disconnected communication path as an unavailable path, and determine the group of the disconnected communication path; In response to the group of the disconnected communication path being the standby group, obtain the number of communication paths in the standby group, and execute a first preset operation based on the obtained number; In response to the group of the disconnected communication path being the available group, obtain the number of communication paths in the standby group, and execute a second preset operation based on the obtained number.
7. The method according to claim 6, wherein The step of, in response to the group of the disconnected communication path being the standby group, obtaining the number of communication paths in the standby group, and executing a first preset operation based on the obtained number includes: For packets in the standby group in response to a disconnected communication path, obtain the number of communication paths in the standby group; In response to the number of communication paths in the standby group being less than 1, send a warning of no redundant path; In response to the number of communication paths in the standby group being equal to 1, clear the warning of no redundant path and send a warning of insufficient redundant path.
8. The method according to claim 6, characterized in that, The steps of obtaining the number of communication paths in the standby group and performing a second preset operation based on the obtained number in response to the packet of the disconnected communication path being the available group include: In response to the packet of the disconnected communication path being the available group, obtain the number of communication paths in the standby group; In response to the number of communication paths in the standby group being equal to 1, switch the communication path of the standby group to the available group, set it to the available state, and send a warning of no redundant path; In response to the number of communication paths in the standby group being greater than 1, calculate the performance of all communication paths in the standby group; Switch the communication path with the highest performance to the available group and set it to the available state; In response to the number of communication paths in the standby group being less than 1, send a warning of no redundant path.
9. The method according to claim 5, wherein The steps of performing a second preset policy in response to the abnormal reason of the communication path being path congestion include: In response to the abnormal reason of the communication path being path congestion, calculate the performance of the congested communication path and all communication paths in the standby group; In response to the communication path with the highest performance being the congested communication path, do nothing; In response to the communication path with the highest performance being a communication path in the standby group, switch the communication path with the highest performance in the standby group to the available group and set it to the available state.
10. The method according to claim 1, characterized in that, The steps of binding the port of the first device to the port of the second device to form a communication path include: Bind the RoCE port of the first device to the RoCE port of the second device to form a communication path; Create a virtual network card for the RoCE port of the first device and a virtual network card for the RoCE port of the second device respectively; Determine whether the first device can communicate with the second device.
11. The method according to claim 10, wherein The steps of binding the RoCE port of the first device to the RoCE port of the second device to form a communication path include: Establish a one-to-one connection between each RoCE port of the first device and each RoCE port of the second device to form a number of communication paths.
12. The method according to claim 10, wherein The steps of determining whether the first device can communicate with the second device include: Configure IPs for the first virtual network card corresponding to the RoCE port of the first device and the second virtual network card corresponding to the RoCE port of the second device respectively; Judge whether the network of the first device can be connected to the network of the second device; In response to the network of the first device being able to be connected to the network of the second device, send the information of the first device to the second device so that the second device configures according to the information of the first device.
13. The method according to claim 12, characterized in that, The first device is the host, the second device is the storage node, and the RoCE ports of the host and the RoCE ports of the storage node are connected by optical fibers. Before configuring IPs for the first virtual network card corresponding to the RoCE port of the first device and the second virtual network card corresponding to the RoCE port of the second device respectively, the method further includes: Configure BOND port binding through the RoCE port of the host to generate a first virtual network card of the host; Configure the BOND port binding through the RoCE port of the storage node to generate a second virtual network card of the storage node.
14. The method according to claim 12, wherein The first device is a host, and the second device is a storage node. The step of sending the information of the first device to the second device so that the second device configures according to the information of the first device includes: The storage node configures host management using the information of the host; The host discovers the storage node using a first command and connects to the storage node using a second command to complete the configuration of the host and the storage node.
15. The method according to claim 14, wherein The storage node configures host management using the information of the host, including: The storage node obtains the unique identifier of the host and configures host management using the unique identifier of the host.
16. The method according to claim 14, wherein The host discovers the storage node using a first command and connects to the storage node using a second command to complete the configuration of the host and the storage node, including: The host discovers the storage node using the nvme discover command and connects to the storage node using the nvme connect command. Among them, the first command includes the nvme discover command, and the second command includes the nvme connect command.
17. The method according to claim 10, wherein The first device is a first storage node, and the second device is a second storage node. The step of determining whether the first device can communicate with the second device includes: Configure IPs for the first virtual network card corresponding to the RoCE port of the first storage node and the second virtual network card corresponding to the RoCE port of the second storage node respectively; Judge whether the network of the first storage node can communicate with the network of the second storage node; In response to the network of the first storage node being able to communicate with the network of the second storage node, create a cluster based on the first storage node and the second storage node.
18. A device for disaster recovery of storage system port binding, characterized in that, The device includes: A binding module configured to bind the ports of the first device and the ports of the second device respectively to form a number of communication paths and calculate the performance of each communication path; A grouping module configured to group the communication paths based on the performance of the communication paths and perform data transmission based on the grouping; An adjustment module configured to, in response to an exception occurring in a communication path, adjust the grouping of the communication paths based on the cause of the exception and the number of communication paths in each group.
19. A computer device, characterized in that, Include: At least one processor; And A memory storing computer instructions that can be run on the processor. When the instructions are executed by the processor, the steps of the method according to any one of claims 1-17 are implemented.
20. A computer non-volatile readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1-17 are implemented.
Citation Information
Patent Citations
Method for main and slave transmission of multimedia video in wireless Ad hoc network
CN101192955A
A network congestion processing method and device
CN109936508A
Storage system port binding disaster recovery method, device, equipment and medium
CN117499205A
Load balancing for multipath group routed flows by re-routing the congested route
US10116567B1
Redundancy and load balancing in remote direct memory access communications
US20130332767A1