An Ethernet storage system, its information notification method, and related apparatus.
By storing partition tables and dynamic tables on the switching devices of the Ethernet storage system, rapid information communication between nodes is achieved, solving the problem of nodes having difficulty obtaining communication partner information and improving service processing capabilities and system reliability.
Patent Information
- Application Number
- CN202310131477.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-12
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2040-06-12
AI Technical Summary
In Ethernet storage systems, nodes struggle to quickly obtain information from other nodes they can communicate with, resulting in insufficient business processing capabilities.
By storing partition tables and dynamic partition tables on the switching device, after receiving a node's join message, the switching device sends a join notification message to nodes in the same partition and stores configuration information on the nodes, thus enabling rapid information communication between nodes.
It improves the service processing capabilities of the Ethernet storage system, avoids waste of network resources, and enhances the reliability and consistency of the system.
Smart Images

Figure CN116192880B_ABST
Abstract
Description
[0001] This application claims priority to Chinese Patent Application No. CN202010537494.2, filed with the State Intellectual Property Office of China on June 12, 2020, entitled "An Ethernet Storage System and Its Information Notification Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field
[0002] This invention relates to the field of Ethernet technology, and in particular to an Ethernet storage system, an information notification method and related apparatus applied in the Ethernet storage system. Background Technology
[0003] Distributed storage is a data storage technology that distributes data across storage nodes in different locations. These storage nodes are interconnected via switching devices to transmit data and exchange information. Nodes that access and use data from storage nodes are called compute nodes. Storage nodes, switching devices, and compute nodes together form a distributed storage system. When the switching device in a distributed storage system is an Ethernet switching device, this application refers to the distributed storage system as an Ethernet storage system. Nodes in an Ethernet storage system need to obtain information about other nodes they can communicate with. Summary of the Invention
[0004] This application provides an Ethernet storage system, an information notification method and related apparatus applied in the Ethernet storage system, so as to enable nodes (including storage nodes and computing nodes) in the Ethernet storage system to quickly obtain information from other nodes that can communicate, thereby improving the service processing capability of the Ethernet storage system.
[0005] This application provides an Ethernet storage system, including a first switching device and a first node, wherein the first switching device is an access device for the first node. The first node is configured to send a join message to the first switching device, the join message including the identifier and configuration information of the first node. The first switching device is configured to: receive the join message; obtain the identifier of the first node based on the join message; determine a first partition corresponding to the first node based on the identifier of the first node; and send a join notification message to a second node belonging to the first partition, the join notification message including the identifier and configuration information of the first node.
[0006] In this application, the first switching device refers to any edge switching device (i.e., access device). When a first node connects to the first switching device, it sends a join message to the first switching device. This join message includes the identifier and configuration information of the first node. Upon receiving the join message, the first switching device determines the first partition to which the first node belongs (the first partition can be one or more partitions), and then sends a join notification message to each second node (the second node can be one or more nodes) belonging to the first partition. This join notification message includes the identifier and configuration information of the first node. In this way, other nodes belonging to the same partition as the first node can obtain the configuration information of the first node and communicate with it.
[0007] Furthermore, the first switching device is also used to send a node notification message to the first node, the node notification message including the identifier and configuration information of the second node; the first node is also used to receive the node notification message, and obtain and store the identifier and configuration information of the second node according to the node notification message.
[0008] In this application, the first switching device also sends configuration information of a second node belonging to the same partition as the first node to the first node. This allows the first node to obtain configuration information of other nodes it can communicate with. This enables communication between the first node and other nodes, improving the service processing capabilities of the Ethernet storage system.
[0009] In one embodiment, the first switching device is used to determine the first partition corresponding to the first node based on a partition table and the identifier of the first node, wherein the partition table records the correspondence between partitions and node identifiers.
[0010] In this application, by storing a partition table on each edge switching device, each edge switching device can quickly identify the second node that belongs to the same partition as the first node and send a join notification message to the second node, thereby improving the efficiency of configuration information transmission and thus improving the service processing capability of the Ethernet storage system.
[0011] In one embodiment, the first switching device is further configured to: record information of the first node in a partition dynamic table; determine the second node as an active node of the first partition based on the partition dynamic table, wherein the partition dynamic table records information of the active node corresponding to each partition.
[0012] In this application, the first switching device also stores a dynamic partition table to record information about active nodes. This allows the first switching device to send configuration information of the first node only to the active nodes recorded in the dynamic partition table, avoiding the waste of network resources caused by sending node configuration information to failed nodes.
[0013] Optionally, the information of the active node includes the information of the access device of the active node. The first switching device is also used to determine the target switching device based on the access device information of each active node in the partition dynamic table, and to detect whether the communication between the first switching device and the target switching device is interrupted.
[0014] In this application, when the Ethernet storage system includes multiple switching devices, in order to avoid the impact of network failures on services, the first switching device also determines the access device corresponding to the active node and uses it as the target switching device. However, detecting whether the communication between the first switching device and the target switching device is interrupted can improve the reliability of services.
[0015] Optionally, when communication between the first switching device and the target switching device is interrupted, the first switching device is further configured to determine the target node corresponding to the target switching device, delete the information of the target node from the partition dynamic table, and send a deletion notification message, the deletion notification message including the identifier of the target node.
[0016] In this application, when communication between the first switching device and the target switching device is interrupted, the first switching device cannot send messages to or receive messages from the nodes connected to the target switching device (i.e., the target nodes). The first switching device then determines all nodes connected to the target switching device as the target nodes from its own partition dynamic table, deletes the information of the target node from the partition dynamic table, and sends a deletion notification message. This application does not restrict the order in which the first switching device deletes the target node's information and sends the deletion notification message. The first switching device can determine other nodes belonging to the same partition as each target node and send deletion notification messages including the identifier of the target node to these other nodes. This deletion notification message can be sent directly to other nodes or via a reflector.
[0017] Optionally, the Ethernet storage system further includes a second switching device, which acts as a reflector. The reflector's function is to forward messages received from one switching device to other switching devices. In one embodiment, the second switching device is used to send the partition table to the first switching device. In this case, the partition table may be configured by the administrator on the second switching device acting as a reflector. Optionally, the first switching device is also used to send the dynamic partition table to the second switching device. In this way, the second switching device can send the dynamic partition table to other switching devices, achieving synchronization of the dynamic partition table across all edge switching devices. Sending the dynamic partition table may involve sending the complete dynamic partition table or only sending update information of the dynamic partition table. The reflector can be deployed on edge switching devices, aggregation switching devices, or backbone switching devices.
[0018] In this application, when an Ethernet storage system includes multiple switching devices, specifying a reflector from among these switching devices enables the synchronization of the partitioned dynamic table across all edge switching devices in the Ethernet storage system. This allows the device closest to the node to perform the task of sending the node's configuration information, improving the efficiency of information notification and saving network bandwidth resources.
[0019] In one implementation, when a join notification message is sent to a second node belonging to the first partition, the first switching device sends the join notification message to the second switching device. When the second switching device is an access device for the second node, it sends the join notification message to the second node. When the second switching device is not an access device for the second node, it sends the join notification message to the second node through a third switching device, and the second node accesses the third switching device. Optionally, the Ethernet storage system further includes a fourth switching device, and the second switching device is further used for:
[0020] Send the join notification message to the fourth switching device; and
[0021] Send the partition dynamic table to the third and fourth switching devices.
[0022] The solution proposed in this application can be applied in various scenarios. Regardless of whether the second switching device acting as a reflector is the access device of the second node, the first switching device can first send the notification message to the second switching device, which then forwards the join notification message. This application enables centralized management of node configuration information, ensuring the consistency of node configuration information stored in the Ethernet storage system.
[0023] Optionally, the first node is further configured to send a leave message to the first switching device, the leave message including the identifier of the first node. The first switching device is further configured to: receive the leave message, obtain the identifier of the first node based on the leave message, determine the first partition based on the identifier of the first node, and send a leave notification message to a third node belonging to the first partition, the leave notification message including the identifier of the first node.
[0024] In this application, before the first node leaves the Ethernet storage system, it sends a departure message to the first switching device so that the first switching device sends the departure message of the first node to the third node that belongs to the same partition as the first node. This can prevent the third node from sending service requests to the first node after the first node leaves, thereby improving the reliability of the service.
[0025] Optionally, the first switching device is further configured to: detect whether the first node is locally reachable; when the first node is locally unreachable, determine the first partition and send a fault notification message to the fourth node belonging to the first partition, the fault notification message including the identifier of the first node.
[0026] In this application, after the first node is locally unreachable, the first switching device sends a fault notification message of the first node to a fourth node belonging to the same partition as the first node. This can prevent the fourth node from sending service requests to the first node after the first node leaves, thereby improving the reliability of the service.
[0027] In this application, the first node, second node, third node, and fourth node are used to distinguish different nodes in different scenarios, but do not specifically point to any particular node. For example, node A that sends the join message at the first moment is the first node; node A that receives the join message from node B at the second moment is the second node; node A that receives the leave message from node C at the third moment is the third node; and node A that receives the fault notification message from node D at the fourth moment is the fourth node.
[0028] A second aspect of this application discloses an information notification method applied to a first switching device in an Ethernet storage system. The first switching device receives a join message sent by a first node, the join message including the identifier and configuration information of the first node, and the first switching device is the access device for the first node. The first switching device obtains the identifier of the first node based on the join message, determines the first partition corresponding to the first node based on the identifier, and sends a join notification message to a second node belonging to the first partition, the join notification message including the identifier and configuration information of the first node.
[0029] Optionally, the first switching device sends a node notification message to the first node, the node notification message including the identifier and configuration information of the second node.
[0030] Optionally, the first switching device determines the first partition corresponding to the first node based on a partition table and the identifier of the first node, wherein the partition table records the correspondence between partitions and node identifiers.
[0031] Optionally, the first switching device also records the information of the first node in a partition dynamic table; and determines the second node as the active node of the first partition based on the partition dynamic table, which records the information of the active node corresponding to each partition.
[0032] Optionally, the information of the active node includes the information of the access device of the active node. The first switching device determines the target switching device based on the information of the access device of each active node in the partition dynamic table, and detects whether the communication between the first switching device and the target switching device is interrupted.
[0033] Optionally, when communication between the first switching device and the target switching device is interrupted, the first switching device further determines the target node corresponding to the target switching device, deletes the information of the target node from the partition dynamic table, and sends a deletion notification message, which includes the identifier of the target node.
[0034] Optionally, the Ethernet storage system further includes a second switching device, which is a reflector. The first switching device also receives the partition table sent by the second switching device and sends the partition dynamic table to the second switching device.
[0035] Optionally, the first switching device sends the join notification message to the second switching device.
[0036] Optionally, the first switching device further receives a leave message sent by the first node, the leave message including the identifier of the first node, determines the first partition based on the identifier of the first node, and sends a leave notification message to a third node belonging to the first partition, the leave notification message including the identifier of the first node.
[0037] Optionally, the first switching device further detects whether the first node is locally reachable. When the first node is locally unreachable, it identifies the first partition and sends a fault notification message to the fourth node belonging to the first partition. The fault notification message includes the identifier of the first node.
[0038] A third aspect of this application provides an information notification method applied to a first node in an Ethernet storage system. The first node sends a join message to a first switching device in the Ethernet storage system. The join message includes the identifier and configuration information of the first node, and the first switching device is an access device for the first node. The first node also receives a node notification message sent by the first switching device. The node notification message includes the identifier and configuration information of a second node, and the second node belongs to the same partition as the first node.
[0039] Optionally, the first node may also acquire and store information about the second node, including the identifier and configuration information of the second node.
[0040] Optionally, the first node receives a deletion message sent by the first switching device, the deletion message including the identifier of the second node; and deletes the information of the second node according to the deletion message.
[0041] Depending on the scenario, the deletion message can be a node leaving message, a deletion notification message, or a fault notification message.
[0042] A fourth aspect of this application provides a switching device. This switching device includes functional modules that perform the information notification method provided in the second aspect or any possible design of the second aspect. This application does not limit the division of functional modules; the functional modules can be divided according to the flow steps of the information notification method in the first aspect, or according to specific implementation needs.
[0043] The fifth aspect of this application provides a node. This node includes a functional module that performs the information notification method provided in the third aspect or any possible design of the third aspect; this application does not limit the division of functional modules, and the functional modules can be divided according to the flow steps of the information notification method in the third aspect, or according to specific implementation needs.
[0044] The sixth aspect of this application provides a host computer. A node runs on this host computer, including a memory, a processor, and a communication interface. The memory stores computer program code and data, and the processor invokes the computer program code and, in conjunction with the data, enables the node to implement the information notification method of the third aspect of this application and any possible design thereof.
[0045] The seventh aspect of this application provides a chip that, when running, can implement the information notification method in the second aspect of this application and any possible design thereof, as well as the information notification method in the third aspect of this application and any possible related application.
[0046] The eighth aspect of this application provides a storage medium storing program code that, when executed, enables a device (switch, server, terminal device, etc.) running the program code to implement the information notification method in the second aspect of this application and any possible design thereof, as well as the information notification method in the third aspect of this application and any possible related application.
[0047] The beneficial effects of aspects two through eight of this application can be found in the description of the beneficial effects of aspect one and its various possible designs, and will not be repeated here. Attached Figure Description
[0048] Figure 1 This is a schematic diagram of the structure of an Ethernet storage system provided in an embodiment of this application;
[0049] Figure 2 This is a schematic diagram of another Ethernet storage system provided in an embodiment of this application;
[0050] Figure 3This is a schematic diagram of another Ethernet storage system provided in an embodiment of this application;
[0051] Figure 4 A flowchart illustrating an information notification method provided in an embodiment of this application;
[0052] Figure 5A and Figure 5B This is a schematic diagram of the partition table structure provided in an embodiment of this application;
[0053] Figure 6A A schematic diagram of the partition dynamic table before the update provided in this application embodiment;
[0054] Figure 6B A schematic diagram of the updated partition dynamic table provided in an embodiment of this application;
[0055] Figure 7A A schematic diagram of a node dynamic table provided in an embodiment of this application;
[0056] Figure 7B A schematic diagram of the node dynamic table after a port failure provided in an embodiment of this application;
[0057] Figure 8 A schematic diagram of the structure of the switching device provided in the embodiments of this application;
[0058] Figure 9 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0059] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0060] In the embodiments of this application, unless otherwise stated, "multiple" means two or more. For example, multiple nodes means two or more nodes. "At least one" means any number, such as one, two, or more. "A and / or B" can be only A, only B, or include A and B. "At least one of A, B, and C" can be only A, only B, only C, or include A and B, include B and C, include A and C, or include A, B, and C. The terms "first," "second," etc., in this application are only used to distinguish different objects and are not used to indicate the priority or importance of objects. The embodiments of this application are used to enable nodes in an Ethernet storage system to quickly obtain information from other nodes belonging to the same partition, improve the service processing capability of the Ethernet storage system, and avoid nodes communicating with abnormal nodes, thereby improving the reliability of the Ethernet storage system. In this application, in order to allow specific nodes / devices to communicate with each other, the group formed by these nodes / devices is called a partition. A partition is similar to a small virtual private network (VPN). To make the purpose, technical solution, and advantages of this application clearer, the application will be described in further detail below with reference to the accompanying drawings.
[0061] This application applies to an Ethernet storage system, which includes multiple nodes and at least one switching device (in this application, the switching device refers to an Ethernet switching device). Each of the multiple nodes is connected to a corresponding switching device (which may be called an edge switching device or a leaf switching device) to publish information to or receive information from the switching device. In terms of form, the nodes in this application can be physical devices or virtual devices deployed on physical devices. When a node is a virtual device, multiple nodes can be hosted on the same physical device. This physical device can be a physical server, workstation, mobile station, general-purpose computer, or other device capable of hosting nodes. Functionally, the nodes in this application can be computing nodes (e.g., servers) or storage nodes (e.g., storage arrays).
[0062] In one implementation, the Ethernet storage system can be Figure 1 The Ethernet storage system 100 shown includes a switching device 110 and multiple nodes 120 connected to the switching device 110. Figure 1 The diagram shows nodes 120a-120d. In another embodiment, the Ethernet storage system can be... Figure 2 The Ethernet storage system 200 shown includes multiple switching devices 210 (switching devices 210a-210c are shown in the figure) and multiple nodes 220 ( Figure 2The diagram shows nodes 220a-220f, each node in which is connected to a corresponding switching device (i.e., an edge switching device), and multiple switching devices 210a-210c are interconnected. In another embodiment, the Ethernet storage system is... Figure 3 The Ethernet storage system 300 shown includes multiple switching devices 310 (switching devices 310a-310d are shown in the figure) and multiple nodes 320 ( Figure 3 The figure shows nodes 320a-320d. Each of these nodes 320a-320d is connected to a corresponding edge switching device, and the multiple switching devices 310a-310d include edge switching devices 310a and 310b and backbone switching devices 310c and 310d. Edge switching devices 310a and 310b communicate through backbone switching devices 310c and 310d. Optionally, the Ethernet storage system may also include a convergence switching device (not shown) for connecting the backbone switching devices and the edge switching devices. When the Ethernet storage system includes multiple switching devices, these multiple switching devices form a switching network (which may be called a fabric), which is transparent to the nodes. For example, Figure 2 The switching equipment 210a-210c in the middle form a switching network. Figure 3 The switching devices 310a-310d in the network form a switching network. In this application, the node connected to the switching device is referred to as the local node of the switching device, and the switching device is referred to as the access device of the local node. For example, Figure 3 In this context, nodes 320a and 320b are local nodes of switching device 310a, and switching device 310b is an access device for nodes 320c and 320d.
[0063] When the Ethernet storage system includes multiple switching devices, at least one of these switching devices is configured as a reflector (i.e., capable of performing the function of a reflector). This reflector acts as a collection point in the switching network, receiving information sent by one switching device and forwarding that information to other switching devices to achieve information synchronization throughout the entire switching network. For example, it can be... Figure 2 The switching device 210c in the middle is configured as a reflector, which will Figure 3 The switching devices 310c and 310d in the middle are configured as reflectors.
[0064] In Ethernet storage systems, to enhance information security and avoid wasting network resources, nodes are divided into different partitions. Each partition contains a group of nodes. A node can belong to multiple partitions. Nodes in an Ethernet storage system can only communicate with other nodes belonging to the same partition. Therefore, it is necessary to advertise node information within the Ethernet storage system so that a node can obtain information about other nodes belonging to the same partition as it.
[0065] The following combination Figures 1-3 as well as Figure 4 This application will now introduce the information notification method provided in its embodiments. For example... Figure 4 The diagram illustrates an information notification method provided in an embodiment of this application, comprising steps S400-S495. In specific implementations, steps S400-S495 can be modified or deleted as needed. Figure 4 The order of any step can be adjusted as needed.
[0066] In step S400, a connection is established between the node and the corresponding switching device. Each node needs to establish a connection with the corresponding switching device to communicate with it when it accesses the network. Figure 4 This example illustrates the concept of a first node connecting to a first switching device and a second node connecting to a third switching device. In actual deployments, the first and second nodes can be... Figure 1 , Figure 2 and Figure 3 Any node in the array.
[0067] In step S405, the first node sends a join message to the first switching device.
[0068] The first node here refers to any node (e.g., Figure 1 Nodes 120a-120d in the middle, Figure 2 Nodes 220a-220f in the middle, Figure 3In the context of nodes 320a-320d, the first switching device is the edge switching device connected to the first node. The join message indicates that the first node wants to join the network to communicate with other nodes. The join message includes the identifier of the first node and its configuration information. The identifier of the first node can be any one or more of the following: Internet Protocol (IP) address, Media Access Control (MAC) address, device name, device number, and device connection relationship. The device connection relationship indicates the identifier of the switching device to which the first node is connected, and may also indicate the port on which the switching device connects to the first node. The configuration information of the first node includes its role and / or protocol-related information. In one implementation, the role of the first node can be a compute node, a storage node, or a general-purpose node (i.e., it can function as both a compute node and a storage node). The protocol-related information of the first node includes the protocols supported by the first node and the ports used by those protocols. When the first node is a storage node, the protocol-related information also includes the target identifier length of the protocol and the target identifier of the protocol. The protocols supported by this first node can include one or more, such as NVMe over RoCE, NVMe over TCP, or iSCSI. The target identifier is used to identify the remote storage target; for example, it can be an NVMe qualified name (NQN) or an iSCSI qualified name (IQN).
[0069] In step S410, the first switching device receives the join message, obtains the identifier of the first node based on the join message, and determines the first partition corresponding to the first node based on the identifier of the first node.
[0070] In one implementation, the first switching device stores a partition table that records the partitions corresponding to each node. Specifically, the partition table records the correspondence between node identifiers and partition identifiers. When the Ethernet storage system includes only one switching device, the partition table includes the zones to which all nodes connected to that switching device belong. When the Ethernet storage system includes multiple switches, the partition table includes the zones to which all nodes connected to all switching devices belong, and the contents of the partition tables on all switching devices are identical. This partition table reflects the network planning; that is, it records all nodes planned to be deployed in the Ethernet storage system. Some of these nodes may have already been activated and established connections with the switching devices (these nodes are called active nodes), while some nodes may not yet be deployed. Active nodes may become invalid due to anomalies or failures during subsequent operation. This partition table can be configured by the network administrator on a reflector and sent to the first switching device by the reflector, or it can be directly configured on the first switching device. The partition table can adopt different structures. For example, Figure 5A and Figure 5B by Figure 3 Taking the Ethernet storage system shown as an example, this application illustrates the structural diagram of the partition table provided in its embodiment. The Ethernet storage system includes nodes 320a-320d. Figure 5A The partitions corresponding to each node are recorded. Node 320a corresponds to partitions zone1 and zone2, node 320b corresponds to zone1 and zone3, node 320c corresponds to zone2, and node 320d corresponds to zone1, zone2, and zone3. That is, each node can correspond to one or more partitions. Figure 5B The system records the nodes corresponding to each partition, and each partition includes multiple nodes. Since each node can correspond to one or more partitions, the first partition in this application actually includes one or more partitions corresponding to the first node. For example, when the first node is 320a, the first partition includes zone1 and zone2.
[0071] In step S420, the first switching device sends a join notification message to the second node belonging to the first partition. The join notification message includes the identifier and configuration information of the first node.
[0072] In one implementation, the second node is all nodes belonging to the first partition except for the first node. For example, if the first node 320a corresponds to partitions zone1 and zone2, then the second nodes determined by the first switching device include nodes 320b, 320c, and 320d.
[0073] Furthermore, the second node is an active node in the first partition. In this case, the first switching device also needs to determine all active nodes in the first partition, excluding the first node, based on the partition dynamic table. This partition dynamic table records information about the active nodes in each partition. Assuming that before 320a sends the join message, Figure 3 If the 320b and 320c are added to the network, then the dynamic partition table on the first switching device is as follows: Figure 6A As shown, since node 320a corresponds to partitions zone1 and zone2, the second nodes determined by the first switching device according to the partition dynamic table include 320b and 320c. The first switching device can send a join notification message to the second node in the following way:
[0074] A. The first switching device is the access device for the second node, and the first switching device directly sends the join notification message to the second node.
[0075] B. The first switching device is not the access device of the second node, and the first switching device is a reflector. The first switching device sends an access notification message to the second node through the third switching device, and the second node accesses the third switching device.
[0076] C. The first switching device is not the access device of the second node, and the first switching device is not a reflector. The first switching device sends the join notification message to the reflector (the second switching device) so that the reflector can directly send the join notification message to the second node, or cause the reflector to send the join notification message to the second node through the third switching device. Figure 4 (As illustrated in this case), the second node connects to the third switching device.
[0077] In step S425, the second node receives the join notification message and obtains and stores the identifier and configuration information of the first node.
[0078] As mentioned earlier, the second node can be connected to the first switching device, or it can be connected to other switching devices. Figure 4 The following example illustrates the first node using the second node accessing the third switching device (non-reflector). After receiving the join notification message through the third switching device, the second node obtains and stores the identifier and configuration information of the first node included in the join notification message.
[0079] In one implementation, the second node stores a node configuration table that records information about the nodes the second node can communicate with. For example, each entry in this node configuration table includes the node's identifier, role, and protocol-related information. The second node can add the identifier and configuration information of the first node to the node configuration table. The entries in this node configuration table may or may not include the partition to which the node belongs. Since a node can belong to one or more partitions, the nodes recorded in the second node's node configuration table can also belong to one or more partitions.
[0080] In step S430, the first switching device sends a node notification message to the first node, which includes the identifier and configuration information of the second node.
[0081] After identifying the second node, the first switching device needs to send the second node's information to the first node. This information includes the second node's identifier and configuration information. When the second node comprises multiple nodes, the first switching device can send the information of all these nodes simultaneously to the first node via a single node notification message, or it can send the information of each node separately to the first node via multiple node notification messages. These multiple nodes can belong to the same partition or different partitions. The first switching device may or may not send the partition identifier of the second node to the first node.
[0082] In step S435, the first node receives a node notification message and obtains and stores the identifier and configuration information of the second node.
[0083] Similar to the second node, the first node obtains the identifier and configuration information of the second node and stores the identifier and configuration information of the second node in the node configuration table.
[0084] In step S440, the first switching device updates the partition dynamic table according to the join message.
[0085] There is no strict order requirement for steps S430 and S440. Step S440 involves the first switching device adding the information of the first node to the partition dynamic table. For example, Figure 3 If node 320a is the first node, and node 320a corresponds to zone1 and zone2, then the updated dynamic partition table is as follows: Figure 6B As shown.
[0086] Furthermore, the first switching device also stores a node dynamic table, which records information about active nodes. This node dynamic table can record the node's identifier and configuration information, the identifier of the switching device (i.e., access device) to which the node is connected, and the port on which the switching device connects to the node (referred to as the access device port). Furthermore, the node dynamic table can also record the node's status, which can be valid or invalid. Valid status indicates that the node can communicate normally, while invalid status indicates that the node or the link to which the node resides has experienced an anomaly or fault, resulting in degraded communication quality or inability to communicate. Furthermore, the node dynamic table can also record the cause of the anomaly or fault. After receiving the join message, the first switching device records the information of the first node in the node dynamic table according to the join message. For example... Figure 7A The diagram shown is a schematic of a node dynamic table provided in an embodiment of this application. This node dynamic table is related to... Figure 6B Correspondingly, information about nodes 320b, 320c, and the newly added 320a is recorded. In another implementation, the information in the node dynamic table can also be recorded in the partition dynamic table. That is, the first switching device only needs to store the partition dynamic table, which records... Figure 6A or Figure 6B and Figure 7A The information shown.
[0087] In step S450, the first switching device sends the partition dynamic table.
[0088] The first switching device sends the partition dynamic table, which can be either the updated partition dynamic table or only the updated content of the partition dynamic table. When there is only one switching device (i.e., the first switching device) in the Ethernet storage system, the first switching device does not need to execute step S450. When the switching network includes multiple switching devices, if the first switching device is a reflector (e.g., switching device 210c), the first switching device sends the updated partition dynamic table to other switching devices (e.g., switching devices 210a and 210b), and the other switching devices store the updated partition dynamic table. If the first switching device is not a reflector (e.g., switching device 210a), the first switching device sends the updated partition dynamic table to a reflector (e.g., switching device 210c), and then the reflector sends the updated partition dynamic table to other switching devices, so that the partition dynamic table is synchronized among all switching devices in the switching network. When multiple reflectors exist in the switching network (e.g., switching devices 310c and 310d), the first switching device (e.g., switching device 310a) sends the partition dynamic table to each of the multiple reflectors, and the multiple reflectors send the partition dynamic table to other switching devices (e.g., switching device 310b). In this case, switching device 310b can store the partition dynamic tables sent by switching devices 310c and 310d respectively. Thus, even if one reflector fails, communication will not be interrupted. Figure 4 In the example of the second access device as a reflector, the method may further include step S450', in which the second access device sends the partition dynamic table to the third access device.
[0089] When the node dynamic table is independent of the partition dynamic table, the process of sending the partition dynamic table described above also applies to sending the node dynamic table. The processes of sending the partition dynamic table and the node dynamic table can be performed simultaneously or separately.
[0090] In step S455, the first node sends a leave message to the first switching device, the leave message including the identifier of the first node.
[0091] In step S460, the first switching device receives the departure message, obtains the identifier of the first node based on the departure message, and determines the first partition based on the identifier of the first node.
[0092] In step S470, the first switching device sends a departure notification message to the third node belonging to the first partition, the departure notification message including the identifier of the first node.
[0093] The method by which the first switching device determines the third node in step S460 can be referred to the description of step S410. The third node in step S470 may be the same as or different from the second node in step S420. For example, after the first node sends a join message, if a node leaves or joins the first partition to which the first node belongs, then the third node determined in step S470 will be different from the second node determined in step S420.
[0094] In step S475, the second node receives the departure notification message and deletes the information of the first node according to the departure notification message.
[0095] In one implementation, the second node receives the departure notification message, retrieves the identifier of the first node from the message, and deletes the first node's information from the node configuration table based on that identifier. Deleting the first node's information includes either deleting the corresponding entry from the node configuration table or invalidating the entry, which includes the first node's identifier and configuration information. After deleting the first node's information, the second node ceases communication with the first node.
[0096] In step S480, the first switching device further identifies a target switching device and detects whether communication with the target switching device is interrupted.
[0097] Optionally, the first switching device treats all switching devices in the switching network except itself as target switching devices, and detects whether the communication with each target switching device is interrupted.
[0098] Optionally, the first switching device determines the target switching device based on the access device information of each active node, and detects whether communication with the target switching device is interrupted. In one implementation, the first switching device obtains the identifier of the active node according to the partition dynamic table, determines the access device information of each active node based on the identifier, and then determines the target switching device based on the access device information of all active nodes. For example, the access devices of all active nodes are added to a set, and then duplicates are removed from the set to obtain a candidate set. Then, the other switching devices in the candidate set besides the first switching device are used as the target switching devices. For example, Figure 7A There are three active nodes 320a, 320b and 320c. Nodes 320a and 320b correspond to switching device 310a, and node 320c corresponds to switching device 310b. Therefore, the target switching device determined by switching device 310a is 310b.
[0099] It should be understood that step S480 can be performed by any switching device. For example, the target switching device determined by switching device 310b is switching device 310a, and the target switching devices determined by switching devices 310c and 310d are switching devices 310a and 310b, respectively.
[0100] In another implementation, the first switching device determines the node corresponding to each partition according to the partition table, determines whether the node is connected to the Ethernet storage system, and if the node is already connected to the Ethernet storage system (i.e. the node is an active node), the switching device corresponding to the node is added to a set, and the set is deduplicated until all nodes of all partitions have been traversed, resulting in a candidate set. Then, the other switching devices in the candidate set except for the first switching device are used as the target switching devices.
[0101] In one implementation, a detection session is established between the first switching device and each target switching device to detect whether communication between the first switching device and the target switching device is interrupted. This detection session may be, for example, a Bidirectional Forwarding Detection (BFD) session or other keep-alive session.
[0102] Step S480 can be executed periodically or when set conditions are met. When communication between the first switching device and the target switching device is uninterrupted, step S480 can continue to be executed. When communication between the first switching device and the target switching device is interrupted, the first switching device executes step S490.
[0103] In step S490, the first switching device performs fault handling on the target switching device.
[0104] The first switching device performs fault handling on the target switching device in the following ways:
[0105] 1. The target switching device is not a reflector.
[0106] The first switching device determines the target node corresponding to the target switching device, deletes the target node's information from the partition dynamic table, and sends a deletion notification message, which includes the identifier of the target node. In one implementation, the node dynamic table can be a sub-table of the partition dynamic table, and the nodes in the partition dynamic table point to the information of that node in the node dynamic table. Therefore, deleting the target node's information from the partition dynamic table includes deleting the node's information from the node dynamic table. In another implementation, the first switching device also needs to delete the target node's information from the node dynamic table.
[0107] In one implementation, the first switching device determines the target node (i.e., the local node of the target switching device) corresponding to the target switching device based on the node dynamic table. Then, the first switching device deletes the information of the target node from the partition dynamic table and the node dynamic table, and publishes a deletion notification message, which includes the identifier of the target node. When there are multiple target nodes, the identifiers of the multiple target nodes can be published through one or more deletion notification messages.
[0108] If the first switch is not a reflector, it sends the deletion notification message to the reflector, which then forwards it to all other switches except the target switch and the first switch. If the first switch is a reflector, it sends the deletion notification message to all other switches except the target switch. If the first switch is connected to a local node belonging to the same partition as the target node, it also sends the deletion notification message to that local node. Each switch that receives the deletion notification message also sends it to the local node connected to that switch that belongs to the same partition as the target node.
[0109] Each switching device that receives the deletion notification message removes the target node's information from its partition dynamic table and node dynamic table. Each node that receives the deletion notification message also removes the target node's information from its node configuration table to avoid communicating with the target node.
[0110] II. The target switching device is a reflector.
[0111] When the target switching device is a reflector, and the first switching device is not a reflector, the first switching device deletes the data received from the target switching device. This data may include one or more of the following: partition tables, dynamic partition tables, dynamic node tables, and other information sent by the target switching device to the first switching device.
[0112] In this embodiment of the application, if a switching device is not a reflector, the switching device records the data received from the reflector and establishes an association between the data and the reflector. In this way, when the communication between the switching device and the reflector is interrupted, the switching device deletes the data received from the reflector.
[0113] Furthermore, the first switching device can also detect whether the first node is locally unreachable. When the first node is locally unreachable, the first switching device determines the first partition to which the first node belongs, and then sends a fault notification message to a fourth node belonging to the first partition. This fourth node can be the same as or different from the nodes included in the second node mentioned above. For example, after the first node sends a join message and joins the Ethernet storage system, if other nodes in the first partition to which the first node belongs leave the Ethernet storage system, the nodes included in the first partition will change. When the first node is locally unreachable, the fourth node included in the first partition will be different from the second node included in the first partition when the first node joined the Ethernet storage system. The fault notification message includes the identifier of the first node. In this application, "locally unreachable" means that data or messages cannot be forwarded to the node through the node's switching device. When communication between the first switching device and the first node is interrupted, or when the route to the first node stored on the first switching device fails, the first switching device determines that the first node is locally unreachable. For example, if a link failure between the first node and the first switching device prevents other nodes from sending messages to the first node through the first switching device, then the first node becomes locally unreachable relative to the first switching device. Communication interruption between the first switching device and the first node could be caused by, for example, the port on the first switching device connecting to the first node failing, or the first switching device not receiving a keep-alive message from the first node within a set time period (e.g., every 5 minutes, or every N set keep-alive message cycles, or at a set time point). The port on the first switching device connecting to the first node failing means that the port used by the first switching device to connect to the first node is in a faulty (e.g., down) state. Several situations can cause the port connecting the first switching device to the first node to fail, such as a faulty port connecting the first switching device to the first node, a faulty port connecting the first node to the first switching device, a cable fault between the first switching device and the first node, a power failure of the first node, a reset of the first node, a faulty optical module of the first switching device, or a priority flow control (PFC) storm on the port, or when checking the packets received on the port, the number of errors in the packets is found to exceed a set threshold. The aforementioned errors may be, for example, cyclic redundancy check (CRC) errors.
[0114] Furthermore, the first switching device can also modify the status of the first node in the node dynamic table to invalid and record the cause of the fault. For example, if node 320a is connected to port 2 of switching device 310a, and port 2 is considered faulty due to a CRC check error, the first switching device can... Figure 7A The node dynamic table shown has been modified to Figure 7B Optionally, the first switching device can delete the information of the first node from the partition dynamic table.
[0115] In one implementation, the first switching device can send the fault notification message to other switching devices in the switching network. The switching device receiving the fault notification message identifies a local node belonging to the same partition as the first node and sends the fault notification message to that local node. There can be one or more local nodes.
[0116] In another implementation, the first switching device determines the first partition to which the first node belongs, and determines the second node belonging to the first partition. Then, the first switching device directly sends a fault notification message to the second node.
[0117] Upon receiving a fault notification message, the node removes the information of the first node from the node configuration table based on the identifier of the first node, in order to avoid service interruption caused by communication with the first node.
[0118] The node join message and node notification message in this embodiment are for descriptive convenience and are message names provided for different objects. In practical applications, the above-mentioned node join message and node notification message can be messages of the same format or different formats, and can be messages of the same type or different types.
[0119] The departure notification message, deletion notification message, and fault notification message in this application embodiment are message names provided for ease of description and to address different scenarios. In practical applications, the aforementioned departure notification message, deletion notification message, and fault notification message can be messages of the same format or different formats, and can be messages of the same type or different types. Furthermore, since the aforementioned departure notification message, deletion notification message, and fault notification message can all cause the receiving node to delete the corresponding node's information, they can be collectively referred to as deletion messages.
[0120] Figure 4 The processing flow of the information notification method in the embodiments of this application is used to illustrate the following: In actual deployment, each switching device can implement all or part of the functions performed by the different switching devices, and each node can implement all or part of the functions performed by the first node and the second node.
[0121] In various embodiments of this application, after a node in the Ethernet storage system establishes a connection with an edge switching device, it proactively sends a join message to the switching device, reporting its configuration information. The switching device can determine the partition to which the node belongs based on its identifier and then send the node's configuration information to other nodes belonging to the same partition. In this way, each node can obtain information about the nodes it can communicate with, improving the service processing efficiency of the Ethernet storage system.
[0122] In the embodiments provided above, the information notification method provided by the embodiments of this application has been described from the perspectives of nodes and switching devices, respectively. It is understood that, in order to achieve the above-mentioned functions, the nodes and switching devices in the embodiments of this application include hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that the functions and steps of the various examples described in the embodiments disclosed in this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art use different methods to implement the described functions, but such implementation should not be considered beyond the scope of this application. The structure of the nodes and switching devices in this application is described below from different perspectives. To implement this application... Figure 4 The method shown in this application embodiment provides a switching device 800. This switching device 800 can be... Figure 4 The first switching device in the system. This switching device 800, besides... Figure 8 In addition to the components shown, other components can be included to achieve more functionality.
[0123] like Figure 8 As shown, the switching device 800 includes a receiving unit 8031, a processing unit 8032, and a transmitting unit 8033. In this embodiment, the receiving unit 8031 is used to perform one or more of steps S405 and S455; the processing unit 8032 is used to perform one or more of steps S410, S440, S460, S480, and S490; and the transmitting unit 8033 is used to perform one or more of steps S420, S430, S450, S450', and S470. The functions of the receiving unit 8031, the processing unit 8032, and the transmitting unit 8033 are described in detail below.
[0124] In one embodiment, the receiving unit 8031 is configured to receive a join message sent by a first node, the join message including the identifier and configuration information of the first node, and the first switching device being the access device of the first node. The processing unit 8032 is configured to obtain the identifier of the first node based on the join message, and determine the first partition corresponding to the first node based on the identifier of the first node. The sending unit 8033 is configured to send a join notification message to a second node belonging to the first partition, the join notification message including the identifier and configuration information of the first node.
[0125] Optionally, the sending unit 8033 is further configured to: send a node notification message to the first node, the node notification message including the identifier and configuration information of the second node.
[0126] Optionally, the switching device further includes a storage unit 8034 for storing a partition table that records the correspondence between partitions and node identifiers. When the first partition corresponding to the first node is determined based on the identifier of the first node, the processing unit 8032 is used to determine the first partition corresponding to the first node based on the partition table and the identifier of the first node.
[0127] Optionally, the storage unit 8034 is further configured to store a partition dynamic table, which records information about the active nodes corresponding to each partition. The processing unit 8032 is further configured to: determine, based on the partition dynamic table, that the second node is an active node of the first partition; and record information about the first node in the partition dynamic table.
[0128] Optionally, the information of the active node includes the information of the access device of the active node, and the processing unit 8032 is further configured to: determine the target switching device according to the access device information of each active node in the partition dynamic table, and detect whether the communication between the first switching device and the target switching device is interrupted.
[0129] Optionally, when communication between the first switching device and the target switching device is interrupted, the processing unit 8032 is further configured to: determine the target node corresponding to the target switching device, delete the information of the target node from the partition dynamic table, and send a deletion notification message, wherein the deletion notification message includes the identifier of the target node.
[0130] Optionally, the Ethernet storage system further includes a second switching device, which is a reflector. The receiving unit 8031 is also used to receive the partition table sent by the second switching device; the sending unit 8033 is also used to send the partition dynamic table to the second switching device.
[0131] Optionally, when sending a join notification message to a second node belonging to the first partition, the processing unit 8032 is used to send the join notification message to the second switching device.
[0132] Optionally, the receiving unit 8031 is further configured to receive a leave message sent by the first node, the leave message including the identifier of the first node; the processing unit 8032 is further configured to determine the first partition based on the identifier of the first node; and the sending unit 8033 is further configured to send a leave notification message to a third node belonging to the first partition, the leave notification message including the identifier of the first node.
[0133] Optionally, the processing unit 8032 is further configured to: detect whether the first node is locally reachable; when the first node is locally unreachable, determine the first partition; the sending unit 8033 is further configured to send a fault notification message to a fourth node belonging to the first partition, the fault notification message including the identifier of the first node.
[0134] The aforementioned receiving unit 8031, processing unit 8032, and transmitting unit 8033 can be implemented in hardware or software. When implemented in software, such as... Figure 8 As shown, the switching device 800 may further include a processor 801, a communication interface 802, and a memory 803. The processor 801, communication interface 802, and memory 803 are connected via a bus system 804. The memory 803 stores program code, which includes instructions that can implement the functions of the receiving unit 8031, the processing unit 8032, and the transmitting unit 8033. The processor 801 can call the program code in the memory 803 to implement the functions of the receiving unit 8031, the processing unit 8032, and the transmitting unit 8033. Furthermore, the switching device 800 may also include a programming interface 808 for writing the program code into the memory 803.
[0135] In this embodiment, the processor 801 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 801 can include one or more processing cores.
[0136] The communication interface 802 is used to communicate with external devices, for example, to implement receiving and / or sending functions in conjunction with the program code in the memory 803. Figure 8 The communication interface 802 in the example is just an example. In practice, the switching device 800 may include multiple communication interfaces to connect to and communicate with multiple different external devices.
[0137] The memory 803 may include read-only memory (ROM) or random access memory (RAM). Any other suitable type of storage device may also be used as memory 803. Memory 803 may include one or more storage devices. Memory 803 may further store an operating system 8034, which supports the operation of the switching device 800.
[0138] In addition to the data bus, the bus system 804 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 804 in the diagram.
[0139] In order to realize this application Figure 4 The method shown in this application embodiment also provides a terminal device 900. The terminal device 900 can be a node (in this case, a physical device) or a host where the node resides (in this case, a virtual device running on a physical device). The terminal device 900 can have... Figure 4 The functions of the first and / or second nodes in the terminal device 900. In addition to this, Figure 9 In addition to the components shown, other components can be included to achieve more functionality.
[0140] like Figure 9As shown, the terminal device 900 includes a transmitting unit 9031, a receiving unit 9032, and a processing unit 9033. In this embodiment, the transmitting unit 9031 is used to execute step S405 or S455, and the receiving unit 9032 is used to execute the receiving behavior in step S420, S430, or S470. The processing unit 9033 is used to execute one or more of steps S425, S435, and S475. The functions of the transmitting unit 9031, the receiving unit 9032, and the processing unit are described in detail below.
[0141] The sending unit 9031 is used to send a join message to the first switching device of the Ethernet storage system. The join message includes the identifier and configuration information of the first node, and the first switching device is the access device of the first node.
[0142] The receiving unit 9032 is used to receive a node notification message sent by the first switching device. The node notification message includes the identifier and configuration information of the second node, and the second node and the first node belong to the same partition.
[0143] Optionally, the processing unit 9033 is used to acquire and store information about the second node, including the identifier and configuration information of the second node.
[0144] Optionally, the receiving unit 9032 is further configured to receive a deletion message sent by the first switching device, the deletion message including the identifier of the second node;
[0145] The processing unit 9033 is also configured to delete the information of the second node according to the deletion message.
[0146] The transmitting unit 9031, the receiving unit 9032, and the processing unit 9033 can be implemented in hardware or software. When implemented in software, such as... Figure 9 As shown, the terminal device 900 may further include a processor 901, a communication interface 902, and a memory 903. The processor 901, communication interface 902, and memory 903 are connected via a bus system 904. The memory 903 stores program code, which includes instructions that implement the functions of the sending unit 9031, the receiving unit 9032, and the processing unit 9033. The processor 901 can call the program code in the memory 903 to implement the functions of the sending unit 9031, the receiving unit 9032, and the processing unit 9033. Furthermore, the terminal device 900 may also include a programming interface 905 for writing the program code into the memory 903.
[0147] In this embodiment, the processor 901 can be a CPU, or it can be other general-purpose processors, DSPs, ASICs, FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor. The processor 901 can include one or more processing cores.
[0148] The communication interface 902 is used to communicate with external devices, for example, to implement receiving and / or sending functions in conjunction with the program code in the memory 903. Figure 9 The communication interface 902 in the example is just an example. In practice, the terminal device 900 may include multiple communication interfaces to connect to and communicate with multiple different external devices.
[0149] The memory 903 may include ROM or RAM. Any other suitable type of storage device may also be used as memory 903. Memory 903 may include one or more storage devices. Memory 903 may further store operating system 9034, which supports the operation of terminal device 900. The second information may be stored in memory 903 or in other memories outside of memory 903.
[0150] In addition to the data bus, the bus system 904 may also include a power bus, a control bus, and a status signal bus. However, for clarity, all buses are labeled as bus system 904 in the diagram.
[0151] The various components of the switching device 800 or terminal device 900 provided in this application are merely exemplary. Those skilled in the art can add or remove components as needed, or divide the function of one component into multiple components. The implementation methods of the various functions of the switching device 800 or terminal device 900 in this application can be referred to... Figure 3 Description of each step.
[0152] Through the above description of the embodiments, those skilled in the art can clearly understand that the switching device 800 or terminal device 900 of this application can be implemented in hardware, or it can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a hardware product or a software product. The hardware product can be a dedicated chip or processor. The software product can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, portable hard drive, etc.), and the software product includes several instructions. When the software product is executed, it can cause a computer device to execute the methods described in the various embodiments of this application.
[0153] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. An Ethernet storage system, characterized by, The first exchange device is a reflector, or the first exchange device is not a reflector, and the second exchange device is a reflector, and the first exchange device is further configured to send the first deletion notification message to other exchange devices through the reflector. The first exchange device is further configured to: detect whether the first node is locally reachable, and when the first node is not locally reachable, determine the first partition, and send a failure notification message to a fourth node belonging to the first partition, the failure notification message comprising an identifier of the first node. The first exchange device is further configured to:
2. The Ethernet storage system of claim 1, wherein, when a condition is met, the first exchange device determines that the first node is not locally reachable, the condition comprising:
3. The Ethernet storage system according to claim 1 or 2, characterized in that, the first exchange device does not receive a keep-alive message from the first node within a set time; or, 4. The Ethernet storage system of claim 3, wherein, a port connected to the first node by the first exchange device is disabled. The first exchange device is further configured to: in the case of a communication interruption between the first exchange device and the target exchange device, send a second deletion notification message to a local node connected to the first exchange device, the second deletion notification message comprising an identifier of a local node connected to the target exchange device, the local node connected to the first exchange device and the local node connected to the target exchange device belonging to the same partition. The first configuration information comprises a role of the first node and / or protocol-related information of the first node. The role of the first node comprises a compute node, a storage node, or a general-purpose node, and the general-purpose node is a compute node and a storage node.
5. The Ethernet storage system according to any one of claims 1-4, wherein, The protocol-related information of the first node comprises a protocol supported by the first node, a port used by the protocol, a target identifier length of the protocol, and a target identifier of the protocol, and the protocol comprises NVMe over RoCE, NVMe over TCP, or iSCSI.
9. The Ethernet storage system of any one of claims 1-8, wherein:
6. The Ethernet storage system according to any one of claims 1-5, wherein, the first exchange device is further configured to send a node notification message to the first node, the node notification message comprising an identifier of the second node and third configuration information of the second node; 7. The Ethernet storage system of claim 6, wherein, 8. The Ethernet storage system according to claim 6 or 7, characterized in that, The first node is further configured to receive the node notification message, and acquire and store the identity of the second node and the third configuration information of the second node according to the node notification message.
10. The Ethernet storage system according to any one of claims 1-8, wherein, The first switching device is configured to determine the first partition corresponding to the first node according to a partition table and the identity of the first node, and the partition table records a correspondence between a partition and a node identity.
11. The Ethernet storage system of claim 10, wherein, The first switching device is further configured to: record information of the first node in a dynamic partition table; and determine, according to the dynamic partition table, that the second node is an active node of the first partition, and the dynamic partition table records information of an active node corresponding to each partition.
12. The Ethernet storage system of claim 11, wherein, The information of the active node includes information of an access device of the active node, and the first switching device is further configured to determine a target switching device according to the information of the access device of each active node in the dynamic partition table, and detect whether communication between the first switching device and the target switching device is interrupted.
13. The Ethernet storage system of claim 12, wherein, when the communication between the first switching device and the target switching device is interrupted, the first switching device is further configured to determine a target node corresponding to the target switching device, the target node being a local node connected by the target switching device, and when the local node connected by the target switching device is deleted, the first switching device is configured to delete information of the target node from the dynamic partition table.
14. The Ethernet storage system of claim 11, wherein, The Ethernet storage system further comprises a third switching device, and the third switching device is a reflector, the third switching device is configured to send the partition table to the first switching device; the first switching device is further configured to send the dynamic partition table to the third switching device.
15. The Ethernet storage system of claim 14, wherein, When sending a join notification message to a second node belonging to the first partition, the first switching device is configured to send the join notification message to the third switching device, and when the third switching device is an access device of the second node, the third switching device is configured to send the join notification message to the second node; when the third switching device is not the access device of the second node, the third switching device is configured to send the join notification message to the second node through a fourth switching device, and the second node accesses the fourth switching device.
16. The Ethernet storage system of claim 15, wherein, The Ethernet storage system further comprises a fifth switching device, and the third switching device is further configured to: send the join notification message to the fifth switching device; and send the dynamic partition table to the fourth switching device and the fifth switching device.
17. The Ethernet storage system of any of claims 1-16, wherein the first node is further configured to send a leave message to the first switching device, and the leave message includes an identity of the first node. The first exchange device is further configured to: receive the leave message, obtain the identity of the first node according to the leave message, determine the first partition according to the identity of the first node, and send a leave notification message to a third node belonging to the first partition, the leave notification message comprising the identity of the first node.
18. The Ethernet storage system of any of claims 1-17, wherein, The first deletion notification message comprises the identity of a local node connected to the target exchange device.
19. An information notification method characterized by comprising: The method applied to a first exchange device in an Ethernet storage system, the method comprising: receiving a join message sent by a first node, the join message comprising an identity of the first node and first configuration information of the first node, the first exchange device being an access device of the first node, the first node and the first exchange device being connected through an Ethernet; sending a join notification message to a second node belonging to a same first partition as the first node, the join notification message comprising the identity of the first node and second configuration information of the first node, the second node and the first exchange device being connected through the Ethernet; in the case of communication interruption between the first exchange device and a target exchange device, deleting a local node connected to the target exchange device, and sending a first deletion notification message to a second exchange device.
20. The method of claim 19, wherein, The first exchange device is a reflector, or the first exchange device is not a reflector and the second exchange device is a reflector, the method further comprising: sending the first deletion notification message to other exchange devices through the reflector.
21. The method of claim 19 or 20, wherein, The method further comprises: detecting whether the first node is locally reachable, and when the first node is not locally reachable, determining the first partition, and sending a failure notification message to a fourth node belonging to the first partition, the failure notification message comprising the identity of the first node.
22. The method of claim 21, wherein, The method further comprises: when a condition is met, the first exchange device determining that the first node is not locally reachable, the condition comprising: the first exchange device not receiving a keep-alive message from the first node within a set time; or a port of the first exchange device connected to the first node being disabled.
23. The method of any one of claims 19-22, wherein, The method further comprises: in the case of communication interruption between the first exchange device and the target exchange device, sending a second deletion notification message to a local node connected to the first exchange device, the second deletion notification message comprising the identity of a local node connected to the target exchange device, the local node connected to the first exchange device and the local node connected to the target exchange device belonging to a same partition.
24. The method of any one of claims 19-23, wherein, The first configuration information comprises a role of the first node and / or protocol-related information of the first node.
25. The method of claim 24, wherein, The role of the first node comprises a compute node, a storage node or a general-purpose node, and the general-purpose node is a compute node and a storage node.
26. The method of claim 24 or 25, wherein, The protocol-related information of the first node comprises a protocol supported by the first node, a port used by the protocol, a target identifier length of the protocol and a target identifier of the protocol, the protocol comprising NVMe over RoCE, NVMe over TCP or iSCSI.
27. The method of any one of claims 19-26, wherein, Further comprising: sending a node notification message to the first node, the node notification message comprising an identity of the second node and third configuration information of the second node.
28. The method of any one of claims 19-26, wherein, The method further comprises: determining the first partition corresponding to the first node according to a partition table and the identity of the first node, the partition table recording a correspondence between a partition and an identity of a node.
29. The method of claim 28, wherein, The method further comprises: recording information of the first node in a dynamic partition table; determining, according to the dynamic partition table, that the second node is an active node of the first partition, the dynamic partition table recording information of an active node corresponding to each partition.
30. The method of claim 29, wherein, The information of the active node comprises information of an access device of the active node, and the method further comprises: determining a target switching device according to the information of the access device of each active node in the dynamic partition table, and detecting whether communication between the first switching device and the target switching device is interrupted.
31. The method of claim 30, wherein, When the communication between the first switching device and the target switching device is interrupted, the method further comprises: determining a target node corresponding to the target switching device, the target node being a local node connected to the target switching device; and the deleting the local node connected to the target switching device comprises deleting information of the target node from the dynamic partition table.
32. The method of claim 29, wherein, The Ethernet storage system further comprises a third switching device, the third switching device being a reflector, and the method further comprises: receiving the partition table sent by the third switching device; sending the dynamic partition table to the third switching device.
33. The method of claim 32, wherein, The sending the join notification message to the second node belonging to the same first partition as the first node comprises sending the join notification message to the third switching device.
34. The method of any one of claims 19-33, wherein, The method further comprises: receiving a leave message sent by the first node, the leave message comprising an identity of the first node; determining the first partition according to the identity of the first node; sending a leave notification message to a third node belonging to the first partition, the leave notification message comprising the identity of the first node.
35. The method of any one of claims 19-34, wherein, The first delete notification message comprises an identity of the local node connected to the target switching device.
36. An information announcement method characterized by comprising: The method applied to a first node of an Ethernet storage system comprises: sending a join message to a first switching device of the Ethernet storage system, the join message comprising an identity of the first node and first configuration information of the first node, the first switching device being an access device of the first node, the first node and the first switching device being connected through an Ethernet; receiving a node notification message sent by the first switching device, the node notification message comprising an identity of a second node and third configuration information of the second node, the second node and the first node belonging to the same partition, the second node and the first switching device being connected through the Ethernet; In a case where communication between the first switching device and a target switching device is interrupted, a first deletion notification message sent by the first switching device is received, the first deletion notification message comprising an identifier of a local node connected to the target switching device, the first node and the local node connected to the target switching device belonging to a same partition.
37. The method of claim 36, wherein, The first configuration information comprises a role of the first node and / or protocol-related information of the first node.
38. The method of claim 37, wherein, The role of the first node comprises a compute node, a storage node or a general node, the general node being a compute node and a storage node.
39. The method of claim 37 or 38, wherein, The protocol-related information of the first node comprises a protocol supported by the first node, a port used by the protocol, a target identifier length of the protocol and a target identifier of the protocol, the protocol comprising NVMe over RoCE, NVMe over TCP or iSCSI.
40. The method of any one of claims 36-39, wherein, The method further comprises: obtaining and storing information of the second node, the information of the second node comprising an identifier of the second node and the third configuration information of the second node.
41. The method of claim 40, wherein, The method further comprises: receiving a deletion message sent by the first switching device, the deletion message comprising an identifier of the second node; and deleting the information of the second node according to the deletion message.
42. A switching device, comprising: The switching device is a first switching device in an Ethernet storage system, and the switching device comprises: a receiving unit configured to receive a join message sent by a first node, the join message comprising an identifier of the first node and first configuration information of the first node, the first switching device being an access device of the first node, and the first node and the first switching device being connected through an Ethernet; a sending unit configured to send a join notification message to a second node belonging to a same first partition as the first node, the join notification message comprising the identifier of the first node and second configuration information of the first node, the second node and the first switching device being connected through the Ethernet; the sending unit is further configured to, in a case where communication between the first switching device and a target switching device is interrupted, delete a local node connected to the target switching device, and send a first deletion notification message to a second switching device.
43. The switch device of claim 42, wherein, The first switching device is a reflector, Alternatively, the first switching device is not a reflector, the second switching device is a reflector, and the sending unit is further configured to send the first deletion notification message to other switching devices through the reflector.
44. The switch device of claim 42 or 43, wherein, The first switching device further comprises a processing unit configured to: detect whether the first node is locally reachable, and determine the first partition when the first node is not locally reachable. the sending unit is further configured to send a failure notification message to a fourth node belonging to the first partition, the failure notification message comprising the identifier of the first node.
45. The switch device of claim 44, wherein, The processing unit is further configured to: determine that the first node is not locally reachable when a condition is met, the condition comprising: the first switching device has not received a keep-alive message from the first node within a set time; or The first exchange device connects a port of the first node to fail.
46. The switch device of claim 44 or 45, wherein, The sending unit is further configured to: In the case that the communication between the first exchange device and the target exchange device is interrupted, send a second deletion notification message to a local node connected by the first exchange device, the second deletion notification message comprising an identification of a local node connected by the target exchange device, the local node connected by the first exchange device and the local node connected by the target exchange device belonging to the same partition.
47. The exchange device of any of claims 44-46, wherein, The first configuration information comprises a role of the first node and / or protocol-related information of the first node.
48. The switch device of claim 47, wherein, The role of the first node comprises a compute node, a storage node or a general node, the general node being a compute node and a storage node.
49. The switch device of claim 47 or 48, wherein, The protocol-related information of the first node comprises a protocol supported by the first node, a port used by the protocol, a target identifier length of the protocol and a target identifier of the protocol, the protocol comprising NVMe over RoCE, NVMe over TCP or iSCSI.
50. The exchange device of any of claims 44-49, wherein, The sending unit is further configured to: send a node notification message to the first node, the node notification message comprising an identification of the second node and third configuration information of the second node.
51. The exchange device of claim 50, wherein: The exchange device further comprises a storage unit configured to store a partition table, the partition table recording a correspondence between a partition and an identification of a node; When the first partition corresponding to the first node is determined according to the identification of the first node, the processing unit is configured to determine the first partition corresponding to the first node according to the partition table and the identification of the first node.
52. The switch device of claim 51, wherein, The storage unit is further configured to store a dynamic partition table, the dynamic partition table recording information of an active node corresponding to each partition; The processing unit is further configured to: determine, according to the dynamic partition table, that the second node is an active node of the first partition; and record information of the first node in the dynamic partition table.
53. The switch device of claim 52, wherein, The information of the active node comprises information of an access device of the active node, and the processing unit is further configured to: determine a target exchange device according to the information of the access device of each active node in the dynamic partition table, and detect whether the communication between the first exchange device and the target exchange device is interrupted.
54. The switch device of claim 53, wherein, When the communication between the first exchange device and the target exchange device is interrupted, the processing unit is further configured to: determine a target node corresponding to the target exchange device, the target node being a local node connected by the target exchange device; and 55. The switch device of claim 54, wherein, when deleting the local node connected by the target exchange device, delete information of the target node from the dynamic partition table. The Ethernet storage system further comprises a third exchange device, the third exchange device being a reflector, The receiving unit is further configured to receive the partition table sent by the third exchange device; The sending unit is further configured to send the dynamic partition table to the third exchange device.
56. The switch device of claim 55, wherein, When sending the join notification message to the second node belonging to the first partition, the processing unit is configured to send the join notification message to the third switching device.
57. The switching device of any of claims 44-56, wherein: The receiving unit is further configured to receive a leave message sent by the first node, the leave message comprising an identity of the first node. The processing unit is further configured to determine the first partition according to the identity of the first node. The sending unit is further configured to send a leave notification message to a third node belonging to the first partition, the leave notification message comprising the identity of the first node.
58. The switch device of any of claims 42-57, wherein, The first deletion notification message comprises an identity of a local node connected to the target switching device.
59. A node, comprising: The node is a first node of an Ethernet storage system, and the node comprises: a sending unit configured to send a join message to a first switching device of the Ethernet storage system, the join message comprising an identity of the first node and first configuration information of the first node, the first switching device being an access device of the first node, and the first node and the first switching device being connected through an Ethernet; a receiving unit configured to receive a node notification message sent by the first switching device, the node notification message comprising an identity of a second node and third configuration information of the second node, the second node and the first node belonging to a same partition, and the second node and the first switching device being connected through the Ethernet; The receiving unit is further configured to receive a first deletion notification message sent by the first switching device in a case where communication between the first switching device and a target switching device is interrupted, the first deletion notification message comprising an identity of a local node connected to the target switching device, and the first node and the local node connected to the target switching device belonging to a same partition.
60. The node of claim 59, wherein, The first configuration information comprises a role of the first node and / or protocol-related information of the first node.
61. The node of claim 60, wherein, The role of the first node comprises a compute node, a storage node, or a general-purpose node, and the general-purpose node is a compute node and a storage node.
62. The node of claim 60 or 61, characterized by The protocol-related information of the first node comprises a protocol supported by the first node, a port used by the protocol, a target identifier length of the protocol, and a target identifier of the protocol, and the protocol comprises NVMe over RoCE, NVMe over TCP, or iSCSI.
63. The node of any of claims 59-62, wherein, Further comprising: a processing unit configured to acquire and store information of the second node, the information of the second node comprising the identity of the second node and the third configuration information of the second node.
64. The node of claim 63, wherein: The receiving unit is further configured to receive a deletion message sent by the first switching device, the deletion message comprising the identity of the second node. The processing unit is further configured to delete the information of the second node according to the deletion message.
65. An Ethernet switching device comprising: The Ethernet switching device is deployed in an Ethernet storage system, and comprises a memory and a processor coupled with the memory, the memory is configured to store computer program code, and the processor is configured to invoke the computer program code to enable the Ethernet switching device to perform the information notification method in any one of claims 19-35.
66. A host, comprising: The host is deployed in an Ethernet storage system, and comprises a memory and a processor coupled with the memory, the memory is configured to store computer program code, and the processor is configured to invoke the computer program code to enable the node to perform the information notification method in any one of claims 36-41.
67. A computer-readable storage medium, characterized in that, The computer readable storage medium stores instructions, and when the instructions run on the processor, the method in any one of claims 19-41 is implemented.
68. A computer program product, characterised in that, The computer readable storage medium stores instructions, and when the instructions run on the processor, the method in any one of claims 19-41 is implemented.
Citation Information
Patent Citations
An Ethernet storage system, its information notification method, and related apparatus.
CN113810439B
Ethernet storage system and information notification method and related device thereof
CN116192879A
Hard zoning on NPIV proxy / NPV devices
US20110022693A1
Network management using port announcements
US20170034008A1