Storage node management methods and electronic devices
By automatically switching the hardware address and receiving and forwarding management requests when the main management controller detects an anomaly from the management controller, the problem of low storage node management efficiency is solved, and efficient and automated storage node management is achieved.
Patent Information
- Application Number
- CN202511236563.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-01
- Publication Date
- 2025-12-02
- Estimated Expiration
- 2045-09-01
AI Technical Summary
Existing technologies have low storage node management efficiency, and manual intervention and complex recovery operations are required when the main management port fails, resulting in low management efficiency.
After detecting an abnormal state of the primary management controller, the hardware address of the secondary management controller is automatically switched to that of the primary management controller. By receiving and forwarding management requests from the secondary management controller, the continuity and efficiency of storage node management are ensured.
It enables automated fault detection and switching for storage node management, improving the continuity of management operations, response speed, and resource utilization, and enhancing system stability and availability.
Smart Images

Figure CN120743200B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, and more specifically, to a method for managing storage nodes and an electronic device. Background Technology
[0002] With the growth of data volume and the increasing demand for data storage, the management of storage nodes is becoming more and more complex and important. Current technologies typically manage storage nodes by configuring two management ports on each node. One port is used for routine management and configuration, while the other serves as a backup in case the primary management port fails.
[0003] However, when the storage node management method provided by the above-mentioned existing technology fails, although theoretically the management function of the storage node can be maintained by switching to the backup management port, in practice this process often requires manual intervention by the operation and maintenance personnel and a series of complex recovery operations, which takes a lot of time and leads to the technical problem of low management efficiency of the storage node.
[0004] Therefore, existing technologies suffer from the technical problem of low management efficiency of storage nodes. Summary of the Invention
[0005] This application provides a storage node management method and electronic device to at least solve the technical problem of low management efficiency of storage nodes.
[0006] According to one embodiment of this application, a method for managing storage nodes is provided, comprising: determining that the master management controller is in an abnormal state when the slave management controller of the storage node set does not receive a status prompt signal sent by the master management controller of the storage node set within a predetermined time window; switching the hardware address of the second management port of the slave management controller to a first hardware address corresponding to the first management port of the master management controller, wherein the first hardware address is bound to a first network address used to access the storage node set; receiving a first request for managing a first storage node in the storage node set through the second management port, wherein the destination network address of the first request is the first network address; and forwarding the first request to the first storage node through the connection link between the second management port and the first storage node, so that the first storage node performs a management operation matching the first request.
[0007] Optionally, receiving a first request for managing a first storage node in the storage node set via the second management port includes: receiving the first request using the first sub-management port when both the first and second sub-management ports in the second management port are in normal condition; or receiving the first sub-data of the first request using the first sub-management port and the second sub-data of the first request using the second sub-management port when both the first and second sub-management ports are in normal condition; or receiving the first request using the second sub-management port when the first sub-management port is in an abnormal condition; wherein the second management port is a logical port including the first sub-management port and the second management port.
[0008] Optionally, forwarding the first request to the first storage node via the connection link between the second management port and the first storage node includes: sending the first request to the first storage node using the connection link between the first sub-management port and the first storage node; or sending first sub-data to the first storage node using the connection link between the first sub-management port and the first storage node, and sending second sub-data to the first storage node using the connection link between the second sub-management port and the first storage node; or sending the first request to the first storage node using the connection link between the second sub-management port and the first storage node.
[0009] Optionally, determining that the master management controller is in an abnormal state when the slave management controller of the storage node set does not receive a status prompt signal from the master management controller of the storage node set within a predetermined time window includes: determining that the master management controller is in an abnormal state when the slave management controller does not receive a status prompt signal from the master management controller through the communication link with the master management controller within N status detection cycles, where N is a positive integer.
[0010] Optionally, before determining that the master management controller is in an abnormal state, if the slave management controller of the storage node set does not receive a status prompt signal sent by the master management controller of the storage node set within a predetermined time window, the method further includes: binding the first sub-management port and the second sub-management port of the slave management controller to obtain a second management port; and binding the third sub-management port and the fourth sub-management port of the master management controller to obtain a first management port.
[0011] Optionally, before determining that the main management controller is in an abnormal state, if the slave management controller of the storage node set does not receive a status prompt signal sent by the master management controller of the storage node set within a predetermined time window, the method further includes: connecting the third sub-management port to the first local area network through the first switch chip of the main management controller, and connecting the fourth sub-management port to the second local area network through the first switch chip, wherein the third sub-management port and the fourth sub-management port are respectively connected to the first switch chip.
[0012] Optionally, before determining that the master management controller is in an abnormal state if the slave management controller of the storage node set does not receive a status prompt signal from the master management controller of the storage node set within a predetermined time window, the method further includes: connecting the first sub-management port to the first local area network through the second switch chip of the slave management controller, and connecting the second sub-management port to the second local area network through the second switch chip, wherein the first sub-management port and the second sub-management port are respectively connected to the second switch chip.
[0013] Optionally, the above-mentioned storage node management method further includes: when the main management controller is in a normal state, receiving a second request for managing the first storage node through a first management port, wherein the destination network address of the second request is the first network address; and forwarding the second request to the first storage node through the connection link between the first management port and the first storage node, so that the first storage node performs a management operation matching the second request.
[0014] According to another embodiment of this application, a storage node management system is provided, including a master management controller, a slave management controller, and a storage node set. The master management controller and the slave management controller are connected to a target general-purpose input / output interface. A first management port of the master management controller establishes a connection link with each node in the storage node set, and a second management port of the slave management controller establishes a connection link with each node. In a normal state, the master management controller receives a second request through its first management port for managing a first storage node in the storage node set. The destination network address of the second request is a first network address bound to a first hardware address corresponding to the first management port. The second request connects to the first storage node through the first management port. The connection link between nodes is also used to send the second request to the first storage node; the management controller is used to determine that the main management controller is in an abnormal state if it does not receive a status prompt signal from the main management controller within a predetermined time window, and switches the hardware address of the second management port of the management controller to the first hardware address, and receives the first request for managing the first storage node through the second management port, wherein the destination network address of the first request is the first network address, and forwards the first request to the first storage node through the connection link between the second management port and the first storage node; the first storage node in the storage node set is used to perform management operations matching the second request, and to perform management operations matching the first request.
[0015] According to another embodiment of this application, a storage node management device is provided, comprising: a determining unit, configured to determine that the master management controller is in an abnormal state when the slave management controller of the storage node set does not receive a status prompt signal sent by the master management controller of the storage node set within a predetermined time window; a switching unit, configured to switch the hardware address of the second management port of the slave management controller to a first hardware address corresponding to the first management port of the master management controller, wherein the first hardware address is bound to a first network address used to access the storage node set; a receiving unit, configured to receive a first request for managing a first storage node in the storage node set through the second management port, wherein the destination network address of the first request is the first network address; and a forwarding unit, configured to forward the first request to the first storage node through the connection link between the second management port and the first storage node, so that the first storage node performs a management operation matching the first request.
[0016] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0017] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0018] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps of the methods described in various embodiments of this application.
[0019] According to the embodiments provided in this application, if the slave management controller of the storage node set does not receive a status notification signal from the master management controller of the storage node set within a predetermined time window, it is determined that the master management controller is in an abnormal state. The hardware address of the slave management controller's second management port is switched to the first hardware address corresponding to the master management controller's first management port, wherein the first hardware address is bound to a first network address used to access the storage node set. A first request for managing a first storage node in the storage node set is received through the second management port, wherein the destination network address of the first request is the first network address. The first request is forwarded to the first storage node through the connection link between the second management port and the first storage node, so that the first storage node performs a management operation matching the first request. By adopting the embodiments of this application, the technical effect of improving the management efficiency of storage nodes is achieved, solving the technical problem of low management efficiency of storage nodes. Attached Figure Description
[0020] Figure 1 This is a flowchart of a storage node management method according to an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of a storage node management method according to an embodiment of this application;
[0022] Figure 3 This is a schematic diagram of another storage node management method according to an embodiment of this application;
[0023] Figure 4 This is a structural block diagram of a storage node management device according to an embodiment of this application. Detailed Implementation
[0024] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] As an optional solution, the specific steps of the above-mentioned storage node management method are as follows: Figure 1 The following are included:
[0027] S102, if the slave management controller of the storage node set does not receive a status prompt signal from the master management controller of the storage node set within a predetermined time window, it is determined that the master management controller is in an abnormal state.
[0028] It should be noted that the above-described storage node management methods can be applied, but are not limited to, storage node management scenarios in centralized storage systems. The primary management controller mentioned above can be, but is not limited to, the management controller primarily responsible for system management and data communication in a multi-controller storage system. It is normally active, handling all management requests and data transmissions. Secondary management controllers: In a multi-controller storage system, these are the other management controllers besides the primary management controller. They are typically in standby or auxiliary status, ready to take over management responsibilities in the event of a failure of the primary management controller.
[0029] Specifically, the aforementioned master management controller can be, but is not limited to, a Chassis Management Controller (CMC), which integrates processing power and network interfaces, enabling it to detect and control all controller nodes within the storage chassis. Correspondingly, the aforementioned slave management controller can also be, but is not limited to, the central management unit in the storage management system, responsible for receiving and processing management requests. Specifically, the aforementioned slave management controller can be, but is not limited to, a Chassis Management Controller (CMC).
[0030] Optionally, the aforementioned predetermined time window may be, but is not limited to, a time interval, used to determine whether a status alert signal sent by the main management controller has been received. If no status alert signal is received within this time window, the main management controller is considered to be in an abnormal state. The aforementioned status alert signal is a signal periodically sent by the main management controller to the slave management controller to indicate its current operating status, including but not limited to normal operation, fault, and maintenance status.
[0031] It should be noted that step S102 above can be implemented, but is not limited to, through a heartbeat link (such as a shared backplane General Purpose Input / Output (GPIO)). Specifically, the master and slave management controllers are configured to share GPIO, allowing them to detect each other's status. Specifically, a set of GPIO interface pairs is configured between the master and slave controllers via the system hardware backplane, with each controller having one GPIO interface. This is typically determined during the system design phase to ensure that GPIO signals can be transmitted directly and quickly between the two management controllers. The master management controller is then configured with a timed task or a dedicated hardware detection mechanism to periodically send a short status signal (also known as a heartbeat signal) to the slave management controller via the GPIO interface. This signal requires no complex software processing; it simply indicates that the master controller is currently operating normally. The slave management controller also has a corresponding mechanism to detect the GPIO interface and receive the status signal from the master management controller. If no signal is received within a preset period, the slave management controller assumes that the master management controller may have encountered a fault or abnormality.
[0032] It should be noted that the first storage node can be, but is not limited to, a node in the storage node set; it is the object that responds to and processes management requests. The storage node set includes multiple storage nodes.
[0033] Furthermore, the aforementioned abnormal states can be used, but are not limited to, to indicate a physical network interface or connection link malfunction, signal interruption, configuration error, or other conditions that prevent normal operation.
[0034] S104, the hardware address of the second management port of the management controller is switched to the first hardware address corresponding to the first management port of the main management controller, wherein there is a binding relationship between the first hardware address and the first network address used to access the storage node set.
[0035] Optionally, the aforementioned second management port may, but is not limited to, be two or more physical network interfaces bound together into a single logical interface using software technology (such as link aggregation) to provide a unified management access point externally. For example, two physical ports, eth0 and eth1, are bound together into a single logical port, bond0, through a bonding mechanism. When external systems access the storage system for management operations, they will interact with the storage system through the single network address of bond0, while internal traffic is distributed between eth0 and eth1 according to a set policy. Accordingly, the aforementioned second management port is similar to the aforementioned first management port, and will not be described in detail in this embodiment.
[0036] It should be noted that the aforementioned hardware address may be, but is not limited to, a Media Access Control Address (MAC), which is the physical address of a network device (such as a network interface card) used for identification and communication at the network layer. The aforementioned first network address may be, but is not limited to, a network address used to access the storage node set, i.e., such as a virtual IP address, bound to the first management port for unified management of access to the storage node set.
[0037] S106, receive a first request through the second management port for requesting management of the first storage node in the storage node set, wherein the destination network address of the first request is the first network address.
[0038] Optionally, the aforementioned first request may be, but is not limited to, an instruction issued by a user account (e.g., an operations and maintenance personnel) or an automated system to perform management operations on the storage node, such as configuration changes, status queries, and troubleshooting. For example, assuming a user needs to change the RAID configuration of a storage node, the first request issued through the management software may include specific parameters and the identifier of the target storage node, such as "change the RAID type of node 1 from RAID5 to RAID10".
[0039] It should be noted that step S106 above can be achieved, but is not limited to, through MAC address shifting technology.
[0040] S108, through the connection link between the second management port and the first storage node, the first request is forwarded to the first storage node so that the first storage node can perform a management operation matching the first request.
[0041] Optionally, the aforementioned connection link may be, but is not limited to, a physical network connection between the management controller and the storage node, such as fiber optic cable, Ethernet cable, etc., for the transmission of data and management requests.
[0042] As an optional example, it can be based on, but is not limited to, such as Figure 2The following example illustrates the steps above:
[0043] like Figure 2 As shown, assume that the storage system includes four storage nodes, namely node 1, node 2, node 3, and node 4, and that there are two CMCs configured in the storage system, namely CMC1 and CMC2. CMC1 includes ports U1 and T1, and CMC2 includes ports U2 and T2. It is necessary to create connection links from node 1, node 2, node 3, and node 4 to U1, T1, U2, and T2 respectively.
[0044] It should be noted that the accuracy of ARP detection can be enhanced by configuring kernel parameters for Address Resolution Protocol (ARP) security, such as enabling arp_validate=1 and arp_announce=2; setting fail_over_mac=1 can prevent Media Access Control Address (MAC) address drift and avoid ARP entry invalidation. The dual CMC management cards synchronize their status via a dedicated heartbeat link (such as a shared backplane General Purpose Input / Output (GPIO)). In the event of a primary CMC failure, automatic switchover to the backup CMC is achieved.
[0045] In this embodiment, if the slave management controller of the storage node set does not receive a status alert signal from the master management controller of the storage node set within a predetermined time window, it determines that the master management controller is in an abnormal state. The hardware address of the slave management controller's second management port is switched to the first hardware address corresponding to the first management port of the master management controller. The first hardware address is bound to a first network address used to access the storage node set. A first request for managing a first storage node in the storage node set is received through the second management port, where the destination network address of the first request is the first network address. The first request is forwarded to the first storage node through the connection link between the second management port and the first storage node, so that the first storage node performs a management operation matching the first request. In other words, using this embodiment, the slave management controller automatically determines the operating status of the master management controller within a predetermined time window by detecting whether it receives a status alert signal from the master management controller. If no signal is received within a preset time, it is determined that the master management controller may be in an abnormal state. This is faster and more accurate than traditional manual detection methods and requires no manual intervention. If the primary management controller is determined to be in an abnormal state, the secondary management controller immediately switches its hardware address (e.g., MAC address), modifying its own second management port hardware address to the hardware address of the primary management controller's first management port. This ensures that external users use a consistent network address (first network address) to access the storage node set, avoiding access interruptions caused by hardware address changes. The secondary management controller forwards management requests to the target storage node (first storage node) through the established connection link, and the storage node executes the management operation matching the request. This avoids the delays and complexities that may exist in traditional manual switching processes, ensuring the continuity and efficiency of management operations. In summary, the embodiments of this application solve the problem of low storage node management efficiency in the prior art, provide an automated fault detection and switching method, improve the continuity of management operations, response speed, and resource utilization, and enhance the stability and availability of the system.
[0046] As an optional approach, receiving a first request for managing a first storage node in the storage node set via the second management port includes:
[0047] When both the first sub-management port and the second sub-management port in the second management port are in normal condition, the first sub-management port is used to receive the first request; or when both the first sub-management port and the second sub-management port are in normal condition, the first sub-management port is used to receive the first sub-data of the first request, and the second sub-management port is used to receive the second sub-data of the first request; or when the first sub-management port is in an abnormal condition, the second sub-management port is used to receive the first request.
[0048] The second management port is a logical port that includes the first sub-management port and the second sub-management port.
[0049] It should be noted that the first sub-management port mentioned above can be, but is not limited to, a User Management Port (U port) or a Technical Operations Port (T port). These two ports are used for user access and system management and maintenance by operations and maintenance personnel, respectively, within the storage device. The second sub-management port mentioned above can be, but is not limited to, a User Management Port (U port) or a Technical Operations Port (T port).
[0050] Specifically, if the first sub-management port is a U port, the second sub-management port is a T port, and vice versa.
[0051] It should be noted that the above normal states can be used, but are not limited to, to indicate that the management port can receive and send data normally without any faults or abnormalities. The above abnormal states can be used, but are not limited to, to indicate that the management port cannot receive or send data normally, possibly due to hardware failure, network problems, or software errors.
[0052] Optionally, the first sub-data and the second sub-data divide the first request data into two parts so that they can be transmitted concurrently through the first sub-management port and the second sub-management port.
[0053] As an optional example, the above steps can be illustrated using, but not limited to, the following examples:
[0054] Suppose there is a system with four storage nodes, each configured with a USB port and a USB-T port. In the system, the first logical management port on the first management controller (CMC1) binds the USB port and the USB-T port together through software-level link aggregation (bond), where the USB port serves as the primary management link and the USB-T port serves as a redundant link.
[0055] Scenario 1: Single-link transmission:
[0056] When the first sub-management port (U port) is in normal working condition, the first request (such as updating the software version on the storage node) will be sent entirely through the U port to the first storage node. The first request initiated by the administrator or automated management software arrives at the U port through bond0 on CMC1 (i.e., the first logical management port), and then the U port transmits all request data to the first storage node. After receiving the request, the node performs the corresponding operation, such as software update.
[0057] Scenario 2: Concurrent Link Transmission:
[0058] With the first sub-management port (U port) functioning normally, the first request can be split into a first sub-data (such as the first half of the request) and a second sub-data (the second half of the request). The first sub-data will be sent via the U port, while the second sub-data will be transmitted concurrently via the second sub-management port (T port). For example, when an administrator needs to check the health status of a large storage cluster, the request can be split into multiple small data packets, with one part sent via the U port and the other via the T port. With both ports operating simultaneously, request processing time can be significantly reduced, improving management efficiency.
[0059] In this embodiment, when both the first and second sub-management ports in the second management port are in normal operation, the first sub-management port is used to receive the first request; or when both the first and second sub-management ports are in normal operation, the first sub-management port is used to receive the first sub-data of the first request, and the second sub-management port is used to receive the second sub-data of the first request; or when the first sub-management port is in an abnormal operation, the second sub-management port is used to receive the first request; wherein, the second management port is a logical port including the first and second sub-management ports. In other words, by designing the second management port as a logical port including the first and second sub-management ports, even if the first sub-management port fails, the second sub-management port can still receive management requests, ensuring the continuity and high availability of system management. Compared with the traditional single-port design, this better addresses network-level faults and avoids management interruptions caused by single-point failures. When both sub-management ports are in normal operation, the first request is divided into first sub-data and second sub-data, which are received through the two sub-ports respectively. This strategy facilitates load balancing, allowing management requests to be distributed across the two sub-ports, preventing overloading of one port and improving overall processing efficiency and response speed. In the event of a failure in the first sub-management port, the system automatically switches to the second sub-management port to receive management requests, simplifying the manual fault recovery process for operations personnel, reducing the risk of human error, and ensuring timely response to management requests, thus improving system stability and operational efficiency. In summary, the embodiments described in this application significantly improve the reliability, efficiency, and security of storage network management, while simplifying the management process and enhancing system flexibility and scalability, making it invaluable for building highly available storage management architectures.
[0060] As an optional solution, forwarding the first request to the first storage node via the connection link between the second management port and the first storage node includes:
[0061] The first request is sent to the first storage node using the connection link between the first sub-management port and the first storage node; or the first sub-data is sent to the first storage node using the connection link between the first sub-management port and the first storage node, and the second sub-data is sent to the first storage node using the connection link between the second sub-management port and the first storage node; or the first request is sent to the first storage node using the connection link between the second sub-management port and the first storage node.
[0062] It should be noted that, in this embodiment, both the first sub-management port and the second sub-management port in the second management port establish connection links with each storage node in the storage node set. These connection links can be, but are not limited to, physical connections used for transmitting data or management commands. They can be direct cable connections or indirect connections through intermediate devices (such as switches), and this embodiment does not impose any limitations on this.
[0063] For example, but not limited to, using RJ45 connectors or Small Form-factor Pluggable (SFP) fiber optic interfaces and cables, the U-port and T-port of the CMC management card can be connected to the management network ports of each controller motherboard (i.e., each storage node), ensuring that each controller is connected to the system through two physical links (U-port and T-port).
[0064] It should be noted that for the specific configuration of the main management controller, please refer to the configuration of the slave management controller; this will not be repeated in this embodiment.
[0065] Alternatively, as an optional example, the hardware framework of the above-described storage node management method may, but is not limited to, referencing... Figure 3 ,like Figure 3 As shown, two switching chips (SW1 and SW2) are mounted on a separate CMC management card. This means the switching chips are integrated within the CMC card. Each switching chip provides two gigabit channels: ETH1 and ETH2; in other words, each card provides both a T-port and a U-port. For physical connections, the T-port and U-port connect to the data center switch (customer's data center switch). Simultaneously, backplane cabling connects to the management ports of storage controller motherboards 1 through 4, allowing a single CMC card to manage all nodes. At the logical / software level, within the CMC's Linux system, eth0 (T-port) and eth1 (U-port) are bound together as bond0. For VLAN isolation, VLAN 10 (U-port) and VLAN 20 (T-port) are defined internally within the switching chip. Figure 3 Taking the interaction via USB as an example, traffic isolation is implemented at the hardware level.
[0066] In this embodiment, the first request is sent to the first storage node using the connection link between the first sub-management port and the first storage node; or the first sub-data is sent to the first storage node using the connection link between the first sub-management port and the first storage node, and the second sub-data is sent to the first storage node using the connection link between the second sub-management port and the first storage node; or the first request is sent to the first storage node using the connection link between the second sub-management port and the first storage node. In other words, using this embodiment, when the first management port fails, the complete first request is sent directly to the target storage node using the second management port. In this mode, the second management port assumes all management request transmission responsibilities, ensuring the continuity of management operations even in the event of partial link failure. The first request is divided into first sub-data and second sub-data, which are sent to the target storage node through the first sub-management port and the second sub-management port, respectively. This mode can improve data transmission efficiency and distribute network load when both sub-management ports are working normally. When the first sub-management port also fails, the second sub-management port is used as a backup transmission path to send the first request to the target storage node. This model provides a multi-layered redundancy design, ensuring that management requests can still reach the target node in the event of multiple failures, thus improving the system's fault tolerance and availability.
[0067] As an optional approach, if the slave management controller of the storage node set does not receive a status alert signal from the master management controller of the storage node set within a predetermined time window, determining that the master management controller is in an abnormal state includes:
[0068] If the slave controller fails to receive any status alert signal from the master controller within N status detection cycles through the communication link with the master controller, it is determined that the master controller is in an abnormal state, where N is a positive integer.
[0069] It should be noted that the status detection period can be, but is not limited to, a periodic detection time interval set in storage system management to detect the status of the Master Management Controller. During this period, the slave management controller will actively or passively check whether it has received a status alert signal from the master management controller.
[0070] Furthermore, the aforementioned communication link can be implemented through the sharing of backplane GPIO between different management controllers as described above, or it can be a network connection path connecting the master management controller and the slave management controller. It can be a physical line (such as fiber optic cable or Ethernet cable) or a logical link (such as an indirect connection through a switch). In this embodiment, no limitations are imposed on the comparison.
[0071] Furthermore, N is a positive integer used to define the threshold for status detection failure. If the management controller fails to receive a status indication signal from the main management controller within N consecutive status detection cycles, the main management controller is determined to be in an abnormal state.
[0072] In this embodiment, if the slave controller fails to receive a status alert signal from the master controller via the communication link within N status detection cycles, it is determined that the master controller is in an abnormal state, where N is a positive integer. In other words, by continuously detecting status alert signals, this embodiment allows for rapid identification of faults when the master controller malfunctions, avoiding prolonged waiting and uncertainty, thereby improving the reliability and response speed of the entire system.
[0073] As an optional approach, before determining that the master management controller is in an abnormal state, if the slave management controller of the storage node set does not receive a status alert signal from the master management controller of the storage node set within a predetermined time window, the following steps are also included:
[0074] Bind the first sub-management port of the management controller to the second sub-management port to obtain the second management port.
[0075] The first management port is obtained by binding the third sub-management port and the fourth sub-management port of the main management controller.
[0076] It should be noted that the first and second sub-management ports can be bound using binding tools, but not limited to these tools. For example, in the CMC management card operating system (such as Linux), the U port (eth0) and T port (eth1) can be bound as logical interface bond0 using the IP LinkCommand (ip link) or NetworkManager Command LineInterface (nmcli) tools, employing the Link Aggregation Control Protocol (LACP) mode (mode=4) and setting the time period miimon=100 (100ms heartbeat detection).
[0077] Furthermore, the binding method between the third and fourth sub-management ports mentioned above can be found in the example above, and will not be repeated in this embodiment.
[0078] It should be noted that the third and fourth sub-management ports within the first management port also establish connection links with each storage node in the storage node set. For specific connection methods, please refer to the connection example of the first and second sub-management ports above; this will not be repeated in this embodiment.
[0079] In this embodiment, the first and second sub-management ports of the slave management controller are bound together to form the second management port; the third and fourth sub-management ports of the master management controller are bound together to form the first management port. In other words, by binding multiple physical ports into a single logical port, link aggregation can be achieved, meaning the bandwidth of multiple sub-ports can be superimposed, providing a higher data transmission rate. Simultaneously, this binding mechanism also provides link redundancy; when one sub-port fails, management requests can automatically switch to another functioning sub-port, thereby improving link reliability and stability.
[0080] As an optional approach, before determining that the master management controller is in an abnormal state, if the slave management controller of the storage node set does not receive a status alert signal from the master management controller of the storage node set within a predetermined time window, the following steps are also included:
[0081] The third sub-management port is connected to the first local area network through the first switch chip of the main management controller, and the fourth sub-management port is connected to the second local area network through the first switch chip. The third and fourth sub-management ports are respectively connected to the first switch chip.
[0082] It should be noted that the first sub-management port mentioned above can be, but is not limited to, a User Management Port (U port) or a Technical Operations Port (T port). These two ports are used for user access and system management and maintenance by operations and maintenance personnel, respectively, within the storage device. The second sub-management port mentioned above can be, but is not limited to, a User Management Port (U port) or a Technical Operations Port (T port).
[0083] It should be noted that a switch chip is a core component of a network device, responsible for forwarding and managing data packets, and capable of functions such as Virtual Local Area Network (VLAN) segmentation and link aggregation.
[0084] For example, but not limited to, two CMC management cards can be installed in the storage chassis, each with a U-port (general management port) and a T-port (dedicated management port), connected to the switch chip via Peripheral Component Interconnect Express (PCIe) or a dedicated backplane. The switch chip is configured with independent VLANs (e.g., VLAN 10 for the U-port and VLAN 20 for the T-port) to achieve U / T port traffic isolation.
[0085] In this embodiment, the third sub-management port is connected to the first local area network (LAN) via the first switch chip of the main management controller, and the fourth sub-management port is connected to the second LAN via the same first switch chip. The third and fourth sub-management ports are respectively connected to the first switch chip. In other words, by using this embodiment, on the one hand, by connecting the sub-management ports to different LANs, even if one LAN fails, the other LAN can still maintain the transmission of management requests, thereby improving the network-level redundancy of the entire management system, reducing the possibility of single points of failure, and enhancing the stability and availability of the system. On the other hand, connecting each sub-management port to an independent LAN helps to isolate and manage network traffic, prevents interference between different management requests, and also facilitates independent configuration and optimization for different network environments.
[0086] As an optional approach, before determining that the master management controller is in an abnormal state, if the slave management controller of the storage node set does not receive a status alert signal from the master management controller of the storage node set within a predetermined time window, the following steps are also included:
[0087] The first sub-management port is connected to the first local area network through the second switch chip of the management controller, and the second sub-management port is connected to the second local area network through the second switch chip, wherein the first sub-management port and the second sub-management port are respectively connected to the second switch chip.
[0088] It should be noted that a switch chip is a core component of a network device, responsible for forwarding and managing data packets, and capable of functions such as Virtual Local Area Network (VLAN) segmentation and link aggregation.
[0089] For example, but not limited to, two CMC management cards can be installed in the storage chassis, each with a U-port (general management port) and a T-port (dedicated management port), connected to the switch chip via Peripheral Component Interconnect Express (PCIe) or a dedicated backplane. The switch chip is configured with independent VLANs (e.g., VLAN 10 for the U-port and VLAN 20 for the T-port) to achieve U / T port traffic isolation.
[0090] In this embodiment, the first sub-management port is connected to the first local area network (LAN) via the second switch chip of the management controller, and the second sub-management port is connected to the second LAN via the second switch chip. The first and second sub-management ports are respectively connected to the second switch chip. In other words, by adopting this embodiment, on the one hand, by connecting the first and second sub-management ports to different LANs (the first LAN and the second LAN), even if one LAN fails, the other LAN can still support management operations, significantly enhancing network redundancy and stability. On the other hand, connecting each sub-management port to an independent LAN helps isolate and manage network traffic, prevents interference between different management requests, and also facilitates independent configuration and optimization for different network environments.
[0091] As an optional solution, the above-mentioned storage node management method also includes:
[0092] When the main management controller is in a normal state, it receives a second request for managing the first storage node through the first management port, wherein the destination network address of the second request is the first network address;
[0093] The second request is forwarded to the first storage node through the connection link between the first management port and the first storage node, so that the first storage node can perform a management operation that matches the second request.
[0094] It should be noted that the second request mentioned above corresponds to the first request. When the main management controller is in a normal state, it is an instruction or data sent by the client or maintenance personnel via the network to request management of the first storage node in the storage node set. The destination network address of the second request matches the first network address bound to the first management port.
[0095] Optionally, the aforementioned connection link is a physical or logical network connection between the first management port and the first storage node, used to transmit management requests and related data. It can be a direct RJ45 or SFP+ cable connection, or an indirect connection through a switch chip. This embodiment does not limit this.
[0096] In this embodiment, when the master management controller is in a normal state, it receives a second request for managing the first storage node through the first management port, wherein the destination network address of the second request is the first network address. The second request is forwarded to the first storage node through the connection link between the first management port and the first storage node, so that the first storage node performs a management operation matching the second request. In other words, by distinguishing between the normal and abnormal states of the master and slave management controllers, this embodiment can flexibly adjust the processing method of management requests in different states, ensuring efficient operation in normal states and providing redundancy and recovery paths in abnormal states, demonstrating the system's high flexibility and adaptability.
[0097] As an optional solution, the above-mentioned storage node management methods also include:
[0098] Intelligent predictive maintenance: While continuously monitoring the health status of all storage nodes in the storage node set through the first logical management port, it predicts potential hardware or software failures by analyzing historical data and real-time performance indicators, and automatically triggers preventive measures.
[0099] Dynamic resource scheduling: Based on the results of predictive maintenance, intelligently adjust the resource allocation in the storage node set, such as migrating data to healthier storage nodes or pre-allocating additional computing resources to nodes that will soon bear more load.
[0100] Optionally, but not limited to, the above steps can be illustrated with examples:
[0101] S1, Data Collection and Analysis: Utilize the first logical management port (such as bond0) of the main management controller to continuously collect performance data (such as CPU utilization, disk I / O, network latency) and health indicators (such as temperature and voltage fluctuations) of storage nodes.
[0102] S2, Fault Prediction Algorithm: Develop or integrate a set of machine learning models that are trained on historical data and can identify early signs of failure and predict the probability of future failures of storage nodes.
[0103] S3, Early Warning and Response Mechanism: When the predicted failure probability of a storage node exceeds a preset threshold, the main management controller generates an early warning message and automatically executes a dynamic resource scheduling strategy.
[0104] S4, Resource Scheduling: Automatically migrates data from storage nodes to healthy nodes, or adjusts the allocation of computing resources to ensure that the overall performance of the storage cluster is not affected by the failure of a single node.
[0105] By employing the embodiments of this application, intelligent analysis, which integrates performance data and health indicators, can detect potential problems earlier and reduce the impact of sudden failures on business operations.
[0106] This embodiment also provides a storage node management device for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0107] As an optional example, the above-mentioned storage node management method can be illustrated using examples, but not limited to the following:
[0108] 1) Regarding hardware configuration, a dual CMC management card deployment is adopted. Specifically, two CMC management cards are installed in the storage chassis, each card having a U-port (general management port) and a T-port (dedicated management port) respectively, connected to the switch chip (such as the Broadcom Trident series) via PCIe or a dedicated backplane. The switch chip is configured with independent VLANs (e.g., VLAN 10 for the U-port and VLAN 20 for the T-port) to achieve U / T port traffic isolation. Furthermore, using a combination of RJ45 or SFP cables, the U-port and T-port of the CMC management card are connected to the management network ports of each controller motherboard, ensuring that each controller accesses the system through two physical links (U-port and T-port).
[0109] 2) For software configuration, a network port bonding and redundancy strategy is adopted. Specifically, in the CMC management card operating system (such as Linux), the U port (eth0) and T port (eth1) are configured to be bonded as logical interface bond0 through the iplink or nmcli tool, using LACP mode (mode=4), and miimon=100 (100ms heartbeat detection).
[0110] 3) For redundant activation, the accuracy of ARP detection is enhanced by enabling the parameters arp_validate=1 and arp_announce=2; MAC address drift is implemented by setting fail_over_mac=1 to avoid ARP entry failure. The dual CMC management cards synchronize their status through a dedicated heartbeat link (such as a shared backplane GPIO), and automatically switch to the backup CMC when the primary CMC fails.
[0111] Furthermore, it should be noted that the hardware architecture can be smoothly expanded to a larger number of controllers (e.g., 8) without requiring network refactoring. By adding switch chips and the Virtual Router Redundancy Protocol (VRRP), it can be further upgraded to a dual-switch redundancy architecture.
[0112] As an optional solution, the management system for the aforementioned storage nodes includes:
[0113] This includes a master management controller, slave management controllers, and a set of storage nodes. The master and slave management controllers are connected to a target general-purpose input / output interface. A connection link is established between the first management port of the master management controller and each node in the set of storage nodes. A connection link is established between the second management port of the slave management controller and each node.
[0114] When the main management controller is in a normal state, it is used to receive a second request for managing a first storage node in the storage node set through a first management port, wherein the destination network address of the second request is a first network address bound to a first hardware address corresponding to the first management port, and is also used to send the second request to the first storage node through the connection link between the first management port and the first storage node.
[0115] If the primary management controller does not receive a status alert signal from the primary management controller within a predetermined time window, and determines that the primary management controller is in an abnormal state, it switches the hardware address of the secondary management port of the primary management controller to the primary hardware address, receives a first request for management of the primary storage node through the secondary management port, wherein the destination network address of the first request is the primary network address, and forwards the first request to the primary storage node through the connection link between the secondary management port and the primary storage node.
[0116] The first storage node in the storage node set is used to perform management operations that match the second request, and to perform management operations that match the first request.
[0117] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0118] Figure 4 This is a structural block diagram of a storage node management device according to an embodiment of this application, such as... Figure 4 As shown, the device includes:
[0119] The determining unit 402 is used to determine that the main management controller is in an abnormal state when the slave management controller of the storage node set does not receive a status prompt signal sent by the main management controller of the storage node set within a predetermined time window.
[0120] The switching unit 404 is used to switch the hardware address of the second management port of the master management controller to the first hardware address corresponding to the first management port of the master management controller, wherein there is a binding relationship between the first hardware address and the first network address used to access the storage node set.
[0121] The receiving unit 406 is configured to receive a first request for managing a first storage node in the storage node set through a second management port, wherein the destination network address of the first request is a first network address.
[0122] The forwarding unit 408 is used to forward the first request to the first storage node through the connection link between the second management port and the first storage node, so that the first storage node can perform a management operation matching the first request.
[0123] Optionally, in this embodiment, the receiving unit is further configured to: receive a first request using the first sub-management port when both the first sub-management port and the second sub-management port in the second management port are in a normal state; or receive the first sub-data of the first request using the first sub-management port and the second sub-data of the first request using the second sub-management port when both the first sub-management port and the second sub-management port are in a normal state; or receive the first request using the second sub-management port when the first sub-management port is in an abnormal state; wherein, the second management port is a logical port including the first sub-management port and the second sub-management port.
[0124] Optionally, in this embodiment, the forwarding unit is further configured to: send a first request to the first storage node using the connection link between the first sub-management port and the first storage node; or send first sub-data to the first storage node using the connection link between the first sub-management port and the first storage node, and send second sub-data to the first storage node using the connection link between the second sub-management port and the first storage node; or send a first request to the first storage node using the connection link between the second sub-management port and the first storage node.
[0125] Optionally, in this embodiment, the determining unit includes: a determining module, used to determine that the main management controller is in an abnormal state when the slave management controller has not received a status prompt signal sent by the main management controller through the communication link with the main management controller within N status detection cycles, where N is a positive integer.
[0126] Optionally, in this embodiment, the above-mentioned device further includes: a first binding unit, used to bind the first sub-management port of the slave management controller and the second sub-management port of the slave management controller to obtain a second management port; and a second binding unit, used to bind the third sub-management port of the master management controller and the fourth sub-management port of the master management controller to obtain a first management port.
[0127] Optionally, in this embodiment, the above-mentioned device further includes: a first local area network access unit, used to connect the third sub-management port to the first local area network through the first switch chip of the main management controller, and to connect the fourth sub-management port to the second local area network through the first switch chip, wherein the third sub-management port and the fourth sub-management port are respectively connected to the first switch chip.
[0128] Optionally, in this embodiment, the above-mentioned device further includes: a second local area network access unit, configured to connect the first sub-management port to the first local area network through the second switch chip of the management controller, and connect the second sub-management port to the second local area network through the second switch chip, wherein the first sub-management port and the second sub-management port are respectively connected to the second switch chip.
[0129] Optionally, in this embodiment, the above-mentioned device further includes: a first receiving unit, configured to receive a second request for managing the first storage node through a first management port when the main management controller is in a normal state, wherein the destination network address of the second request is the first network address; and a first forwarding unit, configured to forward the second request to the first storage node through the connection link between the first management port and the first storage node, so that the first storage node performs a management operation matching the second request.
[0130] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0131] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0132] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0133] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0134] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0135] Embodiments of this application also provide an electronic device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0136] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0137] Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium storing the computer program product, wherein the computer program, when executed by a processor, implements the steps of the methods described in various embodiments of this application.
[0138] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0139] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0140] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for managing storage nodes, characterized in that, include: If a slave management controller of a storage node set does not receive a status notification signal from the master management controller of the storage node set within a predetermined time window, it is determined that the master management controller is in an abnormal state. The master management controller and the slave management controller are connected to a target general-purpose input / output interface. A connection link is established between the first management port of the master management controller and each node in the storage node set. A connection link is established between the second management port of the slave management controller and each node. When the master management controller is in a normal state, it is used to receive a second request for managing a first storage node in the storage node set through the first management port, and also to send the second request to the first storage node through the connection link between the first management port and the first storage node. The hardware address of the second management port is switched to the first hardware address corresponding to the first management port, wherein the first hardware address is bound to a first network address used to access the storage node set. Receiving a first request for managing a first storage node in the storage node set through the second management port, wherein the destination network address of the first request is the first network address, the step of receiving the first request for managing the first storage node in the storage node set through the second management port includes: receiving the first request through the first sub-management port when both the first sub-management port and the second sub-management port in the second management port are in normal condition; or receiving the first sub-data of the first request through the first sub-management port and the second sub-data of the first request through the second sub-management port when both the first sub-management port and the second sub-management port are in normal condition; or receiving the first request through the second sub-management port when the first sub-management port is in an abnormal condition. The first request is sent to the first storage node using the connection link between the first sub-management port and the first storage node; or the first sub-data is sent to the first storage node using the connection link between the first sub-management port and the first storage node, and the second sub-data is sent to the first storage node using the connection link between the second sub-management port and the first storage node; or the first request is sent to the first storage node using the connection link between the second sub-management port and the first storage node.
2. The storage node management method according to claim 1, characterized in that, The second management port is a logical port that includes the first sub-management port and the second sub-management port.
3. The storage node management method according to claim 1, characterized in that, The determination that the master controller is in an abnormal state when the slave controller of the storage node set does not receive a status alert signal from the master controller of the storage node set within a predetermined time window includes: If the slave management controller fails to receive the status alert signal sent by the master management controller through the communication link with the master management controller within N status detection cycles, it is determined that the master management controller is in an abnormal state, where N is a positive integer.
4. The storage node management method according to claim 1, characterized in that, Before determining that the master controller is in an abnormal state, if the slave controller of the storage node set does not receive a status alert signal from the master controller of the storage node set within a predetermined time window, the process further includes: The first sub-management port of the slave management controller and the second sub-management port of the slave management controller are bound together to obtain the second management port; The third sub-management port and the fourth sub-management port of the main management controller are bound together to obtain the first management port.
5. The storage node management method according to claim 4, characterized in that, Before determining that the master controller is in an abnormal state, if the slave controller of the storage node set does not receive a status alert signal from the master controller of the storage node set within a predetermined time window, the process further includes: The third sub-management port is connected to the first local area network through the first switch chip of the main management controller, and the fourth sub-management port is connected to the second local area network through the first switch chip, wherein the third sub-management port and the fourth sub-management port are respectively connected to the first switch chip.
6. The storage node management method according to claim 5, characterized in that, Before determining that the master controller is in an abnormal state, if the slave controller of the storage node set does not receive a status alert signal from the master controller of the storage node set within a predetermined time window, the process further includes: The first sub-management port is connected to the first local area network through the second switch chip of the management controller, and the second sub-management port is connected to the second local area network through the second switch chip, wherein the first sub-management port and the second sub-management port are respectively connected to the second switch chip.
7. The method for managing storage nodes according to any one of claims 1 to 6, characterized in that, The management method for the storage node further includes: When the main management controller is in a normal state, a second request for managing the first storage node is received through the first management port, wherein the destination network address of the second request is the first network address; The second request is forwarded to the first storage node through the connection link between the first management port and the first storage node, so that the first storage node performs a management operation that matches the second request.
8. A management system for storage nodes, characterized in that, This includes a master management controller, slave management controllers, and a set of storage nodes. The master management controller and the slave management controllers are connected to a target general-purpose input / output interface. A first management port of the master management controller establishes a connection link with each node in the set of storage nodes, and a second management port of the slave management controller establishes a connection link with each node. When the main management controller is in a normal state, it is used to receive a second request through the first management port for managing a first storage node in the storage node set, wherein the destination network address of the second request is a first network address bound to a first hardware address corresponding to the first management port; the main management controller is also used to send the second request to the first storage node through the connection link between the first management port and the first storage node; The slave management controller is configured to, if it does not receive a status alert signal from the master management controller within a predetermined time window, determine that the master management controller is in an abnormal state, switch the hardware address of the slave management controller's second management port to the first hardware address, and receive a first request for management of the first storage node through the second management port, wherein the destination network address of the first request is the first network address. Receiving the first request for management of the first storage node through the second management port includes: receiving the first request using the first sub-management port when both the first sub-management port and the second sub-management port of the second management port are in normal states; or receiving the first request using the first sub-management port when both the first sub-management port and the second sub-management port are in normal states. The slave management controller receives the first sub-data of the first request through the first sub-management port and receives the second sub-data of the first request through the second sub-management port; or, if the first sub-management port is in an abnormal state, it receives the first request through the second sub-management port; the slave management controller is further configured to send the first request to the first storage node through the connection link between the first sub-management port in the second management port and the first storage node; or send the first sub-data to the first storage node through the connection link between the first sub-management port and the first storage node, and send the second sub-data to the first storage node through the connection link between the second sub-management port in the second management port and the first storage node; or send the first request to the first storage node through the connection link between the second sub-management port and the first storage node. The first storage node in the set of storage nodes is used to perform management operations matching the second request, and to perform management operations matching the first request.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Node deployment method and device and storage medium
CN113138717A
Storage node configuration method and device and storage medium
CN120295574A