A whole cabinet management method, device, equipment and medium

By sending configuration parameters on the resource exchange node, the child nodes are automatically configured and report their identity information, which solves the problem of topology relationship identification in resource pooled racks, realizes fast and accurate topology discovery and fault diagnosis, and improves the reliability of rack management.

CN115567400BActive Publication Date: 2026-03-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-29
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, how to correctly identify and manage the topological relationships between the ports of resource exchange nodes and their child nodes in a resource pooling rack is a problem that urgently needs to be solved.

Method used

By sending configuration parameters to the child nodes through the first management controller on the resource exchange node, the child nodes are automatically configured and report their identification information. This identification information is used to establish the topology of the entire rack and to manage communication and fault information in conjunction with the storage devices.

Benefits of technology

It enables rapid and accurate identification of the topological relationship between resource exchange nodes and child nodes, improves the reliability and feasibility of whole rack management, and provides fault diagnosis services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115567400B_ABST
    Figure CN115567400B_ABST
Patent Text Reader

Abstract

The application provides a whole cabinet management method, device, equipment and medium, which is applied to a first management controller, the first management controller is deployed on a resource exchange node, and the method comprises the following steps: configuration parameters are respectively sent to N sub-nodes, so that the N sub-nodes are respectively configured according to the respective received configuration parameters, and the identity information of each sub-node is reported to the resource exchange node after the configuration is completed; and the topology relationship of the whole cabinet is established according to the identity information of the N sub-nodes. In the application, the automatic configuration of the sub-nodes and the reporting of the identity information are utilized to realize the topology discovery between the resource exchange node and the sub-nodes, and the fast and accurate identification of the topology relationship between the resource exchange node and the sub-nodes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of server application, and particularly relates to a whole cabinet management method and device, equipment and medium. BACKGROUND

[0002] Server resource pooling technology can bring flexible and elastic resource deployment, improve resource utilization, and more effectively improve the fault repair capability and operation efficiency of the server. The server resource pooling technology is usually deployed in units of whole cabinets, and some important resources are pooled in the whole cabinet, such as CPU pool, memory pool, storage pool, and heterogeneous acceleration pool, and the resource exchange node (Switch node) is used to connect various resource pools in the whole cabinet together to realize the integration and flexible configuration of various resources.

[0003] At present, the resource pooling technology taking the resource exchange node as the core is still in the research stage. For the resource pool whole cabinet, how to correctly identify the topology relationship between the ports of the resource exchange node and the sub-nodes and manage them becomes a problem to be solved. SUMMARY

[0004] In view of the above problems, the embodiments of the present application provide a whole cabinet management method, device, equipment and medium, so as to overcome the above problems or at least partially solve the above problems.

[0005] The first aspect of the embodiments of the present application discloses a whole cabinet management method applied to a first management controller, the first management controller is deployed on a resource exchange node, and the method comprises the following steps:

[0006] configuration parameters are sent to N sub-nodes respectively, so that the N sub-nodes are respectively configured according to the configuration parameters received by themselves, and identity information of the N sub-nodes is reported to the resource exchange node after the configuration is completed;

[0007] According to the identity information of the N sub-nodes, the topology relationship of the whole cabinet is established.

[0008] Optionally, the first management controller and the storage devices deployed on the N sub-nodes respectively establish a communication connection, and the communication connection is realized by the communication link in the connection cable between the N ports of the resource exchange node and the N sub-nodes; the configuration parameters are sent to the N sub-nodes respectively, comprising:

[0009] The control authority of the storage device of each of the N sub-nodes is obtained;

[0010] The configuration parameters of the N sub-nodes are written into the storage device of each of the N sub-nodes respectively, so that the N sub-nodes read the configuration parameters of themselves from the storage device deployed by themselves.

[0011] Optionally, before establishing the topological relationship of the whole cabinet according to the identity information of each of the N sub-nodes, the method further comprises:

[0012] obtaining the identity information reported by the N sub-nodes from the internal LAN of the whole cabinet, wherein the internal LAN of the whole cabinet is a LAN established by the first management controller, the N second management controllers, and TOR network switches in the whole cabinet through network links, and each second management controller is deployed on a sub-node.

[0013] Optionally, after establishing the topological relationship of the whole cabinet, the method further comprises:

[0014] reporting the topological relationship to a management client through an external management network, so that the management client manages the N sub-nodes.

[0015] Optionally, the first management controller and the storage devices deployed on the N sub-nodes are respectively connected through communication links in the connection cables of the N ports of the resource exchange node and the N sub-nodes; the method further comprises:

[0016] when the identity information reported by a sub-node is not received within a preset time length, reading the fault information in the storage device of the sub-node;

[0017] when a fault broadcast of the internal LAN of the whole cabinet is received, reading the fault information in the storage device of the sub-node that has failed;

[0018] reporting the obtained fault information to a management client through an external management network, so that the management client processes the fault.

[0019] Optionally, the N sub-nodes are deployed with Muxes, the first management controller and the storage devices deployed on the N sub-nodes are respectively connected through the Muxes, and the second management controllers are connected with the corresponding storage devices through the Muxes; the method further comprises:

[0020] the first management controller and the second management controller deployed on each sub-node switch control through the Mux, so that the first management controller and the second management controller can respectively perform read and write operations on the storage devices.

[0021] In a second aspect of the embodiments of the application, a whole cabinet management method is disclosed, which is applied to a second management controller deployed on a sub-node, and the method comprises:

[0022] receiving configuration parameters sent by a first management controller;

[0023] configuring according to the received configuration parameters;

[0024] after the configuration is completed, reporting self identity information to the resource exchange node, so that the first management controller establishes the topological relationship of the whole cabinet according to the identity information.

[0025] Optionally, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link; and the receiving of the configuration parameters sent by the first management controller comprises:

[0026] obtaining the control authority of the storage device;

[0027] reading the configuration parameters pre-written by the first management controller into the storage device from the storage device.

[0028] Optionally, the reporting of the self identity information to the resource exchange node after the configuration is completed comprises:

[0029] publishing the self identity information in the internal local area network of the whole cabinet, the internal local area network of the whole cabinet being a local area network established by the first management controller, N second management controllers and TOR network switches in the whole cabinet through network links.

[0030] Optionally, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link, and the method further comprises:

[0031] when a fault occurs in the configuration process or the identity information reporting process of the sub-node, writing the fault information into the storage device of the sub-node, so that the first management controller reads the fault information from the storage device;

[0032] sending a fault broadcast to the internal local area network of the whole cabinet, so that the first management controller reads the fault information in the storage device of the sub-node which has occurred a fault after receiving the fault broadcast.

[0033] In a third aspect of the embodiments of the application, a whole cabinet management method is disclosed, which is applied to a management client, the management client is connected with a TOR network switch in a whole cabinet through an external management network, and the method comprises:

[0034] accessing a first management controller through the external management network to obtain a topological relationship of the whole cabinet, the topological relationship being generated according to the method of the first aspect;

[0035] calling a communication interface on N sub-nodes in the topological relationship by using an internal local area network of the whole cabinet to obtain device information of the N sub-nodes, and managing the N sub-nodes.

[0036] Optionally, the method further comprises:

[0037] obtaining the fault information reported by the first management controller through the external management network, and processing the fault information.

[0038] In a fourth aspect, the embodiment of the present application discloses a whole cabinet management device, which is applied to a first management controller, and the first management controller is deployed on a resource exchange node. The device comprises:

[0039] a sending module, configured to send configuration parameters to N child nodes respectively, so that the N child nodes are configured according to the respective received configuration parameters, and report their own identity information to the resource exchange node after the configuration is completed;

[0040] a recognition module, configured to establish a topology relationship of a whole cabinet according to the respective identity information of the N child nodes.

[0041] In a fifth aspect, the embodiment of the present application discloses a whole cabinet management device, which is applied to a second management controller, and the second management controller is deployed on a child node. The device comprises:

[0042] a receiving module, configured to receive the configuration parameters sent by the first management controller;

[0043] a configuration module, configured to perform configuration according to the received configuration parameters;

[0044] a reporting module, configured to report its own identity information to the resource exchange node after the configuration is completed, so that the first management controller establishes a topology relationship of a whole cabinet according to the identity information.

[0045] In a sixth aspect, the embodiment of the present application discloses a whole cabinet management device, which is applied to a management client, and the management client is connected with a TOR network switch in a whole cabinet through an external management network. The device comprises:

[0046] an access module, configured to access the first management controller through the external management network, so as to obtain a topology relationship of a whole cabinet, and the topology relationship is generated according to the method in the first aspect;

[0047] a management module, configured to call a communication interface on N child nodes in the topology relationship by using an internal local area network of a whole cabinet, so as to obtain device information of the N child nodes, and manage the N child nodes.

[0048] In a seventh aspect, the embodiment of the present application discloses an electronic device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the whole-cabinet management method of the first aspect or the whole-cabinet management method of the second aspect or the whole-cabinet management method of the third aspect when executed.

[0049] In an eighth aspect, the embodiment of the present application discloses a computer readable storage medium, which stores a computer program / instruction, and the computer program / instruction implements the whole-cabinet management method of the first aspect or the whole-cabinet management method of the second aspect or the whole-cabinet management method of the third aspect when executed by a processor.

[0050] The embodiment of the present application has the following advantages:

[0051] In the embodiment of the present application, the first management controller on the resource exchange node in the whole cabinet sends configuration parameters to each sub-node respectively, each sub-node reports its identity information to the resource exchange node after completing the automatic configuration according to the configuration parameters, and then the first management controller establishes the topology relationship of the whole cabinet according to the received identity information. Therefore, the embodiment proposes a highly feasible resource pooling whole-cabinet management method, which realizes the topology discovery between the resource exchange node and the sub-node by using the automatic configuration and identity information reporting of the sub-node, and realizes the rapid and accurate identification of the topology relationship between the resource exchange node and the sub-node. BRIEF DESCRIPTION OF DRAWINGS

[0052] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0053] Figure 1 is a whole-cabinet management method step flow chart provided by the embodiment of the present application and applied to the first management controller;

[0054] Figure 2 is a whole-cabinet management method step flow chart provided by the embodiment of the present application and applied to the second management controller;

[0055] Figure 3 is a whole-cabinet management method step flow chart provided by the embodiment of the present application and applied to the management client;

[0056] Figure 4 is a whole-cabinet management system hardware topology structure schematic diagram provided by the embodiment of the present application;

[0057] Figure 5 is a structural schematic diagram of a whole-cabinet management device applied to a first management controller according to an embodiment of the present application;

[0058] Figure 6 is a structural schematic diagram of a whole-cabinet management device applied to a second management controller according to an embodiment of the present application;

[0059] Figure 7 is a structural schematic diagram of a whole-cabinet management device applied to a management client according to an embodiment of the present application. DETAILED DESCRIPTION

[0060] In order to make the above objectives, characteristics and advantages of the present application more apparent, clear and easy to understand, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0061] Figure 1 The whole-cabinet control method provided by the embodiments of the present application is shown, which is applied to a first management controller, the first management controller is deployed on a resource exchange node, and the method comprises the following steps:

[0062] Step S101: configuration parameters are sent to N sub-nodes respectively, so that the N sub-nodes are respectively configured according to the respective received configuration parameters, and after the configuration is completed, the identity information of each sub-node is reported to the resource exchange node.

[0063] In the embodiment, the whole cabinet is equivalent to a large server, and the whole cabinet includes a resource exchange node (i.e., a Switch node) and a plurality of (N) sub-nodes. The resource exchange node is a core exchange device for resource pooling of the whole cabinet, and can be used for device expansion and resource and information interaction. A first management controller is deployed on the resource exchange node, and the first management controller is used for managing the operation and maintenance of the resource exchange node. Each sub-node is equivalent to a pooled resource (e.g., a CPU pool, a memory pool, a storage pool, and a heterogeneous acceleration pool) in the whole cabinet. A second management controller is deployed on each sub-node, and the second management controller is used for managing the operation and maintenance of the sub-node.

[0064] Since the configuration parameters of each sub-node are different, when the first management controller on the resource exchange node is started, the configuration parameters are sent to each sub-node one by one, wherein the configuration parameters include the management parameters of the local area network in the whole cabinet, and the identification of the corresponding connection port on the resource exchange node. When the sub-node receives the configuration parameters, the sub-node is automatically configured according to the configuration parameters, and the identity information of the sub-node is reported after the configuration is completed. The identity information of each sub-node includes the main capability, management address, device identification, interface identification and the like.

[0065] In an optional embodiment, the first management controller respectively establishes a communication connection with the storage device deployed on the N sub-nodes, and the communication connection is realized through the communication link in the connection cable of the N ports of the resource exchange node and the N sub-nodes; the configuration parameters are respectively sent to the N sub-nodes, including:

[0066] The control authority of the storage device of each of the N sub-nodes is obtained;

[0067] The configuration parameters of the N sub-nodes are respectively written into the storage device of each of the N sub-nodes, so that the N sub-nodes read the self-configuration parameters from the storage device deployed on each sub-node.

[0068] In the embodiment, the storage device is deployed on each sub-node, and the storage device refers to a device or a chip with a data storage function and capable of data read-write operation, for example, EEPROM (Electrically Erasable Programmable Read Only Memory, Electrically Erasable Programmable Read Only Memory). EEPROM is a storage chip that does not lose data after power failure. The communication link refers to a link that can realize the sending and receiving of data between devices. The communication link can be an SMBus link (System Management Bus, System Management Bus). SMBus is a two-wire serial bus, and data can be sent and received between devices through the SMBus link. Further, the first management controller on the resource exchange node establishes a connection with the EEPROM of the sub-node through the SMBus link in the connection cable of the resource exchange node and the sub-node, and realizes the communication between the first management controller and the storage device (EEPROM). That is, the first management controller can perform data write and read operations on the storage device through the communication connection.

[0069] When the first management controller on the resource exchange node is started, the first management controller actively acquires the control right of the storage device of each sub-node, and writes the configuration parameters into the corresponding area of the memory of the sub-node after acquiring the control right, and releases the control right of the storage device of the sub-node after completing the writing operation. In addition, after the first management controller completes the configuration parameter writing operation, the first management controller can notify each sub-node that can start reading the configuration parameters in the storage device by sending a network broadcast.

[0070] In the embodiment, the hardware connection of the resource pool whole-cabinet equipment topology establishment, i.e. management, is performed by using the communication link and the storage device. The first management controller writes the configuration parameters into the storage device, so that the sub-node acquires the configuration parameters through the storage device, and the data interaction between the resource exchange node and the sub-node is realized at a lower cost. Since the cost of the storage device is low, the whole-cabinet management system has higher reliability and expandability.

[0071] Step S102: establishing the topology relationship of the whole cabinet according to the identity information of each of the N sub-nodes.

[0072] In the embodiment, when the first management controller on the resource exchange node receives the identity information reported by each sub-node, the identity information is identified, and the topology relationship of the whole cabinet is established according to the identified identity information, wherein the topology relationship refers to the connection relationship between the resource exchange node and each sub-node. Further, the topology discovery between the resource exchange node and the sub-node is realized, and the fast and accurate identification of the topology relationship between the resource exchange node and the sub-node is realized.

[0073] In an optional embodiment, before establishing the topology relationship of the whole cabinet according to the identity information of each of the N sub-nodes, the method further comprises:

[0074] The identity information reported by the N sub-nodes is acquired from the internal LAN of the whole cabinet, the internal LAN of the whole cabinet is a LAN established by the first management controller, N second management controllers and TOR network switches in the whole cabinet through network links, and each second management controller is deployed on a sub-node.

[0075] In the embodiment, the network link can be a LAN (local area network) link, the first management controller on the resource exchange node and the second management controller on each sub-node are connected to the TOR network switch located at the top of the cabinet through the LAN link, and then the internal LAN connection of the cabinet is established. After the LAN connection is established, the first management controller and the second management controller can exchange network data, and when the sub-nodes complete the configuration according to the configuration parameters, the second management controller sends the identity information of the sub-nodes to the internal LAN of the cabinet, and the first management controller on the resource exchange node obtains the identity information sent by the sub-nodes from the internal LAN of the cabinet.

[0076] The identity information is sent in the form of an LLDPDU (Link Layer Discovery Protocol Data Unit) message, that is, the second management controller encapsulates the identity information (for example, the main capability, management address, device identifier, interface identifier, and the like) of the second management controller in an LLDPDU after completing the configuration, and publishes the LLDPDU in the internal LAN of the cabinet, so that the first management controller receives the LLDPDU message sent by each sub-node from the internal LAN of the cabinet, and establishes the topology relationship in the cabinet according to the identity information carried in the LLDPDU message.

[0077] In an optional embodiment, after the topology relationship of the cabinet is established, the method further comprises:

[0078] The topology relationship is reported to the management client through an external management network, so that the management client manages the N sub-nodes.

[0079] In the embodiment, the TOR network switch located at the top of the cabinet is connected to the external management network and establishes a connection with the management client, so that the management client can manage the cabinet. The first management controller reports the established topology relationship of the cabinet to the management client, and the management client manages each node in the topology relationship, so that the method realizes reliable management of the cabinet based on the topology relationship, and compared with the previous cabinet management technology, a cabinet management scheme with high feasibility is provided.

[0080] In an optional embodiment, the first management controller and the storage device deployed on the N sub-nodes respectively establish a communication connection, and the communication connection is realized through the communication link in the connection cable between the N ports of the resource exchange node and the N sub-nodes; the method further comprises:

[0081] When the identity information reported by the sub-node is not received within a preset time length, reading the fault information in the storage device of the sub-node;

[0082] When the fault broadcast of the internal LAN of the whole cabinet is received, reading the fault information in the storage device of the sub-node that has failed;

[0083] Reporting the obtained fault information to the management client through the external management network, so that the management client processes the fault.

[0084] In the embodiment, when the first management controller does not receive the identity information reported by the sub-node within a preset time length, it indicates that a fault occurs, causing the sub-node to fail to report its identity information in time, or when the fault broadcast of the internal LAN of the whole cabinet is received, it also indicates that the sub-node has failed.

[0085] The fault problem can be a failure in the configuration process of the sub-node, or a failure in the reporting process of the identity information, or a failure of the LAN. After the failure, the second management controller on the sub-node writes the fault information into the storage device of the sub-node, so that the first management controller on the resource exchange node can read the fault information from the storage device of the failed sub-node. Therefore, through the storage devices in the sub-nodes and the reporting of the fault information of the sub-nodes by the first management controller, the management client can process the fault information, so that the whole cabinet management method has a fault diagnosis service, and the reliability of the whole cabinet management method is higher.

[0086] In an optional embodiment, a Mux is deployed on the N sub-nodes, the first management controller establishes a communication connection with the storage devices deployed on the N sub-nodes through the Mux, and the second management controller communicates with the corresponding storage devices through the Mux; the method further comprises:

[0087] The first management controller and the second management controller on each sub-node switch the control right through the Mux, so that the first management controller and the second management controller can respectively read and write the storage device.

[0088] Before the first management controller and the second management controller read and write the storage device of the sub-node, the control right of the storage device needs to be obtained. Therefore, a Mux (multiplexer) is arranged in front of the storage device of each sub-node. The Mux can switch signals as needed. In the embodiment, the Mux is used to switch the control right of the storage device by the first management controller and the second management controller, so as to realize the orderly interaction between the resource exchange node and each sub-node.

[0089] In the embodiment, the first management controller on the resource exchange node in the whole cabinet sends configuration parameters to each sub-node respectively, each sub-node reports its own identity information to the resource exchange node after completing the automatic configuration according to the configuration parameters, and then the first management controller establishes the topology relationship of the whole cabinet according to the received identity information. Therefore, the embodiment provides a high feasible resource pooling whole cabinet management method, which realizes the topology discovery between the resource exchange node and the sub-node by the automatic configuration and identity information reporting of the sub-node, and realizes the fast and accurate identification of the topology relationship between the resource exchange node and the sub-node.

[0090] As shown in Figure 2 According to another aspect of the present application, a whole cabinet management method is provided, which is applied to a second management controller deployed on a sub-node, and the method comprises:

[0091] Step S201: receiving the configuration parameters sent by the first management controller;

[0092] Step S202: configuring according to the received configuration parameters;

[0093] Step S203: reporting its own identity information to the resource exchange node after completing the configuration, so that the first management controller establishes the topology relationship of the whole cabinet according to the identity information.

[0094] In the embodiment, the second management controller is deployed on each sub-node, which is used to manage the operation and maintenance of the sub-node and realize the data interaction with the first management controller on the resource exchange node. Receiving the configuration parameters sent by the first management controller means that each sub-node receives the configuration parameters sent by the first management controller to the sub-node, and the configuration parameters contain the identity information of the corresponding connection port of the resource exchange node and the sub-node, so the configuration parameters of each sub-node are not the same.

[0095] When the second management controller receives the respective configuration parameters, the second management controller automatically configures according to the respective configuration parameters, and sends its own identity information to the first management controller after completing the configuration, wherein the identity information includes the main capability, management address, device identity, interface identity and other information of the sub-node. The first management controller establishes the topology relationship in the whole cabinet according to the received identity information of each sub-node, and then each sub-node realizes the topology discovery between the resource exchange node and the sub-node by the automatic configuration and identity information reporting, and realizes the fast and accurate identification of the topology relationship between the resource exchange node and the sub-node.

[0096] In an alternative embodiment, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link; and the method further comprises:

[0097] acquiring the control right of the storage device;

[0098] reading the configuration parameters previously written into the storage device by the first management controller from the storage device.

[0099] In this embodiment, when the first management controller writes the configuration parameters into the corresponding area of the storage device of the sub-node, the control right of the storage device of the sub-node is released, and the second management controller is notified to read the configuration parameters in the storage device. After receiving the notification, the second management controller acquires the control right of the storage device and reads the configuration parameters therefrom. The first management controller can notify the second management controller to read the configuration parameters in the storage device through network broadcasting.

[0100] In an alternative embodiment, after completing the configuration, the second management controller reports its identity information to the resource exchange node, and the method further comprises:

[0101] publishing the identity information in an internal LAN of the whole cabinet, wherein the internal LAN of the whole cabinet is a LAN established by the first management controller, the N second management controllers and TOR network switches in the whole cabinet through network links.

[0102] In this embodiment, the second management controller on the sub-node and the first management controller on the resource exchange node can exchange network data through the internal LAN of the whole cabinet, and when the second management controller completes the configuration, it publishes its identity information in the internal LAN of the whole cabinet. The identity information can be published in the form of an LLDPDU message, i.e., after completing the configuration, the second management controller encapsulates the identity information of the sub-node (main capability of the sub-node itself, management address, device identifier, interface identifier, etc.) into an LLDPDU message and publishes it in the internal LAN of the whole cabinet, and then the first management controller receives the LLDPDU message in the internal LAN of the whole cabinet.

[0103] In an alternative embodiment, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link, and the method further comprises:

[0104] When a fault occurs in the configuration process or the identity information reporting process of the sub-node, the fault information is written into the storage device of the sub-node, so that the first management controller reads the fault information from the storage device.

[0105] sending a fault broadcast to the internal LAN of the whole cabinet, so that the first management controller reads the fault information in the storage device of the faulty sub-node after receiving the fault broadcast.

[0106] In the embodiment, the second management controller releases the control right of the storage device after writing the fault information into the storage device, and then the first management controller of the resource exchange node can acquire the control right of the storage device of the faulty sub-node again and read the fault information in the storage device when the first management controller receives the fault broadcast or does not receive the identity information reported by the sub-node for a long time.

[0107] As shown in Figure 3 According to still another aspect of the present application, a whole cabinet management method is provided, which is applied to a management client connected with a TOR network switch in a whole cabinet through an external management network, and includes the following steps:

[0108] In step S301, the first management controller is accessed through the external management network to obtain the topology relationship of the whole cabinet, which is generated according to the method of the first aspect.

[0109] In step S302, the communication interface of the N sub-nodes in the topology relationship is called through the internal LAN of the whole cabinet to obtain the device information of the N sub-nodes, and the N sub-nodes are managed.

[0110] In the embodiment, the management client is connected with the TOR network switch through the external management network, and can interact with the first management controller and the second management controller. The management client accesses the first management controller through the internal LAN of the whole cabinet to obtain the topology relationship of the whole cabinet, and then calls the communication interface of the sub-nodes in the topology relationship through the internal LAN of the whole cabinet. The communication interface can be a Redfish interface, and the device information of the sub-nodes is obtained through the Redfish interface of the sub-nodes, and then the node devices are managed.

[0111] In an optional embodiment, the method further includes:

[0112] The fault information reported by the first management controller is obtained through the external management network, and the fault information is processed.

[0113] In the embodiment, the first management controller on the resource exchange node reports the fault information read from the storage device of the fault sub-node to the management client, and the management client processes the fault information after receiving the fault information, wherein the fault information includes parameter configuration error fault, network fault, etc. Thus, the whole cabinet management method has the fault diagnosis service in the case that the sub-node encounters a fault.

[0114] In the embodiment, the hardware connection of the resource pool whole cabinet equipment topology establishment (i.e. management) is implemented by using the communication link and the storage device, the first management controller on the resource exchange node sends the configuration parameters to each sub-node respectively, each sub-node reports its identity information to the resource exchange node after completing the automatic configuration according to the configuration parameters, then the first management controller establishes the topology relationship of the whole cabinet according to the received identity information, and the topology discovery between the resource exchange node and the sub-node is implemented by using the automatic configuration of the sub-node and the identity information reporting, thus the fast and accurate identification of the topology relationship between the resource exchange node and the sub-node is implemented. Since the cost of the storage device is low, the whole cabinet management system has higher reliability and expandability. In addition, the storage device, the first management controller and the second management controller are used to report the fault information of the sub-node, and the fault diagnosis service is provided by the whole cabinet management method when the sub-node encounters a fault.

[0115] Figure 4 The hardware topology structure of the whole cabinet management system provided by the embodiment of the application is shown in FIG. 1. Figure 4 As shown in FIG. 1, the whole cabinet includes one resource exchange node (i.e. Switch node), the first management controller is deployed on the resource exchange node, and the first management controller is used to manage the operation and maintenance of the resource exchange node. There are several sub-nodes in the whole cabinet, each sub-node is equivalent to a pooled resource, the second management controller is deployed on each sub-node, and the second management controller is also used to manage the operation and maintenance of the sub-node; the EEPROM (storage device), Mux and Devices (i.e. other equipment used for business operation) are also deployed on each sub-node. There is also a TOR network switch on the top of the whole cabinet, and the management client is deployed outside the whole cabinet.

[0116] The first management controller deployed on the resource exchange node establishes a communication connection with the EEPROM of the sub-node through the SMBus link (communication link) and Mux in the cable connecting the resource exchange node and the sub-node, and the second management controller deployed on the sub-node establishes a communication connection with the EEPROM through the SMBus link and Mux. The first management controller deployed on the resource exchange node and the second management controller deployed on the sub-node perform control right switching through the Mux in front of the EEPROM, so that they can respectively read and write the EEPROM. In addition, the first management controller deployed on the resource exchange node and the second management controller deployed on the sub-node are respectively connected to the TOR network switch located at the top of the whole cabinet through the network link LAN, to establish a local area network connection, so that the first management controller and the second management controller can exchange network data, such as identity information, network broadcast, etc. The TOR network switch is connected to an external management network, so that a management client can manage the whole cabinet through the external management network.

[0117] In actual application process, when the first management controller on the resource exchange node is started, it will actively acquire the control right of the EEPROM of each sub-node, and perform read and write operations on the EEPROM of each sub-node one by one. The management parameters of the local area network in the whole cabinet, the identification of the corresponding connection port on the resource exchange node, and other configuration information are written into the corresponding area of the EEPROM of the sub-node. After the write operation is completed, the control right of the EEPROM of the sub-node is released. Then the second management controller on each sub-node acquires the control right of the corresponding EEPROM, reads the configuration parameters in the EEPROM, and performs automatic configuration. After the configuration is completed, the main capability, management address, device identification, interface identification and other identity information of the sub-node are encapsulated in the LLDPDU and published in the internal local area network of the whole cabinet, and the control right of the EEPROM is released. After that, the first management controller on the resource exchange node receives the LLDPDU packet information in the internal local area network of the whole cabinet, and establishes the topology relationship of the whole cabinet according to the LLDPDU packet information. The topology relationship can be used to display to the management end or to manage the internal nodes of the whole cabinet. Finally, the management client accesses the first management controller through the external management network, and calls the Redfish interface of each sub-node in the topology relationship through the internal local area network of the whole cabinet to obtain the detailed information of the devices on the sub-nodes, and manages the devices on the sub-nodes.

[0118] Further, when the sub-node is in the automatic configuration process or the LLDPDU message sending error, the second management controller on the sub-node writes the partial fault information into the EEPROM for the first management controller to report the fault information, and releases the EEPROM control right after the fault writing operation is completed; then the first management controller accesses the EEPROM of the sub-node where the fault occurs, reads the fault information in the EEPROM, records the fault information and reports the fault information to the management client through the external management network, and the client processes the fault information after receiving the fault information.

[0119] Figure 5 The whole cabinet management device provided by the embodiment of the application is shown, which is applied to a first management controller deployed on a resource exchange node, and the device comprises:

[0120] The sending module 51 is configured to send configuration parameters to the N sub-nodes respectively, so that the N sub-nodes are configured according to the respective received configuration parameters, and report their own identity information to the resource exchange node after the configuration is completed;

[0121] The identification module 52 is configured to establish a topology relationship of the whole cabinet according to the respective identity information of the N sub-nodes.

[0122] In an optional embodiment, the first management controller and the storage devices deployed on the N sub-nodes are respectively connected in communication, the communication is realized through the communication link in the connection cable between the N ports of the resource exchange node and the N sub-nodes, and the sending module comprises:

[0123] The first permission obtaining module is configured to obtain the control permission of the storage device of each of the N sub-nodes;

[0124] The parameter writing module is configured to write the configuration parameters of the N sub-nodes into the storage devices of the N sub-nodes respectively, so that the N sub-nodes read their own configuration parameters from the respective deployed storage devices.

[0125] In an optional embodiment, the device further comprises:

[0126] The identity obtaining module is configured to obtain the identity information reported by the N sub-nodes from the internal LAN of the whole cabinet, the internal LAN of the whole cabinet is a LAN established by the first management controller, N second management controllers and TOR network switches in the whole cabinet through network links, and each second management controller is deployed on a sub-node.

[0127] In an optional embodiment, the device further comprises:

[0128] The topology reporting module is configured to report the topology relationship to a management client through an external management network, so that the management client manages the N sub-nodes.

[0129] In an optional embodiment, the first management controller establishes a communication connection with a storage device deployed on each of the N sub-nodes, and the communication connection is achieved through a communication link in a connection cable between the N ports of the resource exchange node and the N sub-nodes.

[0130] The first fault reading module is configured to read the fault information in the storage device of the sub-node when the identity information reported by the sub-node is not received within a preset time length.

[0131] The second fault reading module is configured to read the fault information in the storage device of the sub-node that fails when a fault broadcast of an internal LAN of the whole cabinet is received.

[0132] The fault reporting module is configured to report the obtained fault information to the management client through the external management network, so that the management client processes the fault.

[0133] In an optional embodiment, the N sub-nodes are deployed with a Mux, the first management controller establishes a communication connection with a storage device deployed on each of the N sub-nodes through the Mux, and the second management controller communicates with the corresponding storage device through the Mux.

[0134] The authority switching module is configured to switch the control right of the first management controller and the second management controller deployed on each sub-node through the Mux, so that the first management controller and the second management controller can perform read and write operations on the storage device, respectively.

[0135] Figure 6 The whole cabinet management device provided by the embodiment of the application is shown, which is applied to a second management controller, and the second management controller is deployed on a sub-node. The device comprises:

[0136] The receiving module 61 is configured to receive the configuration parameters sent by the first management controller.

[0137] The configuration module 62 is configured to perform configuration according to the received configuration parameters.

[0138] The reporting module 63 is configured to report the identity information of the device to the resource exchange node after the configuration is completed, so that the first management controller establishes the topology relationship of the whole cabinet according to the identity information.

[0139] In an alternative embodiment, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link, and the receiving module comprises:

[0140] The second permission obtaining module is configured to obtain the control permission of the storage device.

[0141] The parameter reading module is configured to read the configuration parameter pre-written by the first management controller into the storage device from the storage device.

[0142] In an alternative embodiment, the reporting module comprises:

[0143] The identity reporting module is configured to publish the identity information of the first management controller in the internal LAN of the whole cabinet, and the internal LAN of the whole cabinet is a LAN established by the first management controller, N second management controllers and TOR network switches in the whole cabinet through network links.

[0144] In an alternative embodiment, the second management controller establishes a communication connection with the storage device deployed on the sub-node through a communication link, and the device further comprises:

[0145] The fault writing module is configured to write fault information into the storage device of the sub-node when a fault occurs in the configuration process or the identity information reporting process of the sub-node, so that the first management controller reads the fault information from the storage device.

[0146] The fault notification module is configured to send a fault broadcast to the internal LAN of the whole cabinet, so that the first management controller reads the fault information in the storage device of the sub-node that has occurred a fault after receiving the fault broadcast.

[0147] Figure 7 The device for managing the whole cabinet is shown, which is applied to a management client, the management client is connected with a TOR network switch in the whole cabinet through an external management network, and the device comprises:

[0148] The access module 71 is configured to access the first management controller through the external management network to obtain the topology relationship of the whole cabinet, and the topology relationship is generated according to the method in the first aspect of the embodiment.

[0149] The management module 72 is configured to call the communication interface on the N sub-nodes in the topology relationship by using the internal LAN of the whole cabinet to obtain the device information of the N sub-nodes and manage the N sub-nodes.

[0150] In an alternative embodiment, the device further comprises:

[0151] The fault processing module acquires the fault information reported by the first management controller through the external management network, and processes the fault information.

[0152] The electronic device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes, the whole-cabinet management method of any of the above embodiments is implemented.

[0153] The computer readable storage medium stores the computer program / instructions, and the computer program / instructions are executed by the processor to implement the whole-cabinet management method of any of the above embodiments.

[0154] Each of the embodiments in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments. The same and similar parts of each embodiment can be referred to each other.

[0155] Although the preferred embodiments of the present application have been described, those skilled in the art can make further changes and modifications to the embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to cover all changes and modifications falling within the scope of the embodiments of the present application.

[0156] Finally, it should be noted that, in this document, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or terminal device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, method, article or terminal device including the element.

[0157] The above describes in detail the whole-cabinet management method, device, equipment and medium provided by the present application. The specific examples are applied to explain the principles and implementation modes of the present application. The above embodiment descriptions are only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed; in conclusion, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A method for managing a complete server rack, characterized in that, Applied to a first management controller, which is deployed on a resource exchange node, the method includes: Configuration parameters are sent to N child nodes respectively, so that the N child nodes can configure themselves according to the configuration parameters they receive, and report their own identity information to the resource exchange node after the configuration is completed; The system obtains the identity information reported by N sub-nodes from the internal LAN of the rack. The internal LAN of the rack is a LAN established by the first management controller, N second management controllers, and the TOR network switch in the rack, which are connected to the rack via network links. Each second management controller is deployed on a sub-node. The topology of the entire cabinet is established based on the identity information of each of the N child nodes.

2. The method according to claim 1, characterized in that, The first management controller establishes communication connections with the storage devices deployed on the N child nodes, respectively. These communication connections are achieved through communication links in the connection cables between the N ports of the resource exchange node and the N child nodes. Configuration parameters are sent to each of the N child nodes, including: Obtain control permissions for the storage devices of each of the N child nodes; The configuration parameters of the N child nodes are written to their respective storage devices, so that the N child nodes can read their own configuration parameters from their respective deployed storage devices.

3. The method according to claim 1, characterized in that, After establishing the topology of the entire rack, the following is also included: The topology is reported to the management client via an external management network so that the management client can manage the N child nodes.

4. The method according to claim 1, characterized in that, The first management controller establishes communication connections with the storage devices deployed on the N child nodes, respectively. These communication connections are achieved through communication links in the connection cables between the N ports of the resource exchange node and the N child nodes. The method further includes: If no identity information is received from the child node within a preset time period, read the fault information from the storage device of that child node; When a fault broadcast is received from the local area network inside the rack, the fault information is read from the storage device of the faulty child node; The acquired fault information is reported to the management client via an external management network, so that the management client can handle the fault.

5. The method according to claim 1, characterized in that, A Mux is deployed on each of the N child nodes. The first management controller establishes communication connections with the storage devices deployed on each of the N child nodes through the Mux, and the second management controller communicates with the corresponding storage devices through the Mux. The method further includes: The first management controller and the second management controller deployed on each child node switch control through the Mux, so that the first management controller and the second management controller can respectively perform read and write operations on the storage device.

6. A method for managing a complete server rack, characterized in that, Applied to a second management controller, which is deployed on a child node, the method includes: Receive configuration parameters sent by the first management controller; Configure according to the received configuration parameters; After completing the configuration, it reports its own identity information to the resource exchange node so that the first management controller can establish the topology of the entire cabinet based on the identity information. After completing the configuration, the entity reports its own identity information to the resource exchange node, including: The system publishes its own identity information in the internal local area network of the rack. The internal local area network of the rack is a local area network established by the first management controller, N second management controllers, and TOR network switches in the rack, which are connected to each other through network links.

7. The method according to claim 6, characterized in that, The second management controller establishes a communication connection with the storage device deployed on the child node via a communication link; receiving configuration parameters sent by the first management controller includes: Obtain control permissions for the storage device; Read the configuration parameters that the first management controller has pre-written into the storage device from the storage device.

8. The method according to claim 6, characterized in that, The second management controller establishes a communication connection with the storage device deployed on the child node via a communication link, and the method further includes: When a child node experiences a failure during configuration or identity information reporting, the failure information is written to the child node's storage device so that the first management controller can read the failure information from the storage device. A fault broadcast is sent to the local area network inside the cabinet so that the first management controller can receive the fault broadcast and read the fault information from the storage device of the faulty child node.

9. A method for managing a complete server rack, characterized in that, The method is applied to a management client, which connects to a TOR network switch within a server rack via an external management network. The first management controller is accessed through the external management network to obtain the topology of the entire cabinet, the topology being generated according to any one of claims 1-5; The communication interfaces on N child nodes in the topology are accessed through the local area network inside the cabinet to obtain device information of the N child nodes and manage the N child nodes.

10. The method according to claim 9, characterized in that, The method further includes: The fault information reported by the first management controller is obtained through the external management network, and the fault information is processed.

11. A cabinet management device, characterized in that, The device, applied to a first management controller deployed on a resource exchange node, comprises: The sending module is used to send configuration parameters to N child nodes respectively, so that the N child nodes can configure themselves according to the configuration parameters they receive, and report their own identity information to the resource exchange node after the configuration is completed; The identity acquisition module is used to obtain the identity identification information reported by N sub-nodes from the internal LAN of the rack. The internal LAN of the rack is a LAN established by the first management controller, N second management controllers, and the TOR network switch in the rack respectively connected through network links. Each second management controller is deployed on a sub-node. The identification module is used to establish the topology of the entire cabinet based on the identification information of each of the N child nodes.

12. A cabinet management device, characterized in that, The device, applied to a second management controller deployed on a child node, comprises: The receiving module is used to receive configuration parameters sent by the first management controller; The configuration module is used to configure based on the received configuration parameters; The reporting module is used to report its own identity information to the resource exchange node after completing the configuration, so that the first management controller can establish the topology of the entire cabinet based on the identity information; The reporting module includes: The identity reporting module publishes its own identity information in the local area network inside the rack. The local area network inside the rack is a network established by the first management controller, N second management controllers, and TOR network switches in the rack, which are connected to each other through network links.

13. A cabinet management device, characterized in that, The device is applied to a management client, which is connected to a TOR network switch within a server rack via an external management network. The device includes: An access module is used to access the first management controller through the external management network to obtain the topology of the entire cabinet, wherein the topology is generated according to any one of the methods described in claims 1-5; The management module is used to utilize the local area network inside the cabinet to call the communication interfaces on N child nodes in the topology to obtain device information of the N child nodes and manage the N child nodes.

14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes, it implements the rack management method as described in any one of claims 1 to 5, or the rack management method as described in any one of claims 6 to 8, or the rack management method as described in any one of claims 9 to 10.

15. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instruction is executed by the processor, it implements the rack management method as described in any one of claims 1 to 5, or the rack management method as described in any one of claims 6 to 8, or the rack management method as described in any one of claims 9 to 10.

Citation Information

Patent Citations

  • Topological structure generating method, device and system for distributed storage system

    CN108900421A

  • Node configuration method and device, distributed system and computer readable medium

    CN115004650A