Server control method, computing platform, computer program product, and device
Through the out-of-band management controller to detect component failures during the server startup phase, the detection lag problem after server startup is solved, early fault identification and rapid response are achieved, and detection efficiency and system stability are improved.
Patent Information
- Application Number
- CN202510838006.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-06-20
AI Technical Summary
In the prior art, the server only conducts component failure detection after starting and entering the operating system, which has obvious lag, resulting in waste of downtime and business continuity.
When the server power is started, the out-of-band management controller broadcasts request packets to each target component based on the PCIe bus, receives response packets, judges the faulty component and controls the stopping of the operating system to realize early fault detection.
Reliance on the operating system is reduced, detection efficiency is improved, detection risk is reduced, fault detection time is shortened, system resource consumption is reduced, and fault resolution is provided to facilitate timely resolution of faults.
Smart Images

Figure CN120353507B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of fault detection technology, and in particular to a server control method, computing platform, computer program product, and device. Background Art
[0002] During server operation, ensuring the proper functioning of every component is crucial. Conventional technologies often detect component failures (or anomalies) only after the server has booted up and entered the operating system, resulting in a significant lag. Once a component failure is detected, the server must be shut down for replacement, which not only wastes time but can also impact business continuity. Summary of the Invention
[0003] The present application provides a server control method, computing platform, computer program product and device to at least solve the problem in the related art that component fault detection has a large lag after the server is started and enters the operating system.
[0004] This application provides a server control method, which is applied to an out-of-band management controller, including:
[0005] When detecting that the server is powered on, broadcasting a request data packet to each target component based on a Peripheral Component Interconnect Express (PCIe) bus connection with each target component of the server; the target component is a component that supports a hardware management protocol based on the PCIe bus; the request data packet is used to request status information of the target component;
[0006] Within a target duration, based on each PCIe bus connection, receiving a response data packet containing status information sent by each target component based on the request data packet;
[0007] When it is determined based on the response data packets and the number of the response data packets that a faulty component exists in each target component in the server, the server is controlled to stop starting, so as to stop starting the operating system of the server.
[0008] The present application also provides a computing platform including at least one server and at least one out-of-band management controller; the out-of-band management controller is installed on the server; and the out-of-band management controller is used to execute the steps of the aforementioned server control method.
[0009] The present application also provides a computer program product, applied to an out-of-band management controller, comprising:
[0010] a broadcast module configured to, upon detecting that the server is powered on, broadcast a request data packet to each target component based on a Peripheral Component Interconnect Express (PCIe) bus connection with each target component of the server; the target component being a component that supports a hardware management protocol based on the PCIe bus; and the request data packet being used to request status information of the target component;
[0011] A receiving module is configured to receive, within a target duration and based on each PCIe bus connection, a response data packet containing status information sent by each target component based on the request data packet;
[0012] The control module is used for controlling the server to stop starting up so as to stop starting up the operating system of the server when it is determined based on each response data packet and the number of response data packets that there is a faulty component in each target component in the server.
[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned server control methods when executing the computer program.
[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned server control methods are implemented.
[0015] The present application also provides another computer program product, including a computer program, which implements the steps of any of the above-mentioned server control methods when executed by a processor.
[0016] Through the present application, when the out-of-band management controller detects that the power of the server is turned on, it broadcasts a request data packet to each target component of the server based on the PCIe bus connection between the target components to perform component detection. There is no need to perform detection after the server's operating system is started. After broadcasting the request data packet, based on the number of response data packets obtained within the target time length and each response data packet, it can be determined whether the server has a faulty component. That is, the fault detection process is earlier than the startup of the operating system, which reduces the dependence on the operating system. It can solve the problem of large lag in component fault detection in related technologies after the server startup is completed and the operating system is entered. The detection process is simple and direct. When a faulty component is found in the server, the server is directly controlled to stop starting to stop the server's operating system from starting, which reduces the risk of inaccurate detection due to operating system failure or instability, improves detection efficiency, reduces the time required for detection and system resource consumption, and provides great convenience for timely solving faults and reducing business interruption time. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flow chart of a method for controlling a server provided in an embodiment of the present application;
[0019] Figure 2 A schematic diagram of an architecture used in a server control method provided in an embodiment of the present application;
[0020] Figure 3 A second flow chart of a method for controlling a server provided in an embodiment of the present application;
[0021] Figure 4 A schematic diagram of the structure of a computer program product provided in an embodiment of the present application. DETAILED DESCRIPTION
[0022] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0023] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0024] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0025] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the server control method depends, the specific application environment architecture or specific hardware architecture is described here.
[0026] An embodiment of the present application provides a method for controlling a server, and the method is described in detail in conjunction with the execution flow of the method for controlling the server.
[0027] The server control method provided in the embodiment of the present application can be executed by the server's out-of-band management controller. The out-of-band management controller is a dedicated hardware chip independent of the server's main CPU (Central Processing Unit) and operating system, and is used to implement out-of-band management (OOB). That is, when the host loses power, the operating system crashes, or the network is interrupted, the server can still be monitored, controlled, and maintained through an independent channel.
[0028] In practical applications, the out-of-band management controller may be a baseboard management controller (BMC).
[0029] See also Figure 1 The present application implements a flow chart of a method for controlling a server, which is applied to an out-of-band management controller and includes the following steps 110, 120, and 130:
[0030] Step 110: When it is detected that the server is powered on, a request data packet is broadcast to each target component of the server based on the PCIe bus connection between the target components; the target component is a component that supports the hardware management protocol based on the PCIe bus; the request data packet is used to request status information of the target component.
[0031] The server in the embodiment of the present application can be any type of server, including but not limited to enterprise-level servers, data center servers, edge computing servers, etc.
[0032] The target component refers to a component that supports a hardware management protocol based on a Peripheral Component Interconnect Express (PCIe) bus. Specifically, the hardware management protocol may be MCTP over PCIe (Management Component Transport Protocol over PCI Express).
[0033] MCTP over PCIe is a core protocol for modern server hardware management. By reusing PCIe's high-performance links, it overcomes the bandwidth bottleneck of traditional management buses (such as I2C). It is particularly suitable for large-scale data centers and heterogeneous computing environments.
[0034] In practical applications, the target components include but are not limited to hard disks, network cards, and memory.
[0035] The out-of-band management controller (BMC) in the embodiment of the present application is a core management unit responsible for initiating interaction processes with each target component.
[0036] The relationship between the BMC and the server is usually one-to-one. In some special scenarios, it can also be one-to-many or many-to-one.
[0037] See also Figure 2 The embodiment of the present application provides an architectural diagram of a server control method. The architectural diagram is only used as an example. In actual applications, other architectures may also be used, and there is no limitation on this.
[0038] exist Figure 2 In the diagram, the system architecture includes a server having an out-of-band management controller. The server has multiple target components, including three hard disks (hard disk 1, hard disk 2, and hard disk 3), two network cards (network card 1 and network card 2), and four memory modules (memory module 1, memory module 2, memory module 3, and memory module 4). These target components all support the hardware management protocol MCTP over PCIE protocol and have corresponding protocol processing modules. The out-of-band management controller and the various target components of the server are connected via a Peripheral Component Interconnect Express (PCIe) bus.
[0039] To promptly detect faulty components in the server, the BMC performs fault detection on each target component in the server when the server is in the startup phase.
[0040] Specifically, when the BMC detects that the server is powered on, it first initializes the MCTP over PCIe protocol stack. When the server boots to a specific stage (such as after the BIOS self-test is completed), the BMC broadcasts a request packet to each target component of the server through the PCIe bus connection between the BMC and the target components. The request packet is used to request the status information of the target component.
[0041] The status information is used to indicate various operating status information of the corresponding target component, including but not limited to at least one of power supply status, storage status, temperature status, certificate and encryption status.
[0042] The request data packet is encapsulated based on a hardware management protocol (such as the MCTP over PCIe protocol), and includes identity information of an out-of-band management controller (such as an IP address), an interaction request instruction (requesting data interaction), and at least part of a local timestamp.
[0043] The BMC interacts with various components in the server by broadcasting request packets without relying on the operating system.
[0044] Step 120: Within the target duration, based on each PCIe bus connection, receive a response data packet containing status information sent by each target component based on the request data packet.
[0045] After each target component receives each request data packet based on the PCIe bus connection between it and the BMC, its internal protocol processing module parses the request data packet. If the target component is working normally, it will respond to the request data packet and generate a response data packet containing its own status information.
[0046] After generating a response data packet, each target component returns the response data packet to the BMC through the corresponding PCIe bus connection.
[0047] After broadcasting each request data packet, the BMC does not always wait for the response of each target component. There is a time limit. Specifically, the BMC will determine whether each target component will respond to each request data packet within the target duration (preset duration, such as 3s). The target duration can be determined by a timer. Within the target duration set by the timer, based on each PCIe bus connection, the BMC receives the response data packet containing status information sent by each target component based on the request data packet.
[0048] The response data packet includes component information and status information. The component information includes but is not limited to the component name (component identification), component model and at least part of the firmware version. The status information includes but is not limited to at least one of power status, storage status, temperature status, certificate and encryption status.
[0049] Of course, some target components may not receive the request data packet due to network delay or other failures. In this case, the target component will not send a response data packet.
[0050] Step 130: When it is determined based on the response data packets and the number of the response data packets that a faulty component exists in each target component in the server, the server is controlled to stop starting, so as to stop starting the operating system of the server.
[0051] The above embodiment has explained that the BMC will receive response data packets returned by each target component based on the PCIe bus connection within the target time. However, some target components may not send response data packets due to failures. Target components that do not send response data packets are faulty components.
[0052] Therefore, it is necessary to count the number of response packets received within the target duration. If the number of response packets is less than the number of target components in the server, it means that there is a faulty component in the server.
[0053] In addition, although some target components have sent response data packets, this does not mean that these target components are not faulty. The response data packets of these target components need to be verified to determine whether these target components are faulty components. If the verification fails, the target component is a faulty component; if the verification passes, the target component is not a faulty component but a normal component.
[0054] When the number of response data packets is equal to the number of target components in the server and all target components are normal components, the server continues to boot and enters the operating system.
[0055] If it is determined that a target component in the server has a faulty component, the BMC will send a control signal to the server to control the server to stop starting, thereby stopping the server's operating system from starting. That is, the target component of the server can be detected as early as the operating system loading stage, so that component failures can be quickly identified in the early stage of server startup. Compared with related technologies, problems can be discovered minutes or even tens of minutes earlier, greatly shortening the time window for fault discovery.
[0056] The control method of the server provided in the embodiment of the present application is applied to an out-of-band management controller. When the out-of-band management controller detects that the power of the server is started, it broadcasts a request data packet to each target component based on the PCIe bus connection between the target components of the server to perform component detection. There is no need to perform detection after the server's operating system is started. After broadcasting the request data packet, it can be determined whether the server has a faulty component based on the number of response data packets obtained within the target time length and each response data packet. That is, the fault detection process is earlier than the startup of the operating system, which reduces the dependence on the operating system. It can solve the problem of large lag in the related technology of performing fault detection on components after the server startup is completed and the operating system is entered. The detection process is simple and direct. When a faulty component is found in the server, the server is directly controlled to stop starting to stop the server's operating system startup, which reduces the risk of inaccurate detection due to operating system failure or instability, improves detection efficiency, reduces the time required for detection and system resource consumption, and provides great convenience for timely solving faults and reducing business interruption time.
[0057] In some embodiments, each request data packet includes component information; and determining, based on the number of each response data packet and each response data packet, whether a faulty component exists in each target component in the server includes:
[0058] Obtaining a target data table; the target data table is a data table that records component information and reference status of each target component in the server;
[0059] In a case where the number of response data packets is less than the total number of target components in the target data table, determining the target components that have not sent the response data packets from the target data table, and determining the target components that have not sent the response data packets as faulty components;
[0060] For the target component that sends the response data packet, verifying the response data packet of each target component based on the component information and reference state information of each target component in the target data table;
[0061] If the response data packet of the target component fails verification, the target component is determined to be a faulty component.
[0062] The BMC in an embodiment of the present application maintains a target data table, also known as an "installed component list." The target data table is a data table maintained by the BMC during power-on initialization. The target data table records component information and reference status data of each target component in the server. The component information includes, but is not limited to, at least part of the component name, component type, component model, and firmware version; the reference status information includes, but is not limited to, at least one of the power supply status under normal operation (e.g., voltage range under normal operation), storage status (e.g., memory usage under normal operation), temperature status (e.g., temperature under normal operation), and certificate and encryption status.
[0063] As explained in the previous examples, some target components may not send response packets due to failure. These target components that do not send response packets are considered faulty components. Therefore, it is necessary to count the number of response packets. If the number of response packets is less than the total number of target components in the target data table, this indicates that some target components in the server have not sent response packets. These target components that have not sent response packets are faulty components.
[0064] In addition, although some target components have sent response data packets, this does not mean that these target components are not faulty. The response data packets of these target components need to be verified to determine whether these target components are faulty components. The response data packets of each target component can be verified based on the component information and reference status information of each target component in the target data table. If the verification fails, the target component is a faulty component.
[0065] In addition, it is worth noting that the status information of different target components may be different. Accordingly, their reference status information is also different, and the verification method may also be different. In practical applications, verification can be performed according to actual conditions.
[0066] The BMC in the embodiment of the present application maintains a target data table that records the component information and reference status of each target component in the server and compares it with the actual response data packet. Based on the number of response data packets and the verification results of each response data packet, it can accurately determine whether there are faulty components in the server and accurately identify each component, thereby greatly improving the accuracy of fault judgment and effectively avoiding the occurrence of misjudgment.
[0067] In some embodiments, the component information includes a component name, a component model, and a firmware version; and determining, from the target data table, a target component that has not sent a response data packet includes:
[0068] For the target component that sends the response data packet, if the target data table includes the component name in the response data packet, or if the target data table does not include the component name in the response data packet but includes the component model and firmware version in the response data packet, the target component is determined to be the component to be verified;
[0069] It is determined that other components in the target data table except the component to be verified are target components that have not sent a response data packet.
[0070] For the target component that sends the response data packet, it can be determined first whether the target data table includes the component name in the response data packet.
[0071] If the target data table includes the component name in the corresponding response data packet, it means that the component name of the target component that sent the response data packet has not been tampered with, and the target component is the component to be verified that needs to be verified later.
[0072] If the target data table does not include the component name in the corresponding response data packet, it means that the component name of the target component that sent the response data packet may have been tampered with. In this case, in order to accurately and comprehensively identify the target component, it is possible to determine whether the component model and firmware version in the response data packet exist in the target data table.
[0073] The part model and firmware version of a component are fundamental metadata in operation and maintenance.
[0074] A part number is a number that uniquely identifies a specific hardware component. It's typically assigned by the manufacturer and contains information about the component's type, specifications, functions, dimensions, and more. Part numbers are commonly used in inventory management, procurement, and repair procedures to quickly and accurately identify hardware components. Part numbers typically consist of a combination of numbers and letters, with the specific format determined by the manufacturer.
[0075] A firmware version is the version number of the firmware (hardware-level software) built into a device or hardware component. Firmware controls and manages basic hardware functions, such as booting, operation, and communication. Firmware versions are often used to indicate different firmware releases and updates to ensure that devices utilize the latest features and fixes. Firmware versions typically consist of numbers, letters, and other characters.
[0076] The part model and firmware version together can also uniquely identify a part.
[0077] If the component model and firmware version exist in the target data table, it means that the target component is the component to be verified later; if the component model and firmware version do not exist in the target data table, it is directly confirmed that the response data packet verification fails and the corresponding target component is a faulty component.
[0078] After determining each component to be verified from the target components that send response data packets, other components in the target data table except the component to be verified can be used as target components that do not send response data packets. These target components that do not send response data packets are faulty components, and alarms can be issued for the faulty components. See the subsequent content for details.
[0079] The embodiment of the present application determines the target component that sends the response data packet based on the component name, component model and firmware version in the response data packet to screen the components to be verified, and then can determine that other components in the target data table except the components to be verified are target components that have not sent response data packets. When the component name is incorrectly modified, each component can still be accurately identified through the component model and component version in the component information, thereby achieving accurate and comprehensive determination of the faulty component, greatly improving the accuracy of fault judgment, and effectively avoiding the occurrence of misjudgment.
[0080] In some embodiments, verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table includes:
[0081] For a first component to be verified among the target components that sent the response data packet, the component name of the first component to be verified is located in the target data table, and the state information of the first component to be verified in the response data packet is compared with the reference state of the first component to be verified in the target data table to obtain a first comparison result;
[0082] If the first comparison result indicates that the status information in the response data packet is inconsistent with the reference status, determining that the response data packet fails verification;
[0083] If the first comparison result indicates that the status information in the response data packet is consistent with the reference status, it is determined that the response data packet passes verification.
[0084] The above embodiment has explained that it is necessary to verify the sent response data packet. For each target component that sends the response data packet, if the component name in the response data packet is located in the target data table, it means that the component name of the target component has not been tampered with. In this case, the target component can be called the first component to be verified.
[0085] The state information of the first component to be verified in the response data packet may be compared with the reference state of the first component to be verified in the target data table to obtain a first comparison result.
[0086] The above embodiments have shown that the operating information of different target components may be different, and the corresponding reference states may also be different. For example, the operating information of the target component memory may be memory usage, and the operating information of the target component hard disk may be temperature status and storage utilization, etc.
[0087] When the first comparison result indicates that the status information in the response data packet is inconsistent with the reference status, it is determined that the response data packet verification fails; when the first comparison result indicates that the status information in the response data packet is consistent with the reference status, it is determined that the response data packet verification passes.
[0088] In an embodiment of the present application, when the component name in the response data packet is located in the target data table, the status information of the first component to be verified in the response data packet is compared with the reference status of the first component to be verified in the target data table, so that the response data packet can be verified accurately and quickly.
[0089] In some embodiments, verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table further includes:
[0090] For a second component to be verified in each target component for which a response data packet is sent, if the component name of the second component to be verified is not located in the target data table, but the component model and component version of the second component to be verified are located in the target data table, comparing the status information of the second component to be verified in the response data packet with the reference status of the second component to be verified in the target data table to obtain a second comparison result;
[0091] If the second comparison result indicates that the state information of the second component to be verified in the response data packet is inconsistent with the reference state of the second component to be verified in the target data table, determining that the response data packet fails verification;
[0092] If the second comparison result indicates that the status information of the second component to be verified in the response data packet is consistent with the reference status of the second component to be verified in the target data table, it is determined that the response data packet passes verification.
[0093] In addition to the first component to be verified mentioned above, there may be a second component to be verified in the target component sending the response data packet. The component name of the second component to be verified is not located in the target data table, but the component type and component version of the second component to be verified are located in the target data table, that is, the status information of the second component to be verified also needs to be verified.
[0094] Similarly, the status information of the second component to be verified in the response data packet can be compared with the reference status of the second component to be verified in the target data table to obtain a second comparison result. When the second comparison result indicates that the status information of the second component to be verified in the response data packet and the reference status of the second component to be verified in the target data table are inconsistent, it is determined that the response data packet verification has failed; when the second comparison result indicates that the status information of the second component to be verified in the response data packet and the reference status of the second component to be verified in the target data table are consistent, it is determined that the response data packet verification has passed.
[0095] In the embodiment of the present application, even if the component name of the target component in the response data packet is tampered with, the target component can also be identified by the component model and firmware version. When it is determined that the target component is the second component to be verified whose component model and component version are located in the target data table, the status information of the second component to be verified in the response data packet is compared with the reference status of the second component to be verified in the target data table. This can accurately and comprehensively identify and verify each component, greatly improving the accuracy of fault judgment and effectively avoiding the occurrence of misjudgment. When the component name is incorrectly modified, it is still possible to accurately judge whether the component is faulty and the specific faulty component based on the component model and firmware version, thereby improving the reliability of server hardware management.
[0096] In some embodiments, verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table further includes:
[0097] For the target component of the response data packet, if the component name of the target component in the response data packet does not exist in the target data table, and the component model and firmware version in the response data packet do not exist, it is determined that the response data packet verification fails.
[0098] For any response data packet, first determine whether the component name in the response data packet is located in the target data table. If not, further determine whether the component model and firmware version in the response data packet are located in the target data table. The two-round judgment process can accurately identify each target component and effectively avoid the occurrence of misjudgment. If the component model and firmware version in the response data packet are also not located in the target data table, it can be directly determined that the verification of the response data packet has failed, and there is no need to verify the operation information, thereby improving the verification efficiency.
[0099] In some embodiments, after broadcasting the request data packet to each target component based on the PCIe bus connection between the target components of the server, the method further includes:
[0100] Start the timer; the duration set by the timer is the target duration;
[0101] Receiving a response data packet containing status information sent by each target component based on the request data packet, including:
[0102] When the timer times out, the reception of the response data packets containing the status information sent by each target component based on the request data packet is stopped.
[0103] The above embodiments have explained that, in addition to the broadcast request data packet, the BMC is not unlimited in receiving each response data packet. Each response data packet needs to be received within a target duration. To this end, a timer can be started and the duration of the timer is set to the target duration. When the timer has not timed out, the response data packet containing status information sent by each target component based on the request data packet is received; when the timer times out, the response data packet containing status information sent by each target component based on the request data packet is stopped. In this way, the accuracy of determining whether the target component is a faulty component can be improved.
[0104] In some embodiments, broadcasting a request packet to each target component based on a PCIe bus connection with each target component of the server includes:
[0105] When the target signal is obtained, the target signal is used to indicate that the basic input and output system self-test of the server is completed and the protocol stack of the hardware management protocol based on the PCIe bus is initialized;
[0106] After determining that the protocol stack initialization is completed, a request data packet is broadcast to each target component based on the PCIe bus connection between the target components of the server.
[0107] After the server in the embodiment of the present application is powered on, the server's Basic Input / Output System (BIOS) performs a self-test. After receiving a target signal indicating the completion of the server's BIOS self-test, the BMC initializes the protocol stack of a PCIe bus-based hardware management protocol (such as MCTP over PCIe). This ensures that the system hardware has been fully initialized and that the protocol stack is initialized when the hardware is stable, avoiding initialization of the protocol stack when the hardware is not yet stable. This improves system stability, avoids resource conflicts, and enables the MCTP protocol stack to operate normally in a reliable hardware environment.
[0108] After the MCTP over PCIe protocol stack is initialized, the BMC broadcasts a request packet to each target component of the server based on the PCIe bus connection between the target components, ensuring that the system is in the optimal state when the management and monitoring functions are started.
[0109] In some embodiments, before broadcasting the request data packet to each target component based on the PCIe bus connection between the target components of the server, the method further includes:
[0110] Acquire information to be encapsulated; the information to be encapsulated includes at least part of the identity information of the out-of-band management controller, the interaction request instruction, and the local timestamp;
[0111] The information to be encapsulated is encapsulated based on the hardware management protocol to obtain the request data packet.
[0112] In the embodiment of the present application, the request data packet is generated by the BMC based on the following method: the information to be encapsulated can be obtained first; the information to be encapsulated includes the identity information of the out-of-band management controller, the interactive request instruction, and at least part of the local timestamp. The identity information can uniquely indicate the out-of-band management controller and can be the identification information of the out-of-band management controller; the interactive request instruction is used to request the return of status information, and the local timestamp is generally used to record the time when the data packet is sent or received at a specific moment, and is used to measure the delay of the data packet from one node to another. Recording the timestamp in the request data packet can help the recipient calculate the time difference from the arrival of the request, thereby estimating the network transmission delay.
[0113] The above-mentioned information to be encapsulated can be encapsulated based on a hardware management protocol (such as MCTP over PCIe) to obtain a request data packet, which can more efficiently manage hardware and monitor the system while providing a more stable and reliable system operating environment. The request data packet is broadcasted to improve device discoverability and simplify the device discovery process.
[0114] In some embodiments, after determining that a faulty component exists in each target component in the server, the method further includes:
[0115] Reporting fault information of the faulty component to the remote management terminal via the network; the fault information includes at least part of the component name, component model, firmware version, and expected location of the component;
[0116] Output alarm information; alarm information is used to instruct the repair or replacement of faulty components.
[0117] In addition to determining the faulty component in the server, the BMC in the embodiment of the present application not only prevents the server from continuing to boot into the operating system to avoid system instability or failure caused by the faulty component, but also reports the fault information of the faulty component to the remote management terminal via the network.
[0118] A remote management terminal is a key management tool that allows administrators to remotely monitor, manage, and control server hardware and operating systems over the network. It can be a remote management tool such as KVM (Keyboard-Video-Mouse).
[0119] Fault information can include component information such as component name, component model, and firmware version, as well as status information and the component's expected location. The expected location of the component refers to the theoretical installation location of the component.
[0120] In addition, an alarm message is output for the faulty component. The alarm message is used to instruct the repair or replacement of the faulty component. The alarm message may also include the above-mentioned fault information.
[0121] Based on the fault information and alarm information, the operation and maintenance personnel of this application can open the server chassis, quickly locate and replace the faulty component. After replacing the component, the server restarts, and the BMC performs the above interaction and detection process again. After all components are working normally, the server successfully enters the operating system.
[0122] After the BMC in the embodiment of the present application determines the faulty component in the server, it reports the fault information of the faulty component to the remote management terminal through the network and outputs alarm information, which can promptly inform the operation and maintenance personnel to solve the fault and improve the reliability of server hardware management.
[0123] In some embodiments, outputting the warning information includes at least one of the following:
[0124] Display alarm information through a display or remote management terminal;
[0125] Play alarm information through voice player;
[0126] Output warning information by controlling the indicator light to flash;
[0127] Output alarm information by controlling the buzzer alarm.
[0128] The embodiments of the present application can output alarm information in a variety of ways, including but not limited to displaying alarm information through a display or remote management terminal; broadcasting alarm information through a voice player; outputting alarm information by controlling the flashing of indicator lights; and outputting alarm information by controlling the alarm of a buzzer, thereby improving the visibility, responsiveness and timeliness of alarms, ensuring that key problems can be discovered and handled as early as possible, and avoiding further expansion of potential risks or failures.
[0129] See also Figure 3 , the embodiment of the present application provides a second flow chart of a method for controlling a server, comprising the following steps:
[0130] Step 301: Detecting that the server is powered on;
[0131] Step 302: Initialize the protocol stack of the hardware management protocol based on the PCIe bus when the target signal is obtained; the target signal is used to indicate that the basic input and output system self-test of the server is completed;
[0132] Step 303: After determining that the protocol stack initialization is complete, broadcast a request data packet to each target component based on the PCIe bus connection between the target components of the server;
[0133] Step 304: Start the timer; the duration set by the timer is the target duration;
[0134] Step 305: If the timer has not timed out, receiving a response data packet containing status information sent by each target component based on the request data packet based on each PCIe bus connection;
[0135] Step 306, when the timer times out, determining the number of response data packets;
[0136] Step 307, obtaining a target data table; the target data table is a data table that records component information and reference status of each target component in the server; execute steps 308 and 309;
[0137] Step 308, if the number of response data packets is less than the total number of target components in the target data table, determine the target components that have not sent response data packets from the target data table, and determine that the target components that have not sent response data packets are faulty components; then execute step 311;
[0138] Step 309 , for the target component that sends the response data packet, verify the response data packet of each target component based on the component information and reference state information of each target component in the target data table;
[0139] Step 310: If the response data packet of the target component fails to pass the verification, the target component is determined to be a faulty component; and step 311 is executed;
[0140] Step 311: Control the server to stop booting, so as to stop the server's operating system from booting;
[0141] Step 312: Report the fault information of the faulty component to the remote management terminal via the network and output an alarm message; the fault information includes at least part of the component name, component model, firmware version, status information, and the expected location of the component; the alarm message is used to instruct the repair or replacement of the faulty component.
[0142] The detailed execution process of the above steps 301 to 312 is described in the above embodiment and will not be repeated here.
[0143] The embodiment of the present application provides a specific embodiment in which a server is installed with three hard drives that support the MCTP over PCIe protocol, two network cards, and four memory modules.
[0144] When the server is powered on and the BIOS self-test completes, the BMC begins initializing the MCTP over PCIe protocol stack. After initialization, the BMC constructs a request packet containing its own identity information (such as the BMC model and serial number), an interaction request (requesting the component to return its status information), and the current timestamp.
[0145] The BMC broadcasts the request data packet through the PCI Express bus.
[0146] After receiving the request data packet, the protocol processing modules of the three hard disks, two network cards and four memory modules parse it respectively.
[0147] After receiving the request packet, hard drive 1 generates a response packet containing its own model (for example, Seagate ST4000DM004), firmware version (CC49), and status information (normal). This response packet is then returned to the BMC via the PCIe bus. Similarly, the other hard drives, network adapters, and memory modules each generate and return a response packet.
[0148] After broadcasting the request packet, the BMC starts a timer, assuming the target duration is 5 seconds. During these 5 seconds, the BMC continuously receives response packets from various components. When the timer expires, the BMC compares the installed components list (which records component information and reference status for three hard drives, two network cards, and four memory modules) with the actual response packets received. If the BMC only receives response packets from two hard drives, two network cards, and four memory modules, it determines that one hard drive has failed. The BMC records information about the failed hard drive, which is expected to be a Seagate ST4000DM004 hard drive located in hard drive slot 2. Simultaneously, the BMC controls the hard drive fault indicator on the server front panel to flash and sends the fault information to the remote management terminal via the network, alerting maintenance personnel that the hard drive in hard drive slot 2 has failed.
[0149] Based on the fault information provided by the BMC, the maintenance personnel opened the server chassis and replaced the hard drive in slot 2. After the replacement, the server restarted. The BMC repeated the aforementioned interaction and detection process. After confirming that all components were responding normally, the server successfully entered the operating system.
[0150] The above examples demonstrate that the present invention can effectively detect component failures during server startup. Applying the MCTP over PCIe protocol to component failure detection during server startup opens up new technical application directions, differing from traditional detection methods at the operating system level. This improves server reliability and operational efficiency.
[0151] In addition, the fault judgment logic based on broadcast interaction and quantity comparison is simple, direct, efficient and accurate, and can quickly locate faulty components, greatly improving the efficiency of server hardware management.
[0152] In addition, the above detection process is applicable to various types of servers, whether they are enterprise-level servers, data center servers or edge computing servers, etc. It can effectively reduce server operation and maintenance costs, improve the overall performance and stability of the server, and has broad market application prospects.
[0153] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0154] An embodiment of the present application also provides a computing platform, which includes at least one server and at least one out-of-band management controller; the out-of-band management controller is installed on the server; the out-of-band management controller is used to execute the process of the control method of the above-mentioned server. The detailed process can be found in the above-mentioned embodiment and will not be repeated here.
[0155] An embodiment of the present application further provides a computer program product applied to an out-of-band management controller.
[0156] See also Figure 4 , the embodiment of the present application provides a structural diagram of a computer program product, including: a broadcast module 410, a receiving module 420 and a control module 430, wherein,
[0157] The broadcast module 410 is configured to, upon detecting that the server is powered on, broadcast a request data packet to each target component of the server based on the PCIe bus connection with the target component; the target component is a component that supports a hardware management protocol based on the PCIe bus; the request data packet is used to request status information of the target component;
[0158] The receiving module 420 is configured to receive, within a target duration and based on each PCIe bus connection, a response data packet containing status information sent by each target component based on the request data packet;
[0159] The control module 430 is configured to control the server to stop starting up, if it is determined based on the response data packets and the number of the response data packets that there is a faulty component in each target component in the server, so as to stop starting up the operating system of the server.
[0160] In some embodiments, the request data packet includes component information; and the computer program product further comprises:
[0161] Processing module for:
[0162] Obtaining a target data table; the target data table is a data table that records component information and reference status of each target component in the server;
[0163] In a case where the number of response data packets is less than the total number of target components in the target data table, determining the target components that have not sent the response data packets from the target data table, and determining the target components that have not sent the response data packets as faulty components;
[0164] For the target component that sends the response data packet, verifying the response data packet of each target component based on the component information and reference state information of each target component in the target data table;
[0165] If the response data packet of the target component fails verification, the target component is determined to be a faulty component.
[0166] In some embodiments, the component information includes component name, component model, and firmware version; the processing module is further configured to:
[0167] For the target component that sends the response data packet, if the target data table includes the component name in the response data packet, or if the target data table does not include the component name in the response data packet but includes the component model and firmware version in the response data packet, the target component is determined to be the component to be verified;
[0168] It is determined that other components in the target data table except the component to be verified are target components that have not sent a response data packet.
[0169] In some embodiments, the processing module is further configured to:
[0170] For a first component to be verified among the target components that sent the response data packet, the component name of the first component to be verified is located in the target data table, and the state information of the first component to be verified in the response data packet is compared with the reference state of the first component to be verified in the target data table to obtain a first comparison result;
[0171] If the first comparison result indicates that the status information in the response data packet is inconsistent with the reference status, determining that the response data packet fails verification;
[0172] If the first comparison result indicates that the status information in the response data packet is consistent with the reference status, it is determined that the response data packet passes verification.
[0173] In some embodiments, the processing module is further configured to:
[0174] For the target component of the response data packet, if the component name of the target component in the response data packet does not exist in the target data table, and the component model and firmware version in the response data packet do not exist, it is determined that the response data packet verification fails.
[0175] In some embodiments, the processing module is further configured to:
[0176] Start the timer; the duration set by the timer is the target duration;
[0177] The receiving module 420 is specifically configured to stop receiving the response data packets containing status information sent by each target component based on the request data packet when the timer times out.
[0178] In some embodiments, the broadcast module 410 is used to:
[0179] When the target signal is obtained, the target signal is used to indicate that the basic input and output system self-test of the server is completed and the protocol stack of the hardware management protocol based on the PCIe bus is initialized;
[0180] After determining that the protocol stack initialization is completed, a request data packet is broadcast to each target component based on the PCIe bus connection between the target components of the server.
[0181] In some embodiments, the computer program product further comprises:
[0182] Encapsulated modules for:
[0183] Acquire information to be encapsulated; the information to be encapsulated includes at least part of the identity information of the out-of-band management controller, the interaction request instruction, and the local timestamp;
[0184] The information to be encapsulated is encapsulated based on the hardware management protocol to obtain the request data packet.
[0185] In some embodiments, the computer program product further comprises:
[0186] A reporting module, configured to report fault information of a faulty component to a remote management terminal via a network; the fault information includes at least part of the component name, component model, firmware version, status information, and expected location of the component;
[0187] The alarm output module is used to output alarm information; the alarm information is used to instruct the repair or replacement of faulty components.
[0188] In some embodiments, the alarm output module is configured to perform at least one of the following:
[0189] Display alarm information through a display or remote management terminal;
[0190] Play alarm information through voice player;
[0191] Output warning information by controlling the indicator light to flash;
[0192] Output alarm information by controlling the buzzer alarm.
[0193] For the description of the features in the embodiments corresponding to the computer program product, please refer to the relevant description of the embodiments corresponding to the server control method, which will not be repeated here.
[0194] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned server control method embodiments.
[0195] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps of any of the above-mentioned server control method embodiments when running.
[0196] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0197] An embodiment of the present application further provides another computer program product, which includes a computer program. When the computer program is executed by a processor, the steps of any of the above-mentioned server control method embodiments are implemented.
[0198] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above-mentioned server control method embodiments are implemented.
[0199] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0200] The above is a detailed introduction to the control method, computing platform, computer program product and device of a server provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only intended to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for controlling a server, characterized in that: Applicable to out-of-band management controllers, including: When detecting that the server is powered on, broadcasting a request data packet to each target component based on a Peripheral Component Interconnect Express (PCIe) bus connection with each target component of the server; the target component is a component that supports a hardware management protocol based on the PCIe bus; the request data packet is used to request status information of the target component; Within a target duration, based on each PCIe bus connection, receiving a response data packet containing status information sent by each target component based on the request data packet; When it is determined based on the response data packets and the number of the response data packets that a faulty component exists in each target component in the server, controlling the server to stop starting, so as to stop starting an operating system of the server; The determining, based on each of the response data packets and the number of the response data packets, that a faulty component exists in each target component in the server includes: Obtaining a target data table; the target data table is a data table that records component information and reference status of each target component in the server; In a case where the number of the response data packets is less than the total number of target components in the target data table, determining the target components that have not sent the response data packets from the target data table, and determining that the target components that have not sent the response data packets are faulty components; For the target component that sends the response data packet, verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table; If the response data packet of the target component fails verification, the target component is determined to be a faulty component.
2. The server control method according to claim 1, wherein: The component information includes a component name, a component model, and a firmware version; and determining the target component that has not sent the response data packet from the target data table includes: For the target component that sent the response data packet, if the target data table includes the component name in the response data packet, or if the target data table does not include the component name in the response data packet but includes the component model and firmware version in the response data packet, determining that the target component is a component to be verified; It is determined that other components in the target data table except the component to be verified are target components that have not sent the response data packet.
3. The server control method according to claim 2, characterized in that: The verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table includes: For a first component to be verified among the target components that sent the response data packet, the component name of the first component to be verified is located in the target data table, comparing the status information of the first component to be verified in the response data packet with the reference status of the first component to be verified in the target data table to obtain a first comparison result; If the first comparison result indicates that the status information in the response data packet is inconsistent with the reference status, determining that the response data packet fails verification; If the first comparison result indicates that the status information in the response data packet is consistent with the reference status, it is determined that the response data packet passes verification.
4. The server control method according to claim 1, wherein: The verifying of the response data packet of each target component based on the component information and reference status information of each target component in the target data table further includes: For the target component that sends the response data packet, if the component name of the target component in the response data packet does not exist in the target data table, and the component model and firmware version in the response data packet do not exist, it is determined that the response data packet verification fails.
5. The server control method according to claim 1, wherein: After broadcasting the request data packet to each target component based on the peripheral component interconnect extended (PCIe) bus connection between the target components of the server, the method further includes: Start a timer; the duration set by the timer is the target duration; The receiving of a response data packet containing status information sent by each target component based on the request data packet comprises: When the timer times out, the receiving of the response data packets containing the status information sent by each target component based on the request data packet is stopped.
6. The server control method according to claim 1, characterized in that: The method of broadcasting a request data packet to each target component based on a peripheral component interconnect (PCIe) bus connection with each target component of the server includes: When a target signal is obtained, the target signal is used to indicate that a basic input / output system self-check of the server is completed and a protocol stack of a hardware management protocol based on a PCIe bus is initialized; After determining that the protocol stack is initialized, a request data packet is broadcast to each target component based on the PCIe bus connection between the target components of the server.
7. The server control method according to claim 1, wherein: Before broadcasting the request data packet to each target component based on the peripheral component interconnect extended (PCIe) bus connection between the target components of the server, the method further includes: Acquire information to be encapsulated; the information to be encapsulated includes at least part of the identity information of the out-of-band management controller, the interaction request instruction, and the local timestamp; The information to be encapsulated is encapsulated based on the hardware management protocol to obtain the request data packet.
8. The server control method according to claim 1, characterized in that: After determining that a faulty component exists in each target component in the server, the method further includes: Reporting fault information of the faulty component to a remote management terminal via a network; the fault information including at least part of the component name, component model, firmware version, status information, and expected location of the component; Outputting an alarm message; the alarm message is used to instruct the repair or replacement of the faulty component.
9. The server control method according to claim 8, characterized in that: The output alarm information includes at least one of the following: Displaying the alarm information via a display or a remote management terminal; Broadcast the alarm information through a voice player; Outputting the warning information by controlling the indicator light to flash; The alarm information is output by controlling the buzzer to alarm.
10. A computing platform, characterized in that: The computing platform includes at least one server and at least one out-of-band management controller; the out-of-band management controller is installed on the server; the out-of-band management controller is used to execute the server control method according to any one of claims 1-9.
11. A computer program product, characterized in that Applicable to out-of-band management controllers, including: a broadcast module configured to, upon detecting that the server is powered on, broadcast a request data packet to each target component of the server based on a Peripheral Component Interconnect Express (PCIe) bus connection with each target component of the server; the target component being a component supporting a hardware management protocol based on the PCIe bus; the request data packet being used to request status information of the target component; A receiving module, configured to receive, within a target duration and based on each PCIe bus connection, a response data packet containing status information sent by each target component based on the request data packet; Processing module for: Obtaining a target data table; the target data table is a data table that records component information and reference status of each target component in the server; In a case where the number of the response data packets is less than the total number of target components in the target data table, determining the target components that have not sent the response data packets from the target data table, and determining that the target components that have not sent the response data packets are faulty components; For the target component that sends the response data packet, verifying the response data packet of each target component based on the component information and reference status information of each target component in the target data table; If the response data packet of the target component fails verification, determining that the target component is a faulty component; The control module is used to control the server to stop starting up so as to stop starting up the operating system of the server when it is determined based on each of the response data packets and the number of the response data packets that there is a faulty component in each target component in the server.
12. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the server control method according to any one of claims 1 to 9 when executing the computer program.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the server control method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Server hardware power-on starting fault troubleshooting method, system and device and medium
CN115470056A
Fault processing method, server system and storage medium
CN119311446A