Control method and apparatus for server, storage medium, and electronic device
Patent Information
- Application Number
- CN202210999614.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-19
- Publication Date
- 2026-08-21
- Estimated Expiration
- 2042-08-19
AI Technical Summary
[0006]本发明实施例提供了一种服务器的控制方法和装置、存储介质及电子设备,以至少解决服务器恢复上电后出现的电力负荷过大的技术问题
[0021]根据本发明实施例的又一方面,还提供了一种电子设备,包括存储器和处理器,上述存储器中存储有计算机程序,上述处理器被设置为通过计算机程序执行上述服务器的控制方法。
Smart Images

Figure CN117640339B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically, to a server control method and apparatus, storage medium, and electronic device. Background Technology
[0002] With the continuous development of cloud computing, the convenience it brings to various industries is becoming increasingly apparent, while the demand for cloud computing and big data is also constantly increasing. To meet this demand, large enterprises in various regions are continuously expanding their data centers, and the number of servers in each data center is also increasing significantly. In the event of a large-scale power outage in a data center, how to ensure that a large number of servers can start normally is an important issue that needs to be considered at this stage.
[0003] In related technologies, it is common practice to set all servers in a data center to power-on auto-start mode to ensure that large-scale servers can be automatically started when the servers are powered on again.
[0004] However, starting the server after power restoration using the above method not only depends on the server's power-on auto-start mode setting being effective, but also the simultaneous power-on of a large number of servers may cause the voltage of the rack-side power control equipment (load switch) to exceed the preset threshold, resulting in the technical problem of excessive power load after the server is restored.
[0005] There is currently no effective solution to the above problems. Summary of the Invention
[0006] This invention provides a server control method and apparatus, storage medium and electronic device to at least solve the technical problem of excessive power load after the server is powered on.
[0007] According to one aspect of the present invention, a server control method is provided, comprising: acquiring a list of servers that have triggered alarms and a set of server status information, wherein the set of server status information includes status information of servers in the server list; when it is determined from the set of server status information that a power outage has occurred in the data center where the servers in the server list are located, and a power failure has occurred in the servers in the server list, when it is detected that the servers in the server list have been restored to power, determining whether the server list includes servers in a first group of servers, wherein the data center includes a first group of servers and a second group of servers, the first group of servers is configured to start automatically after power restoration, and the second group of servers is configured to be prohibited from starting automatically after power restoration; when it is determined that the server list includes a first set of servers in the first group of servers, and a server that has failed to start is detected in the first set of servers, sending a remote start command to the server that failed to start, wherein the remote start command is used to control the server that failed to start to start.
[0008] Optionally, the above method further includes: when it is determined from the server status information set that the data center where the server in the server list is located has experienced a power outage and the server in the server list has lost power, sending a target detection command to the baseboard management controller through the network access address of the baseboard management controller in the server list, wherein the baseboard management controller is set to a static mode, and the network access address of the baseboard management controller in the static mode remains unchanged before the server in the server list loses power and after power is restored; and upon receiving the response information returned by the baseboard management controller in response to the target detection command, determining that the server in the server list has been restored to power.
[0009] Optionally, the above method further includes: when it is determined that the server list includes a first server set in the first group of servers, sending a status acquisition instruction to the baseboard management controller in each server through the network access address of the baseboard management controller in each server in the first server set, wherein the baseboard management controller in each server is set to a static mode, and the network access address of the baseboard management controller in the static mode in each server remains unchanged before the power is cut off and after the power is restored; acquiring the startup status of the baseboard management controller in each server in response to the status acquisition instruction; and determining the servers in the first server set whose startup status is not started as servers that have failed to start.
[0010] Optionally, the above method further includes: when a server in the server list is detected to have resumed power-on, determining whether the server list includes servers in the second group of servers; if it is determined that the server list includes the second set of servers in the second group of servers, remotely controlling the servers in the second set of servers to start up in batches, wherein servers in the second set of servers located on the same rack are started up in batches, servers located on the same rack are connected to the same power control device, and the power control device is used to control the servers located on the same rack to shut down when the power load generated by the servers located on the same rack is greater than or equal to a first preset threshold.
[0011] Optionally, the remote control of the servers in the second server set is performed in batches, including: when the second server set includes multiple servers located on the target rack, dividing the multiple servers located on the target rack into N groups of servers, wherein the power load generated by the simultaneous startup of each group of servers in the N groups of servers is less than a first preset threshold, and N is a positive integer greater than or equal to 2; sending startup instructions to each group of servers in the N groups of servers in N batches, wherein the startup instructions are used to control the startup of the server that receives the startup instructions.
[0012] Optionally, the above-mentioned division of multiple servers located on the target rack into N groups of servers includes: determining the services processed by each server among the multiple servers to obtain a target service set, wherein each service in the target service set is different and is processed by at least one server; and dividing the multiple servers located on the target rack into N groups of servers according to the number of services in the target service set.
[0013] Optionally, the above-mentioned division of multiple servers located on the target rack into N groups of servers based on the number of services in the target service set includes: when the number of services is P and P is less than M, dividing P servers from the multiple servers that process each service in the target service set into a group of servers, wherein the power load generated by the simultaneous startup of M servers located on the target rack is greater than or equal to a first preset threshold, and the power load generated by the simultaneous startup of M-1 servers located on the target rack is less than the first preset threshold, where M is a positive integer greater than or equal to 2, and the resulting group of servers is the group of servers that receives the startup command first among the N groups of servers; or when the number of services is P and P is greater than or equal to M, selecting M-1 servers from the multiple servers as a group of servers, wherein the M-1 servers are respectively used to process M-1 services in the target service set.
[0014] Optionally, the above selection of M-1 servers from multiple servers includes: determining a target type server among the multiple servers, wherein the service processed by the target type server is being processed by an already started server; and selecting M-1 servers from the multiple servers excluding the target type server when the number of servers other than the target type server is greater than or equal to M-1, wherein the selected M-1 servers are the group of servers that received the start instruction first among the N groups of servers.
[0015] Optionally, sending startup instructions to each of the N groups of servers in N batches includes: performing the following operations N times: randomly selecting a group of servers from the N groups of servers that have not yet sent startup instructions, and sending startup instructions to the randomly selected group of servers; or performing the following operations N times: selecting the group of servers with the largest number of servers from the N groups of servers that have not yet sent startup instructions, and sending startup instructions to the selected group of servers.
[0016] Optionally, after obtaining the list of servers that have triggered alarms and the set of server status information, the above method further includes: determining that the data center where the servers in the server list are located has experienced a power anomaly and that the servers in the server list have experienced a power outage when the number of servers in the server list is greater than or equal to a second preset threshold, the target power log indicates that the servers in the server list have experienced a power outage, and the target probe response information indicates that the servers in the server list cannot respond to the target probe command. The set of server status information includes the target power log and the target probe response information.
[0017] According to another aspect of the present invention, a server control device is also provided, comprising: a first acquisition unit, configured to acquire a list of servers that have triggered alarms and a set of server status information, wherein the set of server status information includes status information of servers in the server list; a first processing unit, configured to, when determining, based on the set of server status information, that a power outage has occurred in the data center where the servers in the server list are located, and that the servers in the server list have experienced a power failure, determine whether the server list includes servers from a first group of servers when it is detected that the servers in the server list have been restored to power, wherein the data center includes a first group of servers and a second group of servers, the first group of servers is configured to automatically start after power restoration, and the second group of servers is configured to be prevented from automatically starting after power restoration; and a second processing unit, configured to, when determining that the server list includes a first set of servers from the first group of servers, and detecting that there is a server in the first set of servers that has failed to start, send a remote start command to the server that failed to start, wherein the remote start command is used to control the server that failed to start to start.
[0018] Optionally, the above apparatus further includes: a third processing unit, configured to send a target detection command to a baseboard management controller via the network access address of the baseboard management controller in the server list when it is determined from the server status information set that the data center where the server in the server list is located has experienced a power anomaly and the server in the server list has lost power; wherein the baseboard management controller is set to a static mode, and the network access address of the baseboard management controller in the static mode remains unchanged before the server in the server list loses power and after the server is restored to power; and a fourth processing unit, configured to determine that the server in the server list has been restored to power upon receiving the response information returned by the baseboard management controller in response to the target detection command.
[0019] According to another aspect of the present invention, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, and the computer program is configured to execute the control method of the server described above when it is run.
[0020] According to another aspect of the present invention, a computer program product is also provided, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0021] According to another aspect of the present invention, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to execute the control method of the server through the computer program.
[0022] In this embodiment of the invention, by setting the state of the first group of servers in the data center to automatically start after power-on and the state of the second group of servers to disable automatic startup after power-on, the first and second groups of servers are powered on in a staggered manner when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can start normally, reducing the dependence on the automatic startup setting of servers after power-on and achieving the technical effect of improving the reliability of the first group of servers' automatic startup after power-on. Attached Figure Description
[0023] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:
[0024] Figure 1 This is a schematic diagram illustrating an application scenario of an optional server control method according to an embodiment of the present invention;
[0025] Figure 2 This is a flowchart of an optional server control method according to an embodiment of the present invention;
[0026] Figure 3 This is a schematic diagram of the distribution of servers in an optional data center according to an embodiment of the present invention;
[0027] Figure 4 This is a schematic diagram of an optional server startup mode setting method according to an embodiment of the present invention;
[0028] Figure 5 This is a schematic diagram of an optional BMC static mode setting method according to an embodiment of the present invention;
[0029] Figure 6 This is a schematic diagram of an optional support server remote power-on self-start according to an embodiment of the present invention;
[0030] Figure 7 This is a schematic diagram illustrating an optional method of remotely restarting a support server via a static IP address, according to an embodiment of the present invention.
[0031] Figure 8 This is a schematic diagram of an optional remote power-on startup of a business server according to an embodiment of the present invention;
[0032] Figure 9 This is a schematic diagram of an optional grouping method for a service server according to an embodiment of the present invention;
[0033] Figure 10 This is a schematic diagram of another optional grouping method for a service server according to an embodiment of the present invention;
[0034] Figure 11 This is an overall flowchart of an optional server control method according to an embodiment of the present invention;
[0035] Figure 12 This is a schematic diagram of the structure of an optional server control device according to an embodiment of the present invention;
[0036] Figure 13 This is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. Detailed Implementation
[0037] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0038] First, the terms used or related in the embodiments of the present invention are described as follows. It should be understood that the following description is one interpretation of the terms and not the only interpretation:
[0039] BMC: Baseboard Management Controller, also known as out-of-band, is a dedicated controller used to monitor and manage servers, independent of the server system. Its main function is to facilitate remote management, monitoring, installation, and restart of servers. BMC can start running as soon as it is powered on.
[0040] Static IP: A BMC static IP means that the BMC's management address is set to static mode instead of DHCP mode. In this mode, the IP address, subnet mask, and gateway will not change after the server restarts once configured.
[0041] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0042] According to one aspect of the present invention, a server control method is provided. As an optional implementation, the above-described server control method may be applied, but is not limited to, [examples of other methods]. Figure 1 The application scenarios shown are as follows. In, for example... Figure 1In the illustrated application scenario, the backend system 102 can, but is not limited to, communicate with the server 106 via network 104. The server 106 can, but is not limited to, perform operations on the database 108, such as write or read data operations. The backend system 102 may, but is not limited to, include a human-computer interaction screen, a processor, and a memory. The human-computer interaction screen may, but is not limited to, display a list of servers that have triggered alarms and a set of server status information on the backend system 102. The processor may, but is not limited to, respond to the sending of remote start commands to servers in the first set of servers that failed to start, and send the status information of the servers after responding to the remote start commands to the backend system 102. The memory is used to store relevant processing data, such as the list of servers that have triggered alarms, server status information, and remote start commands.
[0043] As an optional approach, the following steps in the server control method can be executed on the backend system 102: Step S102, obtain a list of servers that have triggered alarms and a set of server status information, wherein the set of server status information includes the status information of the servers in the server list; Step S104, if it is determined from the set of server status information that the data center where the servers in the server list are located has experienced a power outage and the servers in the server list have experienced a power failure, when it is detected that the servers in the server list have been restored to power, determine whether the server list includes the servers in the first group of servers, wherein the data center includes the first group of servers and the second group of servers, the first group of servers is set to start automatically after power restoration, and the second group of servers is set to prevent automatic startup after power restoration; Step S106, if it is determined that the server list includes the first set of servers in the first group of servers, and it is detected that there is a server in the first set of servers that has failed to start, send a remote startup command to the server that failed to start, wherein the remote startup command is used to control the startup of the server that failed to start.
[0044] As an optional example, this embodiment does not limit the execution subject of the above steps S102 to S106. For example, the above steps S102 to S106 can be executed on the background system 102 or the server 106, or they can be partially executed on the background system 102 and partially executed on the computing server that communicates with the server 106.
[0045] By employing the above method, setting the first group of servers in the data center to automatically start after power-on and the second group of servers to disable automatic startup after power-on, the first and second groups of servers can be powered on at staggered times when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can be started normally, reducing the dependence on the automatic startup setting and improving the reliability of the first group of servers' automatic startup after power-on.
[0046] To address the technical problem of excessive power load after the server is powered on again, this invention provides a server control method. Figure 2 This is a flowchart of a server control method according to an embodiment of the present invention, the process including the following steps:
[0047] Step S202: Obtain the list of servers that have issued alarms and the set of server status information, wherein the set of server status information includes the status information of the servers in the server list;
[0048] Step S204: If, based on the server status information set, it is determined that the data center where the server in the server list is located has experienced a power outage and the server in the server list has experienced a power failure, when it is detected that the server in the server list has been restored to power, it is determined whether the server list includes the server in the first group of servers. The data center includes the first group of servers and the second group of servers. The first group of servers is set to start automatically after power restoration, and the second group of servers is set to be prohibited from starting automatically after power restoration.
[0049] To better understand the server control method in the embodiments of this application, a brief introduction is first given to the types of servers included in the data center and the distribution of each type of server.
[0050] like Figure 3 As shown, the data center consists of 10 racks, each containing 10 servers. Racks 1 to 8 each contain 8 business servers (the second group of servers) and 2 support servers. All 10 servers in rack 9 are support servers, and all 10 servers in rack 10 are business servers. The support servers are used to maintain the basic operations and maintenance of the data center, while the business servers are used to process the business data of the data management platform.
[0051] In addition, the first group of servers includes Figure 3 All support servers in the middle, the second group of servers includes Figure 3 All business servers in the system.
[0052] It should be noted that, in order to ensure the stable operation of the data center, servers are usually deployed in different areas according to racks. Servers in different areas (e.g., different cities) use different power supply systems. Therefore, in the event of a power anomaly in the data center, it is necessary to determine the list of servers that have triggered the alarm based on the set of server status information obtained.
[0053] Server status information collection includes, but is not limited to, Figure 3 The last status message sent by all servers before the power outage, the status information of each server in the list of servers that triggered the alarm, etc., where the status information of each server includes, but is not limited to, the server's power status information, network connection status information, etc.
[0054] For example, suppose Figure 3 Servers in racks 1-8 are deployed in city A, while racks 9 and 10 are deployed in city B. In the event of a power outage in racks 9 and 10 in city B, how can the system...? Figure 3 By analyzing the status information of each server, the backend system can determine the list of servers currently experiencing alarms. For example, this list could include the identifiers of all servers in rack 9 and rack 10, the identifiers of all servers in rack 9, or the identifiers of all servers in rack 10.
[0055] For example, in the event of a power anomaly in racks 1 through 8 in city A, by... Figure 3 By analyzing the status information of each server, the backend system can determine the list of servers currently experiencing alarms. For example, this list could include the identifiers of all servers in racks 1-8, or the identifiers of servers in some racks in racks 1-8.
[0056] Obviously, it is easy to understand that the number of racks and the number and type of servers on the racks in this embodiment are merely examples and are not limited thereto. For instance, in real-world applications, the number of servers in a data center could be tens, hundreds, tens of thousands, or even hundreds of thousands. Furthermore, due to differences in architecture and function, support servers and business servers are typically distributed in a cross-location pattern across the racks. For details, please refer to [reference needed]. Figure 3 .
[0057] In practical applications, many server manufacturers currently disable power-on auto-start mode by default when shipping their servers. Therefore, in this embodiment, it is necessary to set the support server to automatically start upon power-on recovery (also known as power-on auto-start). In other words, the support server's system can automatically start after switching from a power-off state to a power-on state.
[0058] There are many types of servers on the market, and the methods for setting up automatic startup on power-on vary between different manufacturers. For example, ... Figure 4 As shown, the BMC's Power policy option and the BIOS's Power restore on AC Loss are examples of settings that typically take effect without a restart, while BIOS settings require a restart to take effect.
[0059] It is easy to understand that in step S202, when the list of servers that have alarmed is obtained, it cannot be determined that a large-scale power failure has occurred in the data center where the servers in the server list are located. This is because, on the one hand, even if the server does not have a power failure but has a network connection failure, it will still send the current "network connection failure" status information to the backend system; on the other hand, if the number of servers in the server list is less than a preset threshold, it cannot be directly determined that a large-scale power failure has occurred in the data center where the servers in the server list are located.
[0060] As an alternative example, after obtaining a list of servers that have issued alarms and a set of server status information, methods for determining that a power outage has occurred in the data center include:
[0061] If the number of servers in the server list is greater than or equal to the second preset threshold, the target power log indicates that a server in the server list has lost power, and the target probe response information indicates that a server in the server list cannot respond to the target probe command, it is determined that the data center where the server in the server list is located has a power anomaly and the server in the server list has lost power. The server status information set includes the target power log and the target probe response information.
[0062] In most cases, power anomalies in data centers are unexpected, and power facility monitoring largely relies on network operators. Therefore, directly detecting large-scale power anomalies in data centers from power facility monitoring is extremely difficult. The following section provides a method for identifying power anomalies in data centers using a specific example.
[0063] For example, if more than 30 servers in the data center experience alarm anomalies within a short period of time, and the BMC power logs of each server indicate that a server in the list of servers experiencing alarm anomalies has lost power, and the ping probe command indicates that the server in the list is inaccessible, then it is determined that a large-scale power anomaly has occurred in the data center.
[0064] The ping command is a commonly used network command in various operating systems (e.g., Linux). It is typically used to test connectivity with a target host, for example, to "ping a machine to see if it's on," or to "try pinging the gateway address 192.168.1.1" when a webpage cannot be opened. It sends ICMP ECHO_REQUEST packets to network hosts and displays the response. This allows you to determine if the target server is accessible based on the output (but this is not absolute). This is because some servers are blocked from ping by firewall settings or kernel parameters, making it impossible to determine the host's status via ping.
[0065] Clearly, the aforementioned 30 servers are merely an example of the second preset threshold and are not a limitation. In practical applications, different second preset thresholds can be determined based on the scale of servers in different data centers.
[0066] For example, this application also provides a data center power monitoring system based on the business layer, which mainly relies on BMC power logs, OS ping reachability, and IPMI commands to report alarms using complex algorithms. The specific judgment logic includes, but is not limited to, the following two methods:
[0067] (1) If the BMC single power log is printed, the business server OSping is unreachable, and the business server power status returns as off, it can be determined that a large-scale power anomaly has occurred in the data center that is not expected, provided that the support server is not affected.
[0068] (2) If the BMC power_drop alarm from the business server and the OS ping unreachable alarm are received without the support server being affected, it means that an unexpected power anomaly alarm has occurred in the data center.
[0069] Large-scale power outage monitoring in data centers allows for rapid detection of anomalies in the server room and the identification of the specific scope of impact, i.e., a list of alarming servers. If supporting servers are affected, the impact may only be limited to OS ping unreachability and single power supply failures.
[0070] By adopting the above method, the system can determine whether a large-scale power anomaly has occurred in the data center where the servers in the server list are located by using a logical algorithm that combines multiple commands such as setting a second preset threshold for the number of servers in the server list, BMC power logs, OS ping reachability, and IPMI commands, thereby improving the accuracy of the determination results.
[0071] Step S206: If it is determined that the server list includes the first server set in the first group of servers, and a server that failed to start is detected in the first server set, a remote start command is sent to the server that failed to start, wherein the remote start command is used to control the server that failed to start to start.
[0072] If a large-scale power outage occurs in the data center where the server list is located, as determined by the above method, the support servers in the first group of servers are all pre-configured to automatically start after power restoration. Therefore, after power restoration, it is necessary to determine whether all the support servers in the first group that experienced power outages have successfully started automatically. If any server in the first set of servers in the first group fails to start, a remote start command is sent to the failed server through the backend system to ensure that the failed support server can start normally. This will be described below with a specific example.
[0073] By adopting the above technical solution, and setting the first group of servers in the data center to automatically start after power-on while setting the second group of servers to disable automatic startup after power-on, the first and second groups of servers can be powered on in a staggered manner when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can be started normally, reducing the dependence on the automatic startup setting and improving the reliability of the first group of servers' automatic startup after power-on.
[0074] As an optional example, the above method also includes:
[0075] If, based on the server status information set, it is determined that the data center where the server in the server list is located has a power anomaly and the server in the server list has a power outage, a target detection command is sent to the baseboard management controller through the network access address of the baseboard management controller in the server list. The baseboard management controller is set to static mode. Before the power outage and after the power is restored to the server in the server list, the network access address of the baseboard management controller in static mode remains unchanged.
[0076] Upon receiving the response information returned by the baseboard management controller in response to the target detection command, the server in the server list is determined to be powered on again.
[0077] The Baseboard Management Controller (BMC) in a server is a small operating system independent of the server's operating system. Its function is to facilitate remote management, monitoring, installation, restart, and other operations of the server.
[0078] When a server in the server list experiences a power outage, if power is restored to the servers in the list, the BMCs on each server in the list will start running immediately upon power connection. Furthermore, each BMC has only one standard RJ45 network port and an independent IP address. Routine maintenance only requires accessing the management page via a browser using the IP:PORT address. Server clusters typically utilize BMC commands for large-scale unattended operations.
[0079] Based on the characteristics of the BMC in the server, it is known that after the servers in the server list are powered on again, theoretically, the BMC in each server in the server list will be the first to be powered on and start running. Therefore, by detecting whether the BMC of each server in the server list has been powered on, it can be determined whether the servers in the server list have been powered on.
[0080] The specific determination methods include: (1) The backend system sends target detection commands, such as ping commands, to each server in the server list through the network access address of the server in the server list; (2) If the network access address of the BMC can be pinged normally, it means that the power of the server has been restored and the network has also been restored. However, if it still cannot be pinged, it is necessary to continue polling the IDC infrastructure system to obtain the power restoration status of the data center. Among them, the access address of the BMC can be, but is not limited to, an IP address or an IP address + subnet mask + gateway.
[0081] It should be noted that, in order to monitor and manage each server, the BMC (Browser Management Console) on the server is usually set to static mode. This is to ensure that the network access address of the BMC on the server does not change before and after a power outage. Specifically, this is achieved by... Figure 5 By setting the "Static Address" command in the background program shown, the network access address of BMC can be set to static mode. After this configuration, the network access address of BMC will not change when the server restarts.
[0082] Therefore, once the BMC in the server is powered on and started running, it can interact with the backend system through a fixed network access address, avoiding the process of waiting for the network access address to be automatically obtained in dynamic mode, thus improving the efficiency of accessing the server.
[0083] Since the role of the support server is to maintain the basic operation and maintenance capabilities of the data center, in the event of a large-scale power outage of the server, it is first necessary to determine whether the support server is affected. If it is affected, the monitoring and alarm module in the background system will retrieve relevant information of the support server, such as in-band and out-of-band IP addresses, out-of-band target accounts and passwords, etc. The following describes this in conjunction with specific implementation examples.
[0084] As an optional example, the above implementation process for detecting whether there is a server that failed to start in the first server set when it is determined that the server list includes the first server set in the first group of servers includes:
[0085] When it is determined that the server list includes the first server set in the first group of servers, a status acquisition instruction is sent to the baseboard management controller in each server through the network access address of the baseboard management controller in each server in the first server set. The baseboard management controller in each server is set to static mode. Before the power is cut off and after the power is restored, the network access address of the baseboard management controller in static mode in each server remains unchanged.
[0086] Obtain the startup status of the baseboard management controller in each server in response to the status acquisition command;
[0087] Servers in the first server set whose startup status is not started are identified as servers that failed to start.
[0088] like Figure 6As shown, when the server list is determined to include the first server set in the first group of servers, the backend system sends a "power status" command to the support server in the first server set through the BMC static IP (network access address) of each server in the first server set in order to obtain the startup status of the support server in the first server set.
[0089] After responding to the "power status" command sent by the backend system, the BMC of each support server in the first server set returns "on" or "off". "On" indicates that the support server has been started, while "off" indicates that although the power has been restored, the support server is still in an off state.
[0090] Therefore, if the returned startup status includes "off", it indicates that a support server has failed to start. In other words, based on the status information returned by each support service to the backend system, it can be determined whether there is a support server that failed to start in the first server set.
[0091] Furthermore, for support servers that fail to start, the system sends the ipmitoolpower on command to them through the background system to start them. After the power on command is sent, it is only necessary to wait for the support server to enter the OS system and confirm that the in-band IP can be pinged to indicate that the startup was successful.
[0092] It should be noted that, to ensure the support server can start normally in the event of a power outage, in addition to setting the support server to power-on auto-start as mentioned above, a static IP address has also been configured for the BMC. The purpose of the static IP address is that, once the server's BMC is powered on, it does not rely on the out-of-band IP allocation in dynamic mode. As long as the BMC powers on normally, remote power-on commands can be sent to the support server from the backend dedicated area. This ensures that the support server can still start normally even if power-on auto-start fails, improving the reliability of the support server starting normally after power restoration.
[0093] As an optional example, the above embodiment describes the startup process of the support server in the server list after power restoration. However, the normal operation of a data center relies on the joint maintenance of both the support server and the business server. Therefore, when the server list includes business servers, it is necessary to ensure that the business servers in the server list can start normally. Specifically, this includes:
[0094] When a server in the server list is detected to have resumed power-on, determine whether the server list includes servers from the second group of servers;
[0095] If the server list is determined to include the second server set in the second group of servers, the servers in the second server set are started in batches by remote control. The servers in the second server set located on the same rack are started in batches. The servers located on the same rack are connected to the same power control device. The power control device is used to control the servers located on the same rack to cut off power when the power load generated by the servers located on the same rack is greater than or equal to a first preset threshold.
[0096] Combination Figure 6 and Figure 7 As can be seen, when a business server in the server list is detected to have regained power, the backend system sends a "power status" command to the business server in the server list through the support server. The business server then returns "on" or "off" to the backend system through the support server. "On" indicates that the business server has started, while "off" indicates that although it has regained power, the business server is still in a non-started state. Therefore, if the returned startup status includes "off," it is determined that there is a business server that failed to start.
[0097] Furthermore, for business servers that fail to start, the backend system uses the `ipmitool power on` command issued by the support server to power them on, and then restarts them. For details, please refer to [reference needed]. Figure 7 As shown in (b).
[0098] It should be noted that, as Figure 3 As shown, data centers are primarily powered in rack units, meaning a power outage only affects that specific rack unit. However, each rack can house tens of thousands of service servers. Therefore, if power is simultaneously restored to all the service servers in a rack, the power control equipment (load switch) connected to that rack can easily reach its peak power level at the moment of power-on, and this peak power level is usually greater than the load switch's first preset threshold. This situation can easily lead to rack tripping.
[0099] Therefore, to resolve the issue of rack tripping caused by simultaneous power-on of servers, it is necessary to consider remote power-on of business servers from a rack perspective. Specifically, this involves remotely controlling the servers in the second server set to start in batches, including:
[0100] In the case where the second server set includes multiple servers located on the target rack, the multiple servers located on the target rack are divided into N groups of servers, wherein the power load generated by each group of servers starting up simultaneously is less than a first preset threshold, and N is a positive integer greater than or equal to 2.
[0101] Startup commands are sent in N batches to each of the N groups of servers. The startup commands are used to control the startup of the servers that receive the startup commands.
[0102] like Figure 3 As shown, assume the second server set includes 10 service servers located on rack 10. These 10 service servers are connected to the same power control device (e.g., a composite switch). Assume the combined power generated by the simultaneous power restoration of the 10 service servers on rack 10 is U. 总 The composite switch connected to rack 10 can withstand a voltage threshold of U. max The instant that 10 business servers simultaneously resume power-on from a power-off state, U 总 It will suddenly rise to its peak, causing U 总 The voltage threshold that the composite switch can withstand is greater than or equal to U. max In this situation, due to the overload operation of the composite switch, a sudden trip may occur, causing the 10 service servers on rack 10 to fail to power on normally.
[0103] To avoid the aforementioned problems, the 10 service servers on rack 10 can be divided into multiple groups. For example, the 10 service servers can be divided into 3 groups, with each group containing 3, 3, and 4 service servers respectively. Obviously, for each group of service servers, the power load generated by all service servers starting up simultaneously is less than the voltage threshold that the composite switch can withstand, U. max .
[0104] After grouping is completed, according to the startup process, one of the servers in groups 1 to 3 is randomly selected to send a startup command. For example, a startup command is first sent to the servers in group 2. After receiving the startup command, the three servers in group 2 are simultaneously powered on. After the servers in group 2 are powered on, startup commands are sent to the servers in group 1 and group 3 in two separate batches. After receiving the startup commands, the three servers in group 1 and the four servers in group 3 are started simultaneously according to the order in which the commands are received.
[0105] Obviously, it's easy to understand that the number of service servers on rack 10 is merely an example and not a limitation. For instance, in real-world applications, the number of service servers on each rack can reach tens of thousands.
[0106] For example, such as Figure 8 As shown, assuming there are 22 servers in the target rack, and the rack is divided into three equal groups (A, B, and C) by slicing and randomization, if the remainder is 1, then group A has one more server; if the remainder is 2, then group A and group B each have one more server.
[0107] In this embodiment, since the number of rack slots is 22, then 22 = 3 * 7 + 1. Therefore, group A includes 8 servers (e.g., slots 1-8), group B includes 7 servers (e.g., slots 9-15), and group C includes 7 servers (e.g., slots 16-22). Simultaneously, each of the three groups of servers (A, B, and C) corresponds to a boot process. That is, according to the boot command issued by the backend system, the 8 servers in group A are started simultaneously to complete the first boot process; or according to the boot command issued by the backend system, the 7 servers in group B are started simultaneously to complete the second boot process, and so on.
[0108] As can be seen, by dividing multiple servers on the target rack into N groups and sending startup commands to each group in N batches, the servers can be powered on at staggered times, avoiding rack tripping issues that would occur if a large number of servers were powered on simultaneously. Furthermore, for each group of servers, simultaneous power-on upon receiving the startup command improves the efficiency of server recovery.
[0109] As can be seen from the analysis of the above embodiments, the number of business servers on each rack and the number of services processed by each group of business servers may be different, and the types of services processed by each business server may also be different. In order to reasonably group multiple servers in the target rack, the following describes the implementation method of grouping multiple servers in combination with the types of services processed by each server and the number of services in the target business set.
[0110] As an optional example, the above divides multiple servers located on the target rack into N groups of servers, including:
[0111] Identify the services handled by each of the multiple servers to obtain a target service set, wherein each service in the target service set is different and is handled by at least one server;
[0112] Based on the number of services in the target service set, the multiple servers located on the target rack are divided into N groups of servers.
[0113] The following section further describes the implementation process of dividing multiple servers located on a target rack into N groups of servers based on the number of services in the target service set, using specific examples:
[0114] When the number of services is P and P is less than M, the P servers that handle each service in the target service set are divided into a group of servers. The power load generated by simultaneously starting M servers located on the target rack is greater than or equal to a first preset threshold, and the power load generated by simultaneously starting M-1 servers located on the target rack is less than the first preset threshold. M is a positive integer greater than or equal to 2. The resulting group of servers is the first group of servers among the N groups to receive the start command; or
[0115] When the number of services is P, and P is greater than or equal to M, select M-1 servers from multiple servers to form a group of servers, where each of the M-1 servers is used to process one of the M-1 services in the target service set.
[0116] Example 1
[0117] like Figure 9 As shown in (a), assuming the target rack is rack 1, the number of services processed by the 10 servers on rack 1 is P=3, and the power load generated by the simultaneous startup of M servers on rack 1 is greater than or equal to the first preset threshold, where M=5, then the power load generated by the simultaneous startup of 4 servers on rack 1 is less than the first preset threshold.
[0118] Assuming servers S1-S3 are used to handle service 1, S4-S6 are used to handle service 2, and S7-S9 are used to handle service 3, then the ways to divide 10 servers into N groups include:
[0119] S91, sequentially select one server from servers S1-S3 (processing service 1), servers S4-S6 (processing service 2), and servers S7-S9 (processing service 3), for example, to obtain the following result: Figure 9 The first group of servers shown in (b) includes S1, S4 and S7;
[0120] S92, following the same method, select one server sequentially from servers S2-S3 (processing business 1), servers S5-S6 (processing business 2), and servers S8-S9 (processing business 3), and obtain the following result: Figure 9 The second group of servers shown in (b) includes S2, S5 and S8;
[0121] S93, the server S3 that processes business 1, the server S6 that processes business 2, the server S9 that processes business 3, and the server S10 are divided into the third group of servers.
[0122] It should be noted that in this embodiment, server S10 serves as a backup server and can be used to handle services 1, 2, and 3. Therefore, when the third group of servers is determined, server S10 can be directly grouped with servers S3, S6, and S9. In this case, since the power load generated when starting all four servers in the third group simultaneously is less than the first preset threshold, the problem of excessive power load after multiple servers in rack 1 are powered on can be avoided.
[0123] Understandably, since server S10 is a backup server, in some special scenarios, server S10 can be powered on but not started, ensuring that the server can be started on demand and reducing resource waste.
[0124] Example 2
[0125] When the number of services mentioned above is P, and P is greater than or equal to M, select M-1 servers from multiple servers, including:
[0126] Identify the target type server among multiple servers, where the business being processed by the target type server is being processed by the already started server;
[0127] If the number of servers other than the target type in a group of servers is greater than or equal to M-1, select M-1 servers from the group of servers other than the target type. The selected M-1 servers are the group of servers that receive the start command first among the N groups of servers.
[0128] like Figure 10 As shown, assuming the target rack is rack 2, the number of services processed by the 10 servers on rack 2 is P=5, and the power load generated by the simultaneous startup of M servers on rack 1 is greater than or equal to the first preset threshold, where M=4, then the power load generated by the simultaneous startup of 3 servers on rack 2 is less than the first preset threshold.
[0129] like Figure 10 As shown in (a), assume that servers S1-S2 are used to process service 1, S3-S4 are used to process service 2, S5-S6 are used to process service 3, S7-S8 are used to process service 4, and S9-S10 are used to process service 5. Also assume that services 4 and 5 are being processed by other servers that have already started.
[0130] Therefore, the ways to divide the 10 servers on rack 2 into N groups include:
[0131] S1001, using a random selection method from the corresponding server positions, one server is selected sequentially from servers S1-S2 (processing service 1), servers S3-S4 (processing service 2), and servers S5-S6 (processing service 3). For example, the result is as follows: Figure 10 The first group of servers shown in (b) includes S1, S4 and S5;
[0132] S1002, following the same method, the server S2 handling service 1, the server S3 handling service 2, and the server S6 handling service 3 are sequentially divided into the second group of servers, such as... Figure 9 (b)
[0133] This is because services 4 and 5 have already been processed by other servers. Therefore, in order to ensure the normal processing of services 1 to 3, servers S1 to S6 are selected to start first. After sending the start command to S1 to S6, servers S7 to S10 can also be started normally by executing step S1003.
[0134] S1003, servers S7-S8 (processing service 4) and servers SS9-S10 (processing service 5) are divided as follows: Figure 10 The servers in groups 3 and 4 are shown in (b).
[0135] It should be noted that the startup command will be sent to each of the four groups of servers in four separate sessions. Each group of servers that receives the startup command will start simultaneously. The power load generated when multiple servers in each group start simultaneously is less than the first preset threshold. This not only avoids the problem of excessive power load after multiple servers in rack 2 are restored to power, but also improves the startup efficiency of the business servers after they are restored to power.
[0136] To ensure that the service server experiencing a power outage can start successfully, after the grouping and startup operations described in Embodiments 1 and 2, it is usually necessary to perform an in-band ping probe on the abnormally powered service server via a ping server. The probe can be performed, but is not limited to, the backend system sending a target probe command to the abnormally powered service server through the support server. For example, a ping command; if the ping is successful, it indicates that the service server has started successfully; if the ping fails, the backend system sends the `ipmitool poweron` command to the abnormally powered service server through the support server to initiate startup. See [example...] for details. Figure 7 As shown in (b).
[0137] By leveraging the interaction between the backend system, support servers, and business servers, the system can automatically restart the grouped business servers that have experienced abnormal power outages. This avoids the problem of missing out due to manual checks for unsuccessful startups, while also reducing manpower and offline communication time, and improving the startup efficiency of servers after abnormal power outages.
[0138] As an optional example, the above-mentioned sending of startup instructions to each of the N groups of servers in N batches includes:
[0139] Perform the following operation N times: randomly select a group of servers from N groups that have not yet sent a startup command, and send a startup command to the randomly selected group of servers; or
[0140] Perform the following operation N times: Select the group of servers with the largest number of servers from among the N groups of servers that have not yet sent a start command, and send a start command to the selected group of servers.
[0141] like Figure 9 As shown, assuming the servers on rack 1 are divided into 3 groups, the specific operation method for sending start commands to these 3 groups of servers includes at least one of the following:
[0142] i) Randomly select one of the three groups of servers. For example, first select the first group of servers and send a start command to the first group of servers. Then randomly select one of the remaining second and third groups of servers. For example, select the third group of servers and send a start command to the third group of servers. Finally, send a start command to the second group of servers.
[0143] The startup command is first sent to the third group of servers, which has the most servers. Then, a group is randomly selected from the remaining first and second groups of servers. For example, if the first group of servers is selected, the startup command is sent to the first group of servers. Finally, the startup command is sent to the second group of servers.
[0144] By sending startup commands to the N groups of servers in different ways, the problem of excessive load caused by powering on a large number of servers at the same time is not only avoided, but also the servers can be selected to start first as needed, improving the flexibility of server control methods.
[0145] To better understand the technical solutions in the embodiments of this application, the following is combined with, for example, Figure 11 The overall flowchart shown is further described below, with the specific steps as follows:
[0146] S1102, Get the list of alarm servers;
[0147] This is mainly obtained through OS ping detection of the Ping server. If more than 30 servers in the data center are detected in a short period of time, it indicates that the data center has experienced a large-scale alarm, which may be due to power, network, or program problems.
[0148] If the abnormal machine BMC reports a power-related error before the alarm, it can be determined to be a power anomaly. If the automated judgment fails, the relevant on-duty personnel will confirm whether it is affected by power offline. If it is confirmed, a remote power-on process will be initiated. However, this situation is rare, and most cases can be judged automatically.
[0149] S1104, Determine if power has been restored;
[0150] The recovery process primarily relies on out-of-band ping status and feedback from the IDC infrastructure system. When the server is powered on, the BMC (Baseband Management Center) recovers first. If the BMC IP can be pinged normally at this time, it means that power has been restored and the network has also been restored. However, if it cannot be pinged, it is necessary to continue polling the IDC infrastructure system to obtain information on the data center's power recovery status.
[0151] S1106, after power has been restored, determine whether the BMC in each server that experienced a power outage is available;
[0152] The determination is based on the most recent BMC probe result before the server crashes. If the probe is normal, it means the server is available; if the probe is abnormal, it means the server is unavailable.
[0153] For servers where BMC is unavailable, proceed to step S1110.
[0154] S1108, if the BMC is confirmed to be available, determine whether the network connection has been restored;
[0155] The methods for determining whether a network connection has been restored include, but are not limited to, three approaches: BMC ping probing, polling of surrounding systems, and manual confirmation.
[0156] S1110, for servers where BMC is unavailable, manual power-on is required on-site;
[0157] The backend system will send a work order to the on-site maintenance personnel indicating that the server's BMC is unavailable. The on-site maintenance personnel will then manually power on the server and will monitor the corresponding indicator lights in the server's BMC during the process, as shown in step S1118. If the server is confirmed to be powered on, this step will be skipped.
[0158] S1112, Determine if the support server is affected;
[0159] In the event of a large-scale power outage, the first step is to determine whether the support server is affected. The method for this determination can refer to the logical algorithm described in the above embodiment, which uses BMC power logs, OSping reachability, and IPMI commands. The system then determines whether the current support server is within the affected area list. If the determination result indicates that the current support server is within the affected area list, the monitoring and alarm system will retrieve relevant information about the support server, including its internal and external IP addresses, and its external username and password.
[0160] S1114, Power on and restart the affected support servers;
[0161] The recovery method mainly involves setting up power-on auto-start (primary) and BMC static mode in the above embodiments (usually used when the support server fails to power-on auto-start). The support server is set to power-on auto-start, and it is usually already in the system before the network is restored, because the network side needs to negotiate between protocols.
[0162] To use BMC in static mode, the current power status is obtained primarily through `Ipmitool power status`. If it's "on", it means the power-on auto-start configuration for the server is active. No further action is needed for out-of-band static IPs; simply wait for the server to boot normally into the OS.
[0163] If the power status returns "off," it means the server is currently powered on but still in a shutdown state. You need to send the `Ipmitool power on` command to the support server via the backend system to power it on. After the `Power on` command is sent, simply wait for the machine to boot into the OS and for its in-band IP address to be pingable to complete the process.
[0164] S1116 uses a background algorithm to remotely start the business server;
[0165] The method for remotely starting the business server can refer to the description of the part in the above embodiments where multiple servers are grouped and remote start instructions are sent in batches, which will not be repeated here.
[0166] It should be noted that, due to the different architectures and functions of the support servers and business servers, in the event of a large-scale power outage in the data center, the remote startup of the support servers will be restored first to ensure the basic operation and maintenance of the data center. Then, after the support servers have been powered on and started up again, the remote startup of the business servers will be restored.
[0167] S1118, Fault Recovery Detection;
[0168] Since the Ping server is redundantly configured across multiple servers in the campus, it is largely unaffected by power outages. Therefore, recovery detection mainly involves in-band pinging of the servers that experienced abnormal power outages using the Ping server servers. If one Ping server can successfully ping the machine, it indicates that the machine has recovered. If the machine cannot be pinged, proceed to step S1120.
[0169] S1120, Exception handling procedure;
[0170] This node mainly refers to the handling of machines that fail to power on remotely. If the power on command directly returns a power failure, or if the in-band IP still cannot be pinged after a certain period of remote power-on, it needs to be transferred to the anomaly handling node for manual assistance in diagnosis. This is mainly done by issuing a work order to the second-line operations and maintenance team for follow-up and handling.
[0171] Through the embodiments provided in this application, by setting the state of the first group of servers in the data center to automatically start after power-on and the state of the second group of servers to disable automatic startup after power-on, the first and second groups of servers are powered on in a staggered manner when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can start normally, reducing the dependence on the automatic startup setting of servers after power-on and achieving the technical effect of improving the reliability of the first group of servers' automatic startup after power-on.
[0172] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, because according to the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.
[0173] According to another aspect of the embodiments of the present invention, as follows is also provided Figure 12 The diagram shows a server control device, the device comprising:
[0174] The first acquisition unit 1202 is used to acquire a list of servers that have issued alarms and a set of server status information, wherein the set of server status information includes the status information of the servers in the server list.
[0175] The first processing unit 1204 is used to determine whether the server list includes servers in the first group of servers when it is detected that the data center where the servers in the server list are located has a power abnormality and the servers in the server list have lost power, based on the server status information set. The data center includes the first group of servers and the second group of servers. The first group of servers is set to start automatically after power is restored, and the second group of servers is set to be prohibited from starting automatically after power is restored.
[0176] To better understand the server control method in the embodiments of this application, a brief introduction is first given to the types of servers included in the data center and the distribution of each type of server.
[0177] like Figure 3 As shown, the data center consists of 10 racks, each containing 10 servers. Racks 1 to 8 each contain 8 business servers (the second group of servers) and 2 support servers. All 10 servers in rack 9 are support servers, and all 10 servers in rack 10 are business servers. The support servers are used to maintain the basic operations and maintenance of the data center, while the business servers are used to process the business data of the data management platform.
[0178] In addition, the first group of servers includes Figure 3 All support servers in the middle, the second group of servers includes Figure 3 All business servers in the system.
[0179] It should be noted that, in order to ensure the stable operation of the data center, servers are usually deployed in different areas according to racks. Servers in different areas (e.g., different cities) use different power supply systems. Therefore, in the event of a power anomaly in the data center, it is necessary to determine the list of servers that have triggered the alarm based on the set of server status information obtained.
[0180] Server status information collection includes, but is not limited to, Figure 3 The last status message sent by all servers before the power outage, the status information of each server in the list of servers that triggered the alarm, etc., where the status information of each server includes, but is not limited to, the server's power status information, network connection status information, etc.
[0181] For example, suppose Figure 3 Servers in racks 1-8 are deployed in city A, while racks 9 and 10 are deployed in city B. In the event of a power outage in racks 9 and 10 in city B, how can the system...? Figure 3By analyzing the status information of each server, the backend system can determine the list of servers currently experiencing alarms. For example, this list could include the identifiers of all servers in rack 9 and rack 10, the identifiers of all servers in rack 9, or the identifiers of all servers in rack 10.
[0182] For example, in the event of a power anomaly in racks 1 through 8 in city A, by... Figure 3 By analyzing the status information of each server, the backend system can determine the list of servers currently experiencing alarms. For example, this list could include the identifiers of all servers in racks 1-8, or the identifiers of servers in some racks in racks 1-8.
[0183] Obviously, it is easy to understand that the number of racks and the number and type of servers on the racks in this embodiment are merely examples and are not limited thereto. For instance, in real-world applications, the number of servers in a data center could be tens, hundreds, tens of thousands, or even hundreds of thousands. Furthermore, due to differences in architecture and function, support servers and business servers are typically distributed in a cross-location pattern across the racks. For details, please refer to [reference needed]. Figure 3 .
[0184] In practical applications, many server manufacturers currently disable power-on auto-start mode by default when shipping their servers. Therefore, in this embodiment, it is necessary to set the support server to automatically start upon power-on recovery (also known as power-on auto-start). In other words, the support server's system can automatically start after switching from a power-off state to a power-on state.
[0185] There are many types of servers on the market, and the methods for setting up automatic startup on power-on vary between different manufacturers. For example, ... Figure 4 As shown, the BMC's Power policy option and the BIOS's Power restore on AC Loss are examples of settings that typically take effect without a restart, while BIOS settings require a restart to take effect.
[0186] It is easy to understand that in step S202, when the list of servers that have alarmed is obtained, it cannot be determined that a large-scale power failure has occurred in the data center where the servers in the server list are located. This is because, on the one hand, even if the server does not have a power failure but has a network connection failure, it will still send the current "network connection failure" status information to the backend system; on the other hand, if the number of servers in the server list is less than a preset threshold, it cannot be directly determined that a large-scale power failure has occurred in the data center where the servers in the server list are located.
[0187] As an alternative example, after obtaining a list of servers that have issued alarms and a set of server status information, methods for determining that a power outage has occurred in the data center include:
[0188] If the number of servers in the server list is greater than or equal to the second preset threshold, the target power log indicates that a server in the server list has lost power, and the target probe response information indicates that a server in the server list cannot respond to the target probe command, it is determined that the data center where the server in the server list is located has a power anomaly and the server in the server list has lost power. The server status information set includes the target power log and the target probe response information.
[0189] In most cases, power anomalies in data centers are unexpected, and power facility monitoring largely relies on network operators. Therefore, directly detecting large-scale power anomalies in data centers from power facility monitoring is extremely difficult. The following section provides a method for identifying power anomalies in data centers using a specific example.
[0190] For example, if more than 30 servers in the data center experience alarm anomalies within a short period of time, and the BMC power logs of each server indicate that a server in the list of servers experiencing alarm anomalies has lost power, and the ping probe command indicates that the server in the list is inaccessible, then it is determined that a large-scale power anomaly has occurred in the data center.
[0191] The ping command is a commonly used network command in various operating systems (e.g., Linux). It is typically used to test connectivity with a target host, for example, to "ping a machine to see if it's on," or to "try pinging the gateway address 192.168.1.1" when a webpage cannot be opened. It sends ICMP ECHO_REQUEST packets to network hosts and displays the response. This allows you to determine if the target server is accessible based on the output (but this is not absolute). This is because some servers are blocked from ping by firewall settings or kernel parameters, making it impossible to determine the host's status via ping.
[0192] Clearly, the aforementioned 30 servers are merely an example of the second preset threshold and are not a limitation. In practical applications, different second preset thresholds can be determined based on the scale of servers in different data centers.
[0193] For example, this application also provides a data center power monitoring system based on the business layer, which mainly relies on BMC power logs, OS ping reachability, and IPMI commands to report alarms using complex algorithms. The specific judgment logic includes, but is not limited to, the following two methods:
[0194] (1) If the BMC single power log is printed, the business server OSping is unreachable, and the business server power status returns as off, it can be determined that a large-scale power anomaly has occurred in the data center that is not expected, provided that the support server is not affected.
[0195] (2) If the BMC power_drop alarm from the business server and the ping unreachable alarm from the OS are received without the support server being affected, it means that an unexpected power anomaly alarm has occurred in the data center.
[0196] Large-scale power outage monitoring in data centers allows for rapid detection of anomalies in the server room and the identification of the specific scope of impact, i.e., a list of alarming servers. If supporting servers are affected, the impact may only be limited to OS ping unreachability and single power supply failures.
[0197] By adopting the above method, the system can determine whether a large-scale power anomaly has occurred in the data center where the servers in the server list are located by using a logical algorithm that combines multiple commands such as setting a second preset threshold for the number of servers in the server list, BMC power logs, OS ping reachability, and IPMI commands, thereby improving the accuracy of the determination results.
[0198] The second processing unit 1206 is configured to send a remote startup command to the server that failed to start when it is determined that the server list includes the first server set in the first group of servers and a server that failed to start is detected in the first server set. The remote startup command is used to control the startup of the server that failed to start.
[0199] If a large-scale power outage occurs in the data center where the server list is located, as determined by the above method, the support servers in the first group of servers are all pre-configured to automatically start after power restoration. Therefore, after power restoration, it is necessary to determine whether all the support servers in the first group that experienced power outages have successfully started automatically. If any server in the first set of servers in the first group fails to start, a remote start command is sent to the failed server through the backend system to ensure that the failed support server can start normally. This will be described below with a specific example.
[0200] By adopting the above technical solution, and setting the first group of servers in the data center to automatically start after power-on while setting the second group of servers to disable automatic startup after power-on, the first and second groups of servers can be powered on in a staggered manner when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can be started normally, reducing the dependence on the automatic startup setting and improving the reliability of the first group of servers' automatic startup after power-on.
[0201] Optionally, the above-mentioned device further includes:
[0202] The third processing unit is used to send a target detection command to the baseboard management controller through the network access address of the baseboard management controller in the server list when it is determined from the server status information set that the data center where the server in the server list is located has a power abnormality and the server in the server list has a power outage. The baseboard management controller is set to a static mode, and the network access address of the baseboard management controller in the static mode remains unchanged before the server in the server list has a power outage and after the power is restored.
[0203] The fourth processing unit is used to determine whether a server in the server list has been powered on upon receiving a response from the baseboard management controller in response to a target detection command.
[0204] Optionally, the above-mentioned device further includes:
[0205] The fifth processing unit is used to send a status acquisition instruction to the baseboard management controller in each server through the network access address of the baseboard management controller in each server in the first server set when it is determined that the server list includes the first server set in the first group of servers. The baseboard management controller in each server is set to a static mode, and the network access address of the baseboard management controller in the static mode in each server remains unchanged before the power is cut off and after the power is restored.
[0206] The second acquisition unit is used to acquire the start-up status of the baseboard management controller in each server in response to the status acquisition command.
[0207] The first determining unit is used to determine the servers in the first server set whose startup status is not started as servers that have failed to start.
[0208] Optionally, the above-mentioned device further includes:
[0209] The second determining unit is used to determine whether the server list includes servers in the second group of servers when it is detected that a server in the server list has resumed power-on.
[0210] The control unit is used to remotely control the servers in the second server set to start in batches when it is determined that the server list includes the second server set in the second group of servers. The servers in the second server set located on the same rack are started in batches, and the servers located on the same rack are connected to the same power control device. The power control device is used to control the servers located on the same rack to cut off power when the power load generated by the servers located on the same rack is greater than or equal to a first preset threshold.
[0211] Optionally, the first control unit mentioned above includes:
[0212] The first processing module is used to divide the multiple servers located on the target rack into N groups of servers when the second server set includes multiple servers located on the target rack, wherein the power load generated by each group of servers starting up at the same time is less than a first preset threshold, and N is a positive integer greater than or equal to 2.
[0213] The sending module is used to send start commands to each of the N groups of servers in N batches. The start commands are used to control the servers that receive the start commands to start.
[0214] Optionally, the first processing module mentioned above includes:
[0215] The first processing submodule is used to determine the services processed by each of the multiple servers to obtain a target service set, wherein each service in the target service set is different and is processed by at least one server.
[0216] The second processing submodule is used to divide multiple servers located on the target rack into N groups of servers according to the number of services in the target service set.
[0217] Optionally, the first processing module mentioned above includes:
[0218] The third processing submodule is used to divide the P servers that handle each service in the target service set into a group of servers when the number of services is P and P is less than M. The power load generated by simultaneously starting M servers located on the target rack is greater than or equal to a first preset threshold, and the power load generated by simultaneously starting M-1 servers located on the target rack is less than the first preset threshold. M is a positive integer greater than or equal to 2. The resulting group of servers is the first group of servers to receive the start command among the N groups of servers; or
[0219] When the number of services is P, and P is greater than or equal to M, select M-1 servers from multiple servers to form a group of servers, where each of the M-1 servers is used to process one of the M-1 services in the target service set.
[0220] Optionally, the first processing module described above further includes:
[0221] The fourth processing submodule is used to determine the target type of server among multiple servers, wherein the business being processed by the target type server is being processed by the already started server;
[0222] The fifth processing submodule is used to select M-1 servers from the multiple servers excluding the target type when the number of servers other than the target type is greater than or equal to M-1. The selected M-1 servers are the group of servers that receive the start command first among the N groups of servers.
[0223] Optionally, the above-mentioned sending module includes:
[0224] The sixth processing submodule is used to perform the following operations N times: randomly select a group of servers from N groups that have not yet sent a start command, and send a start command to the randomly selected group of servers; or
[0225] Perform the following operation N times: Select the group of servers with the largest number of servers from among the N groups of servers that have not yet sent a start command, and send a start command to the selected group of servers.
[0226] Optionally, the above-mentioned device further includes:
[0227] The sixth processing unit is used to determine that the data center where the servers in the server list are located has a power anomaly and the servers in the server list have a power outage when the number of servers in the server list is greater than or equal to a second preset threshold, the target power log indicates that the servers in the server list have a power outage, and the target detection response information indicates that the servers in the server list cannot respond to the target detection command. The server status information set includes the target power log and the target detection response information.
[0228] By applying the aforementioned device to set the first group of servers in the data center to automatically start after power-on and the second group of servers to disable automatic startup after power-on, the first and second groups of servers can be powered on in a staggered manner when restoring power to servers experiencing large-scale power anomalies. This avoids rack tripping caused by simultaneous power-on of a large number of servers and solves the technical problem of excessive power load after server power-on in related technologies. Furthermore, by sending remote startup commands to servers in the first group of servers that failed to start, servers that failed to automatically start after power-on can be started normally, reducing the dependence on the automatic startup setting and improving the reliability of the first group of servers' automatic startup after power-on.
[0229] It should be noted that the embodiments of the server control device described here can refer to the embodiments of the server control method described above, and will not be repeated here.
[0230] According to another aspect of the embodiments of this application, an electronic device for implementing the above-described server control method is also provided, the electronic device being... Figure 13 The terminal device shown is illustrated in this embodiment, which uses this electronic device as an example of a backend device. Figure 13 As shown, the electronic device includes a memory 1302 and a processor 1304. The memory 1302 stores a computer program, and the processor 1304 is configured to execute the steps of any of the above method embodiments through the computer program.
[0231] Optionally, in this embodiment, the aforementioned electronic device may be located in at least one of a plurality of network devices in a computer network.
[0232] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0233] S1, obtain the list of servers that have issued alarms and the set of server status information, wherein the set of server status information includes the status information of the servers in the server list;
[0234] S2, if it is determined from the server status information set that the data center where the server in the server list is located has a power abnormality and the server in the server list has a power outage, when it is detected that the server in the server list has been restored to power, determine whether the server list includes the server in the first group of servers. The data center includes the first group of servers and the second group of servers. The first group of servers is set to start automatically after power restoration, and the second group of servers is set to be prohibited from starting automatically after power restoration.
[0235] S3, if it is determined that the server list includes the first server set in the first group of servers, and a server that failed to start is detected in the first server set, a remote start command is sent to the server that failed to start. The remote start command is used to control the server that failed to start to start.
[0236] Alternatively, as those skilled in the art will understand, Figure 13 The structure shown is for illustrative purposes only. Electronic devices can also be smartphones (such as Android phones, iOS phones, etc.), tablets, PDAs, mobile internet devices (MIDs), PADs, and other target terminals. Figure 13 This does not limit the structure of the aforementioned electronic devices or electronic equipment. For example, electronic devices or electronic equipment may also include components that are more... Figure 13 The more or fewer components shown (such as network interfaces, etc.), or having the same Figure 13 The different configurations shown.
[0237] The memory 1302 can be used to store software programs and modules, such as the program instructions / modules corresponding to the server control method and device in this embodiment. The processor 1304 executes various functional applications and data processing by running the software programs and modules stored in the memory 1302, thereby realizing the aforementioned server control method. The memory 1302 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1302 may further include memory remotely located relative to the processor 1304, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Specifically, the memory 1302 may be used, but is not limited to, to store a list of servers that have alarmed, a set of server status information, remote start instructions, etc. As an example, such as... Figure 13As shown, the memory 1302 may include, but is not limited to, the first acquisition unit 1202, the first processing unit 1204, and the second processing unit 1206 from the control device of the server. Furthermore, it may include, but is not limited to, other module units from the control device of the server, which will not be elaborated upon in this example.
[0238] Optionally, the transmission device 1306 described above is used to receive or send data via a network. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1306 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1306 is a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0239] In addition, the aforementioned electronic device also includes: a display 1308 for displaying the aforementioned virtual resources that are allowed to be pushed; and a connection bus 1310 for connecting the various module components in the aforementioned electronic device.
[0240] In other embodiments, the target terminal or server can be a node in a distributed system, which can be a blockchain system. This blockchain system is formed by connecting multiple nodes through network communication. The nodes can form a peer-to-peer (P2P) network, and any type of computing device, such as a server or terminal, can become a node in the blockchain system by joining this peer-to-peer network.
[0241] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the server control method provided in various optional implementations of the aforementioned server verification processing, wherein the computer program is configured to execute the steps in any of the above-described method embodiments at runtime.
[0242] Optionally, in this embodiment, the computer-readable storage medium described above may be configured to store a computer program for performing the following steps:
[0243] S1, obtain the list of servers that have issued alarms and the set of server status information, wherein the set of server status information includes the status information of the servers in the server list;
[0244] S2, if it is determined from the server status information set that the data center where the server in the server list is located has a power abnormality and the server in the server list has a power outage, when it is detected that the server in the server list has been restored to power, determine whether the server list includes the server in the first group of servers. The data center includes the first group of servers and the second group of servers. The first group of servers is set to start automatically after power restoration, and the second group of servers is set to be prohibited from starting automatically after power restoration.
[0245] S3, if it is determined that the server list includes the first server set in the first group of servers, and a server that failed to start is detected in the first server set, a remote start command is sent to the server that failed to start. The remote start command is used to control the server that failed to start to start.
[0246] Optionally, in this embodiment, those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the target terminal. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0247] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0248] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention.
[0249] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0250] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or the indirect coupling or communication connection of units or modules may be electrical or other forms.
[0251] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0252] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0253] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for controlling a server, characterized in that, include: Obtain a list of servers that have issued alarms and a set of server status information, wherein the set of server status information includes the status information of the servers in the server list; If, based on the server status information set, it is determined that the data center where the server in the server list is located experiences a power outage and the server in the server list experiences a power failure, when it is detected that the server in the server list has been restored to power, it is determined whether the server list includes the server in the first group of servers. The data center includes the first group of servers and the second group of servers. The first group of servers is set to start automatically after power restoration, and the second group of servers is set to be prohibited from starting automatically after power restoration. If it is determined that the server list includes the first server set in the first group of servers, and a server in the first server set has failed to start, a remote start command is sent to the server that failed to start, wherein the remote start command is used to control the server that failed to start to start.
2. The method according to claim 1, characterized in that, The method further includes: If, based on the server status information set, it is determined that the data center where the server in the server list is located experiences a power outage and the server in the server list experiences a power failure, a target detection command is sent to the baseboard management controller through the network access address of the baseboard management controller in the server list. The baseboard management controller is set to a static mode, and the network access address of the baseboard management controller in the static mode remains unchanged before the power failure and after the power is restored to the server in the server list. Upon receiving the response information returned by the baseboard management controller in response to the target detection command, it is determined that the servers in the server list have been restored to power.
3. The method according to claim 1, characterized in that, The method further includes: If it is determined that the server list includes the first server set in the first group of servers, a status acquisition instruction is sent to the baseboard management controller in each server through the network access address of the baseboard management controller in each server in the first server set, wherein the baseboard management controller in each server is set to a static mode, and the network access address of the baseboard management controller in the static mode in each server remains unchanged before the power outage and after the power is restored. Obtain the startup status of the baseboard management controller in each of the servers in response to the status acquisition command; The servers in the first server set whose startup status is not started are identified as the servers that failed to start.
4. The method according to claim 1, characterized in that, The method further includes: When a server in the server list is detected to have resumed power-on, it is determined whether the server list includes a server in the second group of servers; If it is determined that the server list includes the second server set in the second group of servers, the servers in the second server set are started in batches by remote control. The servers in the second server set located on the same rack are started in batches. The servers located on the same rack are connected to the same power control device. The power control device is used to control the servers located on the same rack to cut off power when the power load generated by the servers located on the same rack is greater than or equal to a first preset threshold.
5. The method according to claim 4, characterized in that, The remote control of starting up servers in the second server set in batches includes: In the case where the second server set includes multiple servers located on the target rack, the multiple servers located on the target rack are divided into N groups of servers, wherein the power load generated by each group of servers starting up simultaneously is less than the first preset threshold, and N is a positive integer greater than or equal to 2. Startup commands are sent to each of the N groups of servers in N batches, wherein the startup commands are used to control the servers that receive the startup commands to start.
6. The method according to claim 5, characterized in that, The step of dividing multiple servers located on the target rack into N groups of servers includes: The services processed by each of the plurality of servers are determined to obtain a target service set, wherein each service in the target service set is different and is processed by at least one server; Based on the number of services in the target service set, the multiple servers located on the target rack are divided into the N groups of servers.
7. The method according to claim 6, characterized in that, The step of dividing the plurality of servers located on the target rack into the N groups of servers according to the number of services in the target service set includes: When the number of services is P and P is less than M, the P servers that process each service in the target service set are divided into a group of servers. The power load generated by simultaneously starting M servers located on the target rack is greater than or equal to the first preset threshold, and the power load generated by simultaneously starting M-1 servers located on the target rack is less than the first preset threshold. M is a positive integer greater than or equal to 2. The resulting group of servers is the group of servers that receives the start command first among the N groups of servers; or When the number of services is P and P is greater than or equal to M, M-1 servers are selected from the plurality of servers to form a group of servers, wherein the M-1 servers are used to process M-1 services in the target service set.
8. The method according to claim 7, characterized in that, The step of selecting M-1 servers from the plurality of servers includes: Among the plurality of servers, a server of a target type is identified, wherein the service being processed by the server of the target type is being processed by an already started server; If the number of servers other than the target type among the plurality of servers is greater than or equal to M-1, the M-1 servers are selected from the plurality of servers other than the target type, wherein the selected M-1 servers are the group of servers that first receive the start instruction among the N groups of servers.
9. The method according to claim 5, characterized in that, The process of sending startup commands to each of the N groups of servers in N batches includes: Perform the following operation N times: randomly select a group of servers from the N groups that have not yet sent the startup command, and send the startup command to the randomly selected group of servers; or Perform the following operation N times: Select the group of servers with the largest number of servers from the N groups of servers that have not yet sent the start command, and send the start command to the selected group of servers.
10. The method according to any one of claims 1 to 9, characterized in that, After obtaining the list of servers that have issued alarms and the set of server status information, the method further includes: If the number of servers in the server list is greater than or equal to a second preset threshold, the target power log indicates that a server in the server list has experienced a power outage, and the target detection response information indicates that a server in the server list cannot respond to the target detection command, it is determined that the data center where the servers in the server list are located has experienced a power anomaly and the servers in the server list have experienced a power outage. The server status information set includes the target power log and the target detection response information.
11. A server control device, characterized in that, include: The first acquisition unit is used to acquire a list of servers that have issued alarms and a set of server status information, wherein the set of server status information includes the status information of the servers in the server list; The first processing unit is configured to, when determining, based on the server status information set, that the data center where the servers in the server list are located has experienced a power outage and the servers in the server list have experienced a power failure, determine whether the server list includes servers from the first group of servers when it is detected that the servers in the server list have been restored to power. The data center includes the first group of servers and the second group of servers. The first group of servers is configured to start automatically after power restoration, and the second group of servers is configured to be prevented from starting automatically after power restoration. The second processing unit is configured to send a remote startup instruction to the server that failed to start when it is determined that the server list includes the first server set in the first group of servers and a server that failed to start is detected in the first server set. The remote startup instruction is used to control the startup of the server that failed to start.
12. The apparatus according to claim 11, characterized in that, The device further includes: The third processing unit is used to send a target detection command to the baseboard management controller through the network access address of the baseboard management controller in the server list when it is determined from the server status information set that the data center where the server in the server list is located has a power abnormality and the server in the server list has a power outage. The baseboard management controller is set to a static mode, and the network access address of the baseboard management controller in the static mode remains unchanged before the power outage and after the power is restored to the server in the server list. The fourth processing unit is used to determine, upon receiving the response information returned by the baseboard management controller in response to the target detection command, whether the server in the server list has been powered on again.
13. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein the program is executed by a processor to perform the method of any one of claims 1 to 10.
14. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method described in any one of claims 1 to 10.
15. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to execute the method of any one of claims 1 to 10 through the computer program.
Citation Information
Patent Citations
Data center power supply device control system and method
CN102833083A
Post-power-failure recovery and starting system for ultra-large container vessel
CN104218670A