Cross-architecture server downtime detection method and system
By deploying the cluster network topology detection module, alarm reception module and cross-architecture remote execution command module, combined with the gossip protocol of Consul and the icmp protocol of ping, cross-architecture server downtime detection is realized, solving the server downtime detection problem with huge differences in different hardware and software, and improving the accuracy and applicability of detection.
Patent Information
- Application Number
- CN202510488699.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology lacks cross-architecture server downtime detection methods, and cannot effectively detect downtime failures of cloud servers, bare metal servers and hosts running cloud servers with huge differences in different hardware and software.
Deploy the cluster network topology detection module, alarm reception module, cross-architecture remote execution command module and downtime detection module, and use the consul's gossip protocol and ping's IMP protocol to perform network dual detection to realize cross-architecture downtime detection.
Through dual network detection, the downtime of cloud servers, bare metal servers and hosts running cloud servers are accurately identified, and cross-architecture downtime detection is supported, which improves the reliability and applicability of detection.
Smart Images

Figure CN120342918A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and particularly to a method and system for detecting downtime of cross-architecture servers. Background Art
[0002] In cloud services, a cloud server is a virtual server that runs on a physical server but provides an independent computing environment through virtualization technology. Cloud servers offer high flexibility and can quickly scale resources up or down according to demand.
[0003] A bare metal server is a physical server that runs directly on the hardware without an additional virtualization layer. Due to running directly on the hardware without the overhead of a virtualization layer, bare metal servers can provide extremely high performance. This is extremely ideal for applications that need to process large amounts of data or have strict performance requirements.
[0004] Whether it is a cloud server, a bare metal server, or the host running the cloud server, when a downtime failure occurs, such as abnormal shutdown or restart, it will seriously affect the availability of cloud services. Therefore, it is necessary to detect the downtime failure in a timely manner and perform fault recovery through manual or automated operation and maintenance. For example, when a cloud server fails, operations such as hot migration, cold migration, and evacuation can be considered to migrate it to other hosts; when a bare metal server fails, backup or disaster recovery means can be used to restart the service on another bare machine; when the host running the cloud server fails, the entire host can be evacuated and all cloud servers on the host can be migrated to other hosts together.
[0005] However, there are many types of existing servers. During the adaptation process, it is found that there are huge differences in their hardware and software. For example, some support obtaining the power status through ipmitool, while others do not; some servers support clearing specified out-of-band logs, while others can only clear all out-of-band logs with one key. Therefore, after widely adopting existing servers in the cloud service field, there is a lack of a cross-architecture server downtime detection method to face the various types of servers. Summary of the Invention
[0006] The purpose of the present invention is to provide a method and system for detecting downtime of cross-architecture servers to solve the problems raised in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solutions: A cross-architecture server downtime detection method, deploying a cluster network topology detection module: dividing nodes into master nodes and slave nodes, where the master node is a physical machine and multiple high-availability implementations are set, and there are no type restrictions on the slave nodes; completing the containerized deployment of Consul on the master node and the slave node, running several Consul agents on each node to be bound to the network on their respective nodes respectively, forming an independent network cluster; allowing customization of the number of bindings to the network to form a corresponding cluster, running the Consul agent in server mode on the master node and in client mode on the slave node, and Consul using the gossip protocol to manage the nodes in the cluster; based on the Consul chart provided by the community, changing the client-daemenset and server-configmap, so that a Consul client pod contains several Consul agents that are respectively bound to the networks specified in the configuration file, and using hostNetwork.
[0008] Preferably, it includes deploying an alarm receiving module: deploying an alarm receiving module on the master node, which depends on a custom collector to obtain monitoring metrics, writing a Consul client in the custom collector, and parsing the cluster network topology of different networks by calling the Consul API; allowing customization of alarm rules, when the metrics meet the alarm conditions, relevant network alarms appear on Prometheus, and then using Alertmanager to notify the downtime detection module; the alarm rule is that all networks of a certain slave node are AgentMemberLeft or AgentMemberFailed, and the time lasts for 2 minutes, then a network alarm is triggered and the downtime detection module is automatically called to evaluate again.
[0009] Preferably, it includes a module for deploying cross-architecture remote execution commands: developed based on the open-source configuration management and remote execution tool SaltStack to achieve synchronous execution of remote commands across architectures; using containerized design, the salt-master is managed by deployment, with 3 replicas to achieve high availability and podAntiAffinity is set. The salt-master has 2 containers, namely the keys-init container for cleaning the masterkey on the source node and the salt-master container for starting the salt-master and salt-api services. The health detection mechanism of the salt-master is to call the 8000 port of the local control network for login verification; the salt-slave is managed by daemonset, and its health detection mechanism is to check the tcp connection status of the peer 4505 port and the peer address is the management network address of the node where the salt-master pod is located; when making the image, the root directory of the computing node is mounted to / mnt in the container, and when executing remote commands, the prefix chroot / mnt is added to call the host software commands. The method of writing the command to be executed into a randomly named file is used to ensure that the salt commands do not overwrite in the case of multi-threading; this module supports cross-clusters and cross-architectures, and the salt-slave can be deployed independently. The salt-master node can batch-distribute commands to nodes of different architectures.
[0010] Preferably, it includes a module for detecting downtime: after receiving an alarm, the downtime detection module parses out the name of the slave node; according to the network card metadata information, several nodes that can ping the corresponding network gateway or the network corresponding to the master are selected; using the cross-architecture remote execution command module, the ping command is executed concurrently to detect whether the selected nodes are all unreachable from the slave node to be tested. If they are all unreachable, it indicates that there is a problem with the slave node to be tested; the dual detection of the network is achieved through the gossip protocol of Consul and the icmp protocol of ping. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination; Consul detects the network status of the slave node control network, service network, and storage network every 15s. If all 3 networks of the node to be tested are AgentMemberLeft or AgentMemberFailed among the 4 detection points within 1 minute, an alarm is sent to Prometheus, and Alertmanager notifies the downtime detection module of the alarm. After the downtime detection module parses out the name of the alarm node, network ping detection is performed, and it supports setting the number of nodes for pinging the gateway, the number of detection rounds, and the number of ping packets.
[0011] A system for a cross-architecture server downtime detection method, including a cluster network topology detection module, which is used for:
[0012] The nodes are divided into master nodes and slave nodes. The master nodes are physical machines and multiple master nodes are set up to achieve high availability. There is no type limit for slave nodes;
[0013] Complete the containerized deployment of Consul on the master nodes and slave nodes. Each node runs several Consul agents that are respectively bound to the networks on their respective nodes to form an independent network cluster; Customizing the number of bindings to the network is allowed to form corresponding clusters;
[0014] On the master nodes, the Consul agents run in server mode to maintain the state of Consul; On the slave nodes, they run in client mode for health checks and forwarding queries to the server;
[0015] Based on the Consul chart provided by the community, modify the client-daemenset and server-configmap so that a Consul client pod contains several Consul agents that are respectively bound to the networks specified in the configuration file and use hostNetwork.
[0016] Preferably, it includes an alarm receiving module, which is used for:
[0017] Deployed on the master nodes, it depends on a custom collector to obtain monitoring metrics. A Consul client is written in the custom collector, and the cluster network topologies of different networks are parsed by calling the Consul API;
[0018] Allowing custom alarm rules. When the metrics meet the alarm conditions, relevant network alarms appear on Prometheus, and then the Alertmanager is used to notify the downtime detection module;
[0019] The alarm rule is that all networks of a certain slave node are AgentMemberLeft or AgentMemberFailed and the duration is 2 minutes, then a network alarm is triggered and the downtime detection module is automatically called to evaluate again.
[0020] Preferably, it includes a cross-architecture remote execution command module, which is used for:
[0021] Developed based on the open-source configuration management and remote execution tool SaltStack to achieve cross-architecture synchronous execution of remote commands;
[0022] Adopting containerized design, the salt-master is managed by deployment, and the replicas are set to 3 to achieve high availability, and podAntiAffinity is set. The salt-master is set with 2 containers, namely the keys-init container for cleaning the master key on the source node and the salt-master container for starting the salt-master and salt-api services. The health detection mechanism of the salt-master is to call the 8000 port of the control network of this node for login verification;
[0023] The salt-slave is managed by daemonset, and its health detection mechanism is to check the tcp connection status of the peer 4505 port and the peer address is the management network address of the node where the salt-master pod is located;
[0024] When making the image, mount the root directory of the computing node to / mnt in the container. When executing remote commands, prefix chroot / mnt to call the host software commands, and write the commands to be executed into randomly named files to ensure that salt commands are not overwritten in the case of multi-threading;
[0025] Support cross-cluster and cross-architecture, and the salt-slave can be deployed independently. The salt-master node can batch distribute commands to nodes of different architectures.
[0026] Preferably, it includes a downtime detection module, which is used for:
[0027] After receiving the alarm, parse out the name of the slave node;
[0028] According to the network card metadata information, select several nodes that can ping the corresponding network gateway or the network corresponding to the master;
[0029] Use the cross-architecture remote execution command module to concurrently execute ping commands to detect whether the selected nodes are all unreachable from the slave node to be tested. If they are all unreachable, it indicates that there is a problem with the slave node to be tested;
[0030] Implement double detection of the network through the gossip protocol of consul and the icmp protocol of ping. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination;
[0031] Support setting the number of ping gateway nodes, the number of detection rounds, and the number of ping packets.
[0032] Compared with the prior art, the beneficial effects of the present invention are:
[0033] The cross-architecture server downtime detection method and system proposed by the present invention deploy a cluster network topology detection module, an alarm receiving module, a cross-architecture remote execution command module, and a downtime detection module in sequence. By using the gossip protocol of Consul and the ICMP protocol of Ping, double detection of the network is achieved. The gossip protocol is responsible for triggering, and the ICMP protocol is responsible for further determination. Thus, downtime faults are determined through network detection, and all are containerized deployments without involving specific commands bound to different models. Finally, it supports downtime detection of cross-architecture cloud servers, bare metal servers, and host machines running cloud servers. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to clearly and completely describe the objectives, technical solutions of the present invention, and make the advantages more clearly understood, the following further details the embodiments of the present invention with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are some embodiments of the present invention, rather than all embodiments, and are only used to explain the embodiments of the present invention, not to limit the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0036] Embodiment 1, please refer to Figure 1 , the present invention provides a technical solution: a cross-architecture server downtime detection method, the steps are as follows:
[0037] 1. Deploy a cluster network topology detection module;
[0038] 2. Deploy an alarm receiving module;
[0039] 3. Deploy a cross-architecture remote execution command module;
[0040] 4. Deploy a downtime detection module;
[0041] In the above steps, the first stage and the fourth stage are the cores of the cross-architecture server downtime detection method.
[0042] As a technical solution that needs to be further explained in the present invention, for step 1, it should be noted that: the present invention divides nodes into master nodes and slave nodes. The master nodes are physical machines, and several are set up to achieve high availability; there are no type restrictions for slave nodes, which can be cloud servers, bare metal servers, or host machines running cloud servers.
[0043] The containerized deployment of Consul is completed on the master node and the slave node. Several Consul agents are running on each node, which are respectively bound to the networks on their respective nodes, and finally an independent network cluster is formed. For example, 3 Consul agents are running on each node, which are respectively bound to the control network, the service network, and the storage network. Finally, 3 independent clusters are formed, namely the Consul cluster based on the control network (Control Network), the Consul cluster based on the service network (Tenant Network), and the Consul cluster based on the storage network (Storage Network).
[0044] It is allowed to customize the number of networks to be bound to form the corresponding clusters. Once the network cluster is formed, the network traffic of all nodes' network cards can be detected. On the master node, the Consul agent runs in server mode to maintain the state of Consul; on the slave node, it runs in client mode for health checks and forwarding queries to the server. Consul uses the gossip protocol to manage the nodes in the cluster. If the connection of an agent on a node is found to be disconnected, it will broadcast this message to the entire cluster.
[0045] At the containerization level, based on the Consul chart provided by the community, modify the Consul clients in the client-daemenset to start several containers, that is, several Consul agents are included in a pod of a Consul client, which are respectively bound to the networks specified in the configuration file, and the hostNetwork should be used in this client-daement; modify the server-configmap to also use the hostNetwork. Similarly, 3 Consul agents are running in a pod and running in server mode.
[0046] As a technical solution that needs to be further explained in the present invention, it should be noted in step 2 that: a warning receiving module is deployed on the master node, which depends on a custom collector to obtain monitoring metrics. A Consul client is written in the custom collector. By calling the Consul API, the network topologies of clusters in different networks are parsed. Custom warning rules are allowed. When the metrics meet the warning conditions, relevant network warnings will appear on Prometheus, and further use Alertmanager to notify the downtime detection module.
[0047] It should be noted that the cluster network topology of Consul only plays a triggering role and cannot directly determine whether there is really a problem with the network. The reason is that Consul manages the nodes in the cluster based on the gossip protocol. When the network latency or packet loss rate reaches a threshold, it will cause the status detected by Consul to be abnormal. After actual research, the current status of Consul includes five states: AgentMemberNone, AgentMemberAlive, AgentMemberLeaving, AgentMemberLeft, AgentMemberFailed, and UnKnown. Among them, AgentMemberLeft and AgentMemberFailed can clearly indicate that network alarms need to be triggered, but the truly decisive factor is the deployment of the downtime detection module.
[0048] Therefore, in the present invention, the alarm rule is that if all the networks of a certain slave node are AgentMemberLeft or AgentMemberFailed and the duration is 2 minutes, then the network alarm can be triggered. After the network alarm is triggered, the downtime detection module will be automatically called to evaluate again.
[0049] As a technical solution that needs to be further explained in the present invention, regarding step 3, it should be noted that: a cross-architecture remote execution command module is deployed, and the downtime detection module depends on this module. In the present invention, the interaction between the master node and the slave node includes two methods. One is API call, and the other is remote execution of commands. The former itself supports cross-architecture, while the latter requires design and development. The present invention is developed based on the open-source configuration management and remote execution tool Salt Stack (usually abbreviated as Salt), and realizes the cross-architecture synchronous execution of remote commands, laying a foundation for realizing cross-architecture server downtime detection. Because for different versions of operating systems, the compatibility of software packages is not consistent, so a containerization method is needed to make it compatible with each version of the operating system. Specifically:
[0050] 1) Containerization design, chart management:
[0051] ① The salt-master is managed by deployment, and the replicas are set to 3 to achieve high availability of the salt master, and podAntiAffinity is set to avoid multiple salt-masters being located on the same node
[0052] ② The salt-master sets up two containers. One is keys-init, which is used to clean up the master key that already exists on the source node so that the salt-master can use the same key. If different masters use different master keys, the minion may not be able to communicate correctly with all masters, thus achieving the high availability of the salt-master. The other is salt-master, which is used to start the salt-master and salt-api services.
[0053] ③ The health check mechanism of the salt-master is set to call the 8000 port of the control network of this node for login verification. The 8000 port is the port exposed by the salt-api. If the login is successful, it means that the salt-api can be used normally.
[0054] ④ The salt-slave is managed by the daemonset.
[0055] ⑤ The health check mechanism of the salt-slave is set to check the TCP connection status of the peer on port 4505, and the address of the peer is the management network address of the node where the salt-master pod is located. If the connection is normal, it means that the remote command can be executed normally.
[0056] 2) Image making: Since salt is a remote execution command tool, it is impossible to install all software into the image during containerization. Therefore, the best way is to directly use the software on the host. By mounting the root directory of the computing node to / mnt in the container, and then when executing a remote command in the container, adding the prefix chroot / mnt to the command can achieve the invocation of the software command on the host. In addition, to ensure that the salt command will not be overwritten in the case of multi-threading, the command to be executed is written to a randomly named file, and then each time "chroot / mnt bash + <random file name>" is executed.
[0057] 3) Support cross-cluster and cross-architecture: The salt-slave can be deployed separately. As long as the network communication is normal, salt-slaves with different architectures and different platforms can connect to a specific salt-master. In this way, on the salt-master node, commands can be batch-distributed to nodes with different architectures.
[0058] As a technical solution that needs to be further described in the present invention, it should be noted that for step 4: after receiving the alarm, the downtime detection module first parses the alarm to parse out the name of the slave node. Usually, when deploying nodes, the network card metadata information will be stored on the master node. If there is a problem with the metadata, the Consul network topology cannot run properly, and the abnormal status of the corresponding container will still be detected. Therefore, after the cluster is normally used, the network card metadata of the master node is correct and there is a backup. In addition, some of the different network clusters may deploy gateways, while others do not require gateways.
[0059] After the downtime detection module receives the alarm and parses out the name of the slave node, according to the network card metadata information, it first selects several nodes that can ping the corresponding network gateway or the corresponding network of the master. For example, it is parsed that all networks of the node slave1 are abnormal. According to the configuration file, 2 nodes with normal communication need to be selected. At this time, both slave2 and slave3 can ping the gateway of their own network or the corresponding network of the master node for all networks. Then, slave2 and slave3 are normal nodes. Then, using the cross-architecture remote execution command module, the ping command is executed concurrently to detect whether slave2 and slave3 are both unreachable from slave1. If both are unreachable, it truly indicates that there is a problem with slave1.
[0060] In short, through the above two-layer network detection, using the gossip protocol of Consul and the icmp protocol of ping respectively, double detection of the network is realized. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination, enhancing the reliability of downtime detection. For example, Consul detects the network status of the slave node control network, service network, and storage network every 15s. If in 1 minute, that is, among 4 detection points, all 3 networks of the node to be tested are AgentMemberLeft or AgentMemberFailed, an alarm will be directly sent to Prometheus, and Alertmanager will notify the alarm to the downtime detection module. The downtime detection module parses out the name of the alarm node, uses the remote execution command module to perform network ping detection, selects 2 nodes with a packet loss rate of 0 for pinging the control network, service network, and storage network gateways, and then uses these 2 nodes to ping the node to be tested. If the packet loss rate of all network cards is 100%, the node to be tested is considered to be down; it supports setting the number of nodes for pinging the gateway, the number of detection rounds, and the number of ping packets.
[0061] Embodiment 2, based on Embodiment 1, proposes a system for a cross-architecture server downtime detection method, including a cluster network topology detection module, which is used for: dividing nodes into master nodes and slave nodes, where the master nodes are physical machines and multiple master nodes are set to achieve high availability, and there is no type limit for slave nodes; completing the containerization deployment of Consul on the master nodes and slave nodes, and each node runs several Consul agents respectively bound to the networks on their respective nodes to form an independent network cluster; allowing customization of the number of bindings to the network to form corresponding clusters; on the master nodes, the Consul agent runs in server mode to maintain the state of Consul; on the slave nodes, it runs in client mode to perform health checks and forward queries to the server; based on the Consul chart provided by the community, changing the client-daemenset and server-configmap so that a Consul client pod contains several Consul agents respectively bound to the networks specified in the configuration file and uses hostNetwork.
[0062] Including an alarm receiving module, which is used for: being deployed on the master nodes, depending on a custom collector to obtain monitoring metrics, writing a Consul client in the custom collector, and parsing the cluster network topologies of different networks by calling the Consul API; allowing customization of alarm rules, when the metrics meet the alarm conditions, relevant network alarms appear on Prometheus, and then using Alertmanager to notify the downtime detection module; the alarm rule is that all networks of a certain slave node are AgentMemberLeft or AgentMemberFailed and the duration is 2 minutes, then a network alarm is triggered and the downtime detection module is automatically called to evaluate again.
[0063] It includes a cross-architecture remote execution command module, which is used for: developing based on the open-source configuration management and remote execution tool SaltStack to achieve cross-architecture synchronous execution of remote commands; adopting containerized design, where salt-master is managed by deployment, the replicas are set to 3 to achieve high availability and podAntiAffinity is set, and salt-master has 2 containers, namely the keys-init container for cleaning the masterkey on the source node and the salt-master container for starting the salt-master and salt-api services. The health detection mechanism of salt-master is to call the 8000 port of the local control network for login verification; salt-slave is managed by daemonset, and its health detection mechanism is to check the tcp connection status of the peer 4505 port and the peer address is the management network address of the node where the salt-master pod is located; when making the image, mount the root directory of the computing node to / mnt in the container, and add the prefix chroot / mnt when executing remote commands to call the host software commands. The method of writing the command to be executed into a randomly named file is adopted to ensure that the salt commands do not overwrite in the case of multi-threading; it supports cross-cluster and cross-architecture, and salt-slave can be deployed independently, and the salt-master node can batch distribute commands to nodes of different architectures.
[0064] It includes a downtime detection module, which is used for: after receiving an alarm, parsing out the name of the slave node; according to the network card metadata information, selecting several nodes that can ping the corresponding network gateway or the network corresponding to the master; using the cross-architecture remote execution command module to concurrently execute the ping command to detect whether the selected nodes are all unreachable from the slave node to be tested. If they are all unreachable, it indicates that there is a problem with the slave node to be tested; double network detection is achieved through the gossip protocol of consul and the icmp protocol of ping. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination; it supports setting the number of ping gateway nodes, the number of detection rounds, and the number of ping packets.
[0065] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A cross-architecture server downtime detection method, characterized in that: Deploy the cluster network topology detection module: Divide the nodes into master nodes and slave nodes. The master node is a physical machine and multiple master nodes are set up to achieve high availability. There are no restrictions on the type of slave nodes. Complete the containerized deployment of Consul on the master nodes and slave nodes. Each node runs several Consul agents that are respectively bound to the networks on their own nodes to form an independent network cluster. Allow customizing the number of bindings to the network to form the corresponding cluster. On the master node, the Consul agent runs in server mode, and on the slave node, it runs in client mode. Consul uses the gossip protocol to manage the nodes in the cluster. Based on the Consul chart provided by the community, modify the client-daemenset and server-configmap so that a Consul client pod contains several Consul agents that are respectively bound to the networks specified in the configuration file and use hostNetwork.
2. The cross-architecture server downtime detection method according to claim 1, wherein: Include the deployment of the alarm receiving module: Deploy the alarm receiving module on the master node. This module depends on a custom collector to obtain monitoring metrics. A Consul client is written in the custom collector. The cluster network topology of different networks is parsed by calling the Consul API. Allow customizing the alarm rules. When the metrics meet the alarm conditions, relevant network alarms appear on Prometheus, and then the alertmanager is used to notify the downtime detection module. The alarm rule is that all networks of a certain slave node are AgentMemberLeft or AgentMemberFailed and the duration is 2 minutes, then a network alarm is triggered and the downtime detection module is automatically called to evaluate again.
3. A cross-architecture server downtime detection method according to claim 2, characterized in that: Include the deployment of the cross-architecture remote execution command module: Develop based on the open-source configuration management and remote execution tool SaltStack to achieve synchronous execution of remote commands across architectures. Adopt containerized design. The salt-master is managed by deployment. The replicas are set to 3 to achieve high availability and podAntiAffinity is set. The salt-master is set with 2 containers, namely the keys-init container for cleaning the master key on the source node and the salt-master container for starting the salt-master and salt-api services. The health detection mechanism of the salt-master is to call the 8000 port of the control network of this node for login verification; the salt-slave is managed by daemonset, and its health detection mechanism is to check the tcp connection status of the peer 4505 port and the peer address is the management network address of the node where the salt-master pod is located; when making the image, mount the root directory of the computing node to / mnt in the container. When executing remote commands, prefix them with chroot / mnt to call the host software commands. Write the commands to be executed into randomly named files to ensure that salt commands do not overwrite in the case of multi-threading; this module supports cross-cluster and cross-architecture, and the salt-slave can be deployed independently. The salt-master node can batch distribute commands to nodes of different architectures.
4. A cross-architecture server downtime detection method according to claim 3, characterized in that: It includes a deployment downtime detection module: after receiving an alarm, the downtime detection module parses out the name of the slave node; According to the network card metadata information, select several nodes that can ping the corresponding network gateway or the network corresponding to the master; use the cross-architecture remote execution command module to concurrently execute the ping command to detect whether the selected nodes are all unreachable from the slave node to be tested. If they are all unreachable, it indicates that there is a problem with the slave node to be tested; implement double network detection through the gossip protocol of consul and the icmp protocol of ping. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination; Consul detects the network status of the control network, service network, and storage network of the slave node every 15s. If within 1 minute, all 3 networks of the node to be tested among the 4 detection points are AgentMemberLeft or AgentMemberFailed, an alarm will be sent to prometheus, and alertmanager will notify the downtime detection module of the alarm. After the downtime detection module parses out the name of the alarm node, it performs network ping detection, and supports setting the number of ping gateway nodes, the number of detection rounds, and the number of ping packets.
5. A system for the cross-architecture server downtime detection method according to claim 4, characterized in that: It includes a cluster network topology detection module, which is used for: Divide the nodes into master nodes and slave nodes. The master nodes are physical machines and multiple are set to achieve high availability. There is no type limit for the slave nodes; Complete the containerized deployment of Consul on the master node and slave nodes. Run several Consul agents on each node, which are respectively bound to the networks on their respective nodes to form independent network clusters. Allow customizing the number of bindings to the network to form corresponding clusters. On the master node, Consul agents run in server mode to maintain the state of Consul. On the slave nodes, they run in client mode to perform health checks and forward queries to the server. Based on the Consul chart provided by the community, modify the client-daemenset and server-configmap so that a Consul client pod contains several Consul agents that are respectively bound to the networks specified in the configuration file and use hostNetwork.
6. A system according to claim 5, wherein: It includes an alarm receiving module, which is used for: Deployed on the master node, it depends on a custom collector to obtain monitoring metrics. A Consul client is written in the custom collector, and the network topologies of clusters with different networks are parsed by calling the Consul API. Allow customizing alarm rules. When the metrics meet the alarm conditions, relevant network alarms appear on Prometheus, and then use Alertmanager to notify the downtime detection module. The alarm rule is that all networks of a certain slave node are AgentMemberLeft or AgentMemberFailed and last for 2 minutes, then trigger a network alarm and automatically call the downtime detection module to evaluate again.
7. A system according to claim 6, wherein: It includes a cross-architecture remote execution command module, which is used for: Developed based on the open-source configuration management and remote execution tool SaltStack to achieve synchronous execution of remote commands across architectures. Adopt containerized design. salt-master is managed by Deployment, and the replicas are set to 3 to achieve high availability and set podAntiAffinity. salt-master has 2 containers, namely the keys-init container for cleaning the master key on the source node and the salt-master container for starting the salt-master and salt-api services. The health detection mechanism of salt-master is to call the 8000 port of the control network of this node for login verification. salt-slave is managed by DaemonSet, and its health detection mechanism is to check the TCP connection status of the peer 4505 port and the peer address is the management network address of the node where the salt-master pod is located. When making the image, mount the root directory of the computing node to / mnt in the container. When executing remote commands, add the prefix chroot / mnt to call the host software commands. Ensure that salt commands do not overwrite in the case of multi-threading by writing the commands to be executed into randomly named files. Support cross-cluster and cross-architecture. The salt-slave can be deployed independently, and the salt-master node can batch-distribute commands to nodes with different architectures.
8. A system according to claim 7, wherein: It includes a downtime detection module, which is used for: After receiving an alarm, parsing out the name of the slave node; According to the network card metadata information, selecting several nodes that can ping the corresponding network gateway or the network corresponding to the master; Using the cross-architecture remote execution command module to concurrently execute the ping command to detect whether the selected nodes are all unreachable from the slave node to be tested. If they are all unreachable, it indicates that there is a problem with the slave node to be tested; Implement double detection of the network through the gossip protocol of consul and the icmp protocol of ping. The gossip protocol is responsible for triggering, and the icmp protocol is responsible for further determination; Support setting the number of nodes for pinging the gateway, the number of detection rounds, and the number of ping packets.