System and method for supporting multi-scene cross-platform computing node high availability
The system addresses the limitations of existing cloud computing solutions by implementing automated fault detection and recovery across platforms, enhancing system responsiveness and resource utilization for diverse cloud environments.
Patent Information
- Application Number
- CN202510216737.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing cloud computing solutions for computing node high availability (HA) lack flexibility in handling diverse scenarios, require complex manual operations, do not support nuanced handling of different storage types, and lack automatic alerting and fault analysis, especially in cross-platform environments.
A system and method utilizing an alerting module, cross-platform remote execution module, and multi-scenario evaluation module to handle various fault scenarios, leveraging Telegraf, Prometheus, and Salt Stack for automated fault detection and recovery across different platforms.
Ensures rapid and effective handling of hardware, software, and network faults, enhancing system responsiveness and resource utilization, ensuring business continuity and stability across diverse cloud environments.
Smart Images

Figure CN119996150A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of cloud computing, and in particular to a system and method for supporting high availability of computing nodes across multiple scenarios and platforms. Background Art
[0002] In cloud services, a computing node usually refers to a physical server in the cloud infrastructure, which is used to perform computing tasks and process data, and common forms include virtual machines and containers. These nodes are part of the cloud service provider's resource pool and can be rented to customers to meet their computing needs.
[0003] Compute node high availability (HA) refers to the ability to remain available and operational for a long period of time, even in the face of hardware failures, software errors, network problems, or other types of failures. High availability is a key feature of system design, which ensures that the system can continue to work normally in the face of various failures, thereby reducing downtime, improving user experience and business continuity.
[0004] Many cloud vendors provide high availability features for computing nodes, such as:
[0005] 1. VMware vSphere HA
[0006] A VMware vSphere HA cluster allows a collection of ESXi hosts to work together as a group that provides a higher level of availability for virtual machines than an ESXi host could provide individually. When a vSphere HA cluster is created, one host is automatically elected as the primary host. The primary host communicates with the vCenter Server and monitors the status of all protected virtual machines and the status of the secondary hosts. Different types of host failures can occur, and the primary host must detect and handle the failures appropriately. If the primary host fails, is powered off or placed in standby mode, or is removed from the cluster, a new election is held. After a host failure, restarting a virtual machine on a new host takes into account factors such as the ability to access the virtual machine's files from an active cluster host that can communicate with the primary host over the network, setting the order in which to restart the virtual machines, virtual machine and host compatibility, and resource reservations.
[0007] 2. Inspur InCloud Sphere configuration HA
[0008] InCloud Sphere is an enterprise-level open virtualization solution platform launched by Inspur. Click [Compute Pool] in the homepage menu bar or navigation bar, select a cluster in the navigation bar, and click [Configuration] -> [HA Service] of the current cluster to turn HA on or off. The HA access control policy mainly controls whether other hosts are allowed to take over when a host in the cluster fails. When the HA access control policy is turned on for a business, the user can set the failover host; when the policy is turned off, the system automatically removes the failover host. The failover host is only used to trigger HA to automatically migrate the virtual machine to the destination host after a node fails in the cluster. It is prioritized to run on the failover host. If the failover host does not meet the conditions, it will continue to try to run on the non-failover host. The HA priority supports three levels: high, medium, and low. The default is "medium", which is used to control the startup order of HA for virtual machines in the cluster. After the virtual machine triggers HA, it starts or recovers in the order of high, medium, and low. The HA function of the InCloudSphere system cluster is implemented based on shared storage.
[0009] 3. Huawei Cloud CAR Platform
[0010] The fault handling CAR platform (CloudAutoRemediation) is a platform that provides automated fault handling and fault diagnosis for SRE (Site Reliability Engineering). In terms of fault diagnosis, it discovers faults before users do, shortening the cycle from service interruption to fault recovery in the existing network. In terms of root cause analysis, it supports fault tree / diagnosis process configuration management, converts the experience of operation and maintenance personnel into automated processing capabilities, supports various atomic capability extensions, provides certain customized analysis capabilities, cooperates with big data offline / online analysis capabilities, supports the expansion of fault analysis in the field of machine learning / AI, and integrates the end-to-end processing of the operation and maintenance platform automation, integrates various operation and maintenance automation capabilities such as monitoring, alarm, call chain, API problem discovery, and change automation, covers various public cloud technology stacks and problem location knowledge accumulation; in terms of fault self-healing, it provides automated recovery capabilities after analyzing the root cause of the fault, and provides an end-to-end one-stop solution for faults.
[0011] 4. OpenStack Masakari
[0012] Openstack masakari is an open source high availability solution. Instancemonitor detects virtual machine failures, processmonitor detects service failures, and hostmonitor detects host failures. The host failure detection function provided by hostmonitor realizes the high availability of computing nodes.
[0013] Although many cloud vendors provide high availability solutions for computing nodes, due to closed-source reasons, there is no way to recover services in different ways for different scenarios from the public solutions. In addition, the open source OpenStack Masakari solution has the following problems:
[0014] (1) The operation is complex and involves a lot of manual operations, which are prone to omissions: it is necessary to manually set the number of nodes in the cluster (called segments in masakari) and specify which node is the reserved node. After the evacuation is triggered once, the reserved node’s reserved attribute needs to be set again, otherwise the reserved node cannot be evacuated a second time.
[0015] (2) Does not support detailed processing of multiple scenarios: The currently supported scenario is to use pacemaker and corosync to detect whether the heartbeat between nodes is normal. If it is not normal, business evacuation is performed.
[0016] (3) Few processing methods: Only evacuation is supported, not migration. Evacuation means rebuilding the same virtual machine on another node after a computing node goes down; migration includes cold migration and hot migration. Hot migration allows virtual machines to migrate between two computing nodes without being noticed by the user, that is, the user will not be aware of the migration action. Cold migration is to shut down the virtual machine for migration.
[0017] (4) Using the disable computing service method to allow evacuation may cause disk read and write abnormalities.
[0018] (5) Unable to automatically receive alarms and locate the cause of the fault.
[0019] (6) Unable to perform fine-grained processing for different storage types: Volumes of different storage types should be processed differently. For example, virtual machines mounted with lvm volumes should not be evacuated; virtual machines mounted with ceph volumes need to ensure that the storage network is normal; virtual machines mounted with fcsan volumes should consider whether the FC link is normal. Therefore, for such complex scenarios, Masakari did not consider the different storage types.
[0020] (7) Inactive community: No significant code updates in recent years.
[0021] In real cloud service scenarios, there will be no business scenarios, multiple storage backend scenarios, node downtime scenarios, node stuck scenarios, node restart scenarios, network failure scenarios, etc. Different scenarios need to be sorted and analyzed, and different processing methods are given. Only in this way can we reasonably ensure that the business can be quickly restored in a targeted manner when facing hardware failures, software errors, network problems or other types of failures.
[0022] In addition, with the development of cloud computing, enterprises are increasingly using the cloud. Xinchuang Cloud has the advantages of shielding the complexity and differences of heterogeneous devices and supporting the unified and centralized supply of computing, storage, network and other resources. It is becoming more and more widely used in applications such as government cloud and industry cloud. Therefore, it is also crucial to achieve cross-platform high availability of computing nodes. However, in the known solutions, no specific description is given for the cross-platform function. Summary of the invention
[0023] In response to the needs and deficiencies of current technological development, the present invention provides a system and method for high availability of computing nodes that supports multiple scenarios and cross-platforms.
[0024] In a first aspect, the present invention provides a system that supports high availability of computing nodes in multiple scenarios and across platforms. The technical solutions adopted to solve the above technical problems are as follows:
[0025] A system supporting high availability of computing nodes in multiple scenarios and across platforms, comprising:
[0026] The alarm receiving module is used to obtain fault monitoring indicators through telegraf and custom collectors, report them to prometheus, and use webhooks to push alarms to the multi-scenario evaluation module;
[0027] The cross-platform remote command execution module is developed based on Salt Stack and is used to achieve cross-platform interaction between the control layer and computing nodes through containerized design, image production, and support for cross-cluster and cross-platform.
[0028] The multi-scenario assessment module is used to receive alarms and evaluate and provide treatment suggestions according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure.
[0029] The fault handling module is used to traverse the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario evaluation module, identify the storage type, determine the operations to be performed by the virtual machines, and execute them uniformly to restore the business.
[0030] Optionally, the alarm receiving module involved specifically performs the following operations:
[0031] Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems;
[0032] Report the collected indicator data to prometheus;
[0033] With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
[0034] Further optionally, the cross-platform remote execution command module involved realizes cross-platform interaction between the control layer and the computing node through containerization design, image production and support for cross-cluster and cross-platform. The specific process includes:
[0035] Containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. If the login is successful, it means that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally;
[0036] Image production: Mount the compute node root directory to the container's / mnt. When executing remote commands in the container, add the prefix chroot / mnt to call the host software command. Write the command to be executed into a randomly named file and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads.
[0037] Support cross-cluster and cross-platform: deploy salt-slave separately. When network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to enable salt-master to distribute commands to nodes of different architectures in batches.
[0038] Optionally, after receiving the alarm, the multi-scenario assessment module evaluates and gives processing suggestions in turn according to the seven scenarios of IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows:
[0039] IPMI unreachable scenario judgment: judge whether the IPMI communication of the alarm node is normal. If it is not, subsequent evaluation cannot be performed;
[0040] Judgment of scenario without virtual machine: judge whether there is a virtual machine on the alarm node. If there is no virtual machine on the alarm node, it will not affect the business and no operation is required.
[0041] Local storage node scenario judgment: The local storage node refers to the lvm node. The volumes mounted by all virtual machines of the alarm node are traversed. If a virtual machine mounts an lvm volume, it is determined to be an lvm node. When the lvm node fails, the alarm is sent again.
[0042] Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage virtual machines and centralized storage virtual machines.
[0043] Judgment of node stuck scenario: When the consul network cluster believes that all networks of the alarm node are blocked, and the healthy node cannot ping the networks of the alarm node, the node is considered stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated;
[0044] Node restart scenario judgment: Based on salt, call the IPMI command to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation;
[0045] Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure. Hot migration, alarm or evacuation are performed according to different situations.
[0046] Further optionally, the fault handling module involved specifically performs the following operations:
[0047] Receive operational suggestions from the multi-scenario assessment module;
[0048] Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and evaluation recommendations;
[0049] All pending operations of virtual machines are executed uniformly to achieve business recovery.
[0050] In a second aspect, the present invention provides a method for supporting high availability of computing nodes across multiple scenarios and platforms. The technical solution adopted to solve the above technical problems is as follows:
[0051] A method for supporting high availability of computing nodes in multiple scenarios and across platforms comprises the following steps:
[0052] S1. The alarm receiving module obtains fault monitoring indicators through telegraf and custom collectors, reports them to prometheus, and then uses webhook to push alarms to the multi-scenario evaluation module;
[0053] S2. Develop a cross-platform remote execution command module based on Salt Stack, and realize cross-platform interaction between the control layer and computing nodes through containerized design, image production, and support for cross-cluster and cross-platform.
[0054] After receiving the alarm, the multi-scenario assessment module will assess and provide treatment suggestions according to the seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure.
[0055] S4. The fault handling module traverses the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario assessment module, identifies the storage type, determines the operations to be performed by the virtual machines, and executes them uniformly to restore the business.
[0056] Optionally, the step S1 specifically includes the following process:
[0057] Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems;
[0058] Report the collected indicator data to prometheus;
[0059] With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
[0060] Further optionally, the step S2 involved specifically includes the following process:
[0061] S2.1, containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. If the login is successful, it means that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally;
[0062] S2.2, image production: mount the compute node root directory to the container's / mnt, add the prefix chroot / mnt when executing remote commands in the container to call the host software command; write the command to be executed into a randomly named file, and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads;
[0063] S2.3, support cross-cluster and cross-platform: deploy salt-slave separately. When the network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to realize batch distribution of commands by salt-master to nodes of different architectures.
[0064] Further optionally, step S3 is executed. After receiving the alarm, the multi-scenario assessment module sequentially assesses and gives processing suggestions according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows:
[0065] S3.1, IPMI unreachable scenario judgment: judge whether the alarm node IPMI communication is normal, if not, subsequent evaluation cannot be performed;
[0066] S3.2, No virtual machine scenario judgment: judge whether there is a virtual machine on the alarm node. If there is no virtual machine on the alarm node, it will not affect the business and no operation is required;
[0067] S3.3, local storage node scenario judgment: local storage node refers to lvm node, traverse the volumes mounted by all virtual machines of the alarm node, if there is a virtual machine mounted with lvm volume, it is judged as lvm node, when lvm node fails, send alarm again;
[0068] S3.4, Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage virtual machines and centralized storage virtual machines;
[0069] S3.5, Node stuck scenario judgment: When the consul network cluster believes that all networks of the alarm node are blocked, and the normal node cannot ping the networks of the alarm node, the node is considered to be stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated;
[0070] S3.6, Node restart scenario judgment: Based on salt, call IPMI commands to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation;
[0071] S3.7. Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure, and hot migration, alarm or evacuation are performed according to different situations.
[0072] Further optionally, the step S4 involved specifically includes the following process:
[0073] S4.1. Receive operational suggestions from the multi-scenario assessment module;
[0074] S4.2. Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and the evaluation recommendations;
[0075] S4.3. All pending operations of the virtual machines are executed uniformly to achieve business recovery.
[0076] The system and method of the present invention for supporting high availability of computing nodes in multiple scenarios and across platforms have the following beneficial effects compared with the prior art:
[0077] 1. The present invention can realize the effective transmission of faults, improve the utilization rate of resources, enhance the flexibility and scalability of the system, ensure that the system can make correct responses in various situations, and ensure the continuity and stability of business;
[0078] 2. The present invention realizes high availability of computing nodes. No matter in case of hardware failure, software error or network problem, the system can handle it quickly and effectively, ensuring the normal operation of the business; through the reasonable allocation and management of resources, the utilization rate of resources is improved and the cost is reduced; through the design and implementation of fault monitoring, alarm, remote execution, multi-scenario evaluation and fault handling, the security of the system is ensured;
[0079] 3. The present invention sorts and analyzes no-business scenarios, multi-storage backend scenarios, node downtime scenarios, node stuck scenarios, node restart scenarios, network failure scenarios, etc., and performs targeted business recovery according to different scenarios. Recovery measures include virtual machine evacuation, hot migration, cold migration, alarms, and no operations. At the same time, it is developed based on the open source configuration management and remote execution tool Salt Stack (usually referred to as Salt), which realizes cross-platform synchronous execution of remote commands, lays the foundation for achieving high availability of computing nodes across platforms, and meets the cloud computing support for the Xinchuang Cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] Attached Figure 1 is a module connection block diagram of the first embodiment of the present invention;
[0081] Attached Figure 2 It is a method flow chart of embodiment 2 of the present invention. DETAILED DESCRIPTION
[0082] In order to make the technical solution, the technical problem solved and the technical effect of the present invention more clearly understood, the technical solution of the present invention is clearly and completely described below in conjunction with specific embodiments.
[0083] Embodiment 1:
[0084] Combined with Figure 1 This embodiment proposes a system that supports high availability of computing nodes in multiple scenarios and across platforms, which includes:
[0085] The alarm receiving module is used to obtain fault monitoring indicators through telegraf and custom collectors, report them to prometheus, and use webhooks to push alarms to the multi-scenario evaluation module;
[0086] The cross-platform remote command execution module is developed based on Salt Stack and is used to achieve cross-platform interaction between the control layer and computing nodes through containerized design, image production, and support for cross-cluster and cross-platform.
[0087] The multi-scenario assessment module is used to receive alarms and evaluate and provide treatment suggestions according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure.
[0088] The fault handling module is used to traverse the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario evaluation module, identify the storage type, determine the operations to be performed by the virtual machines, and execute them uniformly to restore the business.
[0089] In this embodiment, the alarm receiving module involved specifically performs the following operations:
[0090] Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems;
[0091] Report the collected indicator data to prometheus;
[0092] With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
[0093] In this embodiment, the cross-platform remote execution command module involved realizes cross-platform interaction between the control layer and the computing node through containerization design, image production and support for cross-cluster and cross-platform. The specific process includes:
[0094] Containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. Successful login indicates that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally;
[0095] Image production: Mount the compute node root directory to the container's / mnt. When executing remote commands in the container, add the prefix chroot / mnt to call the host software command. Write the command to be executed into a randomly named file and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads.
[0096] Support cross-cluster and cross-platform: deploy salt-slave separately. When network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to enable salt-master to distribute commands to nodes of different architectures in batches.
[0097] In this embodiment, after receiving the alarm, the multi-scenario assessment module evaluates and gives processing suggestions in turn according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows:
[0098] IPMI unreachable scenario judgment: judge whether the IPMI communication of the alarm node is normal. If it is not, subsequent evaluation cannot be performed;
[0099] Judgment of scenario without virtual machine: judge whether there is a virtual machine on the alarm node. If there is no virtual machine on the alarm node, it will not affect the business and no operation is required.
[0100] Local storage node scenario judgment: Local storage node refers to LVM node. LVM (Logical Volume Manager) is a tool for managing disk storage space. The method to judge whether a node is an LVM node is to traverse the volumes mounted on all virtual machines of this node. If a virtual machine has an LVM volume mounted, it is judged to be an LVM node. For LVM nodes, if a node fails, the node cannot be shut down or virtual machines cannot be evacuated, because local storage is in the alarm node. If the alarm node is shut down, the virtual machine business mounted with local storage cannot be restored. Therefore, the method adopted is to send an alarm again;
[0101] Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage (specifically ceph storage, ceph is an open source storage platform that provides high-performance, scalable and reliable object, block and file storage solutions) virtual machines and centralized storage (specifically FC storage, a high-speed network technology used to connect computer systems and storage devices) virtual machines;
[0102] Judgment of node stuck scenario: When the consul network cluster believes that all networks of the alarm node are blocked, and the healthy node cannot ping the networks of the alarm node, the node is considered stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated;
[0103] Node restart scenario judgment: Based on salt, call the IPMI command to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation;
[0104] Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure. Hot migration, alarm or evacuation are performed according to different situations.
[0105] It is necessary to add that: ① For business network failure or business network + fc network failure: This combination of failures will affect virtual machines mounted with various types of storage. First, determine whether the virtual machine supports hot migration, and ensure that the service and hot migration network card status are normal. If it is satisfied, hot migration will be performed; if not, the preset configuration item data_net_force_evacuate is used to decide whether to issue an alarm or perform an evacuation operation; ② For storage network failure: This failure only affects virtual machines mounted with ceph type storage. Since hot migration is not supported after the storage network fails (hot migration will be locked), for virtual machines mounted with ceph type storage, the preset configuration item stor_net_force_evacuate is used to decide whether to issue an alarm or perform an evacuation operation; ③ For fc network failure: This failure only affects virtual machines mounted with fc type storage. First determine whether the virtual machine supports hot migration. If it does, hot migrate the fc-type virtual machine. If it does not, decide whether to issue an alarm or evacuate based on the preset configuration item fc_net_force_evacuate. ④ For storage network failure + business network failure or storage network failure + fc network failure or storage network + fc network + business network failure: In this complex failure situation, for virtual machines mounted with ceph type storage, only evacuation can be performed. For virtual machines mounted with fc type storage, if hot migration is supported, hot migration is performed. If hot migration is not supported, shutdown and evacuation are performed.
[0106] In this embodiment, the fault processing module involved specifically performs the following operations:
[0107] Receive operational suggestions from the multi-scenario assessment module;
[0108] Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and evaluation recommendations;
[0109] All pending operations of virtual machines are executed uniformly to achieve business recovery.
[0110] For example, the fault handling module recommends hot migration of fc type virtual machines and no operation on ceph type virtual machines. During the fault handling phase, it is determined that virtual machine A has mounted fc type storage volumes, virtual machine B has mounted ceph type storage volumes, and virtual machine C has mounted both fc type and ceph type storage volumes. In this case, both A and C need to be hot migrated, and no operation is performed on B. The reason is that virtual machine C is hot migrated according to the fc type mounted, but no operation is performed according to the ceph type mounted. In summary, hot migration should be performed because if hot migration is not performed, the virtual machine has already used fc type volumes, and fc network abnormalities will lead to business abnormalities. Hot migration will be done without user perception.
[0111] Embodiment 2:
[0112] Combined with Figure 2 This embodiment proposes a method for supporting high availability of computing nodes in multiple scenarios and across platforms, which includes the following steps:
[0113] S1. The alarm receiving module obtains fault monitoring indicators through telegraf and custom collectors, reports them to prometheus, and then uses webhook to push alarms to the multi-scenario evaluation module. The specific process includes the following:
[0114] Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems;
[0115] Report the collected indicator data to prometheus;
[0116] With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
[0117] S2. Develop a cross-platform remote execution command module based on Salt Stack. Through containerized design, image production, and support for cross-cluster and cross-platform, realize cross-platform interaction between the control layer and computing nodes. The specific process includes the following:
[0118] S2.1, containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. If the login is successful, it means that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally;
[0119] S2.2, image production: mount the compute node root directory to the container's / mnt, add the prefix chroot / mnt when executing remote commands in the container to call the host software command; write the command to be executed into a randomly named file, and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads;
[0120] S2.3, support cross-cluster and cross-platform: deploy salt-slave separately. When the network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to realize batch distribution of commands by salt-master to nodes of different architectures.
[0121] After receiving the alarm, the multi-scenario assessment module will evaluate and provide treatment suggestions according to the seven scenarios of IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows:
[0122] S3.1, IPMI unreachable scenario judgment: judge whether the alarm node IPMI communication is normal, if not, subsequent evaluation cannot be performed;
[0123] S3.2, No virtual machine scenario judgment: judge whether there is a virtual machine on the alarm node. If there is no virtual machine on the alarm node, it will not affect the business and no operation is required;
[0124] S3.3, local storage node scenario judgment: local storage node refers to lvm node, traverse the volumes mounted by all virtual machines of the alarm node, if there is a virtual machine mounted with lvm volume, it is judged as lvm node, when lvm node fails, send alarm again;
[0125] S3.4, Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage virtual machines and centralized storage virtual machines;
[0126] S3.5, Node stuck scenario judgment: When the consul network cluster believes that all networks of the alarm node are blocked, and the normal node cannot ping the networks of the alarm node, the node is considered to be stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated;
[0127] S3.6, Node restart scenario judgment: Based on salt, call IPMI commands to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation;
[0128] S3.7. Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure, and hot migration, alarm or evacuation are performed according to different situations.
[0129] It is necessary to add that: ① For business network failure or business network + fc network failure: This combination of failures will affect virtual machines mounted with various types of storage. First, determine whether the virtual machine supports hot migration, and ensure that the service and hot migration network card status are normal. If it is satisfied, hot migration will be performed; if not, the preset configuration item data_net_force_evacuate is used to decide whether to issue an alarm or perform an evacuation operation; ② For storage network failure: This failure only affects virtual machines mounted with ceph type storage. Since hot migration is not supported after the storage network fails (hot migration will be locked), for virtual machines mounted with ceph type storage, the preset configuration item stor_net_force_evacuate is used to decide whether to issue an alarm or perform an evacuation operation; ③ For fc network failure: This failure only affects virtual machines mounted with fc type storage. First determine whether the virtual machine supports hot migration. If it does, hot migrate the fc-type virtual machine. If it does not, decide whether to issue an alarm or evacuate based on the preset configuration item fc_net_force_evacuate. ④ For storage network failure + business network failure or storage network failure + fc network failure or storage network + fc network + business network failure: In this complex failure situation, for virtual machines mounted with ceph type storage, only evacuation can be performed. For virtual machines mounted with fc type storage, if hot migration is supported, hot migration is performed. If hot migration is not supported, shutdown and evacuation are performed.
[0130] S4. The fault handling module traverses the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario assessment module, identifies the storage type, determines the operations to be performed by the virtual machines, and executes them uniformly to restore the business. The specific process includes the following:
[0131] S4.1. Receive operational suggestions from the multi-scenario assessment module;
[0132] S4.2. Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and the evaluation recommendations;
[0133] S4.3. All pending operations of the virtual machines are executed uniformly to achieve business recovery.
[0134] For example, the fault handling module recommends hot migration of fc type virtual machines and no operation on ceph type virtual machines. During the fault handling phase, it is determined that virtual machine A has mounted fc type storage volumes, virtual machine B has mounted ceph type storage volumes, and virtual machine C has mounted both fc type and ceph type storage volumes. In this case, both A and C need to be hot migrated, and no operation is performed on B. The reason is that virtual machine C is hot migrated according to the fc type mounted, but no operation is performed according to the ceph type mounted. In summary, hot migration should be performed because if hot migration is not performed, the virtual machine has already used fc type volumes, and fc network abnormalities will lead to business abnormalities. Hot migration will be done without user perception.
[0135] In summary, the system and method of the present invention that supports high availability of computing nodes in multiple scenarios and across platforms can achieve effective transmission of faults, improve resource utilization, enhance system flexibility and scalability, ensure that the system can make correct responses in various situations, and ensure business continuity and stability.
[0136] The above specific examples are used to explain the principles and implementation methods of the present invention in detail. These examples are only used to help understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made by technicians in this technical field without departing from the principles of the present invention should fall within the scope of patent protection of the present invention.
Claims
1. A system that supports high availability of computing nodes in multiple scenarios and across platforms, characterized in that: It includes: The alarm receiving module is used to obtain fault monitoring indicators through telegraf and custom collectors, report them to prometheus, and use webhooks to push alarms to the multi-scenario evaluation module; The cross-platform remote command execution module is developed based on Salt Stack and is used to achieve cross-platform interaction between the control layer and computing nodes through containerized design, image production, and support for cross-cluster and cross-platform. The multi-scenario assessment module is used to receive alarms and evaluate and provide treatment suggestions according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The fault handling module is used to traverse the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario evaluation module, identify the storage type, determine the operations to be performed by the virtual machines, and execute them uniformly to restore the business.
2. According to claim 1, a system for supporting high availability of computing nodes in multiple scenarios and across platforms is characterized in that: The alarm receiving module specifically performs the following operations: Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems; Report the collected indicator data to prometheus; With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
3. A system for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 2, characterized in that: The cross-platform remote command execution module realizes cross-platform interaction between the control layer and the computing node through containerized design, image production and support for cross-cluster and cross-platform. The specific process includes: Containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. If the login is successful, it means that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally; Image production: Mount the compute node root directory to the container's / mnt. When executing remote commands in the container, add the prefix chroot / mnt to call the host software command. Write the command to be executed into a randomly named file and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads. Support cross-cluster and cross-platform: deploy salt-slave separately. When network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to enable salt-master to distribute commands to nodes of different architectures in batches.
4. A system for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 3, characterized in that: After receiving the alarm, the multi-scenario assessment module evaluates and gives processing suggestions in turn according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows: IPMI unreachable scenario judgment: judge whether the IPMI communication of the alarm node is normal. If it is not, subsequent evaluation cannot be performed; Judgment of scenario without virtual machine: judge whether there is a virtual machine on the alarm node. If there is no virtual machine on the alarm node, it will not affect the business and no operation is required. Local storage node scenario judgment: The local storage node refers to the lvm node. The volumes mounted by all virtual machines of the alarm node are traversed. If a virtual machine mounts an lvm volume, it is determined to be an lvm node. When the lvm node fails, the alarm is sent again. Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage virtual machines and centralized storage virtual machines. Judgment of node stuck scenario: When the consul network cluster believes that all networks of the alarm node are blocked, and the healthy node cannot ping the networks of the alarm node, the node is considered stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated; Node restart scenario judgment: Based on salt, call the IPMI command to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation; Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure. Hot migration, alarm or evacuation are performed according to different situations.
5. A system for high availability of computing nodes supporting multiple scenarios and cross-platforms according to claim 4, characterized in that: The fault processing module specifically performs the following operations: Receive operational suggestions from the multi-scenario assessment module; Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and evaluation recommendations; All pending operations of virtual machines are executed uniformly to achieve business recovery.
6. A method for supporting high availability of computing nodes in multiple scenarios and across platforms, characterized in that: The steps include: S1. The alarm receiving module obtains fault monitoring indicators through telegraf and custom collectors, reports them to prometheus, and then uses webhook to push alarms to the multi-scenario evaluation module; S2. Develop a cross-platform remote execution command module based on Salt Stack, and realize cross-platform interaction between the control layer and computing nodes through containerized design, image production, and support for cross-cluster and cross-platform. After receiving the alarm, the multi-scenario assessment module will assess and provide treatment suggestions according to the seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. S4. The fault handling module traverses the virtual machines of the faulty nodes according to the processing suggestions of the multi-scenario assessment module, identifies the storage type, determines the operations to be performed by the virtual machines, and executes them uniformly to restore the business.
7. A method for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 6, characterized in that: The step S1 specifically includes the following process: Use telegraf and custom collectors to collect monitoring indicators for hardware failures, software errors, or network problems; Report the collected indicator data to prometheus; With the help of webhook technology, the alarm information in prometheus is pushed to the multi-scenario evaluation module.
8. The method for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 7, characterized in that: The step S2 specifically includes the following process: S2.1, containerized design, chart management: ① Use deployment to manage salt-master, set up 3 replicas to achieve high availability, and configure podAntiAffinity to avoid multiple salt-masters on the same node; ② Set up two containers, keys-init and salt-master, clean up the source node masterkey through the keys-init container, and start the salt-master and salt-api services through the salt-master container; ③ Login verification is performed by calling port 8000 of the control network of this node. If the login is successful, it means that salt-api can be used normally; ④ Daemonset manages salt-slave; ⑤ Check the TCP connection status of port 4505 of the other end. The address of the other end is the management network address of the node where the salt-masterpod is located. If the connection is normal, it means that the remote command can be executed normally; S2.2, image production: mount the compute node root directory to the container's / mnt, add the prefix chroot / mnt when executing remote commands in the container to call the host software command; write the command to be executed into a randomly named file, and execute "chroot / mntbash+<random file name>" to prevent the salt command from being overwritten in multiple threads; S2.3, support cross-cluster and cross-platform: deploy salt-slave separately. When the network communication is normal, salt-slave of different architectures and platforms connect to the specified salt-master to realize batch distribution of commands by salt-master to nodes of different architectures.
9. A method for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 8, characterized in that: Execute step S3. After receiving the alarm, the multi-scenario assessment module evaluates and gives processing suggestions in turn according to seven scenarios: IPMI failure, no virtual machine, local storage node, node downtime, node stuck, node restart, and network failure. The specific process is as follows: S3.1, IPMI unreachable scenario judgment: judge whether the alarm node IPMI communication is normal, if not, subsequent evaluation cannot be performed; S3.2, No virtual machine scenario judgment: judge whether the alarm node has a virtual machine. If the alarm node has no virtual machine, it will not affect the business and no operation is required; S3.3, local storage node scenario judgment: local storage node refers to lvm node, traverse the volumes mounted by all virtual machines of the alarm node, if there is a virtual machine mounted with lvm volume, it is judged as lvm node, when lvm node fails, send alarm again; S3.4, Node downtime scenario judgment: Prioritize node downtime judgment. If downtime occurs, directly evacuate non-lvm distributed storage virtual machines and centralized storage virtual machines; S3.5, Node stuck scenario judgment: When the consul network cluster believes that all networks of the alarm node are blocked, and the normal node cannot ping the networks of the alarm node, the node is considered to be stuck, and the IPMI command is called based on salt to shut down the node, and then the non-lvm virtual machine is evacuated; S3.6, Node restart scenario judgment: Based on salt, call IPMI commands to collect out-of-band logs within 1 hour to determine whether the node restart is caused by hardware failure. If the business is restored after the restart, hot migrate the non-lvm virtual machine; if it is not restored, wait for the set time, and if it is still not restored, continue the subsequent evaluation; S3.
7. Network failure scenario judgment: Cloud service network failure scenarios are subdivided into business network failure or business network + FC network failure, storage network failure, FC network failure, storage network + business network or storage network + FC network or storage network + FC network + business network failure, and hot migration, alarm or evacuation are performed according to different situations.
10. A method for supporting high availability of computing nodes in multiple scenarios and across platforms according to claim 9, characterized in that: The step S4 specifically includes the following process: S4.
1. Receive operational suggestions from the multi-scenario assessment module; S4.
2. Start traversing all virtual machines on the faulty node, identify the storage type mounted on each virtual machine during the traversal process, and determine the operations to be performed on each virtual machine based on the mounted storage type and the evaluation recommendations; S4.
3. All pending operations of the virtual machines are executed uniformly to achieve business recovery.
Citation Information
Patent Citations
KVM virtual machine high availability method and system
CN114816658A
Internet of Things system
CN116368355A
CT cloud and edge cloud security platform
CN118432835A
Cloud platform fault detection and operation and maintenance system, method and device and storage medium
CN118550752A
Application system high availability evaluation method and device, equipment and medium
CN118747147A