Method, device, medium and system for realizing hot plug fault isolation

By monitoring hot-swap events in the device cluster, identifying the fault type and activating redundant hardware devices, switching business processes to the target redundant hardware devices, solving the problem of business capabilities caused by hardware failure or hot-swap in the device cluster, and achieving business continuity and reliability.

CN120336094APending Publication Date: 2025-07-18INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510465534.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In a device cluster, when hardware devices fail or hot plugging occurs, existing solutions cause other hardware devices to be downgraded and operate, resulting in reduced business capabilities.

Method used

By monitoring hot-swap events in the device cluster, identifying the fault type and activating the redundant hardware device, the business process is switched from the failed hardware device to the target redundant hardware device to maintain business uninterrupted, and the business processing capabilities of the target redundant hardware device and the hardware device are the same.

Benefits of technology

It realizes that in the event of hardware equipment failure or hot-swap, maintains business capabilities without degradation, ensures business continuity and reliability in the device cluster, and avoids insufficient performance and bandwidth due to degraded operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336094A_ABST
    Figure CN120336094A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device, a medium and a system for realizing hot plug fault isolation, relates to the technical field of equipment clusters, and aims to switch a business process from fault hardware equipment to target redundant hardware equipment by monitoring and identifying a hot plug event of an equipment cluster so as to keep the business in the equipment cluster uninterrupted and improve the safety of the equipment cluster. Therefore, the service capability is completely switched to the same hardware module, the function of the fault module is not shared by other hardware modules, which is reflected in the same service processing capability of the target redundant hardware equipment and the hardware equipment, so that the technical scheme is different from the degradation operation of the existing scheme, the service capability is not reduced, and the service processing efficiency is improved. Therefore, the problem that when a certain hardware device in a device cluster breaks down or hot plug occurs in an existing scheme, due to the fact that other hardware devices are in degraded operation, the service providing capacity is lowered can be solved, and the service capacity is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of device clusters, and in particular, to a method for implementing hot-plug fault isolation, a device for implementing hot-plug fault isolation, a computer-readable storage medium, and a system for implementing hot-plug fault isolation. Background Art

[0002] In modern data centers, critical services run on highly reliable hardware resources, and usually adopt dual-machine or n+1 redundant configurations. When a hardware resource fails, the system usually takes over the faulty resource through redundant hardware and then starts a fault recovery mechanism to restore the service. However, a large amount of service traffic conversion will occur between the faulty hardware resource and the redundant hardware resource during the service recovery process, resulting in network congestion and ultimately affecting the efficiency of service recovery.

[0003] Moreover, in the existing solutions, when the service of a faulty hardware device is transferred to other hardware devices, a downgrade process is required, and other hardware devices will not handle all the services.

[0004] Therefore, in the existing solutions, when a certain hardware device in the device cluster fails or a hot-plug occurs, since other hardware devices are operating in a degraded mode, the ability to provide services will be reduced. Summary of the Invention

[0005] This application provides a method for implementing hot-plug fault isolation, a device for implementing hot-plug fault isolation, a computer-readable storage medium, and a system for implementing hot-plug fault isolation, so as to at least solve the problem that in the existing solutions, when a certain hardware device in the device cluster fails or a hot-plug occurs, since other hardware devices are operating in a degraded mode, the ability to provide services will be reduced.

[0006] This application provides a method for implementing hot-plug fault isolation, including: when a hot-plug hardware device in the device cluster is powered on or off, monitoring and identifying a hot-plug event of the device cluster; determining a current fault type according to the hot-plug event, where the current fault type includes a hot-removal fault type and a hot-insertion fault type, and the hot-removal fault type includes a hardware device hot-removed without a fault and a hardware device hot-removed with a fault; determining whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; and in the case of determining that it is necessary to activate a redundant hardware device, switching the service process from the faulty hardware device to a target redundant hardware device to keep the service in the device cluster uninterrupted, where the service processing capabilities of the target redundant hardware device and the hardware device are the same.

[0007] The present application also provides a device for implementing hot-plug fault isolation, including: a monitoring module, configured to monitor and identify hot-plug events of a device cluster when a hot-plug hardware device of the device cluster is powered on or off; a first determination module, configured to determine a current fault type according to the hot-plug event, the current fault type including a hot-removal fault type and a hot-insertion fault type, and the hot-removal fault type including hot-removal of a hardware device without a fault and hot-removal of a hardware device with a fault; a second determination module, configured to determine whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; a first processing module, configured to switch a service process from a faulty hardware device to a target redundant hardware device in the case of determining that it is necessary to activate the redundant hardware device, so as to keep the service in the device cluster uninterrupted, and the service processing capabilities of the target redundant hardware device and the hardware device are the same.

[0008] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above methods for implementing hot-plug fault isolation are realized.

[0009] The present application also provides a system for implementing hot-plug fault isolation, including: a fault detection module, configured to monitor and identify hot-plug events of a device cluster when a hot-plug hardware device of the device cluster is powered on or off; a resource takeover module, configured to takeover the resources of a faulty hardware device when the fault detection module detects a hot-plug event, determine a current fault type according to the hot-plug event, the current fault type including a hot-removal fault type and a hot-insertion fault type, and the hot-removal fault type including hot-removal of a hardware device without a fault and hot-removal of a hardware device with a fault, and determine whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; a service transfer module, configured to switch a service process from a faulty hardware device to a target redundant hardware device in the case of determining that it is necessary to activate the redundant hardware device, so as to keep the service in the device cluster uninterrupted, and the service processing capabilities of the target redundant hardware device and the hardware device are the same.

[0010] Through this application, by monitoring and identifying hot-plug events of a device cluster, first determine whether redundant hardware devices need to be activated, and in the case of determining that redundant hardware devices need to be activated, switch the business process from the faulty hardware device to the target redundant hardware device to keep the business in the device cluster uninterrupted, so as to completely switch the business capabilities to the same hardware module, rather than other hardware modules sharing the functions of the faulty module, which is reflected in the same business processing capabilities of the target redundant hardware device and the hardware device, thus being a different technical solution from the degraded operation of the existing solution and not causing a decline in business capabilities. Therefore, it can solve the problem that in the existing solution, when a certain hardware device in the device cluster fails or a hot-plug situation occurs, since other hardware devices operate in a degraded mode, the ability to provide services will decline, and the business capabilities can be maintained. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0012] Figure 1 Schematic diagram of the principle of the hardware architecture of the storage cluster provided by the embodiment of the present application;

[0013] Figure 2 Schematic flowchart of a method for implementing hot-plug fault isolation provided by the embodiment of the present application;

[0014] Figure 3 Block diagram of the structure of a device for implementing hot-plug fault isolation provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.

[0016] It should be noted that in the description of this application, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0017] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.

[0018] In the existing solution, when a certain piece of hardware in the cluster fails (is unplugged), the other hardware operates in a degraded mode, which will lead to a decline in the ability to provide services. For example, if there are four network cards in a node and one network card is unplugged, the remaining three network cards take over the services of the unplugged network card. This method is actually not redundancy but degraded operation, and the performance and bandwidth cannot fully meet the service standards. However, this application can avoid degraded operation.

[0019] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the method for implementing hot-pluggable fault isolation depends, the specific application environment architecture or specific hardware architecture will be described herein.

[0020] As Figure 1 shown, the hardware architecture of the storage cluster includes multiple storage clusters, each storage cluster includes multiple nodes, each node includes redundant hardware and hot-pluggable hardware, and the control steps for the redundant hardware and hot-pluggable hardware involve fault detection, fault status management, and redundant switching control. The storage cluster interacts with user management, and user management refers to a visual graphical interface.

[0021] When using modules between nodes or between clusters as redundant modules, fault recovery must be achieved. In a hot-pluggable fault isolation system, once the hardware device of the faulty node is repaired or replaced and the system confirms that the device has returned to normal, it is necessary to quickly float back the IP address and service link (such as FC link) that were previously transferred to the redundant device to the node where the fault has been recovered. At the same time, if PCIe was once used as a temporary cluster interconnection means, then it is also necessary to switch the cluster interconnection method back to a more efficient or more business-demand-compliant RDMA over Converged Ethernet (RoCE) interconnection. The following are the glossary explanations and specific processes of this process:

[0022] IP address drift: In the scenarios of hot plugging or hardware failures, the IP address of the original node will be transferred to the same or similar network interfaces of the redundant nodes to maintain the continuity of network services. When the original node recovers from the failure, the process of "drifting back" the IP address from the redundant node to the original node is called IP address drift back to the original node.

[0023] FC link: That is, Fiber Channel link, which is a high-speed network technology mainly used in data storage systems such as SAN (Storage Area Network). In the hot plugging fault isolation scenario, the FC link may need to be switched back from the redundant node to the node that has recovered from the failure to restore the storage service of the original node.

[0024] PCIe (Peripheral Component Interconnect Express): A high-speed serial computer expansion bus standard used to connect high-performance subsystems to the computer motherboard. PCIe can be used as an interconnection method between cluster nodes, especially in scenarios that require quick response switching in a short time.

[0025] RoCE (RDMA over Converged Ethernet): The implementation of RDMA (Remote Direct Memory Access) over Ethernet. RoCE uses Ethernet hardware to provide remote direct memory access, thus enabling efficient data transfer between servers or cluster nodes. Compared with traditional Ethernet communication, RoCE has lower latency and higher throughput.

[0026] The SAN (Storage Area Network) configuration refers to the process of setting up and parameter adjustment of storage devices, servers, network connections, and corresponding management software in the SAN environment to ensure the efficiency, security, and reliability of data storage.

[0027] Specific process: First, the system needs to confirm that the faulty hardware device has returned to normal through hardware status feedback (such as PCIe link up signal, link quality check, etc.) or software detection mechanisms, including hardware function verification and performance testing. The fault management module notifies the current redundant takeover node (or cluster) that the hardware device of the faulty node has recovered and it is necessary to start preparing to drift back the IP address and service links. On the redundant takeover node, the resource takeover process is responsible for drifting back the IP address. The process of drifting back the IP address by the resource takeover process may include: releasing the binding of the old IP address on the redundant node; re-binding the IP address on the node that has recovered from the failure, including activation of network interfaces, update of network configuration, etc., and updating the routing table to ensure that data packets can be correctly forwarded to the node that has recovered from the failure.

[0028] FC Link Switching: The switching of the FC link involves re - establishing the FC connection of the original node. It is necessary to restart the FC controller or driver on the fault - recovered node, establish an FC connection with the storage system, which may include steps such as re - authentication and re - configuration of the FC address, update the SAN configuration to ensure that data streams can be accurately transmitted through the FC link.

[0029] Switching back from PCIe to RoCE: During the hot - plug fault isolation period, if PCIe is used as the temporary inter - cluster interconnection method, then after the fault recovery, it is necessary to re - configure the cluster interconnection mechanism. First, close or terminate the temporary communication link on PCIe, and then re - activate the RoCE communication protocol on the fault - recovered node. Update the network configuration between clusters, including RoCE addresses, port settings, etc., to ensure that cluster communication can seamlessly switch back to RoCE; perform network - layer fault detection and performance testing to ensure the stability of the RoCE link.

[0030] During the entire drifting - back and switching process, the system notification module updates the status in real - time and displays it to the user through a visualization interface to ensure that the user is aware of the current recovery progress and status. Finally, after confirming that all drifting - back and switching operations are completed, the service recovery module starts to officially take over the service. It is necessary to confirm that all network links and resources have been correctly configured, check the service status to ensure data consistency, and notify the upper - layer applications and services to inform that the service has been restored to the fault - recovered node.

[0031] Regarding the use of hardware redundancy processing: Build a redundant hardware topology for each hot - pluggable hardware on a single node. For example, install two network cards on each node, one network card is the service network card, and the other network card is the backup network card. When the service network card fails or is replaced and upgraded, use the backup network card to provide services. In scenarios with high cost requirements, the network card of another node in the cluster can be used as the backup network card. For example, if there are four network cards in a network card and two network ports can meet the service requirements, then the remaining two network ports can be used as the backup network cards for other nodes. When the network card of a certain node is fault - pulled out, this network card undertakes the service of the pulled - out network card.

[0032] Embodiments of this application provide a method for implementing hot - plug fault isolation, as Figure 2 shown. This method includes the following steps:

[0033] Step S201: When the hot - pluggable hardware device in the device cluster is powered on or off, monitor and identify the hot - plug event of the device cluster;

[0034] The device cluster can be a distributed storage cluster.

[0035] Among them, step S201 is carried out by the built-in hot-plug detection module. The hot-plug detection module can monitor the power-on and power-off states of hardware devices in the cluster in real time, capture hot-plug events in a timely manner, and provide immediate hardware status information for subsequent fault type judgment and redundant takeover decision-making. This monitoring mechanism ensures that the system can quickly respond to hot-plug operations and avoid the risk of service interruption.

[0036] Step S202: Determine the current fault type according to the hot-plug event. The current fault type includes hot-removal fault type and hot-insertion fault type. The hot-removal fault type includes hot-removal of hardware device without fault and hot-removal of hardware device with fault.

[0037] Among them, step S201 is carried out by the built-in redundant switching control module. The redundant switching control module can adopt different coping strategies according to different fault types and fault levels. For example, for hot-removal with fault, it may be necessary to immediately activate the redundant hardware device, while for hot-removal without fault, it may only be necessary to update the status and adjust resources, which improves the pertinence and efficiency of fault handling. Subdividing the hot-removal fault type into hot-removal of hardware device without fault and hot-removal with fault makes the fault management more refined, enables differential coping strategies to be adopted for different scenarios, and improves the flexibility and pertinence of fault handling.

[0038] Step S203: Determine whether it is necessary to activate the redundant hardware device according to the current fault type and the fault level of the current fault type.

[0039] Among them, step S203 involves the collaborative work of the fault detection module and the redundant switching control module. According to the severity of the fault and the running state of the cluster, it intelligently judges whether it is necessary to immediately activate the redundant hardware device to prevent service interruption or data loss. By activating the target redundant hardware device, the rapid switching of the business process is realized. Even when a hardware device fails or a hot-plug operation occurs, the service in the device cluster can be kept uninterrupted, ensuring the continuity of the service and the consistency of the data, and improving the reliability and availability of the cluster.

[0040] Step S204: In the case of determining that it is necessary to activate the redundant hardware device, switch the business process from the faulty hardware device to the target redundant hardware device to keep the service in the device cluster uninterrupted. The business processing capabilities of the target redundant hardware device and the hardware device are the same.

[0041] In the above steps, by monitoring and identifying the hot plug and unplug events of the device cluster, first determine whether to suddenly activate redundant hardware devices. And in the case of determining that redundant hardware devices need to be activated, switch the business process from the faulty hardware device to the target redundant hardware device to keep the business in the device cluster uninterrupted, so as to completely switch the business capabilities to the same hardware module, rather than having other hardware modules share the functions of the faulty module. This is reflected in the fact that the business processing capabilities of the target redundant hardware device and the hardware device are the same, so it is a different technical solution from the degraded operation of the existing solution and will not cause a decline in business capabilities. Therefore, it can solve the problem that in the existing solution, when a certain hardware device in the device cluster fails or a hot plug and unplug situation occurs, since other hardware devices operate in a degraded mode, the ability to provide services will decline, and the business capabilities can be maintained.

[0042] In addition, the present application has a complete fault handling link from hardware interruption to driver re-binding; it realizes zero perception of fault isolation, fault transfer, and service recovery; the hierarchical architecture design supports flexible adaptation to different operating systems and hardware platforms. It fundamentally solves other faults caused by hot plug and unplug faults, and isolates layer by layer from fault detection, fault status handling, redundant switching, and fault recovery, isolating both the hot plug and unplug faults and preventing the business from stopping or even crashing due to hot plug and unplug.

[0043] In an embodiment provided by the present application, switching the business process from the faulty hardware device to the target redundant hardware device in step S204 includes: determining the target redundant hardware device at least according to the current fault type and node load conditions, where the device cluster includes multiple nodes, and each node respectively includes at least one hot plug and unplug hardware device; using resource management and scheduling technologies to switch the business process from the faulty hardware device to the target redundant hardware device.

[0044] Specifically, by analyzing the fault type and node load conditions, the most suitable target redundant hardware device is intelligently selected to ensure the rationality of resource allocation, avoid resource waste and performance bottlenecks; the multi-node configuration of the device cluster can ensure that when a fault occurs, the cluster can achieve rapid switching of business processes and load balancing through cross-node resource scheduling, avoid overloading of a single node, and enhance the overall stability of the cluster; resource management and scheduling technologies are used to switch the business process from the faulty hardware device to the target redundant hardware device, ensuring that the business continuity is not affected by hot-plug events. At the same time, through the data consistency check mechanism, data loss or damage during the business switching process is avoided; the rapid activation of the target redundant hardware device and the seamless switching of the business process effectively isolate the impact of the faulty hardware device. At the same time, through the fault recovery mechanism, once the faulty hardware device is repaired, it can quickly resume its normal operation and reduce the time of business interruption; through the multi-node selection strategy and resource takeover scheduling technology, not only the efficiency of fault handling is improved, but also the overall processing capacity and efficiency of the device cluster are enhanced through optimized resource allocation and load balancing, and the operation and maintenance costs are reduced.

[0045] In an embodiment provided by the present application, using resource management and scheduling technologies to switch the business process from the faulty hardware device to the target redundant hardware device includes: after activating the redundant hardware device, updating the IO path of the faulty hardware device to the IO path of the target redundant hardware device; or, updating the resource mapping table to map the resources of the faulty hardware device to the target redundant hardware device, where the resources of the faulty hardware device include PCI resources, memory mapping, interrupt vectors, and DMA channels, so that the business process transitions from the faulty hardware device to the target redundant hardware device.

[0046] Specifically, by updating the IO path configuration, all IO operations originally directed to the faulty hardware device are automatically redirected to the target redundant hardware device, ensuring the continuity of the business process during the hardware switchover and avoiding business interruptions caused by IO path changes; update the resource mapping table to immediately take over and map key resources such as the PCI resources, memory mapping, interrupt vectors, and DMA channels of the faulty hardware to the target redundant hardware device. This not only accelerates the resource takeover process but also ensures the coherence of data transmission and the correctness of business logic, reducing the risks of data loss and business latency; the update of the resource mapping enables the system resources originally allocated to the faulty hardware to be fully utilized by the target redundant hardware, avoiding resource idleness or waste, and at the same time ensuring the efficient reallocation of resources to support the uninterrupted operation of the business process; by updating the resource mapping table and IO path, data consistency is ensured during the hardware device switchover, avoiding problems such as data transmission interruptions or data loss. In addition, the immediate transition of the business process guarantees service continuity, allowing users to continuously access data and perform business operations without being affected by hardware failures; once the faulty hardware device is repaired, through reverse resource mapping updates and IO path switching, the normal service of the faulty hardware device can be quickly restored, simplifying the recovery process of the faulty hardware device and improving the efficiency of hardware fault repair; the hardware redundancy switch achieved through resource mapping updates and IO path switching improves the system's ability to respond to and recover from hardware failures, enhancing the overall stability and reliability of the system.

[0047] The IO path refers to the path for data transmission between the operating system and the hardware device. In a computer system, all input and output operations must be completed through specific IO paths, which usually include device drivers, device files, hardware interfaces, etc. When a device is replaced or fails, updating the IO path is a key step to ensure that the new or redundant device can take over the tasks of the original device, including remapping the device file and updating the driver, etc., to ensure that the upper-layer software can correctly interact with the new hardware device.

[0048] PCI resources refer to the resources used by devices that communicate with the CPU through the PCI (Peripheral Component Interconnect) bus. The PCI bus is a common standard interface in computer hardware that allows the CPU to perform high-speed data exchange with various peripheral devices. PCI resources include address space, interrupt request lines, clock signals, etc., which are allocated to devices by the operating system during device initialization. When a device is hot-plugged or fault-isolated, these PCI resources need to be reallocated and configured to ensure the normal operation of the device.

[0049] Memory mapping refers to mapping the physical address space of a hardware device into the system's virtual memory address space, enabling the device to perform data read and write operations as if accessing memory. For hot-pluggable devices, their memory mapping information may change with device plugging and unplugging. Therefore, when a device fails or is replaced, the memory mapping needs to be updated to map the memory area of the faulty device to the redundant device to ensure data continuity and consistency.

[0050] An interrupt vector is the address used by the operating system to point to a specific interrupt handler when processing an interrupt request. When a device undergoes hot-plugging or fault isolation, the interrupt vector needs to be updated to ensure that the correct interrupt handler can respond to interrupt requests generated by the redundant device, thus avoiding interrupt handling chaos and data loss.

[0051] A DMA channel is a direct data transfer channel between a device and memory. It allows data to be transferred directly between the device and memory without CPU intervention. The efficient utilization of DMA channels is particularly important for I / O-intensive operations (such as disk read and write, network data transfer). Reconfiguring the DMA channel to the redundant device during hot-plugging or fault isolation is a key step in ensuring data transfer efficiency and business continuity.

[0052] In an embodiment provided by the present application, determining a target redundant hardware device at least according to the current fault type and the node load condition includes: parsing fault information, where the fault information includes the current fault type, the ID of the faulty hardware device, the faulty IO path, and the address, and identifying the faulty hardware device and the hardware device status of the faulty hardware device according to the fault information, where the hardware device status includes the running status and the health status of the faulty hardware device; evaluating the node load condition in real time, where the node load condition includes CPU utilization rate, memory occupancy, and disk I / O performance, and determining the load data of the current node and other nodes except the current node according to the node load condition to analyze the real-time processing capacity of the nodes; inputting the fault information, the node load condition, and the load difference between the current node and other nodes except the current node into a neural network model as the input of the neural network model; processing the fault information, the node load condition, and the load difference between the current node and other nodes by using the neural network model to obtain the output of the neural network model, where the output of the neural network model includes a target confidence level, and the target confidence level is the confidence level that the redundant hardware device has the ability to take over the services of the faulty hardware device; determining that the redundant hardware device has the ability to take over the services of the faulty hardware device when the target confidence level is greater than or equal to the confidence level threshold; determining that the redundant hardware device does not have the ability to take over the services of the faulty hardware device when the target confidence level is less than the confidence level threshold; checking the resource mapping condition of the redundant hardware device to ensure that the resources of the redundant hardware device can match the resources of the faulty hardware device when it is determined that the redundant hardware device has the ability to take over the services of the faulty hardware device, where the resource mapping condition includes the hardware device driver program, the hardware device ID, and the IO address space of the hardware device; taking over all the resources of the faulty hardware device, where all the resources of the faulty hardware device include the hardware device driver resources, the hardware device resources, and the hardware device IO resources, and determining that the redundant hardware device is the target redundant hardware device, and visually displaying the hardware device status of the faulty hardware device.

[0053] Specifically, disk I / O performance refers to the efficiency and ability of disk input / output (I / O) operations, and is a key indicator for measuring disk read / write speed and data processing capacity. The confidence level threshold can be 85%. For a specific usage scenario of determining a target redundant hardware device at least according to the current fault type and the node load condition: In a data center, a highly available storage cluster is operated, which is composed of multiple computing nodes, and each node is equipped with at least one redundant hardware device. Suppose in a hot plug event, the system detects that a hot plug hard disk device of node A fails, and immediate action needs to be taken to avoid data loss and service interruption.

[0054] Fault Information Analysis: Fault Type: Hot Plug Fault Type, specifically a hard disk fault. Faulty Hardware Device ID: The unique identifier of the hard disk. Faulty IO Path and Address: Records the IO access path of the hard disk and the current mount point to locate the specific scope affected by the fault. Real-time CPU Utilization: The CPU utilization rate of Node A is 75%, and the CPU utilization rates of other nodes such as B and C are 50% and 60% respectively. Memory Occupancy: The memory occupancy of Node A is 80%, and those of Nodes B and C are 60% and 70% respectively. Disk I / O Performance: The disk I / O performance indicator of Node A was at a normal level before the fault, and the performance dropped sharply after the fault. Input the fault information (hard disk fault, ID, IO path), node load conditions (CPU utilization, memory occupancy, disk I / O), and load differences into the neural network model; the model analyzes the possibility of whether the faulty hardware device can be taken over by redundant hardware devices through pre-trained weights and outputs the target confidence level. If the target confidence level is higher than the confidence level threshold (e.g., 85%), it indicates that there is a high possibility that the redundant hardware device (such as a redundant hard disk of Node B) can take over the faulty hard disk. Conduct a resource mapping check on the redundant hardware device to confirm that its driver is compatible, the device ID is the same as the faulty device ID, and the IO address space matches; after confirming the resource match, immediately activate the redundant hard disk to take over all resources of the faulty hard disk, including its driver, device resources, and IO resources; the entire process is visually displayed through the user management interface, and users can view the status of the faulty hard disk in real time (such as fault identification, resource takeover status, redundant hardware activation progress, etc.) and the real-time load conditions of the cluster to understand the overall health status and fault isolation progress.

[0055] Beneficial effects brought by a specific usage scenario of determining a target redundant hardware device at least according to the current fault type and node load conditions: The neural network model can accurately evaluate the probability of the redundant hardware device taking over the faulty hardware device, improving the intelligence and accuracy of fault handling; When a fault occurs, it can quickly identify and activate the matching redundant hardware device, greatly shortening the service recovery time and avoiding data loss and service interruption; By evaluating the node load conditions in real time, it intelligently selects the node with a lower load as the source of the redundant device, avoiding overloading of a single node and ensuring balanced allocation and efficient utilization of resources; The output of the neural network model is used as the basis for the redundant takeover decision, realizing the automation of the fault detection and redundant activation processes, reducing the workload of the operation and maintenance personnel, and improving the efficiency and accuracy of fault handling; The visual monitoring provided by the user management interface enables the operation and maintenance personnel to intuitively understand the fault status and resource takeover progress, which is beneficial to fault tracking and subsequent operation and maintenance operations, enhancing the transparency and manageability of the system; Before the redundant hardware device is activated, through resource mapping inspection, it is ensured that the resources of the faulty hardware device can be seamlessly docked, avoiding the risk of secondary faults caused by resource mismatch and improving the stability of the system; By taking over the resources of the faulty hardware device, including drivers, device resources, and IO resources, it ensures data consistency and integrity, avoiding data confusion or inconsistency during the fault switchover process and enhancing the data processing ability of the system.

[0056] In an embodiment provided by the present application, the method further includes: When each hardware device is configured with a redundant hardware device, after the faulty hardware device is restored or replaced, it is determined that no service recovery operation is required; If there is a hardware device without a configured redundant hardware device, after the faulty hardware device is restored or replaced, the service is restored to the original hardware device.

[0057] Specifically, when all hardware devices are configured with redundant hardware devices, even when a fault occurs, the system can immediately take over the service through the redundant device without waiting for the recovery or replacement of the faulty device. This means that the system can tolerate the impact of hardware faults to the greatest extent, improving the continuity and reliability of the service, avoiding the risk of additional delays or interruptions caused by service recovery operations, and ensuring the smoothness of the business process and the consistency of user operations; For the case where there is no redundant hardware device configured, the strategy of directly performing service recovery after the faulty hardware device is restored or replaced is clarified, simplifying the fault recovery process, reducing the decision-making burden of the operation and maintenance personnel in the fault recovery operation, and improving the efficiency and success rate of fault recovery; This strategy dynamically adapts to the redundant configuration status of the hardware devices in the device cluster. When the hardware device and its redundant device are both in normal working state, no service recovery operation is required, saving system resources and time overhead. Otherwise, necessary service recovery is performed, reflecting the flexibility and intelligence of the strategy.

[0058] In an embodiment provided by the present application, switching the service process from a faulty hardware device to a target redundant hardware device includes: receiving and responding to a user preset operation, obtaining the hardware device information of the hardware device to be unplugged. The user preset operation indicates that the user manually marks the hardware device to be replaced through a graphical user interface or a command line method and triggers a hot plug operation request. The hardware device information includes the resources, current status, and relevance to the service of the hardware device; starting a resource takeover process to switch the service process from the faulty hardware device to the target redundant hardware device by using resource management and scheduling technologies, and generating a status update notification to visually display the progress and status of replacing the hardware device.

[0059] Specifically, the user can directly participate in the scheduling of the hot plug operation through a graphical user interface (GUI) or a command line interface (CLI) to manually select and mark the hardware device to be replaced. This active user participation enhances the flexibility of system management, allowing the user to independently determine the timing and method of hardware replacement according to business requirements and maintenance plans; before the hot plug operation, the system can accurately obtain the detailed resource information of the faulty hardware device, including its address on the PCI bus, memory mapping, interrupt vector, etc., as well as the current status of the device and its relevance to the service. Through this information, the resource takeover process can precisely switch the service from the faulty hardware device to the target redundant hardware device, ensuring business continuity while avoiding resource conflicts and data chaos; the status update notification generated by the system can visually display the progress and status of replacing the hardware device in real time. For operation and maintenance personnel, this means being able to immediately understand the progress of the hot plug operation, including the resource takeover status, the activation degree of the redundant device, the completion of the service switch, etc., so as to better monitor and manage the hot plug process and ensure the smooth progress of the operation.

[0060] In an embodiment provided by the present application, the method further includes: after generating the status update notification, performing an uninstall operation on the hardware device to uninstall the hardware device from the system. The uninstall operation includes releasing the hardware device resources and unregistering the hardware device driver.

[0061] Specifically, by uninstalling the hardware device driver, it is ensured that there are no longer driver instances for the faulty hardware in the system, which is an important guarantee for system security and stability. Uninstalling the driver can also avoid problems after system restart that may be caused by residual drivers, such as duplicate driver loading and device recognition errors; the resource release and driver uninstallation of the faulty hardware device can reduce the memory occupancy and CPU burden of the system, improving the overall performance. Especially for resource-intensive application scenarios, such as large-scale data processing and high-performance computing, this operation is particularly important and helps the system maintain an efficient operating state; the uninstallation operation eliminates the traces of the faulty hardware device in the system configuration, simplifies the system configuration, and improves the convenience of management. This is especially true for complex cluster environments and multi-device management systems, which can avoid configuration confusion, reduce the workload of operation and maintenance, and improve the efficiency of configuration management.

[0062] In the scenario of hardware upgrade and iteration, the scenarios of non-violent hot plugging and violent hot plugging can be realized. The scenario of violent hot plugging is the processing logic of the above method. It is also possible to actively uninstall the module to be upgraded and specify which hardware is needed for takeover, which needs to be implemented in user management. A user management interface can be implemented, or only terminal commands can be implemented for operation. When a hardware needs to be replaced, mark the hardware to be replaced as about to be hot unplugged from the user management interface. At this time, obtain the resources of the module to be unplugged, and then perform redundant switching, and then uninstall the hardware from the system. After the user completes the hardware replacement, load the new hardware into the system where the device cluster is located.

[0063] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0064] This application also provides a specific implementation manner of a method for implementing hot plugging fault isolation:

[0065] Fault status management: The fault status management module is responsible for monitoring the status of all hardware devices in the system, especially paying attention to the changes in hardware devices caused by hot plugging operations. After the kernel senses the loss of a hardware device and triggers a hardware device interrupt, the fault status management module starts and runs, determines whether the change in the hardware device is a hot unplug or a hot plug, classifies and processes hardware device faults and hot plugging events, and ensures that faults are promptly identified and enter the isolation process.

[0066] Redundancy Switching Control: The redundancy switching control module is the key to hot-plug fault isolation. It is responsible for transferring services from faulty hardware devices to redundant hardware devices. When the fault status management module confirms that a hot-plug event is a fault, the redundancy switching control module is called to take over the resources applied for by the faulty hardware device in the kernel and user states, including bus resources and memory resources, to ensure service continuity.

[0067] Service Transfer: For in-node service transfer, it directly switches to the redundant hardware device. For cross-node or cross-cluster service takeover, the most suitable redundant node is selected through inter-cluster and inter-node communication, which may be accompanied by a short delay.

[0068] System Notification: The system notification module, as an optional function, is used to display the service status of the current hardware device to the user, enhancing system transparency. After the redundancy switching control is completed, the system notification module updates the status information and presents it to the user, informing the progress and status of the hardware device replacement to improve the user experience.

[0069] Recovery Processing: The recovery processing module is responsible for reactivating the original faulty hardware device and restoring the service process after the hardware device fault is recovered. In the polling mechanism where the faulty node and the redundant node run simultaneously, the faulty node polls the isolation module status, and the redundant node polls the fault recovery instructions from the cluster. When the faulty node receives the recovery instruction and confirms that the fault has been eliminated, the redundant node completes the last I / O processing and exits, and the faulty node takes over the service again to ensure the continuity and consistency of data processing.

[0070] User Management and Pre-Notification: The user manually marks the hardware device that needs to be replaced through the graphical user interface or command line, triggering the preset hot-plug operation. After the user's preset operation, the resource takeover process starts, intelligently schedules resources, switches the service from the marked hardware device to the target redundant device, and the system notification module updates and displays the progress and status of the hardware device replacement in real time to ensure the user's real-time monitoring of the replacement process.

[0071] Fault Detection and Judgment: The fault detection module, as part of the kernel interrupt processing, is used to identify hot-plug events. After the hardware device loss triggers an interrupt, the fault detection module in the bottom half of the kernel runs to determine whether it is a hot-plug or hot-insert, and treats the hardware device fault as a hot-plug scenario; by registering the fault detection module as a hook function of the hardware device management subsystems such as PCI, it ensures an immediate response when a hardware device interrupt occurs and quickly starts the fault status management.

[0072] Resource takeover and switching: The fault management module takes over all resources of the faulty hardware device, such as IP, routing, unsent data, etc., and switches the service from the faulty hardware device to the redundant hardware device. Node-level takeover uses inter-node communication to select a suitable redundant hardware device, while cluster-level takeover completes the rapid takeover of resources through inter-cluster communication. After the takeover is completed, the system notification module updates the status information to reflect the resource takeover situation for easy user monitoring.

[0073] Fault recovery and service recovery: The faulty node periodically polls the status register to confirm whether the fault has been eliminated, and the redundant node polls the recovery instructions from the cluster. After the fault recovery module confirms that the hardware device of the faulty node has recovered, it notifies the redundant node to stop processing the service, and the faulty node resumes the service processing task. During the service takeover and switching process, data caching technology is used to ensure data consistency and avoid data loss or service interruption during the switching process.

[0074] An embodiment of the present application also provides a device for implementing hot-pluggable fault isolation, such as Figure 3 shown. The device includes:

[0075] A monitoring module 31, configured to monitor and identify hot-plug events of the device cluster when a hot-pluggable hardware device in the device cluster is powered on or off; a first determination module 32, configured to determine the current fault type according to the hot-plug event, where the current fault type includes a hot-removal fault type and a hot-insertion fault type, and the hot-removal fault type includes hot-removal of a hardware device without a fault and hot-removal of a hardware device with a fault; a second determination module 33, configured to determine whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; a first processing module 34, configured to switch the service process from the faulty hardware device to the target redundant hardware device in the case of determining that it is necessary to activate the redundant hardware device, so as to keep the service in the device cluster uninterrupted, and the service processing capabilities of the target redundant hardware device and the hardware device are the same.

[0076] In an embodiment provided by the present application, the first processing module includes a first determination sub-module and a first processing sub-module. The first determination sub-module is configured to determine the target redundant hardware device at least according to the current fault type and the node load situation, where the device cluster includes multiple nodes, and each node includes at least one hot-pluggable hardware device; the first processing sub-module is configured to switch the service process from the faulty hardware device to the target redundant hardware device by using resource management and scheduling technologies.

[0077] In an embodiment provided by the present application, the first processing sub-module includes: a second processing sub-module, configured to update the IO path of the faulty hardware device to the IO path of the target redundant hardware device after activating the redundant hardware device; or update the resource mapping table to map the resources of the faulty hardware device to the target redundant hardware device, where the resources of the faulty hardware device include PCI resources, memory mapping, interrupt vectors, and DMA channels, so that the service process transitions from the faulty hardware device to the target redundant hardware device.

[0078] In an embodiment provided by the present application, the first determination sub-module includes: a third processing sub-module, a fourth processing sub-module, a fifth processing sub-module, a sixth processing sub-module, a second determination sub-module, a seventh processing sub-module, and an eighth processing sub-module. The third processing sub-module is configured to parse the fault information, where the fault information includes the current fault type, the ID of the faulty hardware device, the faulty IO path, and the address, and identify the faulty hardware device and the hardware device status of the faulty hardware device according to the fault information. The hardware device status includes the operating status and health status of the faulty hardware device. The fourth processing sub-module is configured to evaluate the node load situation in real time. The node load situation includes CPU utilization, memory occupancy, and disk I / O performance, and determine the load data of the current node and other nodes except the current node according to the node load situation to analyze the real-time processing ability of the nodes. The fifth processing sub-module is configured to input the fault information, the node load situation, and the load difference between the current node and other nodes except the current node into a neural network model as the input of the neural network model. The sixth processing sub-module is configured to process the fault information, the node load situation, and the load difference between the current node and other nodes by using the neural network model to obtain the output of the neural network model. The output of the neural network model includes a target confidence level, where the target confidence level is the confidence level that the redundant hardware device has the ability to take over the services of the faulty hardware device. The second determination sub-module is configured to determine that the redundant hardware device has the ability to take over the services of the faulty hardware device when the target confidence level is greater than or equal to the confidence level threshold; and determine that the redundant hardware device does not have the ability to take over the services of the faulty hardware device when the target confidence level is less than the confidence level threshold. The seventh processing sub-module is configured to check the resource mapping situation of the redundant hardware device to ensure that the resources of the redundant hardware device can match the resources of the faulty hardware device when it is determined that the redundant hardware device has the ability to take over the services of the faulty hardware device. The resource mapping situation includes the hardware device driver, the hardware device ID, and the IO address space of the hardware device. The eighth processing sub-module is configured to take over all the resources of the faulty hardware device. All the resources of the faulty hardware device include hardware device driver resources, hardware device resources, and hardware device IO resources, and determine that the redundant hardware device is the target redundant hardware device, and visually display the hardware device status of the faulty hardware device.

[0079] In an embodiment provided by the present application, the device further includes a third determination module and a second processing module. The third determination module is configured to determine that no service recovery operation is required after a faulty hardware device is recovered or replaced when each hardware device is configured with a redundant hardware device; the second processing module is configured to, if there is a hardware device without a redundant hardware device configured, recover the service to the original hardware device after the faulty hardware device is recovered or replaced.

[0080] In an embodiment provided by the present application, the first processing module includes: a ninth processing sub-module and a tenth processing sub-module. The ninth processing sub-module is configured to receive and respond to a user preset operation, and obtain the hardware device information of the hardware device to be unplugged. The user preset operation indicates that the user manually marks the hardware device to be replaced through a graphical user interface or a command line method, and triggers a hot unplug operation request. The hardware device information includes the resources, current status, and relevance to the service of the hardware device; the tenth processing sub-module is configured to start a resource takeover process to switch the service process from the faulty hardware device to the target redundant hardware device by using resource management and scheduling technologies, and generate a status update notification to visually display the progress and status of replacing the hardware device.

[0081] In an embodiment provided by the present application, the device further includes: a third processing module, configured to perform an uninstall operation on the hardware device after generating the status update notification to uninstall the hardware device from the system. The uninstall operation includes releasing the hardware device resources and unregistering the hardware device driver.

[0082] For the description of the features in the corresponding embodiment of the device for implementing hot plug fault isolation, reference may be made to the relevant description in the corresponding embodiment of the method for implementing hot plug fault isolation, which will not be elaborated here one by one.

[0083] An embodiment of the present application further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above XX method embodiments.

[0084] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any of the above XX method embodiments when running.

[0085] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disc, and other media that can store a computer program.

[0086] An embodiment of the present application further provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments for realizing hot-plug fault isolation.

[0087] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the method embodiments for realizing hot-plug fault isolation.

[0088] An embodiment of the present application further provides a system for realizing hot-plug fault isolation, including: a fault detection module, configured to monitor and identify hot-plug events of a device cluster when a hot-plug hardware device in the device cluster is powered on or off; a resource takeover module, when the fault detection module detects a hot-plug event, configured to take over the resources of the faulty hardware device, determine the current fault type according to the hot-plug event, the current fault type includes a hot-removal fault type and a hot-insertion fault type, the hot-removal fault type includes hot-removal of a hardware device without a fault and hot-removal of a hardware device with a fault, and determine whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; a service transfer module, configured to, when it is determined that it is necessary to activate a redundant hardware device, switch the service process from the faulty hardware device to the target redundant hardware device to keep the service in the device cluster uninterrupted, and the service processing capabilities of the target redundant hardware device and the hardware device are the same.

[0089] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0090] The above has introduced in detail a method for implementing hot-plug fault isolation, a device for implementing hot-plug fault isolation, a computer-readable storage medium, and a system for implementing hot-plug fault isolation provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can still be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A method for implementing hot plug fault isolation, characterized in that, Including: When a hot-pluggable hardware device in a device cluster is powered on or off, monitor and identify hot-plug events of the device cluster; Determine a current fault type according to the hot-plug event, where the current fault type includes a hot-removal fault type and a hot-insertion fault type, and the hot-removal fault type includes hot-removal of a hardware device without a fault and hot-removal of a hardware device with a fault; Determine whether it is necessary to activate a redundant hardware device according to the current fault type and the fault level of the current fault type; In the case of determining that it is necessary to activate the redundant hardware device, switch the service process from the faulty hardware device to the target redundant hardware device to keep the service in the device cluster uninterrupted, and the service processing capabilities of the target redundant hardware device and the hardware device are the same.

2. The method for realizing hot plugging fault isolation according to claim 1, characterized in that Switching the service process from the faulty hardware device to the target redundant hardware device includes: Determine the target redundant hardware device at least according to the current fault type and the node load condition, where the device cluster includes multiple nodes, and each of the nodes respectively includes at least one of the hot-pluggable hardware devices; Use resource management and scheduling technologies to switch the service process from the faulty hardware device to the target redundant hardware device.

3. The method for implementing hot plug fault isolation according to claim 2, wherein Using resource management and scheduling technologies to switch the service process from the faulty hardware device to the target redundant hardware device includes: After activating the redundant hardware device, update the IO path of the faulty hardware device to the IO path of the target redundant hardware device; Alternatively, update the resource mapping table to map the resources of the faulty hardware device to the target redundant hardware device, where the resources of the faulty hardware device include PCI resources, memory mapping, interrupt vectors, and DMA channels, so that the service process transitions from the faulty hardware device to the target redundant hardware device.

4. The method for implementing hot plug fault isolation according to claim 2, wherein Determining the target redundant hardware device at least according to the current fault type and the node load condition includes: Analyze fault information, where the fault information includes the current fault type, the ID of the faulty hardware device, the faulty IO path, and the address, and identify the faulty hardware device and the hardware device state of the faulty hardware device according to the fault information, and the hardware device state includes the running state and the health state of the faulty hardware device; Evaluate the node load condition in real time, where the node load condition includes CPU utilization rate, memory occupancy, and disk I / O performance, and determine the load data of the current node and other nodes except the current node according to the node load condition to analyze the real-time processing capabilities of the nodes; Input the fault information, the node load condition, and the load difference between the current node and other nodes except the current node into a neural network model as the input of the neural network model. The neural network model processes the fault information, the node load condition, and the load difference between the current node and other nodes to obtain the output of the neural network model. The output of the neural network model includes a target confidence level, and the target confidence level is the confidence level that the redundant hardware device has the business capability to take over the failed hardware device; When the target confidence level is greater than or equal to the confidence level threshold, it is determined that the redundant hardware device has the business capability to take over the failed hardware device; when the target confidence level is less than the confidence level threshold, it is determined that the redundant hardware device does not have the business capability to take over the failed hardware device; When it is determined that the redundant hardware device has the business capability to take over the failed hardware device, check the resource mapping situation of the redundant hardware device to ensure that the resources of the redundant hardware device can match the resources of the failed hardware device. The resource mapping situation includes the hardware device driver program, the hardware device ID, and the IO address space of the hardware device; Take over all the resources of the failed hardware device. The all resources of the failed hardware device include hardware device driver resources, hardware device resources, and hardware device IO resources, and determine that the redundant hardware device is the target redundant hardware device, and visually display the hardware device status of the failed hardware device.

5. The method for implementing hot plug fault isolation according to claim 1, characterized in that, The method further includes: When each of the hardware devices is configured with a redundant hardware device, after the failed hardware device is restored or replaced, it is determined that no business recovery operation is required; If there is a hardware device without the redundant hardware device configured, after the failed hardware device is restored or replaced, the service is restored to the original hardware device.

6. The method for implementing hot plug fault isolation according to claim 1, characterized in that, Switching the business process from the failed hardware device to the target redundant hardware device includes: Receiving and responding to a user preset operation to obtain the hardware device information of the hardware device to be unplugged. The user preset operation indicates that the user manually marks the hardware device to be replaced through a graphical user interface or a command line method and triggers a hot plug operation request. The hardware device information includes the resources, the current state, and the relevance to the service of the hardware device; Starting a resource takeover process to switch the business process from the failed hardware device to the target redundant hardware device by using resource management and scheduling technologies, and generating a status update notification to visually display the progress and status of replacing the hardware device.

7. The method for implementing hot plug fault isolation according to claim 6, wherein The method further includes: After generating the status update notification, perform an uninstall operation on the hardware device to uninstall the hardware device from the system. The uninstall operation includes releasing the hardware device resources and unregistering the hardware device driver.

8. A device for realizing hot plugging fault isolation, characterized in that Including: A monitoring module for monitoring and identifying hot plug events of the device cluster when the hot plug hardware device of the device cluster is powered on or off; A first determination module, configured to determine a current fault type according to the hot plug and unplug event, where the current fault type includes a hot unplug fault type and a hot plug fault type, and the hot unplug fault type includes hot unplug without hardware device fault and hot unplug with hardware device fault; A second determination module, configured to determine whether to activate redundant hardware devices according to the current fault type and the fault level of the current fault type; A first processing module, configured to, when it is determined that the redundant hardware devices need to be activated, switch the service process from the faulty hardware device to the target redundant hardware device to keep the service in the device cluster uninterrupted, where the target redundant hardware device and the hardware device have the same service processing capabilities; 9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, where when the computer program is executed by a processor, the steps of the method for implementing hot plug and unplug fault isolation as described in any one of claims 1 to 7 are implemented.

10. A system for implementing hot-swap fault isolation, characterized in that, Including: A fault detection module, configured to monitor and identify the hot plug and unplug events of the device cluster when the hot plug and unplug hardware device of the device cluster is powered on or off; A resource takeover module, when the fault detection module detects the hot plug and unplug event, is configured to take over the resources of the faulty hardware device, determine the current fault type according to the hot plug and unplug event, where the current fault type includes a hot unplug fault type and a hot plug fault type, and the hot unplug fault type includes hot unplug without hardware device fault and hot unplug with hardware device fault, and determine whether to activate redundant hardware devices according to the current fault type and the fault level of the current fault type; A service transfer module, configured to, when it is determined that the redundant hardware devices need to be activated, switch the service process from the faulty hardware device to the target redundant hardware device to keep the service in the device cluster uninterrupted, where the target redundant hardware device and the hardware device have the same service processing capabilities.