Device Online Maintenance Method, Device, Equipment and Computer Readable Storage Medium

The online maintenance method for electrical equipment ensures continuous system operation by identifying device types, performing fault tolerance adjustments, and reintroducing devices post-repair, thus reducing maintenance costs and downtime.

CN114358340BActive Publication Date: 2025-07-15SUZHOU INOVANCE CONTROL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111683411.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-31
Publication Date
2025-07-15
Estimated Expiration
2041-12-31

AI Technical Summary

Technical Problem

In electrical equipment running parallel, when equipment failure causes shutdown or replacement, the entire system needs to be powered off, resulting in high maintenance costs.

Method used

When a device failure is detected, determine the device type, and make fault-tolerant adjustments, remove the faulty equipment, and maintain the system operation until the maintenance is completed, confirming that the equipment to be put into operation is normal through the maintenance process before connecting to the system.

Benefits of technology

Without affecting the normal operation of the system, repair or replacement of faulty equipment is realized, reducing maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114358340B_ABST
    Figure CN114358340B_ABST
Patent Text Reader

Abstract

The present invention discloses an on-line maintenance method, device, equipment and computer-readable storage medium for equipment. When a device failure occurs in a parallel system, the present invention first determines the device type, and then performs corresponding fault-tolerant operations on the system for different device types, so that the system can still maintain a normal operating state after the failure occurs; by timely removing the faulty device from the system, the faulty device can be repaired or replaced without affecting the normal operation of the system; after completing the repair or replacement and confirming that the device to be put into operation passes the inspection process including a buffer detection process and / or an operation detection process preset, the device to be put into operation is reconnected to the system, so that during the entire process from the occurrence of the failure to the troubleshooting and reconnection of the devices in the parallel system, the normal operation of the system will not be disturbed, that is, the system can always maintain the power-on state for normal production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electrical equipment control, and particularly to an on-line equipment maintenance method, device, equipment and computer-readable storage medium. Background Art

[0002] In the current field of electrical equipment control, electrical equipment may be applied to parallel operation applications. In this case, if a device in the parallel operation fails and shuts down, and machine replacement or repair is required, the entire parallel operation system needs to be powered off until the maintenance staff repairs or replaces the faulty device and then restores the parallel operation state of the entire system. Since normal production cannot be carried out during fault troubleshooting, relatively high maintenance costs are incurred. Summary of the Invention

[0003] The main object of the present invention is to provide an on-line equipment maintenance method, device, equipment and computer-readable storage medium, aiming to solve the technical problem of relatively high maintenance costs in the existing fault troubleshooting method for parallel operation.

[0004] To achieve the above object, the present invention provides an on-line equipment maintenance method, which is applied to a parallel system. The method includes:

[0005] When a device failure is detected in the parallel system, determine the device type of the faulty device;

[0006] According to the device type, perform corresponding fault tolerance adjustment on the parallel system to maintain the operation of the parallel system, and remove the faulty device from the parallel system;

[0007] Until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation passes a preset inspection process, connect the device to be put into operation to the parallel system, where the inspection process includes a buffer detection process and / or an operation detection process.

[0008] Optionally, the device type includes a host type and a slave type. The step of performing corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system and removing the faulty device from the parallel system includes:

[0009] If the device type of the faulty device is the host type, change the device type of the faulty device to the slave type, and select a new host from several slaves in the parallel system except the faulty device;

[0010] Maintain the operation of the parallel system based on the new host, and remove the faulty device from the parallel system.

[0011] Optionally, the step of selecting a new host from several slave machines in the parallel operation system except the faulty device includes:

[0012] Obtain the status information and number information of the several slave machines, and select the new host according to the status information and number information.

[0013] Optionally, the step of, until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation after maintenance passes a preset inspection process, then connecting the device to be put into operation to the parallel operation system includes:

[0014] Until a device new access instruction is received, determine the device to be put into operation after repair or replacement;

[0015] Execute the inspection process on the device to be put into operation, and connect the device to be put into operation to the parallel operation system after the inspection passes, wherein the DC power switch of the device to be put into operation is not closed during the inspection process.

[0016] Optionally, the step of executing the inspection process on the device to be put into operation and connecting the device to be put into operation to the parallel operation system after the inspection passes includes:

[0017] Judge whether the buffer function of the device to be put into operation is normal;

[0018] If so, judge whether the device to be put into operation can operate normally;

[0019] If so, determine that the device to be put into operation passes the inspection;

[0020] Close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel operation system.

[0021] Optionally, the step of judging whether the buffer function of the device to be put into operation is normal includes:

[0022] Determine that both the AC power switch and the DC power switch of the device to be put into operation are opened;

[0023] Close the AC power switch, and judge whether the actual DC side voltage value of the device to be put into operation is consistent with the standard DC side voltage value;

[0024] If so, determine that the buffer function of the device to be put into operation is normal;

[0025] If not, determine that the buffer function of the device to be put into operation is abnormal.

[0026] Optionally, the device type includes a host type and a slave type. The step of performing corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system and removing the faulty device from the parallel system includes:

[0027] If the device type of the faulty device is the slave type, notify the host in the parallel system to reassign tasks to a number of slaves except the faulty device to maintain the operation of the parallel system, and remove the faulty device from the parallel system.

[0028] In addition, to achieve the above object, the present invention also provides a device online maintenance device, which is provided in the parallel system. The device includes:

[0029] A faulty device determination module, configured to determine the device type of the faulty device when detecting a device fault in the parallel system;

[0030] A system fault tolerance adjustment module, configured to perform corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system, and remove the faulty device from the parallel system;

[0031] A device reconnection module, configured to, until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation under maintenance passes a preset inspection process, connect the device to be put into operation to the parallel system, where the inspection process includes a buffer detection process and / or an operation detection process.

[0032] In addition, to achieve the above object, the present invention also provides a device parallel maintenance device, which includes: a memory, a processor, and a device online maintenance program stored on the memory and executable on the processor. When the device online maintenance program is executed by the processor, the steps of the above-mentioned device online maintenance method are implemented.

[0033] In addition, to achieve the above object, the present invention also provides a computer-readable storage medium, on which a device online maintenance program is stored. When the device online maintenance program is executed by a processor, the steps of the above-mentioned device online maintenance method are implemented.

[0034] In addition, to achieve the above object, the present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the above-mentioned device online maintenance method are implemented.

[0035] When a device failure occurs in the parallel operation system, the present invention first determines the device type, and then performs corresponding fault tolerance operations on the system for different device types, enabling the system to maintain a normal operating state after the failure occurs; by promptly removing the faulty device from the system, it is possible to repair or replace the faulty device without affecting the normal operation of the system; by reconnecting the device to be put into operation to the system after completing the repair or replacement and confirming that the device to be put into operation has passed the inspection process including a buffer detection process and / or an operation detection process, during the entire process from the occurrence of the fault to the troubleshooting and reconnection of the devices in the parallel operation system, the normal operation of the system will not be disturbed, that is, the system can always carry out normal production, thus solving the technical problem of the high maintenance cost of the existing troubleshooting method for parallel operation. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention;

[0037] Figure 2 is a schematic flowchart of the first embodiment of the device online maintenance method of the present invention;

[0038] Figure 3 is a schematic flowchart of a specific embodiment of the device online maintenance method of the present invention;

[0039] Figure 4 is a schematic diagram of the functional modules of the device online maintenance device of the present invention.

[0040] The realization, functional features and advantages of the object of the present invention will be further described in conjunction with the embodiments with reference to the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0042] In the current field of electrical equipment control, electrical equipment may be applied to application scenarios of parallel operation. In this scenario, if a device in the parallel operation fails and causes a shutdown, and machine replacement or repair is required, the entire parallel operation system needs to be powered off until the staff repairs or replaces the faulty device and then restores the parallel operation state of the entire system. Since normal production cannot be carried out during the troubleshooting period, it results in a high maintenance cost.

[0043] To solve the above problems, the present invention provides a method for online maintenance of equipment. When a device failure occurs in a parallel system, first determine its device type, and then perform corresponding fault tolerance operations on the system for different device types, so that the system can still maintain a normal operating state after the failure occurs; by promptly removing the faulty device from the system, it is possible to repair or replace the faulty device without affecting the normal operation of the system; after completing the repair or replacement and confirming that the device to be put into operation passes a preset inspection process including a buffer detection process and / or an operation detection process, then reconnect the device to be put into operation to the system, so that during the entire process from the occurrence of the failure to the reconnection of the faulty device after troubleshooting in the parallel system, the normal operation of the system will not be disturbed, that is, the system can always carry out normal production, thus solving the technical problem of the relatively high maintenance cost of the existing fault troubleshooting method for parallel operation.

[0044] As Figure 1 shown, Figure 1 is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment of the present invention.

[0045] As Figure 1 shown, the device online maintenance device may include: a processor 1001, such as a CPU, a user interface 1003, a network interface 1004, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. Optionally, the memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0046] Those skilled in the art can understand that Figure 1 the device structure shown in

[0047] As Figure 1 shown, the memory 1005 (provided in the parallel system), as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a device online maintenance program.

[0048] In Figure 1In the device shown, the network interface 1004 is mainly used to connect to the background server and communicate with the background server for data; the user interface 1003 is mainly used to connect to the client (user side) and communicate with the client for data; and the processor 1001 can be used to call the device online maintenance program stored in the memory 1005 and perform the following operations:

[0049] When it is detected that there is a device failure in the parallel system, determine the device type of the faulty device;

[0050] According to the device type, perform corresponding fault tolerance adjustments on the parallel system to maintain the operation of the parallel system, and remove the faulty device from the parallel system;

[0051] Until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation after maintenance passes the preset inspection process, then connect the device to be put into operation to the parallel system, where the inspection process includes a buffer detection process and / or an operation detection process.

[0052] Further, the device type includes a host type and a slave type, and the step of performing corresponding fault tolerance adjustments on the parallel system according to the device type to maintain the operation of the parallel system and removing the faulty device from the parallel system includes:

[0053] If the device type of the faulty device is the host type, change the device type of the faulty device to the slave type, and select a new host from several slaves in the parallel system other than the faulty device;

[0054] Maintain the operation of the parallel system based on the new host, and remove the faulty device from the parallel system.

[0055] Further, the step of selecting a new host from several slaves in the parallel system other than the faulty device includes:

[0056] Obtain the status information and number information of the several slaves, and select the new host according to the status information and number information.

[0057] Further, the step of until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation after maintenance passes the preset inspection process, then connecting the device to be put into operation to the parallel system includes:

[0058] Until a device new access instruction is received, determine the device to be put into operation after repair or replacement;

[0059] Execute the maintenance process on the device to be put into operation, and connect the device to be put into operation to the parallel operation system after the maintenance is passed. During the maintenance process, the DC power switch of the device to be put into operation is not closed.

[0060] Further, the step of executing the maintenance process on the device to be put into operation and connecting the device to be put into operation to the parallel operation system after the maintenance is passed includes:

[0061] Judge whether the buffer function of the device to be put into operation is normal;

[0062] If so, judge whether the device to be put into operation can operate normally;

[0063] If so, determine that the device to be put into operation passes the maintenance;

[0064] Close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel operation system.

[0065] Further, the step of judging whether the buffer function of the device to be put into operation is normal includes:

[0066] Determine that both the AC power switch and the DC power switch of the device to be put into operation are switched off;

[0067] Close the AC power switch, and judge whether the actual DC side voltage value of the device to be put into operation is consistent with the standard DC side voltage value;

[0068] If so, determine that the buffer function of the device to be put into operation is normal;

[0069] If not, determine that the buffer function of the device to be put into operation is abnormal.

[0070] Further, the device type includes a host type and a slave type. The step of performing corresponding fault tolerance adjustment on the parallel operation system according to the device type to maintain the operation of the parallel operation system and removing the faulty device from the parallel operation system includes:

[0071] If the device type of the faulty device is the slave type, notify the host in the parallel operation system to reassign tasks to several slaves except the faulty device to maintain the operation of the parallel operation system, and remove the faulty device from the parallel operation system.

[0072] Based on the above hardware structure, an embodiment of the device online maintenance method of the present invention is proposed.

[0073] Refer to Figure 2 , Figure 2Schematic flowchart of the first embodiment of the on-line maintenance method for the device of the present invention. The on-line maintenance method for the device is applied to a parallel operation system, and the method includes:

[0074] Step S10, when a device failure is detected in the parallel operation system, determine the device type of the faulty device;

[0075] In this embodiment, the parallel operation system refers to a system formed by the parallel operation of multiple electrical devices (including but not limited to inverters, rectifiers, energy storage converters). Generally, a parallel operation system includes a main machine and multiple slave machines. The main machine is used to send control instructions to each slave machine, and each slave machine operates according to the control instructions sent by the main machine. It should be noted that in actual situations, one or more slave machines may fail in the parallel operation system, the main machine may fail, or both the main machine and the slave machines may fail. Therefore, the device type of the faulty device may be the main machine type and / or the slave machine type.

[0076] Specifically, since each device in the parallel operation system monitors its own operating status, and each device also sends its own status information to other devices in the system (for example, a heartbeat message can be sent), when a device in the system stops operating due to a failure, the system can promptly determine the faulty device and thus know the device type of the faulty device.

[0077] Step S20, according to the device type, perform corresponding fault tolerance adjustments on the parallel operation system to maintain the operation of the parallel operation system, and remove the faulty device from the parallel operation system;

[0078] Step S30, until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation passes a preset inspection process, connect the device to be put into operation to the parallel operation system, where the inspection process includes a buffer detection process and / or an operation detection process.

[0079] In this embodiment, the fault tolerance adjustment strategy corresponding to the main machine type is different from the fault tolerance adjustment strategy corresponding to the slave machine type. When only a device of the slave machine type or only a device of the main machine type fails in the system, the system can adopt the fault tolerance adjustment strategy corresponding to the main machine type or the slave machine type; when both the main machine and the slave machines in the system fail, the system can adopt the above two fault tolerance adjustment strategies simultaneously. The inspection process includes at least a buffer detection process and / or an operation detection process, and other inspection processes can also be added according to actual needs.

[0080] Specifically, for the host type, the fault tolerance adjustment strategy can be: select a device from the slave devices in the system as the new host device to replace the faulty host and continue to send instructions, and at the same time (or afterwards) remove the faulty host from the system; for the slave type, since directly removing it will not affect the system, it can be directly removed from the system, or inform the host in the system, and the host will reassign the tasks to be executed to other normally operating slaves, and at the same time (or afterwards) remove the faulty slave from the system, so as to timely assign the tasks that should have been executed by the faulty slave to other slaves to complete, thus avoiding delays in the task execution progress.

[0081] After removing the faulty device, the faulty device can be repaired or replaced. After the repair or replacement is completed, the repaired device or the newly replaced device can be prepared as a device to be put into use and connected to the original parallel system. It should be noted that before connecting the device to be put into use to the system, usually for the consideration of system stability and security, it is also necessary to inspect it to confirm that there are no abnormalities in the performance of each function of the device to be put into use.

[0082] This embodiment provides a method for online maintenance of devices. When a device fails in a parallel system, the method for online maintenance of devices first determines its device type, and then performs corresponding fault tolerance operations on the system for different device types, so that the system can still maintain a normal operating state after the failure occurs; by timely removing the faulty device from the system, it is possible to repair or replace the faulty device without affecting the normal operation of the system; after the repair or replacement is completed and it is confirmed that the device to be put into use passes the inspection process including a buffer detection process and / or an operation detection process, and then reconnect the device to be put into use to the system, so that during the entire process from the occurrence of the fault to the troubleshooting and reconnection of the devices in the parallel system, the normal operation of the system will not be disturbed, that is, the system can always carry out normal production, thus solving the technical problem of the high maintenance cost of the existing fault troubleshooting method for parallel operation.

[0083] Further, based on the above Figure 2 shown in the first embodiment, the second embodiment of the method for online maintenance of devices of the present invention is proposed. In this embodiment, the device type includes a host type and a slave type, and step S20 includes:

[0084] Step S21, if the device type of the faulty device is the host type, change the device type of the faulty device to the slave type, and select a new host from several slaves in the parallel system except the faulty device;

[0085] Step S22, maintain the operation of the parallel system based on the new host, and remove the faulty device from the parallel system.

[0086] In this embodiment, when the system detects that the host has a fault and cannot operate normally, first, the faulty host needs to be changed to a slave, and at the same time, a new host needs to be selected from among the slaves in the system, and then the faulty host that has been demoted to a slave is removed from the system.

[0087] For the selection strategy of the new host, it can be flexibly set according to the actual situation. Specifically, it can be: randomly selecting a slave from all the slaves in the system with normal operating status as the new host; or selecting the slave that best meets the criteria from all the slaves in the system with normal operating status according to certain parameter criteria as the new host.

[0088] In this embodiment, further, when the host has a fault, a new host is quickly competed from the currently running slaves, so as to ensure that the remaining slaves can continue to receive correct control instructions, and further ensure the stable operation of the system.

[0089] Further, the step of selecting a new host from several slaves in the parallel system except the faulty device in step S21 includes:

[0090] Step S211, obtaining the status information and number information of the several slaves, and selecting the new host according to the status information and number information.

[0091] In this embodiment, the status information refers to the information reflecting the operating status of the device, which can be specifically divided into the normal operating status and the abnormal status. The number information refers to the device number uniquely corresponding to each device.

[0092] Specifically, for the selection of the new host, the system can obtain the status information of each slave device, screen out the slave devices with normal operating status according to these status information, and then select the slave device with the largest (or smallest) device number from these slave devices with normal operating status as the new host.

[0093] Further, step S30 includes:

[0094] Step S31, until a device new access instruction is received, determining the device to be put into use after maintenance or replacement;

[0095] Step S32, performing the inspection process on the device to be put into use, and connecting the device to be put into use to the parallel system after the inspection passes, where the DC power switch of the device to be put into use is not switched on during the inspection process.

[0096] In this embodiment, the device new access instruction can be initiated by relevant maintenance personnel to the system. After the maintenance personnel complete the maintenance of the faulty device, they send a device new access instruction to the system. After receiving this instruction, the system determines the device to be put into operation that needs to be accessed, and then performs maintenance on it according to the preset maintenance process. Specifically, it can perform maintenance on each function of the device to be put into operation. If it is confirmed that each function is normal, it can be connected to the system. The specific operation of connecting to the system includes closing the DC power switch of the device to be put into operation.

[0097] It should be noted that before the end of the maintenance, the DC power switch of the device to be put into operation is always in the off state to ensure that the operating state of the device to be put into operation will not affect the system and ensure that the system can always operate stably.

[0098] Further, based on the above second embodiment, a third embodiment of the device online maintenance method of the present invention is proposed. In this embodiment, step S32 includes:

[0099] Step S321, determine whether the buffering function of the device to be put into operation is normal;

[0100] Step S322, if it is, then determine whether the device to be put into operation can operate normally;

[0101] Step S323, if it is, then determine that the device to be put into operation passes the maintenance;

[0102] Step S324, close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel system.

[0103] In this embodiment, the buffering function is mainly verified by changing the opening and closing state of the AC power switch of the device to be put into operation. The system determines whether its buffering function is normal by changing the opening and closing state of its AC power switch; if it fails, it means that there is a functional abnormality in the device to be put into operation. At this time, it cannot be connected to the system and further troubleshooting is needed. Specifically, the system can output an abnormal prompt message on the display interface to prompt the maintenance personnel that the maintenance fails; if it is normal, the system continues to determine whether the device to be put into operation can operate normally. If the device to be put into operation can operate normally at this time, it can be determined that the device to be put into operation passes the maintenance, and it can be connected to the system at this time; if the device to be put into operation cannot operate normally (which can be judged by the operation error message), it means that the device to be put into operation has not yet passed the maintenance process, and it cannot be connected to the system at this time.

[0104] Further, step S321 includes:

[0105] Step S3211, determine that both the AC power switch and the DC power switch of the device to be put into operation are switched off;

[0106] Step S3212: Close the AC power switch and determine whether the actual DC side voltage value of the device to be connected is consistent with the standard DC side voltage value;

[0107] Step S3213: If it is, determine that the buffer function of the device to be connected is normal;

[0108] Step S3214: If not, determine that the buffer function of the device to be connected is abnormal.

[0109] In this embodiment, the system first ensures that the AC power supply and DC power supply of the device to be connected are both switched off (i.e., not powered on), then closes the AC power switch (i.e., powers on), and determines whether the actual DC side voltage value of the device to be connected is consistent with the standard DC side voltage value at this time to detect the buffer function of the device to be connected; if they are consistent, the system can determine that the buffer function of the device to be connected is normal, and specifically, the system can output a prompt message of normal function on the display interface; if they are not consistent, the system determines that the buffer function verified this time is abnormal, and specifically, it can also output a prompt message of abnormal function on the display interface.

[0110] Further, the device types include host type and slave type, and step S20 includes:

[0111] Step A: If the device type of the faulty device is the slave type, notify the host in the parallel system to reassign tasks to several slaves except the faulty device to maintain the operation of the parallel system, and remove the faulty device from the parallel system.

[0112] In this embodiment, if a slave device in the system fails, the system can notify the host of this failure situation. After learning about it, the host can reassign the tasks that need to be executed next to the remaining slaves, transfer or allocate the tasks originally executed by the faulty slave to the remaining normally executing slaves to ensure that all tasks can proceed normally, and then remove the faulty slave from the system for separate maintenance.

[0113] As a specific embodiment, as Figure 3 shown.

[0114] After the system enters the maintenance operation mode for the equipment to be put into operation, it will first trip the DC power switch and AC power switch of the equipment to be put into operation. After confirming that the tripping is completed, it will close the AC power switch. After confirming that the closing is completed, it will judge whether the DC side voltage value of the equipment to be put into operation is consistent with the standard value at this time to judge whether the pre-charging is completed (that is, to detect the buffer function). If it is inconsistent, it will prompt the maintenance failure through the human-machine interface (HMI, Human Machine Interface) and trip the AC power switch. If it is consistent, it will determine that the pre-charging is completed, and continue to judge whether this equipment to be put into operation can operate normally. If it cannot operate normally, it will prompt the maintenance failure through the human-machine interface. If it can operate normally, it will determine that the maintenance process is passed and prompt the maintenance success through the human-machine interface. After the maintenance is successful, the system will first trip the AC power switch of the equipment to be put into operation to exit the maintenance operation mode, and then enter the normal operation mode after the tripping is completed. In this mode, the DC power switch and AC power switch of the equipment to be put into operation will be closed to officially connect to the system, and the equipment can operate normally after being connected to the system.

[0115] As Figure 4 shown, the present invention also provides an on-line maintenance device for equipment. The device is arranged in a parallel system, and the device includes:

[0116] A faulty equipment determination module 10, configured to determine the equipment type of the faulty equipment when detecting that there is equipment failure in the parallel system;

[0117] A system fault tolerance adjustment module 20, configured to perform corresponding fault tolerance adjustment on the parallel system according to the equipment type to maintain the operation of the parallel system, and remove the faulty equipment from the parallel system;

[0118] An equipment reconnection module 30, configured to, until the maintenance of the faulty equipment is completed, if it is determined that the equipment to be put into operation to be maintained passes a preset maintenance process, connect the equipment to be put into operation to the parallel system, where the maintenance process includes a buffer detection process and / or an operation detection process.

[0119] Optionally, the equipment type includes a host type and a slave type, and the system fault tolerance adjustment module 20 includes:

[0120] A host fault tolerance unit, configured to, if the equipment type of the faulty equipment is the host type, change the equipment type of the faulty equipment to the slave type, and select a new host from several slaves other than the faulty equipment in the parallel system;

[0121] A first equipment removal unit, configured to maintain the operation of the parallel system based on the new host and remove the faulty equipment from the parallel system.

[0122] Optionally, the host re-selection unit is further configured to:

[0123] Obtain the status information and number information of the plurality of slave machines, and select the new host according to the status information and number information.

[0124] Optionally, the device reconnection module 30 includes:

[0125] A new instruction receiving unit, configured to determine the device to be put into operation after maintenance or replacement until a device new access instruction is received;

[0126] A device maintenance unit, configured to execute the maintenance process on the device to be put into operation, and connect the device to be put into operation to the parallel system after the maintenance passes, wherein the DC power switch of the device to be put into operation is not closed during the maintenance process.

[0127] Optionally, the device maintenance unit is further configured to:

[0128] Judge whether the buffer function of the device to be put into operation is normal;

[0129] If so, judge whether the device to be put into operation can operate normally;

[0130] If so, determine that the device to be put into operation passes the maintenance;

[0131] Close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel system.

[0132] Optionally, the step of judging whether the buffer function of the device to be put into operation is normal includes:

[0133] Determine that both the AC power switch and the DC power switch of the device to be put into operation are opened;

[0134] Close the AC power switch, and judge whether the actual DC side voltage value of the device to be put into operation is consistent with the standard DC side voltage value;

[0135] If so, determine that the buffer function of the device to be put into operation is normal;

[0136] If not, determine that the buffer function of the device to be put into operation is abnormal.

[0137] Optionally, the device type includes a host type and a slave type, and the system fault tolerance adjustment module 20 includes:

[0138] The slave fault tolerance unit is used to notify the host in the parallel system to reassign tasks to several slaves except the faulty device if the device type of the faulty device is the slave type, so as to maintain the operation of the parallel system, and remove the faulty device from the parallel system.

[0139] The present invention also provides a device parallel maintenance device.

[0140] The device parallel maintenance device includes a processor, a memory, and a device online maintenance program stored on the memory and executable on the processor. When the device online maintenance program is executed by the processor, the steps of the device online maintenance method as described above are implemented.

[0141] Wherein, the method implemented when the device online maintenance program is executed can refer to the various embodiments of the device online maintenance method of the present invention, which will not be elaborated here.

[0142] The present invention also provides a computer-readable storage medium.

[0143] The device online maintenance program is stored on the computer-readable storage medium of the present invention. When the device online maintenance program is executed by a processor, the steps of the device online maintenance method as described above are implemented.

[0144] Wherein, the method implemented when the device online maintenance program is executed can refer to the various embodiments of the device online maintenance method of the present invention, which will not be elaborated here.

[0145] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the device online maintenance method as described above are implemented.

[0146] Wherein, the method implemented when the computer program is executed can refer to the various embodiments of the device online maintenance method of the present invention, which will not be elaborated here.

[0147] It should be noted that in this article, the terms "including", "comprising" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or system including that element.

[0148] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.

[0149] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.

[0150] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.

Claims

1. An on-line maintenance method for a device, characterized in that, The method for online maintenance of the device is applied to a parallel system, and the method includes: When a device failure is detected in the parallel system, determine the device type of the faulty device; According to the device type, perform corresponding fault tolerance adjustment on the parallel system to maintain the operation of the parallel system, and remove the faulty device from the parallel system; Until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation after maintenance passes a preset inspection process, connect the device to be put into operation to the parallel system, where the inspection process includes a buffer detection process and / or an operation detection process; The step of connecting the device to be put into operation to the parallel system until the maintenance of the faulty device is completed and if it is determined that the device to be put into operation after maintenance passes a preset inspection process includes: Until a device new access instruction is received, determine the device to be put into operation after repair or replacement; Execute the inspection process on the device to be put into operation, and connect the device to be put into operation to the parallel system after the inspection passes. During the inspection, the DC power switch of the device to be put into operation is not closed. Determine whether the buffer function of the device to be put into operation is normal; if so, determine whether the device to be put into operation can operate normally; if so, determine that the device to be put into operation passes the inspection; close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel system.

2. The method for online maintenance of the device according to claim 1, characterized in that, The device type includes a host type and a slave type. The step of performing corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system and removing the faulty device from the parallel system includes: If the device type of the faulty device is the host type, change the device type of the faulty device to the slave type, and select a new host from several slaves in the parallel system except the faulty device; Maintain the operation of the parallel system based on the new host, and remove the faulty device from the parallel system.

3. The on-line maintenance method of the device according to claim 2, characterized in that, The step of selecting a new host from several slaves in the parallel system except the faulty device includes: Obtain the status information and number information of the several slaves, and select the new host according to the status information and number information.

4. The method for online maintenance of the device according to claim 1, characterized in that, The step of determining whether the buffer function of the device to be put into operation is normal includes: Determine that both the AC power switch and the DC power switch of the device to be put into operation are opened; Close the AC power switch, and determine whether the actual DC side voltage value of the device to be put into operation is consistent with the standard DC side voltage value; If so, determine that the buffer function of the device to be put into operation is normal; If not, determine that the buffer function of the device to be put into operation is abnormal.

5. The method for online maintenance of the device according to claim 4, characterized in that, The device type includes a host type and a slave type. The step of performing corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system and removing the faulty device from the parallel system includes: If the device type of the faulty device is a slave type, notify the host in the parallel system to reassign tasks to several slaves except the faulty device to maintain the operation of the parallel system, and remove the faulty device from the parallel system.

6. An on-line maintenance device for a device, characterized in that, The device is provided in a parallel system, and the device includes: A faulty device determination module, configured to determine the device type of the faulty device when detecting a device fault in the parallel system; A system fault tolerance adjustment module, configured to perform corresponding fault tolerance adjustment on the parallel system according to the device type to maintain the operation of the parallel system, and remove the faulty device from the parallel system; A device reconnection module, configured to, until the maintenance of the faulty device is completed, if it is determined that the device to be put into operation to be maintained passes a preset inspection process, connect the device to be put into operation to the parallel system, where the inspection process includes a buffer detection process and / or an operation detection process; The device reconnection module is further configured to, until a device new connection instruction is received, determine the device to be put into operation after repair or replacement; execute the inspection process on the device to be put into operation, and connect the device to be put into operation to the parallel system after the inspection passes, where the DC power switch of the device to be put into operation is not closed during the inspection process, determine whether the buffer function of the device to be put into operation is normal; if so, determine whether the device to be put into operation can operate normally; if so, determine that the device to be put into operation passes the inspection; close the DC power switch of the device to be put into operation to connect the device to be put into operation to the parallel system.

7. A parallel maintenance device for equipment, characterized in that, The device parallel inspection device includes: a memory, a processor, and a device online maintenance program stored on the memory and executable on the processor, and when the device online maintenance program is executed by the processor, the steps of the device online maintenance method according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program, and when the computer program is executed by a processor, the steps of the device online maintenance method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Uninterrupted power supply system

    CN206313520U

  • Fault-tolerant computer system with / config filesystem

    EP0433979A2

  • Control method of UPS module connected in parallel and system thereof

    TW200419884A

  • Fault-tolerant computer system with online recovery and reintegration of redundant components

    US5295258A