Fault processing method based on double-control system, double-control system and storage medium
Through the troubleshooting method of the dual-control system, the CPU fault is monitored by separate motherboards and circuits, and the reset command is generated to restart the system, solving the problem of downtime of the storage product when the CPU fails, and improving the reliability and stability of the system.
Patent Information
- Application Number
- CN202510897129.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing storage products are inefficient in detection and recovery in the event of CPU failure, resulting in system downtime and data corruption, especially in case of catastrophic errors at the hardware level.
The dual control system is adopted, including a separate first motherboard and a second motherboard. The target device is monitored by the first circuit arranged on the first motherboard, and the fault device is determined in response to changes in the control signal, and a reset command is generated to restart the dual control system when it is not communicable, realizing full-link fault processing.
It improves the reliability and stability of storage products, avoids downtime problems caused by target equipment failure, realizes rapid response from signal-level fault detection to system recovery, and enhances the system's fault tolerance and high availability.
Smart Images

Figure CN120492209A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a fault handling method based on a dual-control system, a dual-control system, and a storage medium. Background Art
[0002] When storage products like storage area networks, network-attached storage, and distributed storage systems process massive amounts of data read and write requests, their control devices, such as the central processing unit (CPU), undertake core tasks such as metadata management, cache scheduling, and protocol conversion. As the core control unit of storage products, the CPU's reliability is directly related to data security and service continuity.
[0003] During the implementation of the embodiments of the present application, it was found that although the CPU will attempt to repair an error when it detects one, the detection and recovery efficiency is low. Even when a catastrophic error occurs at the CPU hardware level, the CPU cannot continue to work, resulting in a direct system crash and data corruption. Summary of the Invention
[0004] In view of the above problems, the present application provides a fault handling method based on a dual-control system, a dual-control system, a fault handling device based on a dual-control system, a computer-readable storage medium, and a computer program product.
[0005] According to the first aspect of the present application, a fault handling method based on a dual-control system is provided, wherein the dual-control system includes a separate first main board and a second main board, and the fault handling method is applied to a first circuit arranged on the first main board; the fault handling method based on the dual-control system includes: in response to receiving a change in a control signal, determining a target device in which a fault occurs in the dual-control system based on identification information of the control signal; and generating a reset instruction when it is determined that the target device is located on the second main board and cannot communicate with the second circuit arranged on the second main board, so as to restart the dual-control system in response to the reset instruction.
[0006] The second aspect of the present application provides a dual-control system, including: a first mainboard, on which a first circuit is arranged; a second mainboard separately from the first mainboard, on which a second circuit is arranged, and a communication link is provided between the first circuit and the second circuit; wherein the first circuit is used to: in response to a change in a received control signal, determine a target device that has failed in the dual-control system based on identification information of the control signal; and generate a reset instruction when it is determined that the target device is located on the second mainboard and cannot communicate with the second circuit, so as to restart the dual-control system in response to the reset instruction.
[0007] A third aspect of the present application further provides a fault handling device based on a dual-control system, the dual-control system comprising a first separate mainboard and a second separate mainboard, wherein the fault handling device based on the dual-control system is disposed in a first circuit disposed on the first mainboard. The fault handling device based on the dual-control system comprises: a determination module configured to, in response to receiving a change in a control signal, determine a target device in the dual-control system where a fault has occurred based on identification information of the control signal; and a first generation module configured to, upon determining that the target device is located on the second mainboard and is unable to communicate with a second circuit disposed on the second mainboard, generate a reset instruction to restart the dual-control system in response to the reset instruction.
[0008] The fourth aspect of the present application further provides a computer-readable storage medium having a computer program or instructions stored thereon, which implements the steps of the above-mentioned fault handling method when the above-mentioned computer program or instructions are executed by a processor.
[0009] The fifth aspect of the present application further provides a computer program product, comprising a computer program or instructions, which implement the steps of the above-mentioned fault handling method when executed by a processor.
[0010] According to the embodiments of the present application, since the devices in the dual-control system are monitored by the first circuit arranged on the first mainboard, it is possible to promptly determine the faulty target device when the control signal changes, and then by determining whether the first circuit arranged on the first mainboard and the second circuit arranged on the second mainboard can communicate, it is possible to restart the dual-control system through other node recovery mechanisms other than the target device, so that the faulty target device returns to normal, avoiding downtime problems caused by target device failures, and avoiding the risk of the dual-control system being unable to achieve redundancy due to direct downtime of the target device. In addition, a full-link processing mechanism is implemented from signal-level fault detection to dual-control system recovery, which can be applied to the multi-node whole machine environment of storage products. Through the redundant design of the dual-control system, dual-control management of storage product failures is achieved, thereby improving the reliability of the dual-control system. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above contents and other objects, features and advantages of the present application will become more apparent through the following description of the embodiments of the present application with reference to the accompanying drawings.
[0012] Figure 1 The diagram shows an application scenario of a fault handling method and device based on a dual-control system according to an embodiment of the present application.
[0013] Figure 2 A flowchart of a fault handling method based on a dual-control system according to an embodiment of the present application is shown.
[0014] Figure 3AA schematic diagram of generating a reset instruction according to an embodiment of the present application is shown.
[0015] Figure 3B A schematic diagram of generating a reset instruction according to another embodiment of the present application is shown.
[0016] Figure 3C A schematic diagram of generating log information according to an embodiment of the present application is shown.
[0017] Figure 4 A schematic diagram of generating log information according to a specific embodiment of the present application is shown.
[0018] Figure 5 A schematic diagram of generating feedback information according to an embodiment of the present application is shown.
[0019] Figure 6A The figure shows a structural block diagram of a dual control system according to an embodiment of the present application.
[0020] Figure 6B A structural block diagram of a dual control system according to another embodiment of the present application is shown.
[0021] Figure 6C A structural block diagram of a dual-control system according to another embodiment of the present application is shown.
[0022] Figure 6D A structural block diagram of a dual control system according to another embodiment of the present application is shown.
[0023] Figure 7 A structural block diagram of a fault handling device based on a dual-control system according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0024] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the present application. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present application. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present application.
[0025] The terms used herein are only for describing specific embodiments and are not intended to limit the present application. The terms "comprise," "include," etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0027] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0028] During the implementation of the embodiments of the present application, it was discovered that when the CPU detects a recoverable error, such as a single-bit error, the location of the error can be marked and the error can be repaired through a series of hardware and software mechanisms, so that users are less likely to perceive the occurrence of such errors and reduce the impact on users. However, if recoverable errors occur frequently, the system administrator may need to check related hardware components, such as memory modules, and perform maintenance or replacement, which results in low detection and recovery efficiency. In addition, due to the differences between terminal services and hardware themselves, even recoverable errors may cause the CPU to crash, and thus the dual-control system to crash.
[0029] When the CPU detects an unrecoverable error, such as a bad memory block or bus error, it can isolate the fault, for example by isolating the bad memory block or reducing the bus frequency, to maintain system operation. However, unrecoverable errors can lead to degraded system performance, data loss, or system instability. While the system can maintain operation by isolating the faulty area, such as marking a bad memory block and switching to an alternate path, hardware maintenance or replacement will ultimately be required.
[0030] In addition, if a more serious failure occurs in the CPU, such as a catastrophic error at the hardware level, the CPU cannot continue to work, which will cause the system to crash directly and data to be damaged.
[0031] In view of this, an embodiment of the present application provides a fault handling method based on a dual-control system, wherein the dual-control system includes a separate first main board and a second main board, and the fault handling method based on the dual-control system is applied to a first circuit arranged on the first main board; the fault handling method based on the dual-control system includes: in response to receiving a change in a control signal, determining a target device where a fault occurs in the dual-control system based on identification information of the control signal; when it is determined that the target device is located on the second main board and cannot communicate with the second circuit arranged on the second main board, generating a reset instruction to restart the dual-control system in response to the reset instruction.
[0032] Figure 1 The diagram shows an application scenario of a fault handling method and device based on a dual-control system according to an embodiment of the present application.
[0033] like Figure 1 As shown, an application scenario 100 according to this embodiment may include a terminal device 101 and a dual-control system 102. A network may be used as a medium for providing a communication link between the terminal device 101 and the dual-control system 102. The network may include various connection types, such as wired or wireless communication links or optical fiber cables.
[0034] The user can use the terminal device 101 to interact with the dual control system 102 via the network to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0035] The terminal device 101 may be any electronic device having a display screen and supporting web browsing, including but not limited to a smart phone, a tablet computer, a laptop computer, a desktop computer, and the like.
[0036] The dual-control system 102 may include a separate first mainboard and a second mainboard. The first mainboard and the second mainboard may be control mainboards with exactly the same layout. The hardware design of storage products focuses on storage density and scalability. In order to ensure that data is stably and reliably stored in the hard disk, the hardware will adopt a redundant design, such as the redundant design of the first mainboard and the second mainboard in the dual-control system 102. As the main and backup mainboards, when the main mainboard fails, it switches to the backup mainboard in milliseconds to ensure read and write continuity. However, in order to ensure stable and reliable operation of the CPU of the storage product and to effectively monitor and restore the CPU in a timely manner when a serious error occurs, the first mainboard of the dual-control system 102 may be equipped with a first circuit. The second mainboard may be equipped with a second circuit. The first circuit and the second circuit communicate with each other. Both the first circuit and the second circuit can monitor devices arranged on the first mainboard and / or the second mainboard, such as the CPU.
[0037] It should be noted that the fault handling method based on the dual-control system provided in the embodiment of the present application can generally be executed by the first circuit of the first mainboard arranged on the dual-control system 102. Accordingly, the fault handling device based on the dual-control system provided in the embodiment of the present application can generally be arranged in the first circuit of the first mainboard arranged on the dual-control system 102. The fault handling method based on the dual-control system provided in the embodiment of the present application can also be executed by a circuit or circuit cluster that is different from the first circuit of the first mainboard arranged on the dual-control system 102 and can communicate with the terminal device 101 and / or the dual-control system 102. Accordingly, the fault handling device based on the dual-control system provided in the embodiment of the present application can also be arranged in a circuit or circuit cluster that is different from the first circuit of the first mainboard arranged on the dual-control system 102 and can communicate with the terminal device 101 and / or the first circuit of the first mainboard arranged on the dual-control system 102.
[0038] It should be understood that Figure 1 The number of terminal devices and dual-control systems in the embodiment is merely illustrative. Any number of terminal devices and dual-control systems may be provided as required.
[0039] The following will be based on Figure 1 The scene described by Figures 2 to 5 The fault handling method based on the dual control system of the embodiment of the application is described in detail.
[0040] Figure 2 The flowchart of the fault handling method based on the dual-control system according to an embodiment of the present application is schematically shown.
[0041] like Figure 2 As shown, the fault handling method based on the dual-control system of this embodiment includes operations S210 to S220. The fault handling method based on the dual-control system can be applied to the first circuit arranged on the first mainboard.
[0042] In operation S210 , in response to receiving a change in the control signal, a target device in which a fault occurs in the dual-control system is determined based on identification information of the control signal.
[0043] In operation S220 , if it is determined that the target device is located on the second mainboard and cannot communicate with the second circuit disposed on the second mainboard, a reset instruction is generated to restart the dual-control system in response to the reset instruction.
[0044] In the embodiment of the present application, faults may include recoverable faults and unrecoverable faults. A recoverable fault or an unrecoverable fault occurring in a target device may cause the target device to shut down, and a target device shutdown may cause the dual-control system to shut down. The main causes of a recoverable fault occurring in a target device may be, for example, the following (1) to (5).
[0045] (1) Poor heat dissipation leads to excessively high temperatures. When the target device is running at high load for a long time, if the cooling system cannot meet the heat dissipation capacity, the temperature exceeding the threshold will trigger the protection mechanism, directly causing the target device to shut down.
[0046] (2) Unstable or insufficient power supply. If the power conversion efficiency is low, the voltage fluctuates, or the power is insufficient, especially in high-load scenarios of the target device, power outages may occur due to abnormal power supply.
[0047] (3) Hardware compatibility issues. For example, the motherboard does not match the target device, or the external frequency is set too high during overclocking, resulting in unstable operation of the target device.
[0048] (4) Hardware quality issues, such as inferior components. For example, the quality of power supplies and radiators does not meet the standards, affecting the normal operation of the target equipment.
[0049] (5) Software or system issues, such as driver conflicts / system file corruption, malware infection, etc. Driver conflicts / system file corruption can be caused by incompatible drivers or damaged system files, which may cause resource conflicts or overload on the target device. Malware infection can be caused by viruses or Trojans that occupy a large amount of target device resources, causing downtime.
[0050] When an unrecoverable failure occurs on the target device, such as a catastrophic error at the hardware level, the target device will also crash and become unable to operate normally.
[0051] A control signal can indicate a target device failure. For example, control signals include, but are not limited to, catastrophic error (CATERR) signals and thermal protection signals. Triggering a CATERR signal typically causes a system shutdown to prevent data loss or corruption. Administrators can perform hardware maintenance or replacement based on error logs. A change in a control signal can refer to a level change. For example, a change in the CATERR signal from a high level to a low level indicates a CATERR signal change.
[0052] The identification information of the control signal can be used to indicate whether the target device is arranged on the first main board or the second main board.
[0053] The first and second mainboards of the dual-control system may both be equipped with target devices, and the identification information of the target devices on each mainboard may be different. Based on the identification information, it can be determined which mainboard is equipped with the target device that has failed.
[0054] Figure 3A A schematic diagram of generating a reset instruction according to an embodiment of the present application is shown.
[0055] like Figure 3AAs shown, a first circuit 311 is disposed within the first mainboard 310. A second circuit 321 and a target device 322 are disposed within the second mainboard 320. A device redundant with the target device may also be disposed within the first mainboard 310 (not shown). Both the first circuit 311 and the second circuit 321 can detect a device disposed within the first mainboard 310 or a target device 322 disposed within the second mainboard 320. When the first circuit 311 and / or the second circuit 321 detects a change in control information, it is determined based on the identification information that the target device 322 has failed. If the target device 322 has failed, the first circuit 311 determines that the link with the second circuit 321 is disconnected, meaning that the first circuit 311 and the second circuit 321 cannot communicate. Consequently, the second circuit 321 cannot function normally, requiring the first circuit 311 to generate a reset instruction 330.
[0056] The reset instruction 330 may include an instruction capable of triggering a reset of a control signal of the target device 322 , thereby restarting the dual-control system.
[0057] The target device 322 may be used to process data read and write requests sent by the terminal device, and may include, for example, but not limited to, a CPU.
[0058] Since the devices in the dual-control system are monitored by the first circuit arranged on the first mainboard, the faulty target device can be promptly identified when the control signal changes. Furthermore, by determining whether the first circuit arranged on the first mainboard and the second circuit arranged on the second mainboard can communicate, the dual-control system can be restarted through the recovery mechanism of other nodes other than the target device, so that the faulty target device can be restored to normal, avoiding downtime problems caused by target device failures and the risk of the dual-control system being unable to achieve redundancy due to direct downtime of the target device. In addition, a full-link processing mechanism is implemented from signal-level fault detection to dual-control system recovery, which can be applied to the multi-node whole machine environment of storage products. Through the redundant design of the dual-control system, dual-control management of storage product failures can be achieved, thereby improving the reliability of the dual-control system.
[0059] According to another embodiment of the present application, the fault handling method based on the dual control system may include the above Figure 2 In addition to the operations S210 to S220 shown, the process may further include: controlling the second circuit to generate a reset instruction when it is determined that the target device is located on the second mainboard and can communicate with the second circuit.
[0060] Figure 3B A schematic diagram of generating a reset instruction according to another embodiment of the present application is shown.
[0061] like Figure 3BAs shown, in this embodiment, the method in which the first circuit 311 determines that the target device 322 fails is the same as that described above. Figure 3A The method is the same as in , so I will not repeat it here.
[0062] When it is determined that the target device 322 fails, the first circuit 311 determines that the link with the second circuit 321 is connected, that is, the first circuit 311 and the second circuit 321 can communicate. Therefore, the second circuit 321 can work normally and the second circuit 321 can generate a reset instruction 330.
[0063] Because the circuit and the faulty device (the target device) are located on the same motherboard, the circuit generates the reset command. This means the local circuit triggers a control information reset, enabling a faster response to faults, reducing reliance on peer circuits and improving system reliability and stability. Furthermore, this local circuit triggers a control information reset, allowing for faster fault isolation or recovery, avoiding system downtime caused by communication delays or peer failures, thereby enhancing the system's fault tolerance and high availability.
[0064] According to another embodiment of the present application, the fault handling method based on the dual control system may include the above Figure 2 In addition to operations S210 to S220, the following operations may also be included: if the target device is determined to be located on the first mainboard, generating a fault notification indicating that the target device has failed. The fault notification is sent to a first controller located on the first mainboard, so that the first controller performs the following operations based on the fault notification: In response to the fault notification, monitoring the storage information of the target device based on a pre-stored adaptation file to obtain a monitoring result. If the monitoring result indicates that the storage information is abnormal, generating log information based on a preset alarm mechanism.
[0065] In this embodiment, the fault notification may include identification information of the target device and fault information.
[0066] The first controller and the target device can communicate via an interface (such as Inter-Integrated Circuit (I2C) or Platform Environment Control Interface (PECI)). The first controller collects the target device's storage information in real time. It's important to note that the design of the hardware link between the first controller and the target device should take into account signal stability and anti-interference capabilities.
[0067] The pre-stored adaptation file may be a file pre-configured in the first controller. The first controller may parse the adaptation file to obtain content information in the adaptation file. Upon receiving a fault notification, the first controller may monitor the stored information of the target device collected in real time to obtain a monitoring result.
[0068] The preset alarm mechanism can be used to instruct the first controller to execute a predetermined alarm action when a preset condition is met, such as generating a log message to record the specific information of the event in detail. The monitoring result indicates that the stored information is abnormal when the preset condition is met.
[0069] Since the first circuit sends a fault notification to the first controller after determining that a fault has occurred in the target device, the first controller can accurately analyze the fault and generate log information when receiving the fault notification, which is conducive to subsequent analysis of the cause of the fault and evaluation of the effectiveness of the system response, providing an important basis for system optimization and maintenance, thereby improving the reliability and maintainability of the system.
[0070] Exemplarily, the storage information may come from multiple registers of the target device. In response to the fault notification, the storage information of the target device is monitored based on a pre-stored adaptation file to obtain a monitoring result. This may include: determining, based on the types of the multiple registers, a monitoring strategy for the storage information of any of the multiple registers; parsing association information associated with the monitoring strategy from the pre-stored adaptation file; and comparing the association information with the storage information to obtain a monitoring result.
[0071] The monitoring strategies for the storage information of registers are different depending on the type of register.
[0072] The types of registers may include, but are not limited to, general-purpose registers and flag registers.
[0073] For a general register, the monitoring strategy may be to monitor whether the value corresponding to the storage information of the general register is within a predetermined range. The associated information associated with the monitoring strategy may be a predetermined range of values corresponding to the storage information matching the general register.
[0074] For example, a predetermined range of values corresponding to storage information matching a general-purpose register can be parsed from a pre-stored adaptation file. Upon receiving a fault notification, if the value corresponding to the storage information of the general-purpose register collected in real time is within the predetermined range, the monitoring result indicates that the storage information is normal; otherwise, the monitoring result indicates that the storage information is abnormal.
[0075] For a flag register, the monitoring strategy may be to monitor whether the combination of flag bits in the stored information of the flag register meets a predetermined rule. The associated information associated with the monitoring strategy may be a predetermined rule corresponding to the combination of flag bits in the stored information that matches the flag register.
[0076] For example, a predetermined rule corresponding to a combination of flag bits in the storage information matching the flag register can be parsed from a pre-stored adaptation file. When a fault notification is received, if the combination of flag bits in the storage information of the flag register collected in real time meets the predetermined rule, the monitoring result indicates that the storage information is normal; otherwise, the monitoring result indicates that the storage information is abnormal.
[0077] Since the first controller determines different monitoring strategies for the storage information of the register according to different register types, it can achieve refined monitoring of the storage information and fault warning, thereby improving system reliability and maintenance efficiency.
[0078] In another example, the first controller may also be configured to send the log information to a second controller disposed on a second mainboard so that the second controller can perform redundant storage.
[0079] Figure 3C A schematic diagram of generating log information according to an embodiment of the present application is shown.
[0080] like Figure 3C As shown, the first mainboard 310 is provided with a first circuit 311, a first controller 312, and a target device 322. The second mainboard 320 is provided with a second circuit 321 and a second controller 323. The second mainboard 320 may also be provided with a device redundant with the target device (not shown). Both the first circuit 311 and the second circuit 321 can detect the target device 322 or the device within the first mainboard 310. When the first circuit 311 and / or the second circuit 321 detect a change in control information, they determine, based on the identification information, that the target device 322 has failed. If the target device 322 has failed, the first circuit 311 generates a fault notification and sends it to the first controller 312. After receiving the fault notification, the first controller 312 monitors the stored information of the target device. If the monitoring results indicate an abnormality in the stored information, the first controller 312 generates and stores log information based on a preset alarm mechanism. The first controller 312 also sends the log information to the second controller 323, enabling the second controller 323 to perform redundant storage upon receipt of the log information.
[0081] Sending log information to the peer controller for redundant storage ensures that log information is fully preserved when local storage fails, thereby improving data reliability and security, and facilitating subsequent fault analysis and recovery operations.
[0082] In yet another example, the first controller 312 may also be configured to feed back log information to the terminal device so that the terminal device can maintain the dual-control system based on the log information.
[0083] Log information may include but is not limited to: fault timestamp, target device fault type, fault register, fault description and other information.
[0084] Figure 4 A schematic diagram of generating log information according to a specific embodiment of the present application is shown.
[0085] like Figure 4 As shown, in this embodiment, the first mainboard 310 and the second mainboard 320 in the dual-control system are separated by a backplane 410, and the target device is the first CPU 411, the first circuit is the first digital integrated circuit (first CPLD) 412, the first controller is the first baseboard controller (first BMC) 413, the second circuit is the second digital integrated circuit (second CPLD) 422, the second controller is the second baseboard controller (second BMC) 423, and the device installed in the second mainboard 320 is the second CPU 421 as an example.
[0086] When a fault occurs in the first CPU 411, a CATERR signal is simultaneously reported to the first CPLD 412 and the second CPLD 422. After the first CPLD 412 and the second CPLD 422 detect that the CATERR signal transitions from a high level to a low level, both the first CPLD 412 and the second CPLD 422 generate CPLD log records and distinguish whether the fault occurred in the first CPU 411. The first CPLD 412 notifies the first BMC 413 of the fault, and the second CPLD 422 notifies the second BMC 423 of the fault. After receiving the fault notification of the first CPU 411, the first BMC 413 obtains the error record of the first CPU 411, such as the error status and error information of the first CPU 411, via a System Management Bus (SMBus) signal. The error record is then sent to the second BMC 423 via a redundant link to create a similar log record, making it easier for users to query and understand CPU fault information.
[0087] A heartbeat link can exist between the first CPLD 412 and the second CPLD 422 to mutually check whether the link is normal. Because a CPU failure may cause the entire motherboard to malfunction, the first CPLD 412 can be monitored via the Universal Asynchronous Receiver / Transmitter (UART) link of the second CPLD 422 to determine whether the first CPLD 412 and the second CPLD 422 are able to communicate.
[0088] For example, if the UART link is communicating with each other, it means that the first CPLD 412 and the second CPLD 422 can communicate with each other. In this case, only the first CPLD 412 needs to trigger the first CPU 411 to reset the level and restart the dual-control system. If the UART link is not connected, it means that the first CPLD 412 cannot communicate with each other. In this case, the second CPLD 422 needs to trigger the first CPU 411 to reset the level and restart the dual-control system.
[0089] The dual control of the present application can be a storage system composed of two identical control motherboards for redundancy design of the storage product. Each control motherboard has an independent fault management system, and the management system between the dual controls has a redundant link design, which can achieve data interaction and mutual monitoring of the dual control systems, so that the fault management system between the dual controls can manage the dual control systems and data redundancy at the same time. The present application can be used as a recovery mechanism in the scenario where a single control motherboard of a storage device goes down, avoiding the risk of the control motherboard not being redundant due to a direct downtime. The present application evolves the system downtime recovery solution from passive recovery to active recovery.
[0090] According to an embodiment of the present application, the fault handling method also includes: when it is determined that the number of restarts of the dual-control system is greater than a predetermined number, and the control signal changes again within a predetermined time after the dual-control system is restarted, generating feedback information indicating that the fault of the target device is an irreparable fault; when it is determined that the number of restarts is less than or equal to the predetermined number, and the control signal does not change within a predetermined time after the dual-control system is restarted, generating feedback information indicating that the fault of the target device is a repairable fault.
[0091] Figure 5 A schematic diagram of generating feedback information according to an embodiment of the present application is shown.
[0092] like Figure 5 As shown, in this embodiment, generating feedback information may include operations S501 to S505.
[0093] In operation S501 , the dual-control system is restarted, and the number of restarts is recorded.
[0094] In operation S502, it is determined whether the control signal changes within a predetermined time period after the dual-control system is restarted. If it is determined that the control signal changes within the predetermined time period after the dual-control system is restarted, operation S504 is performed. If it is determined that the control signal does not change within the predetermined time period after the dual-control system is restarted, operation S503 is performed.
[0095] In operation S503 , feedback information indicating that the fault of the target device is a repairable fault is generated.
[0096] In operation S504, it is determined whether the number of restarts of the dual-control system is less than or equal to the predetermined number. If it is determined that the number of restarts of the dual-control system is less than or equal to the predetermined number, operation S501 is repeated. If it is determined that the number of restarts of the dual-control system is greater than the predetermined number, operation S505 is performed.
[0097] In operation S505 , feedback information indicating that the fault of the target device is an unrepairable fault is generated.
[0098] The reservation duration and number of reservations may be experience values.
[0099] By recording the number of times the dual-control system restarts and whether the anomaly recurs after a restart, it's possible to effectively determine whether a fault is repairable. If the anomaly persists after a restart, it indicates the fault is likely temporary or automatically recoverable. If the anomaly persists, even after a restart, it suggests hardware damage or a deeper software issue, requiring further diagnosis and repair. This feedback mechanism helps quickly identify the nature of the fault, improves system maintenance efficiency, and reduces unnecessary repair costs.
[0100] Based on the above-mentioned fault handling method based on the dual control system, this application also provides a dual control system. Figure 6A to Figure 6D The system is described in detail.
[0101] Figure 6A The figure shows a structural block diagram of a dual control system according to an embodiment of the present application.
[0102] like Figure 6A As shown, in this embodiment, the dual-control system 102 may include a first mainboard 310 and a second mainboard 320 provided separately from the first mainboard 310. The first mainboard 310 may be provided with a first circuit 311. The second mainboard 320 may be provided with a second circuit 321 and a target device 322. A communication link is provided between the first circuit 311 and the second circuit 321.
[0103] The first circuit 311 is used to respond to changes in the received control signal, determine the target device 322 that has failed in the dual-control system 102 based on the identification information of the control signal, and generate a reset instruction when it is determined that the target device 322 is located on the second mainboard 320 and cannot communicate with the second circuit 321, so as to restart the dual-control system 102 in response to the reset instruction.
[0104] Figure 6B A structural block diagram of a dual control system according to another embodiment of the present application is shown.
[0105] like Figure 6B As shown, in this embodiment, the dual control system 102 includes the above Figure 6AIn addition to the components shown, the first mainboard 310 of the dual-control system 102 is also provided with a first controller 312 .
[0106] The first controller 312 is used to: respond to the fault notification sent by the first circuit 311, monitor the storage information of the target device 322 based on the pre-stored adaptation file, obtain the monitoring results, and when it is determined that the monitoring results indicate that the storage information is abnormal, generate log information based on the preset alarm mechanism, and send the log information to the second mainboard 320.
[0107] Figure 6C A structural block diagram of a dual-control system according to another embodiment of the present application is shown.
[0108] like Figure 6C As shown, in this embodiment, the dual control system 102 includes the above Figure 6B In addition to the figure, the second mainboard 320 of the dual-control system 102 is also provided with a second controller 323 .
[0109] The second controller 323 is configured to receive log information from the first controller 312 and perform redundant storage on the log information.
[0110] Figure 6D A structural block diagram of a dual control system according to another embodiment of the present application is shown.
[0111] like Figure 6D As shown, in this embodiment, the dual-control system may include a first control mainboard 610, a second control mainboard 620 separate from the first control mainboard 610, and a backplane 410 for separating the first control mainboard 610 from the second control mainboard 620. The first control mainboard 610 and the second control mainboard 620 are both equipped with a CPU, a BMC, a CPLD, a Basic Input / Output System (BIOS), a Peripheral Component Interconnect Express (PCIe), and a Dual In-line Memory Module (DIMM).
[0112] The CPU can be equipped with registers (REG), general-purpose input / output ports (GPIO), advanced error reporting (AER) and control and status registers (CSR).
[0113] The CPU can have basic reliability, availability, and maintainability technologies, support fault handling and fault tolerance, handle hardware errors and collect and record error information, and have the ability to manage BIOS, PCIe, and DIMM hardware failures.
[0114] The BMC is the core of the fault location system, responsible for collecting, summarizing, and analyzing faults, and presenting log information to users through a web management interface.
[0115] The CPLD can be interconnected with interfaces such as the CPU, DIMM, and PCIe devices, obtain the status registers of hardware modules to capture hardware abnormalities, and then interconnect with the BMC to transmit fault notifications.
[0116] The BIOS can collect and locate faults of the CPU, DIMM, and PCIe devices, and provide the fault location results to the BMC.
[0117] The interfaces and protocols used in the dual-control system may include an interface for connecting low-speed devices (Low PinCount, LPC), a one-way communication protocol (Platform Environment Control Interface, PECI), PCIe, an interface for asynchronous serial communication (Universal Asynchronous Receiver / Transmitter, UART), I2C, SMBus, and an interface for high-speed communication (Serializer / Deserializer, SerDes).
[0118] Both the CPLD and BMC have redundant link designs. Fault notifications can be transmitted to both BMCs for monitoring and log generation to prevent fault information loss.
[0119] The dual-control system of this application can serve as a fault management system, consisting of a CPU, CPLD, BMC management, BIOS, and protocol interfaces and links that work together and interact. After a system failure occurs, it provides comprehensive fault information collection, faulty node reporting, redundant link monitoring, and faulty redundant link reporting. Furthermore, this fault management system runs on the BIOS and BMC, independent of the operating system, and remains operational at all times. Therefore, it can monitor the system at all times.
[0120] Based on the above-mentioned fault handling method based on the dual control system, this application also provides a fault handling device based on the dual control system. Figure 7 The device is described in detail.
[0121] Figure 7A structural block diagram of a fault handling device based on a dual-control system according to an embodiment of the present application is shown.
[0122] like Figure 7 As shown, the fault handling device 700 based on the dual-control system of this embodiment includes a determination module 710 and a first generation module 720 .
[0123] The determination module 710 is configured to determine the target device that has failed in the dual-control system based on the identification information of the control signal in response to receiving a change in the control signal. In one embodiment, the determination module 710 may be configured to execute the operation S210 described above, which will not be described in detail here.
[0124] The first generation module 720 is configured to generate a reset instruction, if it is determined that the target device is located on the second mainboard and cannot communicate with the second circuit disposed on the second mainboard, so as to restart the dual-control system in response to the reset instruction. In one embodiment, the first generation module 720 can be configured to perform the operation S220 described above, which will not be further described here.
[0125] According to an embodiment of the present application, the dual-control system-based fault handling device 700 further includes a control module configured to control the second circuit to generate a reset instruction when determining that the target device is located on the second mainboard and can communicate with the second circuit.
[0126] According to an embodiment of the present application, the fault handling device 700 based on the dual-control system further includes: a second generation module and a sending module. The second generation module is used to generate a fault notification indicating that a fault has occurred in the target device when it is determined that the target device is located on the first mainboard. The sending module is used to send the fault notification to the first controller disposed on the first mainboard, so that the first controller performs the following operations based on the fault notification: in response to the fault notification, the storage information of the target device is monitored based on a pre-stored adaptation file to obtain a monitoring result; and when it is determined that the monitoring result indicates that the storage information is abnormal, log information is generated based on a preset alarm mechanism.
[0127] According to an embodiment of the present application, the storage information comes from multiple registers of a target device. In response to a fault notification, the storage information of the target device is monitored based on a pre-stored adaptation file to obtain a monitoring result. The monitoring result may include the following operations: determining a monitoring strategy for the storage information of any of the multiple registers based on the types of the multiple registers; parsing association information associated with the monitoring strategy from the pre-stored adaptation file; and comparing the association information with the storage information to obtain a monitoring result.
[0128] According to an embodiment of the present application, the first controller is further configured to send log information to a second controller disposed on a second mainboard so that the second controller performs redundant storage.
[0129] According to an embodiment of the present application, the dual-control system-based fault handling device 700 further includes: a third generation module and a fourth generation module. The third generation module is configured to generate feedback information indicating that the fault of the target device is an unrepairable fault if it is determined that the number of restarts of the dual-control system is greater than a predetermined number and that the control signal changes again within a predetermined time period after the dual-control system is restarted. The fourth generation module is configured to generate feedback information indicating that the fault of the target device is a repairable fault if it is determined that the number of restarts is less than or equal to the predetermined number and that the control signal does not change within a predetermined time period after the dual-control system is restarted.
[0130] According to embodiments of the present application, any multiple modules in the determination module 710 and the first generation module 720 can be combined into a single module, or any one of them can be split into multiple modules. Alternatively, at least part of the functionality of one or more of these modules can be combined with at least part of the functionality of other modules and implemented in a single module. According to embodiments of the present application, at least one of the determination module 710 and the first generation module 720 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware through any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the determination module 710 and the first generation module 720 can be at least partially implemented as a computer program module that, when executed, can perform the corresponding functionality.
[0131] This application also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of this application is implemented.
[0132] According to an embodiment of the present application, a computer-readable storage medium may be a non-volatile computer-readable storage medium, such as, but not limited to, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present application, a computer-readable storage medium may include one or more memories other than the ROM described above.
[0133] The embodiments of the present application also include a computer program product, which includes a computer program containing program code for executing the method shown in the flowchart. When the computer program product is run in a computer system, the program code is used to enable the computer system to implement the method provided in the embodiments of the present application.
[0134] When the computer program is executed by the processor, the above functions defined in the system / device of the embodiment of the present application are performed. According to the embodiment of the present application, the system, device, module, unit, etc. described above can be implemented by a computer program module.
[0135] In one embodiment, the computer program may be stored on a tangible storage medium, such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal over a network medium, downloaded and installed via a communication component, and / or installed from a removable medium. The program code contained in the computer program may be transmitted using any suitable network medium, including but not limited to wireless, wired, or any suitable combination thereof.
[0136] In such an embodiment, the computer program can be downloaded and installed from a network via the communication portion, and / or installed from a removable medium. When the computer program is executed by the processor, the above-mentioned functions defined in the system of the embodiment of the present application are performed. According to the embodiment of the present application, the systems, devices, means, modules, units, etc. described above can be implemented by computer program modules.
[0137] According to an embodiment of the present application, the program code for executing the computer program provided by the embodiment of the present application can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).
[0138] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of the boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0139] Those skilled in the art will appreciate that the features described in the various embodiments of this application may be combined and / or coupled in various ways, even if such combinations or couplings are not explicitly described in this application. In particular, the features described in the various embodiments of this application may be combined and / or coupled in various ways without departing from the spirit and teachings of this application. All such combinations and / or couplings fall within the scope of this application.
[0140] The embodiments of the present application have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present application. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present application, those skilled in the art may make various substitutions and modifications, and these substitutions and modifications should all fall within the scope of the present application.
Claims
1. A fault handling method based on a dual control system, characterized in that: The dual-control system includes a first and a second separate mainboard, and the fault handling method is applied to a first circuit arranged on the first mainboard; The fault handling method includes: In response to receiving a change in a control signal, determining a target device in the dual-control system where a fault occurs based on identification information of the control signal; When it is determined that the target device is located on the second mainboard and cannot communicate with the second circuit disposed on the second mainboard, a reset instruction is generated to restart the dual-control system in response to the reset instruction.
2. The fault handling method according to claim 1, characterized in that: The fault handling method further includes: When it is determined that the target device is located on the second mainboard and can communicate with the second circuit, the second circuit is controlled to generate the reset instruction.
3. The fault handling method according to claim 1, characterized in that: The fault handling method further includes: generating a fault notification indicating that a fault has occurred in the target device if it is determined that the target device is located on the first mainboard; The fault notification is sent to a first controller disposed on a first mainboard, so that the first controller performs the following operations based on the fault notification: In response to the fault notification, based on a pre-stored adaptation file, monitoring the storage information of the target device to obtain a monitoring result; When it is determined that the monitoring result indicates that the stored information is abnormal, log information is generated based on a preset alarm mechanism.
4. The fault handling method according to claim 3, characterized in that: The storage information comes from a plurality of registers of the target device; In response to the fault notification, monitoring the storage information of the target device based on a pre-stored adaptation file to obtain a monitoring result includes: determining, according to types of the plurality of registers, a monitoring strategy for storage information of any register among the plurality of registers; parsing association information associated with the monitoring strategy from the pre-stored adaptation file; The associated information is compared with the stored information to obtain the monitoring result.
5. The fault handling method according to claim 3, characterized in that: The first controller is further configured to: The log information is sent to a second controller disposed on a second mainboard so that the second controller performs redundant storage.
6. The fault handling method according to any one of claims 1 to 5, characterized in that: The fault handling method further includes: If it is determined that the number of restarts of the dual-control system is greater than a predetermined number, and the control signal changes again within a predetermined time period after the dual-control system is restarted, generating feedback information indicating that the fault of the target device is an unrepairable fault; When it is determined that the number of restarts is less than or equal to the predetermined number and the control signal does not change within the predetermined time period after the dual-control system is restarted, feedback information indicating that the fault of the target device is a repairable fault is generated.
7. A dual control system, characterized in that: The dual control system includes: a first mainboard, wherein the first mainboard is provided with a first circuit; a second mainboard provided separately from the first mainboard, the second mainboard being provided with a second circuit, and a communication link being provided between the first circuit and the second circuit; Wherein, the first circuit is used for: In response to a change in the received control signal, determining a target device in the dual-control system where a fault occurs based on identification information of the control signal; When it is determined that the target device is located on the second mainboard and cannot communicate with the second circuit, a reset instruction is generated to restart the dual-control system in response to the reset instruction.
8. The dual control system according to claim 7, characterized in that: The first mainboard is also provided with a first controller; The first controller is used for: In response to the fault notification sent by the first circuit, based on a pre-stored adaptation file, monitoring the storage information of the target device to obtain a monitoring result; When it is determined that the monitoring result indicates that the storage information is abnormal, log information is generated based on a preset alarm mechanism, and the log information is sent to the second mainboard.
9. The dual control system according to claim 8, characterized in that: The second mainboard is further provided with a second controller; The second controller is used to receive log information from the first controller and perform redundant storage on the log information.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the fault handling method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Fault repair method and system based on dual-control storage
CN107967195A
Equipment fault monitoring method and device, equipment and readable storage medium
CN114356708A
Server hardware power-on starting fault troubleshooting method, system and device and medium
CN115470056A
Fault memory positioning method, system and device, computer equipment and storage medium
CN117234771A
Cited By
Parallel interface module and data physical link fault tolerance method thereof
CN122132220A