Hardware-based non-transitory memory device, method, and apparatus for a self-healing control system

By automatically detecting and repairing the failures of logic blocks and memory blocks in the process control system, the automatic control loss caused by transient failures is solved, and self-healing without manual intervention is achieved to ensure the continuous and normal operation of the system.

CN114253225BActive Publication Date: 2025-08-05HONEYWELL INTERNATIONAL INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111079170.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-23
Filing Date
2021-09-15
Publication Date
2025-08-05
Estimated Expiration
2041-09-15

AI Technical Summary

Technical Problem

When existing process control systems face transient failures, they require manual intervention to restore redundant operations, resulting in the loss of automatic control, and traditional methods cannot effectively solve the problem of memory online bus failure.

Method used

By monitoring the logic blocks and memory blocks, the fault condition is automatically detected, the affected subset is determined, and it is exchanged into a second logic or memory block that can perform the corresponding functions, the fault block is disabled, and the corresponding actions are automatically repaired and performed to achieve self-healing.

Benefits of technology

It realizes the automatic recovery of redundant operation of the process control system without manual intervention, repairs transient failures, and ensures that the system continues to operate normally.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253225B_ABST
    Figure CN114253225B_ABST
Patent Text Reader

Abstract

The present invention is entitled “Self-Healing Process Control System.” One embodiment of the present invention is one or more hardware-based non-transitory memory devices storing computer-readable instructions that, when executed by one or more processors disposed in a computing device, cause the computing device to: monitor logic blocks and memory blocks to detect a fault condition; determine a subset of logic blocks or memory blocks affected by the fault condition; and perform at least one action on the logic blocks and memory blocks.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application incorporates by reference in its entirety application serial number 16 / 377,237, filed on April 7, 2019 (the “HIVE Patent”), entitled “CONTROL HIVE ARCHITECTURE ENGINEERING EFFICIENCY FOR AN INDUSTRIAL AUTOMATION SYSTEM,” and application serial number 16 / 502,692, filed on July 3, 2019 (the “Redundant Patent”), entitled “REDUNDANT CONTROLLERS OR INPUT-OUTPUT GATEWAYS WITHOUT DEDICATED HARDWARE.” Background Art

[0003] Process control systems are used across a wide range of industries to achieve consistent, efficient, and safe production levels that would be unattainable with manual control alone. Process control systems are widely used in industries such as oil refining, pulp and paper manufacturing, chemical processing, and power plants. For example, a process control system can be located near the plant floor and manipulate valves, start motors, and operate chillers or mixing units as needed.

[0004] Process control systems are required to diagnose control device failures to prevent unmanaged issues such as inappropriate control actions. One example is a watchdog timer. A watchdog timer can take action, such as placing the machine in its defined fault state, if the timer expires due to a hardware or, in some cases, software failure. For example, it might initiate the closing of a valve, stop a motor, or maintain the last controlled state of the equipment. Maintenance engineers can replace the faulty component that caused the watchdog timer to expire with a best-fit replacement unit. Software is used to refresh the watchdog timer, letting the electronics know that the processor and its executable code are still functioning properly. This is a traditional approach that requires manual human intervention to restore full operation, including redundancy of the control electronics.

[0005] A problem arises when there is a transient fault, such as a bus fault on the memory line between the program and the processor. In this case, a software refresh of the watchdog timer will not solve the problem. One solution to this problem is to use redundancy, but when the fault occurs, the control device is in a non-redundant operating mode, so that a second fault will result in a loss of automatic control. It is preferable to self-heal the faulty device and restore full redundancy (if redundant) or restore / maintain automatic operation even when non-redundant. Summary of the Invention

[0006] One specific implementation is a method comprising: monitoring logic blocks and memory blocks to detect a fault condition; determining a subset of logic blocks or memory blocks affected by the fault condition; swapping the subset of logic blocks or memory blocks with a second logic block or second memory block that is capable of performing the functions of the subset of logic blocks or memory blocks; disabling the subset of logic blocks or memory blocks; automatically repairing the subset of logic blocks or memory blocks; and performing at least one action with respect to the subset of logic blocks or memory blocks.

[0007] Another specific implementation is directed to a device comprising: a monitoring module configured to monitor logic blocks and storage blocks to detect a fault condition; a fault location module configured to determine a subset of logic blocks or storage blocks affected by the fault condition; a swap module configured to swap the subset of logic blocks or storage blocks into a highly integrated virtual environment (HIVE) capable of performing functions of the subset of logic blocks or storage blocks; an orchestrator module configured to disable the subset of logic blocks or storage blocks; and a repair module configured to perform at least one action with respect to the subset of logic blocks or storage blocks.

[0008] Another specific implementation is directed to one or more hardware-based non-volatile memory devices that store computer-readable instructions that, when executed by one or more processors disposed in a computing device, cause the computing device to: monitor logic blocks and memory blocks to detect fault conditions; determine a subset of logic blocks or memory blocks affected by the fault conditions; and perform at least one action on the logic blocks and memory blocks. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 It is a simplified block diagram of an industrial automation system.

[0010] Figure 2 It is a simplified block diagram of a self-healing control system.

[0011] Figure 3 It is a simplified block diagram of a self-healing control system.

[0012] Figure 4 It is a simplified block diagram of a self-healing control system.

[0013] Figure 5 is a flow chart illustrating current usage of a self-healing process control system.

[0014] Figure 6 is a flow chart illustrating current usage of a self-healing process control system.

[0015] Figure 7 is a flow chart illustrating current usage of a self-healing process control system. DETAILED DESCRIPTION

[0016] Figure 1 An exemplary industrial automation system 100 according to the present disclosure is shown. Figure 1 As shown, the system 100 includes various components that facilitate the production or processing of at least one product or other material. For example, the system 100 is used herein to facilitate the control of components in one or more plants 101a to 101n. Each plant 101a to 101n represents one or more processing facilities (or one or more portions thereof), such as one or more manufacturing facilities for producing at least one product or other material. Generally speaking, each plant 101a to 101n can implement one or more processes and can be referred to individually or collectively as a process system. A process system generally represents any system or portion thereof that is configured to process one or more products or other materials in some manner. In Figure 1 In FIG. 1 , system 100 is implemented using the Purdue model of process control. However, it should be noted that an architecture that conforms to the Purdue model is not required. A flat architecture with substantially all of its components on one communication network may also be used.

[0017] In addition, some circuit components and designs allow hardware reset. This can include a complete power reset of the selected subsystem at any level (see, for example, levels 1-5 of the Purdue model or other process control systems). By way of example, as electronic design component geometries have shrunk and are arranged in multiple layers, new categories of faults are resistant to classic online correction (e.g., memory parity EDAC and ECC), but are transient faults that do not reappear once corrected. One example is bit latching, where one or more RAM bits latch and remain in an incorrect and uncorrectable state without a complete hardware circuit power cycle. Another example is a soft error in the gate or memory of a field programmable logic (FPL), field programmable gate array (FPGA), or system-on-chip device (SoC), which also remains in an incorrect and uncorrectable state without a complete hardware circuit power cycle.

[0018] In previous designs, such faults would force the equipment into a fault state requiring manual service (e.g., physical equipment removal and replacement). As process control and other industrial automation equipment become not only more integral to manufacturing operations, but also physically remote from human central control rooms, such transient faults must be repaired in an autonomous and self-healing manner at any level of any plant 101a-101n, including machine controllers 114, cell controllers 122, plant controllers 130, or any other controller.

[0019] Figure 2 is a simplified block diagram of the self-healing control system 200 . Figure 2 The above mentioned transient fault categories can be corrected in a manner that does not require human intervention. Figure 2 The self-healing control system 200 is coupled to a subset of the system 205 and the rest of the system 225. The subset of the system 205 includes a logic block 280 and a memory block 290. A fault condition may occur in either or both of the logic block 280 and the memory block 290. For simplicity, the fault condition may be referred to as occurring in the subset of the system 205, but this means that the fault condition occurs in one or both of the logic block 280 and the memory block 290.

[0020] The self-healing control system includes a monitoring module 235. In operation, monitoring module 235 waits until a fault occurs in subset 205 of the system. For example, subset 205 of the system may enter a faulty state and have a stuck bit. In this case, monitoring module 235 detects the presence of a fault in the system. Fault localization module 240 determines that subset 205 of the system has entered a faulty state. In this example, multiple additional subsets 206 of the system have not entered a faulty state. The task of fault localization module 240 is to identify which of the subsets has entered a faulty state. Subsets 205 and 206 of the system include circuitry, software, memory, and / or firmware, where a stuck bit is possible in some error states.

[0021] When one of the system's subsets 205 or 206 enters a fault condition (in this case, assuming subset 205 is faulty), the repair module corrects the fault condition and returns the system to a normal state. For example, repair module 230 can automatically reboot subset 205 from its stored program memory. In this way, the fault can be cleared by powering on the memory subsystem of the faulty subset 205 of the system. In another example, repair module 230 can repair the faulty node in the system's subset 205 with the stuck bit by resetting the memory and restoring database 299. Once repair module 230 repairs the faulty node, it can return the system to an operational state.

[0022] The control system 200 may also write to a fault log 250 and an event log 260. In one example, when the control system 200 automatically repairs the subset 205, it also updates the fault log 250. The fault log 250 may be sent to the database 299 so that there is a record of every action taken by the control system 200. In the example where the control system 200 performs successive repairs on the subset 205 of the system and / or the same fault, it updates the event log 260. The event log 260 may be used by a human operator to understand that a portion of the overall system has degraded behavior and may need to take additional action.

[0023] Figure 3is a simplified block diagram of the self-healing control system 300 . Figure 3 The above mentioned transient fault categories can be corrected in a manner that does not require human intervention. Figure 3 The self-healing control system 300 is coupled to a primary device or system 305 and an auxiliary device or system 306. The primary device or system 305 and the auxiliary device or system 306 are coupled to the rest of the system 325. The self-healing control system includes a monitoring module 235. In operation, the monitoring module 235 waits until a failure occurs in a subset of the system. In this example, it is assumed that the primary device or system 305 enters a failed state.

[0024] For example, the primary device or system 305 has a stuck position. In this case, the monitoring module 235 detects that there is a fault in the system. The fault location module 240 determines that the primary device or system 305 has entered a fault state. Because the primary device or system 305 is in a fault state, the switching module 345 replaces the primary device or system 305 with an auxiliary device or system 306. The orchestrator module 350, which can distribute workloads based on the availability and / or health of the nodes, disables the primary device or system 305 so that it is no longer used by the rest of the system when it is in a fault state. In this way, the auxiliary device or system 306 becomes the primary device, and the primary device or system 305 is taken offline for repair. The auxiliary device or system 306 is able to replace the functionality of the failed node and allow the system to continue normal operation.

[0025] Repair module 230 is configured to repair master device or system 306 while it is offline. Repair module 230 may, for example, force master device or system 305 to power cycle the memory area where the fault is identified. In one example, repair module 330 may repair the faulty node using a user command to reset or repair the faulty node. In other examples, the node may be repaired automatically by repair module 330. This eliminates the need for someone to physically interact with the node, which is particularly useful when the control node is located in a processing area and it is desirable not to have to travel to the node. This may also eliminate the need to open the cabinet enclosure, which may require a work permit.

[0026] In another example, the master device or system 305 can be automatically rebooted from its stored program memory. Thus, the fault can be cleared by the repair module 230 powering on the memory subsystem of one of the failed nodes. In another example, the repair module 230 can repair the failed node with a stuck bit by resetting the memory and restoring the database 399. The error correction module 355 can also be used. For example, once the repair module 330 reloads the database 399, the error correction module 355 can use internal checkpoints and diagnose whether the flash memory and database checksums are intact. Once the repair module 330 repairs the failed node, it can return the system to an operational state, with the repaired failed node acting as a secondary subsystem and the current replacement node acting as a primary node. For more detailed information on how a redundant system can operate, see the redundancy patent, which is omitted here for brevity.

[0027] The control system 300 may also write to the fault log 250 and the event log 260. In one example, when the control system 300 automatically repairs a fault, it also updates the fault log 250. The fault log 250 may be sent to the database 399 so that there is a record of every action taken by the control system 300. In the example where the control system 300 performs consecutive repairs and / or the same fault, it updates the event log 260. The event log 260 may be used by a human operator to understand that a portion of the overall system is exhibiting degraded behavior and may require additional action.

[0028] Figure 4 is a simplified block diagram of the self-healing control system 400 . Figure 4 The self-healing control system 400 is coupled to a subset 410 of the system and an HIVE 420. The subset 410 of the system is coupled to the rest of the system 430. The self-healing control system 400 includes a monitoring module 235. In operation, the monitoring module 235 waits until a fault occurs in the subset 410 of the system. For example, the subset 410 of the system may have a stuck bit. In this case, the monitoring module 235 detects that a fault exists in the system. The fault location module 240 determines that the subset 410 of the system has entered a fault state.

[0029] In one example, HIVE functional block 405 includes cooperating control nodes 440, 450, and 460. A node with an uncorrectable fault will have its workload shunted as per the HIVE redundancy design and then automatically rebooted with a power cycle of the memory region with the bit fault. This returns the failed node to operational mode. The orchestrator module 350 can then distribute the workload back to the subset 410 of the system as needed by the HIVE. For more details on how the HIVE can operate, please refer to the HIVE patent, which is omitted here for the sake of brevity. However, it should be noted that the repaired node 410 can return to an operational state as a HIVE component and will generally not receive the same workload it had before the fault condition occurred. In a manner similar to that described with respect to Figure 3 In the manner described, once repaired, the failed primary device can be maintained as a secondary device. In the case of HIVE, once repaired, the failed node can be incorporated into HIVE and will present a new state that is different from the old state before the node failed.

[0030] To this end, the repair module 230 is configured to repair the subset 410 of the system while it is offline. The control system 400 may also write to the fault log 250 and the event log 260. In one example, when the control system 400 automatically repairs a fault, it also updates the fault log 250. The fault log 250 may be sent to the historical database 499 so that there is a record of every action taken by the control system 300. In the example where the control system 400 performs consecutive repairs and / or the same fault, it updates the event log 260. The event log 260 may be used by human operators to understand that a portion of the overall system is exhibiting degraded behavior and may require additional action.

[0031] Figure 5 is a flow chart illustrating the current use of a self-healing process control system. At step 500, the system determines whether a fault condition exists. Typically, the system will be operating normally, and any nodes within the system that could have a transient fault category (such as a stuck bit) will not have this condition. Therefore, step 500 will continue until a fault is detected. At step 510, once the system detects a fault, it uses logic to determine where the fault is within the larger system.

[0032] At step 520, the faulty subset of the system is disabled. At step 530, the system determines whether the subset of the system has been repaired. This may occur, for example, via a repair module or other mechanism that isolates the faulty subset of the system and automatically causes it to be repaired. One way this may occur is by forcing a power cycle on the memory area where the fault was identified. In another example, the repair module may use a user command to reset or repair the faulty node to repair the faulty node. In yet another example, one of the faulty nodes may be automatically rebooted from its stored program memory. In this way, the fault can be cleared by the repair module powering on the memory subsystem of one of the faulty nodes. In yet another example, the repair module may repair the faulty node using a stuck bit by resetting the memory and restoring the database. Once the system is repaired, actions are taken, which may include updating the fault log and / or event log at step 540.

[0033] Figure 6 is a flow chart illustrating the current use of a self-healing process control system. At step 600, the system determines whether a transient fault exists, such as a soft error, a stuck bit in field programmable logic (FPL), a stuck bit in a field programmable gate array (FPGA), or a stuck bit in a system-on-chip device (SoC). Thus, step 600 continues until a fault is detected in the primary system. Once the system detects a fault, the system uses logic at step 610 to swap a secondary node with the faulty primary node having the transient fault.

[0034] When a secondary node is swapped for a failed primary node, the primary node is disabled at step 620. At step 640, at least one action is performed on the failed node. In one example, this includes a memory reset and / or database operation on the failed node. The database operation can be, for example, a synchronization with the database of the active primary device. Thereafter, at step 640, the system determines whether the failed node has been repaired. Step 640 is repeated until the failed node is repaired, in which case, at step 650, the repaired failed node is used as a secondary node for the current primary node. In other examples, the repaired failed node can become the primary node. Other schemes are also possible.

[0035] Figure 7is a flow chart illustrating current use of a self-healing process control system. At step 700, the system determines whether a node has a stuck or latched bit. Step 700 is repeated until a fault is detected. Once the system detects a stuck or latched bit in a node, the orchestrator can assign the workload of the failed node to a HIVE at step 710. At step 720, the orchestrator can disable the failed node. At step 730, the system performs memory operations and / or database operations on the failed node. This can include, for example, forcing a power cycle in the memory region where the fault was identified, rebooting from its stored programming memory, resetting the memory and restoring the database, and other options.

[0036] At step 740, the system determines whether the failed node has been repaired. If not, step 740 is repeated. After the failed node is repaired, in one example, it can be used as part of the HIVE in the future and workloads can be assigned as needed by the orchestrator. It should be noted that after the failure condition is repaired, the repaired node typically has a different workload than before and is integrated into the HIVE. It may no longer be necessary for the repaired node to continue operating with the same workload it had before the failure.

[0037] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.

Claims

1. One or more hardware-based non-transitory memory devices storing computer-readable instructions that, when executed by one or more processors provided in a computing device included in a self-healing control system, cause the computing device to: monitoring logic blocks (280) and memory blocks (290) included in the subset of the self-healing control systems to detect a stuck bit fault condition associated with the subset of the self-healing control systems; determining the subset of the logic blocks (280) or the memory blocks (290) affected by the fault condition, wherein nodes including the logic blocks or the memory blocks affected by the fault condition are determined to be faulty nodes in a faulty state; exchanging the logic block or the subset of the memory blocks with a second logic block or a second memory block, the second logic block or the second memory block being capable of performing the functions of the logic block or the subset of the memory blocks; disabling the subset of the logic blocks or the storage blocks when the faulty node is in the faulty state; Automatically repairing the subset of the logic blocks or the memory blocks affected by the fault condition, wherein the repaired subset of logic blocks or memory blocks returns the failed node to an operational mode, wherein the automatic repair comprises at least one of the following: providing a power cycle on a memory region in the memory block affected by the fault condition, Reset by the repair module to repair the faulty node based on user commands, rebooting the failed node from the memory block to clear power to the memory subsystem of one of the failed nodes through the repair module, and Repairing the failed node by resetting the storage block and restoring the database; as well as performing at least one action with respect to the subset of the logic block (280) and the memory block (290), including updating a fault log and / or an event log to record each action taken by the self-healing control system, wherein performing at least one action causes the computing device to: disabling the logic block (280) and the memory block (290); repairing the logic block (280) and the memory block (290); and The logic block (280) and the memory block (290) are enabled.

2. The one or more hardware-based non-transitory memory devices of claim 1, wherein the memory device is selected from the group consisting of: a field programmable logic, a field programmable gate array, or a system-on-chip device.

3. The one or more hardware-based non-volatile memory devices of claim 2, wherein the failure condition is selected from the group consisting of: a soft error, a stuck bit in the field programmable logic, the field programmable gate array, or the system-on-chip device.

4. A method for a self-healing control system, comprising: monitoring logic blocks (280) and memory blocks (290) included in the subset of the self-healing control systems to detect a stuck bit fault condition associated with the subset of the self-healing control systems; determining the subset of the logic blocks (280) or the memory blocks (290) affected by the fault condition, wherein nodes including the logic blocks or memory blocks affected by the fault condition are determined to be faulty nodes in a faulty state; as well as exchanging the subset of the logic block (280) or the memory block (290) for a second logic block or a second memory block, the second logic block or the second memory block being capable of performing the functions of the subset of the logic block (280) or the memory block (290); disabling the subset of the logic blocks (280) or the memory blocks (290); Automatically repairing the subset of the logic blocks (280) or the memory blocks (290) affected by the fault condition, wherein the repaired subset of logic blocks or memory blocks returns the failed node to an operational mode, wherein the automatic repair comprises at least one of: providing a power cycle on a memory region in the memory block affected by the fault condition, Reset by the repair module to repair the faulty node based on user commands, rebooting the failed node from the memory block to clear power to the memory subsystem of one of the failed nodes through the repair module, and Repairing the failed node by resetting the storage block and restoring the database; as well as performing at least one action with respect to the subset of the logic block (280) or the memory block (290), including updating a fault log and / or an event log to record each action taken by the self-healing control system, Wherein performing at least one action comprises: disabling the logic block (280) and the memory block (290); repairing the logic block (280) and the memory block (290); and The logic block (280) and the memory block (290) are enabled.

5. The method according to claim 4, further comprising: designating the logic block (280) and the storage block (290) as a master system; as well as The second logic block and the second storage block are designated as a secondary system.

6. The method of claim 5, wherein the performing at least one action further comprises maintaining the secondary system as the primary system after the subset of the logic blocks (280) or the memory blocks (290) are repaired.

7. The method of claim 4, wherein the logic block (280) and the memory block (290) are incorporated into one or more devices selected from the group consisting of: field programmable logic, field programmable gate array, or system on chip devices.

8. A device for a self-healing control system, comprising: a monitoring module configured to monitor logic blocks (280) and memory blocks (290) included in the subset of the self-healing control system to detect a stuck bit fault condition associated with the subset of the self-healing control system; a fault localization module configured to determine a subset of the logic blocks (280) or the memory blocks (290) affected by the fault condition, wherein nodes including the logic blocks or memory blocks affected by the fault condition are determined to be faulty nodes in a faulty state; and a switching module configured to switch the subset of the logic block (280) or the storage block (290) to a highly integrated virtual environment HIVE, the HIVE being capable of performing the functions of the subset of the logic block (280) or the storage block (290); an orchestrator module configured to disable the subset of the logic blocks (280) or the memory blocks (290); and A repair module, wherein the repair module is configured to: Automatically repairing the subset of the logic blocks or the memory blocks affected by the fault condition, wherein the repaired subset of logic blocks or memory blocks returns the failed node to an operational mode, wherein the automatic repair comprises at least one of the following: providing a power cycle on a memory region in the memory block affected by the fault condition, Reset by the repair module to repair the faulty node based on user commands, rebooting the failed node from the memory block to clear power to the memory subsystem of one of the failed nodes through the repair module, and Repairing the failed node by resetting the storage block and restoring the database; as well as performing at least one action with respect to the subset of the logic block (280) or the memory block (290), including updating a fault log and / or an event log to record each action taken by the self-healing control system, wherein performing at least one action causes the repair module to: disabling the logic block (280) and the memory block (290); repairing the logic block (280) and the memory block (290); and The logic block (280) and the memory block (290) are enabled.

9. The apparatus of claim 8, wherein the at least one action comprises automatically repairing the subset of the logic blocks (280) or the memory blocks (290).

Citation Information

Patent Citations

  • Redundant controllers or input-output gateways without dedicated hardware

    US11481282B2

  • Control hive architecture engineering efficiency for an industrial automation system

    US20200319623A1

  • Memory redundancy and recovery from uncorrectable errors

    GB2404261A