Method for operating an automated system with running redundancy

By cached synchronous data in the second failsafe subsystem in the redundant automated system and confirming that there is no fault-free post-processing, the problem of synchronous data error caused by the failure of the first failsafe subsystem is solved, and the reliability and availability of the system are improved.

CN114253766BActive Publication Date: 2025-08-05SIEMENS AG
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111100916.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-22
Filing Date
2021-09-18
Publication Date
2025-08-05
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

In a redundant automation system, when the first failsafe subsystem fails, synchronous data transmission errors may occur, resulting in problems with the second failsafe subsystem and failure of the system control.

Method used

The synchronous data is cached in the second fail-safe subsystem and post-processing is carried out through the failure-free message confirmation to ensure the accuracy of the synchronous data and avoid system control interference caused by wrong data.

Benefits of technology

It reduces the system response time, improves the reliability and availability of the system, avoids control interference caused by wrong data, and ensures the stable operation of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114253766B_ABST
    Figure CN114253766B_ABST
Patent Text Reader

Abstract

The present application relates to a method for operating a redundant automation system (100) for controlling a technical process, wherein a second fail-safe subsystem (2) is operated redundantly for a first fail-safe subsystem (1), a faulty second fail-safe subsystem (2) is used, synchronization data (SD) are first buffered in the second subsystem (2), and in the case of a detected no-fault situation, the first fail-safe subsystem (1) sends a no-fault message (FFOK) to the second fail-safe subsystem (2), whereupon the second fail-safe subsystem receives the no-fault message (FFOK) with a no-fault certificate (FFQ) and first processes the buffered synchronization data (SD).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a method for operating a redundant automation system for controlling a technical process, wherein a first fail-safe subsystem is operated using a first control program divided into first program segments, wherein the technical process is controlled by the first fail-safe subsystem, and a second fail-safe subsystem is operated redundantly using a second control program divided into second program segments, wherein the first fail-safe subsystem generates and evaluates events using its first control program, influences the execution sequence of first program segments in the first control program depending on the events that occur, and provides synchronization data, provided with an index reflecting the respective program segment, to the second fail-safe subsystem per program segment based on the generated or occurring events, and provides raw data to an output component, wherein the raw data is initially retained and not yet written to the output component, wherein the first subsystem, by processing the respective first program segment, precedes the second subsystem with respect to the index with respect to the processing of the respective second program segment of the second subsystem. Background Art

[0002] In the automation sector, there is an increasing demand for high-availability solutions (H systems) that are suitable for minimizing possible downtimes of technical equipment. H systems used in automation are characterized by two or more subsystems in the form of automation devices or computing systems, coupled to each other via a synchronous connection. In principle, both subsystems have readable and / or writable access to peripheral units connected to the H system. One of the two subsystems is dominant with respect to the peripherals connected to the system. This means that outputs to or for peripheral units are processed by only one of the two subsystems, meaning that this subsystem acts as the master or takes over the master function.

[0003] In addition to the reliability of automation systems, they often also have to implement safety-critical functions in automation technology. Such automation systems are referred to below as HF systems. The two redundantly operating subsystems (HF-CPU1 and HF-CPU2) can now also process safety-related control programs. Furthermore, the control programs communicate via safety-related protocols with the likewise safety-related input and output modules (F-IO). Summary of the Invention

[0004] In the context of the present invention, synchronization data is used to synchronize events occurring on the first subsystem with the second subsystem. This allows the second subsystem to immediately take over the execution of the controlled technical process in the event of a failure in the first subsystem. Without knowledge of events occurring on the first subsystem, the execution sequence of program segments on the second subsystem cannot be maintained. Only when the second subsystem is aware of events on the first system can it execute the required program sequence without disrupting the method for controlling the technical process.

[0005] Furthermore, within the meaning of the present invention, a safe HF-CPU' or a corresponding fail-safe subsystem can diagnose faults during the processing of its control program and subsequently put the device or the corresponding subsystem into a safe state.

[0006] The present invention addresses the following problem when implementing a redundantly operating automation system based on two fail-safe subsystems (two HF CPUs). The first fail-safe subsystem operates as the leading HF CPU and detects a fault (e.g., caused by a hardware failure) in its local processing. It then deactivates itself to take advantage of the fault. However, a further fault can occur during the already ongoing transmission of synchronization data from the first fail-safe subsystem to the second fail-safe subsystem. In other words, the leading HF CPU provides the trailing HF CPU with incorrect synchronization data. In response, the trailing HF CPU or the second fail-safe subsystem then also detects a problem in its processing after a short time and deactivates. The problem is that the two now-deactivated subsystems (two HF CPUs) lead to a loss of control of the system. However, the fault response of the trailing HF CPU is solely due to the synchronization of the faulty data. Nevertheless, the trailing HF CPU can continue to control the process without any problems.

[0007] Accordingly, the object of the present invention is to provide a method for operating a redundant automation system having a first fail-safe subsystem and a second fail-safe subsystem, which ensures that, in the event of a failure of the first fail-safe subsystem, no erroneous synchronization data is provided to the second fail-safe subsystem.

[0008] According to the method, the object is achieved by first buffering synchronization data in the second subsystem, wherein the first fail-safe subsystem first sends the provided raw data with an index of the corresponding first program segment to the second fail-safe subsystem, and the second fail-safe subsystem confirms this by issuing a certificate to the first fail-safe subsystem, in the first fail-safe subsystem, a fault check is carried out at the end of the corresponding program segment, using the fault check to check the error-free process of the first control program in the corresponding program segment, in the case of a detected fault, the first fail-safe subsystem is deactivated, and the control of the technical process is performed by the second fail-safe subsystem, in the case of a detected error, the first fail-safe subsystem sends a error-free message to the second fail-safe subsystem, then the second fail-safe subsystem signs the error-free message with the error-free certificate and first processes the buffered synchronization data of the second program segment with the matching index, and using the received error-free certificate, the first raw data is written to the output component via the first fail-safe subsystem.

[0009] The first fail-safe subsystem can also be regarded as the preceding system, and the second fail-safe subsystem can be regarded as the subsequent system accordingly. If the output of the raw data is now executed by the preceding system, then the output is first signed for by the subsequent system. This ensures that the output will not be affected when the preceding system fails. The subsequent system stores the synchronization data locally in, for example, an IIFO memory, but does not process it first. Only after a fault check has been performed to confirm the trouble-free processing of the corresponding program segment, the preceding system performs the output of the raw data to the peripheral device. The fault check is usually performed on the preceding system at the end of each program segment (for example, segment n).

[0010] If the check determines that there are no errors, the preceding system notifies the following system with a special message ("F-check OK"). This message prompts the following system to process the previously stored synchronization data and, for its part, also perform an F-check at the end of segment n.

[0011] The preceding system outputs the raw data to the peripheral device via the following system according to the fault-free certificate. Preferably, the next section is started only when the following system also completes section n and it is signed by the preceding system using the cycle certificate.

[0012] In this case, it is advantageous if the second subsystem sends a cycle certificate to the first subsystem, the first subsystem confirming that the second program section associated with the index was successfully processed without errors.

[0013] In the event that a fault is detected on the first fail-safe subsystem, the second fail-safe subsystem rejects all synchronization data, which were stored after the last fault-free certificate, and takes over control of the process for standalone operation.

[0014] If, for example, the first subsystem displays a fault in the next segment n+1, it immediately deactivates itself. The second subsystem now rejects all synchronization data after the last "F-check passed" message in the FIFO memory. Consequently, the second subsystem switches to independent operation. Consequently, the previously subsequent HF-CPU, i.e., the second subsystem, now becomes the dominant subsystem and now executes segment n+1, resulting in a correspondingly longer reaction time in the event of a fault in the preceding system. Some of this longer reaction time can be reduced by the following solution. The advantage is that, in the second subsystem, synchronization data is immediately processed using a second program segment with a matching index, regardless of a fault-free message. In the event that a fault-free message reaches the second fail-safe subsystem, a safety backup of the program state is also stored in the memory image. In the event that a fault-free message is missing and the first subsystem fails, the last backup of the program state is downloaded from the memory image, and program processing continues using this program state, with the second fail-safe subsystem controlling the technical process.

[0015] Response time is optimized by the fact that the subsequent system immediately processes the received synchronization data without waiting for the "F Check Passed" message. However, to avoid delays in the transmission of errors from the preceding system to the subsequent system, as soon as the subsequent system receives the "F Check Passed" message, it stores the current state, i.e., the memory image of the program running approximately synchronously. Locally storing the program state via the memory image is very efficient, while the subsequent system simultaneously deletes the memory image of the previous segment. As soon as the subsequent system reaches the end of a program segment, it confirms this to the preceding system. This allows the next program segment to be started immediately. This speeds up the system flow and thus reduces response time.

[0016] If the subsequent system detects a fault in the preceding system, it loads the last backed-up memory image and processes the process values / events independently in a separate operation. Accordingly, the already received and processed synchronization data from step n+2 is rejected, as it could contain potentially erroneous instructions.

[0017] It has also proven advantageous if the synchronization data can be transferred from the preceding system to the succeeding system asynchronously in time. This decouples the processing power of the preceding system from the communication bandwidth available for event synchronization, which can conflict with the increasing imbalance between the increasing processing power of the relevant processors and the increasing number of communication processors.

[0018] Due to the time-asynchronous communication between the preceding and succeeding systems, it is possible to create a high-availability automation system even using slower communication connections. This means that a communication connection with poor transmission bandwidth or response time can also be set up, which is also used by other communication participants and thus does not conflict with the two participants provided for the purpose of synchronization. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The present invention, its configuration and advantages are explained in detail below based on the accompanying drawings which illustrate exemplary embodiments of the present invention.

[0020] Figure 1 It is a redundant automation system based on existing technology,

[0021] Figure 2 is a sequence of steps according to the method for a redundant automation system in a first alternative embodiment,

[0022] Figure 3 is a sequence of steps according to a method for a redundant automation system according to a second alternative embodiment,

[0023] Figure 4 This is the sequence for a redundant automation system according to a third alternative configuration. DETAILED DESCRIPTION

[0024] Figure 1 A redundant automation system 100 for controlling a technical process is shown. According to the prior art, a first fail-safe subsystem 1 and a second fail-safe subsystem 2 are coupled via a communication channel 5 for synchronizing data. The first fail-safe subsystem 1 and the second fail-safe subsystem 2 are each coupled to a peripheral device 3 via a bus 4. The peripheral device 3 has an output component IO-Dev.

[0025] according to Figure 1 A disadvantage of the known redundant automation system 100 of the prior art is that faulty synchronization data can be transmitted when synchronizing the second fail-safe subsystem 2 via the first fail-safe subsystem 1. The second fail-safe subsystem can also be disrupted by these faulty synchronization data.

[0026] Figure 2A first method-based solution is shown in order to circumvent the problem of operating the second fail-safe subsystem 2 with faulty synchronization data.

[0027] As according to Figure 1 As in the prior art, a first fail-safe subsystem 1 is communicatively connected to a second fail-safe subsystem 2 and exchanges synchronization data SD. A first control program P1, divided into first program sections P1n, exists in the first fail-safe subsystem 1 and runs accordingly. A second control program P2, divided into second program sections P2n, runs redundantly with the first control program P1 in the second fail-safe subsystem 2.

[0028] The first fail-safe subsystem 1 generates and evaluates events using its first control program P1. These program- and process-dependent events influence the execution sequence of the first program sections P1n in the first control program P1. To ensure that the second fail-safe subsystem 2 is aware of these influencing events, the first fail-safe subsystem 1 provides synchronization data SD to the second fail-safe subsystem 2 per program section, which is indexed by n and reflects the corresponding program section.

[0029] In the first fail-safe subsystem 1, raw data A1 is provided to the output component 10-Dev via a first control program P1. The raw data A1 is initially retained and not yet written to the output component 10-Dev. The first fail-safe subsystem 1 typically precedes the second fail-safe subsystem 2 by index n in relation to the processing of the corresponding second program segment of the second fail-safe subsystem 2 by processing the corresponding first program segment P1n.

[0030] In order to prevent the second fail-safe subsystem 2 from operating with faulty synchronization data SD, the synchronization data SD are initially buffered in the second fail-safe subsystem 2 in the second memory area SB2 .

[0031] Since a fault check FP is performed at the end of the corresponding program segment in the first fail-safe subsystem 1, the error-free progress of the first control program P1 can be notified in the corresponding program segment. If a fault is detected, the first fail-safe subsystem 1 sends a fault-free message FFOK to the second fail-safe subsystem 2. The second fail-safe subsystem then accepts the fault-free message FFOK with the fault-free certificate FFQ and reads the previously buffered synchronization data SD from the second memory area SB2 for data processing. The synchronization data is processed using the second program segment P2n matching the index n. Upon receipt of the fault-free certificate FFQ, the first fail-safe subsystem 1 writes the raw data A1 to the output component IO-Dev. Accordingly, a fault check FP is periodically introduced in the vertical time curve relative to the first fail-safe subsystem 1. Program segments P1n and P1n+1 are respectively executed. It is important to note that although the raw data A1 has been transmitted in step 20 and is available in the second fail-safe subsystem 2, it has not yet been written to the process component or output component IO-Dev in step 21.

[0032] according to Figure 2 In the method described in FIG. 1 , no fault occurs in the first program section P1n, but a fault condition 22 occurs in the subsequent program section P1n+1 at the end of the fault check FP. In the event of a detected fault, the first fail-safe subsystem 1 is deactivated. The first fail-safe subsystem enters the STOP 23 state, and the control of the technical process is performed by the second fail-safe subsystem 2. Consequently, the second fail-safe subsystem 2 enters standalone operation 24 and therefore does not download the synchronization data, which are assumed to be faulty, from the second memory area SB2.

[0033] according to Figure 3 In an alternative method, the original data A1 and the output certificate AQ are already exchanged between the fault detection unit FP, but the original data A1 is not written to the process via step 21 until the fault-free message FFOK is sent and the corresponding fault-free certificate FFQ is received.

[0034] In terms of the response time of the entire system, Figure 4 method is considered to be an efficient method.

[0035] It is proposed that the synchronization data SD in the second fail-safe subsystem 2 be immediately processed using the second program segment P2n matching index n, independently of the error-free message FFOK. For better illustration, the second fail-safe subsystem 2 is divided into a processor area 2a and a memory area 2b. The step data processing 25 makes it clear that the incoming synchronization data SD for the second program segment P2n is immediately processed in the processor area 2a. In parallel, the raw data A1 is provided to the second fail-safe subsystem 2. The second fail-safe subsystem accordingly sends an output certificate Aq and a cycle certificate ZQ2. An F check FP now occurs in the first fail-safe subsystem 1, confirming the error-free nature of the first program segment P1n. Due to this error-free nature, the cached raw data A1 is now written to the process, and the error-free message FFOK is transmitted to the second fail-safe subsystem 2. This triggers the backup of the program state PM in the memory image SA in the second fail-safe subsystem 2. The memory image n is backed up in step 40. In the event of the absence of the error-free message FFOK and a failure of the first fail-safe subsystem 1 , the last backup of the program state Pn- 1 is downloaded from the memory image SA and program processing continues in this program state, and the control of the technical process is carried out by the second fail-safe subsystem 2 .

Claims

1. A method for operating a redundant automation system (100) for controlling a technical process, wherein: A first fail-safe subsystem (1) operates with a first control program (P1) divided into first program segments (P1n), wherein the control of the technical process is performed by the first fail-safe subsystem (1), and a second fail-safe subsystem (2) operates redundantly with a second control program (P2) divided into second program segments (P2n), wherein the first fail-safe subsystem (1) generates and evaluates events with the first control program (P1) of the first fail-safe subsystem, influences the execution sequence of the first program segments (P1n) in the first control program (P1) according to the occurrence of the events, and based on the generated or the occurrence of the event, synchronization data (SD) provided with an index (n) reflecting the corresponding program segment is provided to the second fail-safe subsystem (2) per program segment, and raw data (A1) is provided to the output component (IO-Dev), wherein the raw data (A1) is initially retained and not yet written to the output component (IO-Dev), wherein the first fail-safe subsystem (1) precedes the second fail-safe subsystem (2) with respect to the index (n) of the processing of the corresponding second program segment of the second fail-safe subsystem (2), with respect to the processing of the corresponding first program segment (P1n), characterized in that The synchronization data (SD) is first cached in the second fail-safe subsystem (2), the first fail-safe subsystem (1) first sends the provided original data (A1) provided with the index (n) of the corresponding first program section (P1n) to the second fail-safe subsystem (2), and the second fail-safe subsystem (2) signs the original data by issuing a certificate (AQ) to the first fail-safe subsystem (1). In the first fail-safe subsystem (1), a fault check (FP) is performed at the end of the corresponding program section, with which the error-free running of the first control program (P1) is checked in the corresponding program section. In the event of a detected fault, the first fail-safe subsystem (1) is deactivated and the control of the technical process is performed by the second fail-safe subsystem (2), In the case of a detected error-free state, the first fail-safe subsystem (1) sends a error-free message (FFOK) to the second fail-safe subsystem (2), whereupon the second fail-safe subsystem receives the error-free message (FFOK) with an error-free certificate (FFQ) and processes the synchronization data (SD) of the first buffered second program section (P2n) matching the index (n). Upon receiving the free-fault certificate (FFQ), first raw data (A1) is written to the output component (IO-Dev) via the first fail-safe subsystem (1).

2. The method according to claim 1, wherein The second fail-safe subsystem (2) sends a cycle certificate (ZQ2) to the first fail-safe subsystem (1), which confirms that the second program section (P2n) associated with the index (n) was processed successfully and without errors.

3. The method according to claim 1 or 2, wherein: In the event that a fault is detected on the first fail-safe subsystem (1), the second fail-safe subsystem (2) rejects all the synchronization data (SD) stored after the last fault-free certificate (FFQ) and takes over the control of the process for standalone operation.

4. The method according to claim 1, wherein The synchronization data (SD) in the second fail-safe subsystem (2) are immediately processed independently of the fault-free message (FFOK) using the second program section (P2n) matched to the index (n), and For the case that the no-fault message (FFOK) reaches the second fail-safe subsystem (2), the program status (P n ) are additionally backed up to the storage image (SA), In the event that the no-fault message (FFOK) is missing and the first fail-safe subsystem (1) fails, a program state (P) is loaded from the memory image (SA). n-1 ) and continue program processing using the program state, and perform the control of the technical process by means of the second fail-safe subsystem (2).

Citation Information

Patent Citations

  • Method for operating a redundant automation system

    CN103377083B

  • Method for operating an automation system

    CN104571078A