Real-time data communication method and system for safety instrument system with fault self-healing capability
By employing parallel path detection and dynamic path decision algorithms, the communication interruption problem in SIS systems during faults or congestion is solved, achieving seamless switching and efficient data transmission. This meets the deterministic and availability requirements of SIS systems for communication, and improves the system's resilience and communication efficiency.
Patent Information
- Application Number
- CN202511907186.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-03-20
AI Technical Summary
Existing safety instrumented systems (SIS) cannot dynamically select the optimal communication path when faced with equipment failure or network congestion, resulting in communication interruptions and failing to meet the requirements of high availability and real-time performance.
By employing a parallel path detection mechanism and a dynamic path decision algorithm, the system generates secure data frames with timestamps and criticality level identifiers, monitors network status in real time, selects the optimal transmission path, and seamlessly switches to a backup path in the event of a failure, ensuring uninterrupted transmission of critical data.
It enables the dynamic and intelligent selection of the optimal communication path in complex industrial network environments, ensuring seamless transmission of critical security data, meeting the stringent requirements of SIS systems for communication determinism and availability, and improving the system's resilience and communication efficiency.
Smart Images

Figure CN121711291A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of communication in industrial automation control systems, and in particular to a real-time data communication method and system for safety instrumented systems. Background Technology
[0002] Safety Instrumented Systems (SIS) are critical systems for ensuring industrial production safety, and reliable and timely data communication between internal safety stations is essential. Several existing technologies attempt to address the communication security and reliability issues in SIS systems.
[0003] For example, patent CN110798480A proposes a secure data packet format that detects data errors by including fields such as version number, sequence number, and CRC checksum, thus ensuring data integrity and security. However, this scheme mainly focuses on data link layer verification and lacks dynamic management of network layer communication paths. When a switch fails or a link is interrupted in the network, this scheme cannot guarantee continuous connectivity of the communication link, which may lead to a complete communication interruption and fail to meet the stringent high availability requirements of SIS systems.
[0004] For example, the patent with publication number CN113852568B solves the data congestion problem in many-to-one communication and achieves load balancing through queue management and A / B network port staggered transmission mechanism. However, the communication path of this solution is static or preset, and its "load balancing" is more about scheduling the pressure on the sending end than actively optimizing and tolerating the underlying network communication path. Once the preset physical link (such as the switch or network cable corresponding to network A or network B) fails, its communication reliability will drop sharply. Summary of the Invention
[0005] To address the technical problem that existing technologies cannot dynamically and intelligently select the optimal communication path in complex industrial network environments when facing equipment failures or network congestion, this invention proposes a real-time data communication method for safety instrumented systems with fault self-healing capabilities. When a communication path failure or performance degradation is detected, it can seamlessly and quickly switch to a backup path, thereby ensuring uninterrupted transmission of critical safety data and meeting strict real-time requirements to guarantee deterministic latency and extremely high availability of safety data communication.
[0006] To achieve the above objectives, the technical solution of the present invention is implemented as follows: a real-time data communication method for a safety instrumented system with fault self-healing capability, applied to a transmitting control station, comprising the following steps: Step S101: Generate a security data frame based on the user data to be sent. The header of the security data frame includes a timestamp and a criticality level identifier. Step S102: Send a path probe frame to the destination control station through a parallel path probing mechanism to obtain the current status information of multiple available paths in the network. The current status information includes path delay, jitter, and packet loss rate. Step S103: Based on the current status information and the criticality level identifier, select the optimal transmission path for the security data frame using a dynamic path decision algorithm; wherein, the dynamic path decision algorithm prioritizes low-latency and low-jitter paths for high-criticality security data frames, and selects paths with lower loads for low-criticality security data frames. Step S104: Send the secure data frame to the destination control station through the selected optimal transmission path; Step S105: Monitor the communication quality of the optimal transmission path in real time. If the communication quality is lower than the preset threshold, trigger the seamless switching mechanism, immediately select a new path from the backup path set and switch to continue transmitting subsequent secure data frames.
[0007] Preferably, the user data includes: process variables, interlock status, operator instructions, equipment diagnostic information, and safety-critical events. The user data is generated or received by the controller of the control station. The generation of the secure data frame is performed by the controller's communication service program or the protocol processing unit of the communication module. The secure data frame encapsulates user data as the payload and adds a message header before the payload. The secure data frame also includes necessary fields: sequence, source / destination identifier, and CRC checksum.
[0008] Preferably, the timestamp is the precise time when the data is generated or prepared for transmission; The criticality level is categorized into high criticality, medium criticality, and low criticality. High criticality directly triggers or executes safety interlock functions, involving core instructions and status signals related to personnel, equipment, or environmental safety, including emergency stop instructions, safety interlock trigger signals, and heartbeat / survival messages. Medium criticality data participates in routine process control, important status monitoring, and system synchronization, including routine process control instructions and important status synchronization information. Low criticality data is used for monitoring, recording, and maintenance, including non-real-time log uploads, diagnostic information, and configuration information.
[0009] Preferably, the path management unit initiates a parallel path detection mechanism and simultaneously sends path detection frames to the destination control station; the path management unit is a logical functional unit integrated into the communication module, used to maintain a dynamic path status table and execute detection, decision-making, and switching logic; The parallel path probing mechanism is triggered before each path selection for highly critical / medium critical security data frames, or proactively triggered at fixed intervals. The path management unit simultaneously sends lightweight path probing frames to the destination control station through all available physical network ports. Each path probing frame contains a sending timestamp T_send. Upon receiving the path probing frame, the destination station immediately replies with an acknowledgment frame containing a receiving timestamp T_receive and calculates the processing delay. Upon receiving the acknowledgment frame, the sending station calculates: Path delay = (T_arrival - T_send) - known processing delay, where T_arrival is the time the sending station received the acknowledgment; Jitter = the absolute value of the difference between the current delay and the historical average delay; Packet loss rate = (Number of probe frames sent - Number of acknowledgment frames received) / Number of probe frames sent. The calculated delay, jitter, and packet loss rate are then updated in the dynamic path status table.
[0010] Preferably, for data of high criticality, the latency threshold is set to 1 / 3 to 1 / 2 of the system's maximum allowable communication latency; paths that meet the latency threshold are low-latency paths; the jitter threshold is set to be much smaller than the latency threshold, and paths whose latency is stable within the fluctuation range of the latency threshold are low-jitter paths; load indicators are compared among all available paths, and the lightest one is selected. The load is assessed indirectly and relatively through one or more of the following methods: recent average latency trend, local send queue depth, packet loss rate, path probe frame queuing time, and bandwidth utilization. The input to the dynamic path decision algorithm is the state set of all available paths and the criticality level identifier of the security data frame. It determines whether the criticality level indicates high criticality. If it is high criticality, it filters out paths from the state set of available paths whose path delay and jitter are both within the tolerance range, and selects the path with the lowest path delay. If it is medium or low criticality, it prioritizes the path load and selects the path with the lightest load.
[0011] Preferably, the set of available paths is a dynamic database maintained by the path management unit through active probing and network discovery. During system initialization or network topology changes, the path management unit identifies all possible physical paths from the current control station to the destination control station based on static routing configuration or through a secure routing discovery protocol, forming an initial list of available paths. The path management unit periodically sends lightweight probe frames to the destination station through all available paths. Based on the response from the destination station, it calculates and updates the status information of each path in real time. The status set of available paths is the set of associations between each path in the list of available paths and its latest status information. With network support, the path management unit automatically discovers and records multiple network paths to the destination station by running a secure industrial routing protocol. The path management unit sends lightweight path probe frames to the destination control station through all identified paths. Upon receiving the frames, the destination station immediately replies with a timestamp response frame. The sending station calculates and continuously updates the status information of each path based on the response frames, including: current latency, jitter, packet loss rate, and load assessment metrics. The path delay threshold refers to the maximum allowable one-way transmission delay for a single path, which must be less than the total communication delay budget. The jitter threshold refers to the maximum allowable range of latency fluctuations to ensure time determinism, and is 10% to 30% of the path latency threshold.
[0012] Preferably, the path management unit sends the output result based on the dynamic path decision algorithm. The output result includes the logical identifier of the optimal transmission path and is also associated with all the feature information of the optimal transmission path in the dynamic path status table. The path management unit queries its maintained path-port mapping table based on the logical identifier of the optimal transmission path, delivers the secure data frame to be sent to the underlying network protocol stack for final encapsulation, and sets the destination address of the data link layer to the next-hop address corresponding to the optimal transmission path. The underlying network sends complete secure data frames to the network through the designated physical port; in continuous communication, the path management unit maintains a short-lived path binding state for the same data stream or session.
[0013] Preferably, the real-time monitoring method in step S105 includes: carrying the transmission sequence number and timestamp in the service data frame, analyzing the return time of the application layer confirmation frame replied by the receiver, calculating the application layer round-trip time of the service message as a direct reflection of the path quality; statistically analyzing the packet loss rate and out-of-order rate of data frames on the current path; comparing the actual service flow data with the underlying path status information obtained by the parallel detection mechanism, and making a comprehensive judgment. The preset threshold for communication quality is 120%-150% of the path delay threshold; When the delay of service packets exceeds the preset threshold for communication quality multiple times in a row, the seamless handover mechanism is triggered. The update of the dynamic path status table includes: after each parallel probe, the corresponding path's table entry is updated with new delay, jitter, and packet loss rate data; when a seamless handover event occurs, the status of the path to be switched is marked as "performance degraded" or "unavailable" and its priority is reduced; after the seamless handover is completed, the status of the newly enabled path is marked as "in use"; if a negative response is received from the receiving station, the packet loss or error count of the corresponding path is updated.
[0014] A real-time data communication method for a safety instrumented system with fault self-healing capability, applied to a receiving control station, includes the following steps: Step S201: Receive a security data frame from the transmitting control station; Step S202: Perform security verification on the secure data frame, including CRC verification, sequence number check, and timestamp freshness check; Step S203: If the verification is successful, the user data is written to the operation data area; if the verification fails, an error log is recorded and a negative response is sent to the source control station. Step S204: Receive the path probe frame sent by the sending control station and immediately reply with a response frame containing the local receiving timestamp, which is used by the source control station to calculate the path delay.
[0015] A real-time data communication system for a safety instrumented system with fault self-healing capability includes at least two SIS control stations interconnected via industrial Ethernet. Each SIS control station includes a controller and a communication module, wherein: The controller is used to execute the real-time data communication method of the safety instrumented system with fault self-healing capability applied to the transmitting control station, or to execute the real-time data communication method of the safety instrumented system with fault self-healing capability applied to the receiving control station. The communication module integrates a path management unit, which is used to execute a parallel path detection mechanism, a dynamic path decision algorithm, and a seamless switching mechanism. The communication module is also configured with multiple physical network ports for simultaneously connecting to different network switches to build a redundant network. Multiple SIS control stations are interconnected through two or more industrial Ethernet switches to form a redundant network topology; each SIS control station's communication module has at least two independent physical network ports, which are connected to different switches respectively.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: Dynamic path optimization and fault self-healing: Through parallel path detection and dynamic path decision-making algorithms, this invention no longer statically relies on the A / B dual network, but can perceive the network status in real time, actively select the best path, and achieve seamless switching in milliseconds when a fault occurs, which greatly improves the availability and resilience of the communication system.
[0017] Deterministic real-time communication: By marking data frames with criticality levels and combining them with real-time latency / jitter information of the path, it can be ensured that the most critical security commands are always transmitted through the fastest path, meeting the stringent requirements of SIS systems for communication determinism.
[0018] Enhanced synergy with existing security mechanisms: This invention does not replace existing secure data packet formats and verification mechanisms, but rather adds intelligent management of the network path layer, forming a comprehensive secure communication solution that extends from data content to transmission path.
[0019] Efficient network resource utilization: Through dynamic path decision algorithms, data streams of different priorities can be rationally distributed to different network paths, avoiding local network congestion and achieving true network load balancing. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a communication flowchart of the sending control station in an embodiment of the present invention.
[0022] Figure 2 This is a flowchart of the dynamic path decision algorithm in an embodiment of the present invention.
[0023] Figure 3 This is a timing diagram of the seamless switching mechanism in an embodiment of the present invention.
[0024] Figure 4 This is a schematic diagram of the system network topology provided for an embodiment of the present invention.
[0025] Figure 5 This is a schematic diagram of the communication process of the receiving control station provided in an embodiment of the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] Example 1 like Figure 1 As shown, the present invention provides a data communication method between safety instrumented systems with fault self-healing capability, applied to a transmitting control station, comprising: Step S101: Generate a security data frame based on the user data to be sent. The header of the security data frame includes a timestamp and a criticality level identifier.
[0028] In SIS, user data originates from the logical operation results of the controller, and may specifically include: process variables (such as gas concentration, temperature sensor readings), interlock status, operator instructions, equipment diagnostic information, and safety-critical events (such as emergency shutdown trigger conditions). This data is generated or received by the controller in the control station (such as the CPU of a safety PLC). The generation of safety data frames is executed by the controller's communication service program or the protocol processing unit of the communication module. The generation process includes: 1) Encapsulate user data: Use the above user data as the payload.
[0029] 2) Add a security header: Add a header before the payload. The header must contain at least the following: Timestamp: The precise time when data is generated or prepared for transmission (usually taken from the controller's high-precision clock), used by the receiver for "freshness checks" to prevent replay attacks or process expired data. The format can be IEEE 1588 PTP time or a millisecond-level countdown after system power-on.
[0030] 3) Criticality Level Indicator: Used to indicate the real-time nature and importance of the data in this frame. For example, it can be divided into: Highly critical: Emergency stop (ESD) commands, safety interlock trigger signals, heartbeat / survival messages. Requires extremely low latency and jitter.
[0031] Key elements: routine process control instructions and important status synchronization information.
[0032] Low criticality: Non-real-time log uploads, diagnostic information, and configuration information.
[0033] Other necessary fields include: sequence number (used to detect packet loss and out-of-order delivery), source / destination identifier, and CRC checksum (covering the message header and payload to ensure integrity).
[0034] Step S102: Send a path probe frame to the destination control station through a parallel path probing mechanism to obtain the current status information of multiple available paths in the network. The status information includes path delay, jitter, and packet loss rate. The path management unit initiates a parallel path detection mechanism and simultaneously sends path detection frames to the destination control station. The path management unit is a logical functional unit integrated into the communication module (such as a dedicated communication chip or FPGA). It can be a resident service of the software driver layer or part of the hardware logic. Its core responsibility is to maintain a dynamic path state table and execute detection, decision-making, and switching logic.
[0035] The specific implementation of the parallel path probing mechanism: Trigger: Trigger proactively before each time a path needs to be selected for a high / medium critical data frame, or at a fixed period (e.g., every 10-100ms).
[0036] Sending probe frames: The path management unit sends probe frames through all available physical network ports (such as...). Figure 4 The dual network ports simultaneously (or nearly simultaneously) send a lightweight path probe frame containing a sending timestamp T_send to the destination control station.
[0037] Status information collection: After receiving a probe frame, the destination station immediately replies with an acknowledgment frame, containing the reception timestamp T_receive, or the directly calculated processing delay. Upon receiving the acknowledgment, the sending station calculates: Path delay = (T_arrival - T_send) - Known processing delay, where T_arrival is the time the sending station received the acknowledgment. Jitter = Absolute value (or statistical standard deviation) of the difference between the current delay and the historical average delay. Packet loss rate = (Number of probe frames sent - Number of acknowledgment frames received) / Number of probe frames sent, based on a time window.
[0038] Status information storage: Update the calculated latency, jitter, and packet loss rate to the dynamic path status table.
[0039] Step S103: Based on the status information and criticality level identifier, select the optimal transmission path for the secure data frame using a dynamic path decision algorithm; wherein, the dynamic path decision algorithm prioritizes low path latency and low jitter paths for high-criticality data frames, and selects paths with lower loads for low-criticality data frames.
[0040] The dynamic path decision algorithm is the core of this invention for achieving "deterministic real-time communication" and "efficient resource utilization," specifically producing the following synergistic technical effects: Ensuring the determinism of safety-critical business operations: By prioritizing and consistently allocating low-latency, low-jitter physical paths for high-critical data (such as emergency stop commands), the timeliness and stability of such command transmission are ensured, directly meeting the millisecond-level deterministic requirements of SIS for the most stringent security functions.
[0041] Achieve intelligent network load balancing: By guiding medium / low-critical data (such as logs and non-real-time status) to select lighter paths, non-critical traffic is prevented from blocking critical paths, thus optimizing bandwidth utilization from a network-wide perspective and improving overall communication efficiency.
[0042] Achieving dynamic matching of Service Level (QoS) and communication resources: The algorithm intelligently associates the "criticality level" of data with the "real-time status" of the network, so that data streams of different importance automatically obtain the corresponding transmission service quality. This is an adaptive and refined resource scheduling strategy that goes beyond static A / B path allocation.
[0043] Basis for setting: All thresholds are determined by working backward from the "overall safety response time requirements of the Safety Instrumented System (SIS)". For example, if the system requires the total time from sensor malfunction to actuator action to not exceed 100ms, then the allowable delay in the communication link may be required to be 10-20ms.
[0044] "Low latency" and latency threshold refer to the one-way propagation delay of a path. For data of high criticality, the latency threshold is typically set to 1 / 3 to 1 / 2 of the system's maximum allowable communication latency to allow for margin. Example: If the system requires end-to-end communication latency <10ms, the path latency threshold can be set to <3ms. A path that meets this condition is a "low latency path".
[0045] "Low jitter" and jitter threshold refer to the range of latency fluctuation. To ensure determinism, the jitter threshold is usually set much smaller than the latency threshold. Example: It can be set to <1ms or 20-30% of the latency threshold (e.g., 0.5ms). The path where the latency stabilizes within this fluctuation range is the "low jitter path".
[0046] The metric for "low load" is a relative comparison, not an absolute threshold. When making a decision, the algorithm compares the "load metrics" across all available paths and selects the lightest. Load can be assessed indirectly and relatively in one or more of the following ways: Recent average latency trend: Under non-congestion conditions, latency is positively correlated with load, so choose the path with the gentlest historical latency growth trend.
[0047] Probe frame queuing time: The path management unit can estimate the waiting time of probe frames in the local port's transmission queue.
[0048] Bandwidth utilization statistics: If the network device supports it, port utilization information can be obtained as a reference through protocols such as SNMP.
[0049] Determining the tolerance range: The above latency and jitter thresholds are system-level parameters that can be pre-configured in the configuration file of the path management unit during system integration based on network scale, switch performance, cable length, etc., or can be issued and adjusted online by the host computer engineering tools.
[0050] like Figure 2 As shown, the core logic of the dynamic path decision algorithm is as follows: Input: The state set of all available paths, the criticality level of the security data frame.
[0051] The set of available paths is a dynamic database maintained by the path management unit through proactive probing and network discovery. Its acquisition and construction process is as follows: Path discovery: During system initialization or network topology changes, the path management unit identifies all possible physical paths from the current control station to the destination control station based on static route configuration or through secure route discovery protocols (such as industrial security variants based on OSPF or RSTP). Figure 4 The paths shown (path 1, path 2, path 3) form an initial list of available paths.
[0052] Status information collection: This is the core task of the parallel path probing mechanism (step S102). The path management unit periodically sends lightweight probe frames to the destination station through all available paths. Based on the response from the destination station, the status information of each path is calculated and updated in real time, including: current instantaneous latency, statistically derived jitter (such as standard deviation), recent packet loss rate, and load assessment metrics.
[0053] Set Formation: The state set of all available paths is the associated set of each path in the above list of available paths and its latest state information. This set exists in the memory of the path management unit in the form of a data table or data structure, and serves as the direct input to the dynamic path decision algorithm.
[0054] Determine if the criticality level is high: If it is high criticality, filter out paths from the path set whose path delay and jitter are both within the tolerance range, and then select the path with the lowest path delay from them.
[0055] If the network is classified as medium or low critical, the load on the path should be taken into account first, and the path with the lightest load should be selected to optimize the overall network efficiency.
[0056] In this invention, the path set refers to the list of available physical paths and their associated real-time status information set. Its acquisition and construction is a dynamic process that includes initialization and continuous updates, as detailed below: Initial setup: During system startup or network topology changes, an initial list of available paths is created in the following ways: Static configuration: based on, for example Figure 4 The network physical topology shown (e.g., control station 101 → switch A → control station 102 is path 1; control station 101 → switch B → control station 102 is path 2) is pre-configured by the engineer during the system configuration phase, with all legal communication paths pre-defined.
[0057] Dynamic discovery: When network support is available, the path management unit can automatically discover and record multiple network paths to the destination station by running secure industrial routing protocols (such as deterministic variants of RSTP).
[0058] Dynamic maintenance of status information (core): After the initial list is established, the path management unit actively and periodically obtains the real-time performance data of each path through a parallel path detection mechanism.
[0059] The specific process is as follows: The path management unit simultaneously sends lightweight probe frames to the destination control station through all identified paths. Upon receiving the frames, the destination station immediately replies with a timestamp-containing response frame. Based on this, the sending station calculates and continuously updates the status information of each path, including: current latency, jitter (latency standard deviation), packet loss rate, and load assessment metrics. This constantly updated status information, together with the path identifiers, constitutes the status set of all available paths upon which the algorithm's decisions depend. This set typically exists within the path management unit as a dynamically updated internal data table or data structure.
[0060] Tolerance ranges, latency, and jitter thresholds are the quantitative boundaries for achieving deterministic real-time communication and seamless triggering switching. Their settings follow a configurable principle of "top-down decomposition based on system safety objectives." All communication performance thresholds are derived from the Safety Integrity Level (SIL) and overall safety response time budget required by the Safety Instrumented System (SIS). A SIL2 safety loop requires a total time ≤ 150ms from hazard detection to actuator action. This time budget is decomposed into components such as sensors, controller logic processing, communication networks, and actuators. If the total latency budget allocated to the communication network is ≤ 20ms, then this 20ms is the top-level constraint for setting all communication-related thresholds.
[0061] The latency threshold (used for filtering highly critical data paths) refers to the maximum allowable one-way transmission latency for a single path. Its value must be less than the total communication latency budget, allowing for margins in other processing stages. Setting method: Path latency threshold = Total system communication latency budget × Weighting coefficient (e.g., 0.5-0.7). For example, if the total budget is 20ms, to ensure critical instructions are transmitted, the latency threshold can be set to 5ms to 10ms. Only paths with latency consistently below this threshold are considered low-latency paths suitable for transmitting highly critical data.
[0062] The jitter threshold (used for filtering highly critical data paths) refers to the maximum allowable range of latency fluctuations to ensure time determinism. It is typically set to 10% to 30% of the latency threshold. For example, if the latency threshold is 5ms, the jitter threshold can be set to 0.5ms or 1ms. Paths with jitter exceeding the jitter threshold are considered unstable and unsuitable for transmitting highly critical data.
[0063] The tolerance range refers to the range within which the real-time latency and jitter measurements of a certain path simultaneously meet the aforementioned latency threshold and jitter threshold. This is a "hard" quality threshold for selecting highly critical data transmission paths.
[0064] The criticality level is classified based on the function and timeliness requirements of the data in security control, and its scope is defined as follows: High Criticality Level: Directly triggers or executes Safety Interlock Functions (SIFs), involving core commands and status signals related to personnel, equipment, or environmental safety. Loss or delay of this data will directly lead to safety risks. Examples include Emergency Stop (ESD) commands, safety interlock trigger input signals (such as over-limit trip signals), safety output commands (such as closing hazardous valves), and heartbeat / survival messages between controllers.
[0065] Medium-criticality level: Data involved in routine process control, critical status monitoring, and system synchronization. Its transmission reliability requirements are high, but the requirement for extremely low latency is less stringent than for highly critical data. Examples: Important process variable setpoints, equipment mode switching commands, non-urgent process alarm signals, and batch control procedure commands.
[0066] Low Criticality Level: Data not involved in real-time closed-loop control or safety interlocks, used for monitoring, recording, maintenance, etc. Timeliness requirements for transmission are less stringent. Examples: Historical data recording, diagnostic log uploads, performance statistics reports, and non-urgent configuration parameter reading and writing.
[0067] The classification method is automatically assigned by the controller of the sending control station when generating data frames, based on the predefined attributes of the data points (determined during the engineering configuration phase), and encoded in the header of the security data frame.
[0068] In this invention, "load" is a relative metric used for comparing paths, not an absolute measure. Load assessment aims to select the path with the lowest probability of congestion for medium- and low-critical data, achieving network traffic balance. Path load assessment is typically based on comparisons using one or more of the following indirect metrics: Detecting frame latency trends: The recent average latency of the path shows an upward trend, indicating that its load may be increasing.
[0069] Local transmit queue depth: The path management unit monitors the instantaneous or average depth of the transmit queue for the corresponding physical port. A deeper queue depth generally indicates more severe congestion and heavier load at the path's egress point.
[0070] Packet loss rate metric: A periodic increase in packet loss rate caused by non-physical faults is a typical manifestation of network congestion.
[0071] Bandwidth utilization (if available): Port traffic statistics obtained from network devices via methods such as Simple Network Management Protocol (SNMP).
[0072] When selecting paths for medium- and low-critical data, the algorithm compares the above-mentioned load metrics of all available paths and selects the path with the lightest overall load or the relatively least idle, thereby "diverting" non-critical traffic from paths that may carry critical traffic and optimizing overall network efficiency.
[0073] Step S104: Send the secure data frame to the destination control station through the selected optimal transmission path.
[0074] The execution entity is the path management unit. Integrated within the control station's communication module, the path management unit is the executor of the dynamic path decision algorithm and the controller for network packet forwarding. It sends the output results based on the dynamic path decision algorithm. This output not only includes the logical identifier of the "optimal path" but also associates all characteristic information of that path in the dynamic path status table, especially its corresponding physical exit (network port) and next-hop address.
[0075] After receiving the "optimal path" identifier selected for a certain secure data frame, the path management unit operates as follows: Address and port resolution: The path management unit queries its maintained path-port mapping table based on the optimal path identifier. This table is created during system initialization and records each logical path (e.g., ...). Figure 4 The local physical network port (such as network port A and network port B) corresponding to path 1 and path 2 in the network, and the address of the next-hop network device (such as the MAC address / IP address of switch A or switch B).
[0076] Frame encapsulation and direction: The secure data frame to be sent is delivered to the underlying network protocol stack for final encapsulation. During this process, the path management unit or driver layer instructs the protocol stack to explicitly specify that the data frame should be sent from the queried specific physical network port. Simultaneously, the destination address of the data link layer is set to the next-hop address corresponding to this path.
[0077] Sending: The underlying network driver command sends the complete data frame to the network through the specified physical port. For example, if the algorithm selects path 1 (via switch A), the data frame will be sent from the port connected to switch A; if path 2 (via switch B) is selected, it will be sent from the port connected to switch B.
[0078] Session continuity maintenance (for seamless handover): In continuous communication, the path management unit maintains a brief path binding state for the same data stream or session. This means that before a handover is triggered, subsequent data frames to the same destination station will be sent using the current optimal path by default until the next periodic probe and decision refreshes the path, or a quality degradation is detected to trigger a handover.
[0079] The technical essence of this step lies in mapping intelligent decisions based on real-time network status and business criticality to redundant physical network hardware in a timely and accurate manner, enabling selective use of physical links. Unlike existing technologies that employ static, polling, or simple rule-based A / B port scheduling, this invention dynamically and optimally selects based on real-time perceived network conditions (latency, jitter, load) before transmission. This process of intelligent decision-making followed by precise execution ensures that each transmission utilizes the optimal physical channel available. This step is a natural continuation and final execution of parallel probing and dynamic decision-making. Without this step to precisely apply the decision results to the physical transmission port, the aforementioned intelligent decisions would not be effective. Simultaneously, it provides a real-time traffic observation target for real-time monitoring (step S105).
[0080] Step S105: Monitor the communication quality of the optimal path in real time. If the communication quality is lower than a preset threshold, trigger the seamless switching mechanism, immediately select a new path from the backup path set and switch to continue transmitting subsequent secure data frames.
[0081] Meanwhile, the path monitoring unit continues to operate, immediately triggering path switching and updating the path status table upon detecting a problem. This mechanism enables switching before performance degradation or in a sub-healthy state. Its core effects are: Zero service interruption awareness: Proactive switching occurs when communication path performance degrades but is not completely interrupted, avoiding service interruptions caused by traditional redundant protocols (such as STP with second-level convergence), ensuring high continuity of SIS control flow. Deterministic recovery: The switching time (ΔT) is extremely short and deterministic (e.g., <10ms), making it possible to include communication interruption time in the system's security response time budget, meeting the stringent requirements of SIL level for communication reliability. Enhanced system resilience: Compared to simple physical redundancy, it increases the ability to withstand soft faults such as performance degradation, improving the overall system resilience.
[0082] The path monitoring unit is a logical functional submodule (software thread or hardware logic block) within the path management unit, which continuously tracks the actual business packet performance of the currently used optimal path.
[0083] Monitoring methods include: Active measurement: By including the transmission sequence number and timestamp in the service data frame and analyzing the return time of the application layer acknowledgment (ACK) frame from the receiver, the application layer round-trip time of the service message is calculated as a direct reflection of path quality. Passive analysis: The packet loss rate (based on discontinuous sequence numbers) and out-of-order rate of data frames on the current path are statistically analyzed. Cross-validation: The measured data of the above service flow are compared with the underlying path status information obtained by the parallel probing mechanism for comprehensive judgment.
[0084] The preset threshold acts as a "warning line" for triggering handover, typically more lenient but more sensitive than the "admission threshold" (latency / jitter threshold in step S103) used for path decision-making. Setting principle: For example, it can be set to 120%-150% of the path decision latency threshold. If the decision latency threshold is 5ms, the handover monitoring threshold can be set to 6-7.5ms. Handover is triggered when the service packet latency exceeds this threshold for N consecutive times (e.g., 3 times). This avoids erroneous handovers caused by momentary network jitter and ensures rapid response even when performance continues to deteriorate.
[0085] The path status table is the core data structure of the path management unit. Updates occur as follows: Periodic updates: After each parallel probe, the corresponding path's entries are updated with new latency, jitter, and packet loss rate data; Event-triggered updates: When a switchover event occurs: the status of the path to be switched (e.g., path 1) is marked as "performance degraded" or "unavailable," and its priority is reduced. After the switchover is complete: the status of the newly enabled path (e.g., path 2) is marked as "in use." Upon receiving a negative acknowledgment (NAK): If a negative acknowledgment (NAK) is received from the receiving station, the packet loss or error count of the corresponding path may be updated.
[0086] like Figure 3 As shown, the seamless switching mechanism is implemented as follows: At time T1, the data frames sent by the control station, including data frame 1 and data frame 2, are transmitted normally through path 1.
[0087] At time T2, the sending end detects a sudden increase in the delay of path 1 (exceeding the threshold).
[0088] At time T3, the transmitting end decides to switch to path 2. This process does not require waiting for path 1 to be completely interrupted, and the switching time (ΔT) is extremely short, ensuring the continuity of the data stream. Subsequent data frames are transmitted to the receiving control station via path 2 starting from time T3.
[0089] like Figure 3 As shown, at time T2, a sudden increase in latency is detected: the path monitoring unit analyzes the time interval between the transmission of data frame 1 and the receipt of its ACK, and finds that the round-trip time significantly exceeds the preset threshold, thus determining that the performance of path 1 has deteriorated. At time T3, a decision switch is made: the path management unit immediately calls the dynamic path decision algorithm, but at this time the input state set may have excluded path 1 (or assigned it a very low weight), thus quickly selecting path 2 as the new path. Subsequently, the path management unit updates the internal forwarding rules, pointing the egress port of subsequent data frames (data frame 3 and later) to the physical network interface corresponding to path 2.
[0090] An ACK (acknowledgment frame) is an application-layer acknowledgment message returned by the receiving control station to the sending control station after successfully receiving and verifying a data frame. It is crucial for achieving reliable transmission and measuring end-to-end latency. Data frame 3 at time T3 is the first service data frame sent via the newly selected path (path 2) after the handover decision is completed. Figure 3 The ACK indicates acknowledgment of the data frame (such as frame 1 or frame 2) previously transmitted via path 1. After the handover, the sending station begins transmitting frame 3 via path 2 and expects to receive the corresponding ACK via path 2.
[0091] Example 2 like Figure 5 As shown, the present invention provides a data communication method between safety instrumented systems with fault self-healing capability, applied to a receiving control station, comprising: Step S201: Receive a security data frame from the source control station; The source control station refers to the Safety Instrumented System (SIS) control station that initiated this communication, i.e., the control station that executes the sending method in Example 1. Figure 4 In the topology, if control station 101 sends data to 102, then 101 is the source control station.
[0092] Step S202: Perform security verification on the secure data frame, including CRC verification, sequence number check and timestamp freshness check; CRC check: The receiving station recalculates the cyclic redundancy check value for the received frame data (excluding the CRC field) using the same polynomial as the sending station, and compares it with the CRC value carried at the end of the frame. This ensures that data has not been tampered with or corrupted during transmission, guaranteeing data integrity and serving as a fundamental security barrier.
[0093] Sequence number checking verifies whether the sequence number in the current frame is equal to the sequence number in the previous frame plus 1. Effect: Detects packet loss and out-of-order arrival events. Discontinuous sequence numbers indicate potential network congestion or path problems.
[0094] Timestamp freshness check: The timestamp in the data frame is compared with the local high-precision clock of the receiving station. If the time difference exceeds a preset valid time window (e.g., 100ms), it is determined to be an expired frame. Effectively defending against replay attacks and ensuring that the system only processes the latest and valid real-time data is crucial to meeting the functional safety requirements of SIS.
[0095] Step S203: If the verification is successful, the user data is written to the operation data area; if the verification fails, an error log is recorded and a negative response is sent to the source control station. User data: This refers to the payload of the secure data frame after all security headers (timestamp, sequence number, CRC, etc.) have been stripped away; it contains the original control commands, process variables, or status information. It is extracted from the communication protocol stack after security verification and decapsulation. Writing to the computation data area: The receiving station's controller writes the successfully verified user data into its internal shared memory or input image area, accessible to secure applications. Subsequently, the security control logic program (such as a program conforming to the IEC 61131-3 standard) periodically reads this data from this area for computation.
[0096] Step S204: Receive the path probe frame sent by the source control station and immediately reply with a response frame containing the local receiving timestamp, which is used by the source control station to calculate the path delay.
[0097] Reception and Reply: The receiving station's path management unit listens for specific probe frames. Upon receipt, it immediately interrupts or prioritizes processing, extracts the sending timestamp T_send from the probe frame, and records the local receiving time T_receive. A lightweight reply frame is generated, containing at least the receiving time T_receive and the timestamp T_send (echo). Delay Calculation and Usage: After receiving this reply, the sending station can calculate the total timeout = T_arrival - T_send based on its own precise clock. If the receiving station's processing time T_processing is available, the one-way propagation delay can be accurately calculated as ≈ (total timeout - _processing) / 2. This calculation result is used to update the sending station's dynamic path state table, serving as the core input for path decision-making.
[0098] Example 3 like Figure 4 As shown, the present invention provides a data communication system between safety instrumented systems, including at least two SIS control stations interconnected via industrial Ethernet. Each SIS control station includes a controller and a communication module, wherein: The controller is used to execute the method steps applied to the transmitting control station as described in Embodiment 1, or to execute the method steps applied to the receiving control station as described in Embodiment 2; the execution of the controller is manifested on two levels: Application Logic Layer: The controller runs the security control application, responsible for generating user data to be sent and assigning criticality levels to the data according to security logic. Simultaneously, it also obtains received user data from the communication module for secure computation.
[0099] Service Invocation Layer: The controller issues commands such as sending data and requesting path information by calling the standardized service interface (API) provided by the communication module. The specific communication processes, such as sending, probing, decision-making, and switching, are autonomously completed by the path management unit within the communication module. This is an architecture that enables division of labor and collaboration.
[0100] The communication module integrates a path management unit, which is used to execute the parallel path detection mechanism, dynamic path decision algorithm and seamless switching mechanism; the communication module is also configured with multiple physical network ports for simultaneously connecting to different network switches to build a redundant network.
[0101] Multiple SIS control stations (101, 102, 103) are interconnected through two or more industrial Ethernet switches (switch A, switch B) to form a redundant network topology. Each SIS control station's communication module has at least two independent physical network ports, each connected to a different switch.
[0102] The communication module of the integrated path management unit embeds intelligent path management functionality into the communication hardware or firmware, achieving decoupling between communication and control. The controller focuses on security logic, while the communication module focuses on network optimization, improving system efficiency and reliability. Multiple physical ports can connect to different switches: such as... Figure 4 As shown, this constructs complete redundancy at both the physical and link layers. A single switch failure or link interruption will not cause any control station to lose connectivity, providing a solid physical foundation and abundant path selection for intelligent path management at the upper layers (network and application layers). Synergistic effect: Physical redundancy provides multiple paths, while intelligent path management achieves "choosing the right path and quickly switching to another." The combination of these two achieves a synergistic effect greater than the sum of its parts (1+1>2), comprehensively ensuring extremely high availability and determinism of communication from both physical and logical perspectives. This is the core value of the system-level innovation of this invention.
[0103] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A real-time data communication method for a safety instrumented system with fault self-healing capability, characterized in that, When applied to a transmitting control station, the following steps are included: Step S101: Generate a security data frame based on the user data to be sent. The header of the security data frame includes a timestamp and a criticality level identifier. Step S102: Send a path probe frame to the destination control station through a parallel path probing mechanism to obtain the current status information of multiple available paths in the network. The current status information includes path delay, jitter, and packet loss rate. Step S103: Based on the current status information and the criticality level identifier, select the optimal transmission path for the security data frame using a dynamic path decision algorithm; wherein, the dynamic path decision algorithm prioritizes low-latency and low-jitter paths for high-criticality security data frames, and selects paths with lower loads for low-criticality security data frames. Step S104: Send the secure data frame to the destination control station through the selected optimal transmission path; Step S105: Monitor the communication quality of the optimal transmission path in real time. If the communication quality is lower than the preset threshold, trigger the seamless switching mechanism, immediately select a new path from the backup path set and switch to continue transmitting subsequent secure data frames.
2. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 1, characterized in that, The user data includes: process variables, interlock status, operator instructions, equipment diagnostic information, and safety-critical events. The user data is generated or received by the controller of the control station. The generation of the secure data frame is performed by the controller's communication service program or the protocol processing unit of the communication module. The secure data frame encapsulates user data as the payload and adds a message header before the payload. The secure data frame also includes necessary fields: sequence, source / destination identifier, and CRC checksum.
3. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 2, characterized in that, The timestamp is the precise time when the data was generated or prepared for transmission; The criticality level is categorized into high criticality, medium criticality, and low criticality. High criticality directly triggers or executes safety interlock functions, involving core instructions and status signals related to personnel, equipment, or environmental safety, including emergency stop instructions, safety interlock trigger signals, and heartbeat / survival messages. Medium criticality data participates in routine process control, important status monitoring, and system synchronization, including routine process control instructions and important status synchronization information. Low criticality data is used for monitoring, recording, and maintenance, including non-real-time log uploads, diagnostic information, and configuration information.
4. The real-time data communication method for a safety instrumented system with fault self-healing capability according to any one of claims 1-3, characterized in that, The path management unit initiates a parallel path detection mechanism and simultaneously sends path detection frames to the destination control station. The path management unit is a logical functional unit integrated into the communication module, used to maintain a dynamic path status table and execute detection, decision-making, and switching logic. The parallel path probing mechanism is triggered before each path selection for highly critical / medium critical security data frames, or proactively triggered at fixed intervals. The path management unit simultaneously sends lightweight path probing frames to the destination control station through all available physical network ports. Each path probing frame contains a sending timestamp T_send. Upon receiving the path probing frame, the destination station immediately replies with an acknowledgment frame containing a receiving timestamp T_receive and calculates the processing delay. Upon receiving the acknowledgment frame, the sending station calculates: Path delay = (T_arrival - T_send) - known processing delay, where T_arrival is the time the sending station received the acknowledgment; Jitter = the absolute value of the difference between the current delay and the historical average delay; Packet loss rate = (Number of probe frames sent - Number of acknowledgment frames received) / Number of probe frames sent. The calculated delay, jitter, and packet loss rate are then updated in the dynamic path status table.
5. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 4, characterized in that, For data of high criticality, the latency threshold is set to 1 / 3 to 1 / 2 of the system's maximum allowable communication latency; paths that meet the latency threshold are low-latency paths; the jitter threshold is set to be much smaller than the latency threshold, and paths whose latency is stable within the fluctuation range of the latency threshold are low-jitter paths; among all available paths, the load index is compared, and the lightest one is selected. The load is assessed indirectly and relatively through one or more of the following methods: recent average latency trend, local send queue depth, packet loss rate, path probe frame queuing time, and bandwidth utilization. The input to the dynamic path decision algorithm is the state set of all available paths and the criticality level identifier of the security data frame. It determines whether the criticality level indicates high criticality. If it is high criticality, it filters out paths from the state set of available paths whose path delay and jitter are both within the tolerance range, and selects the path with the lowest path delay. If it is medium or low criticality, it prioritizes the path load and selects the path with the lightest load.
6. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 5, characterized in that, The set of available paths is a dynamic database maintained by the path management unit through active probing and network discovery. During system initialization or network topology changes, the path management unit identifies all possible physical paths from the current control station to the destination control station based on static routing configuration or through a secure routing discovery protocol, forming an initial list of available paths. The path management unit periodically sends lightweight probe frames to the destination station through all available paths. Based on the response from the destination station, it calculates and updates the status information of each path in real time. The status set of available paths is the set of associations between each path in the list of available paths and its latest status information. With network support, the path management unit automatically discovers and records multiple network paths to the destination station by running a secure industrial routing protocol. The path management unit sends lightweight path probe frames to the destination control station through all identified paths. Upon receiving the frames, the destination station immediately replies with a timestamp response frame. The sending station calculates and continuously updates the status information of each path based on the response frames, including: current latency, jitter, packet loss rate, and load assessment metrics. The path delay threshold refers to the maximum allowable one-way transmission delay for a single path, which must be less than the total communication delay budget. The jitter threshold refers to the maximum allowable range of latency fluctuations to ensure time determinism, and is 10% to 30% of the path latency threshold.
7. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 5 or 6, characterized in that, The path management unit sends the output results based on the dynamic path decision algorithm. The output results include the logical identifier of the optimal transmission path and also associate all the feature information of the optimal transmission path in the dynamic path status table. The path management unit queries its maintained path-port mapping table based on the logical identifier of the optimal transmission path, delivers the secure data frame to be sent to the underlying network protocol stack for final encapsulation, and sets the destination address of the data link layer to the next-hop address corresponding to the optimal transmission path. The underlying network sends complete secure data frames to the network through the designated physical port; in continuous communication, the path management unit maintains a short-lived path binding state for the same data stream or session.
8. The real-time data communication method for a safety instrumented system with fault self-healing capability according to claim 7, characterized in that, The real-time monitoring method in step S105 includes: carrying the transmission sequence number and timestamp in the service data frame, analyzing the return time of the application layer confirmation frame replied by the receiver, calculating the application layer round-trip time of the service message as a direct reflection of the path quality; statistically analyzing the packet loss rate and out-of-order rate of data frames on the current path; comparing the actual service flow data with the underlying path status information obtained by the parallel detection mechanism, and making a comprehensive judgment. The preset threshold for communication quality is 120%-150% of the path delay threshold; When the delay of service packets exceeds the preset threshold for communication quality multiple times in a row, the seamless handover mechanism is triggered. The update of the dynamic path status table includes: after each parallel probe, the corresponding path's table entry is updated with new delay, jitter, and packet loss rate data; when a seamless handover event occurs, the status of the path to be switched is marked as "performance degraded" or "unavailable" and its priority is reduced; after the seamless handover is completed, the status of the newly enabled path is marked as "in use"; if a negative response is received from the receiving station, the packet loss or error count of the corresponding path is updated.
9. The real-time data communication method for a safety instrumented system with fault self-healing capability according to any one of claims 1-3, 5, 6, and 8, characterized in that, When applied to a receiving control station, the following steps are included: Step S201: Receive a security data frame from the transmitting control station; Step S202: Perform security verification on the secure data frame, including CRC verification, sequence number check, and timestamp freshness check; Step S203: If the verification is successful, write the user data into the calculation data area; If the verification fails, an error log is recorded and a negative response is sent to the source control station; Step S204: Receive the path probe frame sent by the sending control station and immediately reply with a response frame containing the local receiving timestamp, which is used by the source control station to calculate the path delay.
10. A real-time data communication system for a safety instrumented system with fault self-healing capability, characterized in that, Includes at least two SIS control stations interconnected via industrial Ethernet, each SIS control station including a controller and a communication module, wherein: The controller is configured to execute the real-time data communication method for a safety instrumented system with fault self-healing capability applied to a transmitting control station as described in any one of claims 1-8, or to execute the real-time data communication method for a safety instrumented system with fault self-healing capability applied to a receiving control station as described in claim 9. The communication module integrates a path management unit, which is used to execute a parallel path detection mechanism, a dynamic path decision algorithm, and a seamless switching mechanism. The communication module is also configured with multiple physical network ports for simultaneously connecting to different network switches to build a redundant network. Multiple SIS control stations are interconnected through two or more industrial Ethernet switches to form a redundant network topology; each SIS control station's communication module has at least two independent physical network ports, which are connected to different switches respectively.
Citation Information
Patent Citations
Data communication method and system between safety instrument systems
CN110798480A
Data transmission method and safety instrument system
CN113852568B