Event Stream Processing Failover for Data Integrity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computing systems face challenges in maintaining data processing integrity and reliability, particularly in event stream processing environments where failures can lead to service interruptions and significant data loss, especially in mission-critical operations like manufacturing or drilling.
Innovation Solution
An Event Stream Processing (ESP) system with failover capabilities that allows seamless and rapid switching between active and standby nodes, ensuring continuous data processing without service interruption or data loss, even in the event of node failures.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional event stream processing systems are used without failover capabilities, then the system structure is simpler, but the reliability deteriorates due to service interruptions and data loss during node failures
Solution Approach 1:
The system is divided into multiple independent nodes (active node and standby nodes) that can operate autonomously. Each node maintains its own event stream processing capability, allowing the system to segment functionality across multiple units to ensure continuity during failures
Solution Approach 2:
Standby nodes are pre-configured and maintained in readiness before failures occur. The system performs preliminary actions by establishing standby nodes with synchronized data and state information, enabling immediate failover without waiting for failure detection and system reconfiguration
2Reliability
If rapid failover switching is implemented between active and standby nodes, then service continuity is improved, but the complexity of node coordination and state synchronization increases
Solution Approach 1:
Standby nodes are created as copies of the active node, maintaining identical data structures, state information, and processing logic. This copying approach simplifies coordination by ensuring the standby node is already configured and ready, requiring only activation rather than complex reconfiguration during failover
Solution Approach 2:
The system implements continuous monitoring and synchronization mechanisms where the active node's state is continuously fed back to standby nodes. This feedback loop ensures that standby nodes remain synchronized with the active node, enabling seamless transitions while maintaining system consistency
3Loss of information
If multiple standby nodes are maintained for failover purposes, then data loss is minimized during failures, but the resource consumption and system complexity increase
Solution Approach 1:
Different standby nodes serve different purposes or cover different failure scenarios. The system assigns specific roles or data subsets to different standby nodes, optimizing resource usage by having each node specialized for particular failover scenarios rather than maintaining identical full replicas
Data Source
AI summary
A computing device quantifies an expected benefit from a calibrated coefficient of variation (CV) and/or a calibrated service level (SL). The target optimization model determines a number and a time a new requisition is placed for an item at each node of the plurality of nodes. A validation time value is updated using an incremental time value and the process is repeated until the validation time value is greater than or equal to a stop time.


