Rail transit signal system based on off-site four-system redundant architecture and disaster recovery method
Patent Information
- Application Number
- CN202610758556.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-08-21
AI Technical Summary
[0003]现有容灾方案多采用单点冗余或本地备份的方式,虽能提供基本的系统保护,但在发生区域性灾难(如站点断电、火灾)时,仍然存在失效风险
[0023]In this embodiment, by deploying subsystems at two separate sites, with each site's object control system containing two redundant systems, a four-system redundancy architecture is constructed in a different location. This ensures that even if a single site fails entirely, the two systems at the other site can still take over control completely, reducing the risk of single-point failure. When the primary object control system experiences an internal failure, it performs a failover itself without the need for a computer interlocking system, achieving millisecond-level seamless switching and ensuring uninterrupted service. When the primary object control system experiences a complete failure, either the first or second computer interlocking system detects and executes a system-level failover, transferring control to the backup site. This fault type-based hierarchical processing avoids unnecessary remote switching delays while ensuring disaster recovery capabilities under extreme failures, making the overall fault recovery time controllable. Furthermore, an output interlocking mechanism is implemented by setting a controllable switch, ensuring that only one control system has output permissions at any given time, thereby improving system security.
Smart Images

Figure CN122607399A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of rail transit technology, and in particular relates to a rail transit signaling system and disaster recovery method based on a four-system redundancy architecture in a different location. Background Technology
[0002] In the field of rail transit, which has extremely high requirements for safety and reliability, a failure in the core business system could lead to serious operational disruptions or even safety accidents. Therefore, disaster recovery and redundancy mechanisms are needed to ensure safety.
[0003] Existing disaster recovery solutions mostly adopt single-point redundancy or local backup methods, which can provide basic system protection, but there is still a risk of failure when regional disasters occur (such as site power outages or fires). Summary of the Invention
[0004] This application provides a rail transit signaling system and disaster recovery method based on a geographically distributed four-system redundancy architecture, which can improve the disaster recovery capability of the rail transit signaling system.
[0005] In a first aspect, embodiments of this application provide a rail transit signaling system based on a geographically dispersed four-system redundancy architecture, comprising: The first subsystem, deployed at the first physical site, includes a first computer interlocking system and a first object control system. The first object control system includes a first trackside control module and a second trackside control module. The second subsystem is deployed at the second physical site, which is geographically separated from the first physical site. The second subsystem includes a second computer interlocking system and a second object control system. The second object control system includes a third trackside control module and a fourth trackside control module. One of the first, second, third, and fourth trackside control modules is the primary trackside control module, and the other three are backup trackside control modules. The primary trackside control module is used to perform trackside equipment control tasks. Each trackside control module includes a controllable switch, which is used to control the output of the trackside control module to be turned on or off. The primary object control system is used to monitor the health status of its own trackside control modules. In the event of a failure of its primary trackside control module, the primary trackside control module is switched to an unavailable state, and another trackside control module it includes is switched to a new primary trackside control module. The primary object control system is the object control system that includes the primary trackside control module in the first object control system and the second object control system. The first and second computer interlocking systems are communicatively connected to the first and second object control systems, and are used to monitor the health status of each object control system. In the event of a complete failure of the primary object control system, one trackside control module in the backup object control system is switched to the new primary trackside control module, and the controllable switch of the new primary trackside control module is set to the ON state, while the controllable switches of other trackside control modules besides the new primary trackside control module are set to the OFF state. The backup object control system is the object control system in the first and second object control systems other than the primary object control system.
[0006] In one embodiment of the first aspect, the first object control system and the second object control system are respectively connected to the first computer interlocking system and the second computer interlocking system through independent communication networks.
[0007] In one embodiment of the first aspect, the overall machine failure includes at least one of the following: Internal bus dual-network failure, dual-system motherboard failure, external communication dual-network failure, circuit breaker failure.
[0008] In one embodiment of the first aspect, the first computer interlocking system and the second computer interlocking system detect the communication status with the first object control system and the second object control system through periodic heartbeat detection, and receive fault codes sent by the first object control system and the second object control system to determine whether the first object control system and the second object control system have malfunctioned and the type of malfunction.
[0009] In one embodiment of the first aspect, the primary object control system is further configured to report a switching event to the first computer interlocking system and the second computer interlocking system when the primary trackside control module is switched to an unavailable state and another trackside control module contained therein is switched to a new primary trackside control module. The switching event is used to indicate that the primary trackside control module has been switched.
[0010] In one embodiment of the first aspect, the first computer interlocking system or the second computer interlocking system is configured to activate a reverse delay protection mechanism in the event of a complete failure in the primary object control system, to perform the following processing within a first delay window: Verify that the standby control system is in an available state; If it is determined that the standby object control system is in an available state, switch one of the trackside control modules in the standby object control system to the new primary trackside control module, and switch the standby object control system to the new primary object control system. Set the controllable switch of the new primary trackside control module to the ON state, and set the controllable switches of other trackside control modules other than the new primary trackside control module to the OFF state.
[0011] In one embodiment of the first aspect, the window length of the first delay window is determined based on at least one of the following: The duration of communication interruption between the object control system and the computer interlocking system, the duration of local fault handling in the object control system, the processing time of the switch machine board, the maximum duration of a single operation of the switch machine board, and the maximum duration of communication between the object control system and the switch machine board.
[0012] In one embodiment of the first aspect, each trackside control module includes a switch machine control unit, a data acquisition unit, and a processing unit; The acquisition unit includes multiple acquisition nodes, which are used to acquire the position information of the same turnout in parallel to obtain multi-source acquisition data; The acquisition unit is also used to transmit multi-source acquired data to the processing unit; The processing unit is used to generate equipment status information based on multi-source acquired data; and to synchronize the equipment status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system; wherein, the equipment status information is used to indicate the working status of the trackside equipment.
[0013] In one embodiment of the first aspect, the processing unit is configured to perform consistency verification on the multi-source acquired data; if the multi-source acquired data is consistent, generate equipment status information based on the multi-source acquired data, and synchronize the equipment status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
[0014] In one embodiment of the first aspect, the processing unit is configured to initiate a delayed confirmation mechanism in the event of inconsistency between multi-source acquired data, to perform the following processing within a second delay window: Receive data from multiple sources; Perform consistency verification on each received multi-source data collection; If the received multi-source acquisition data is consistent, the equipment status information generated based on the multi-source acquisition data will be synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system; Alternatively, if the received multi-source acquisition data is inconsistent, the device status information is generated based on the multi-source acquisition data using the majority voting principle. If consistent multi-source acquisition data is not received within the second delay window, the final device status information is determined based on the device status information generated during the delay period. The final device status information is then synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
[0015] In one embodiment of the first aspect, the processing unit is configured to determine the device status information with the highest proportion among the device status information generated within the second delay window as the final device status information if consistent multi-source acquisition data is not received within the second delay window.
[0016] In one embodiment of the first aspect, the first computer interlocking system includes a first control system and a second control system; The second computer interlocking system includes a third control system and a fourth control system; One of the first, second, third, and fourth control systems is the primary control system, and the other three are backup control systems. Rail transit signaling systems also include intelligent arbitration platforms; The primary control system is used to synchronize status information and output signals to the backup control system, so that the backup control system is in hot standby mode. The intelligent arbitration platform is used to monitor the health status of each control system. If a failure is detected in the primary control system, the primary control system is switched to an unavailable state, and the backup control system in the primary computer interlocking system is switched to the new primary control system. Alternatively, if a complete failure is detected in the primary computer interlocking system, one of the backup control systems in the backup computer interlocking system is switched to the new primary control system. The primary computer interlocking system is the computer interlocking system to which the primary control system belongs, and the backup computer interlocking systems are the computer interlocking systems in the first and second computer interlocking systems other than the primary computer interlocking system.
[0017] In one embodiment of the first aspect, an intelligent arbitration platform is used to acquire multidimensional information of each control system and calculate the multidimensional information to determine the health status of the control system.
[0018] In one embodiment of the first aspect, the control system is further configured to maintain data version number information, perform version verification on received data based on the data version number information, reject data with duplicate versions, and identify whether data loss has occurred.
[0019] Secondly, embodiments of this application provide a disaster recovery method applied to the rail transit signaling system based on a geographically dispersed four-system redundancy architecture, as described in the first aspect. The method includes: The primary object control system monitors the health status of its own trackside control modules. If the primary trackside control module in the primary object control system fails, the primary trackside control module is switched to an unavailable state, and another trackside control module in the primary object control system is switched to a new primary trackside control module. The health status of each object control system is monitored through the first or second computer interlocking system. In the event of a complete failure of the primary object control system, one trackside control module in the backup object control system is switched to the new primary trackside control module, and the controllable switch of the new primary trackside control module is set to the on state, while the controllable switches of other trackside control modules other than the new primary trackside control module are set to the off state.
[0020] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; The processor implements redundancy methods as described in the second aspect when executing computer program instructions.
[0021] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the redundancy method as described in the second aspect.
[0022] Fifthly, embodiments of this application provide a computer program product in which instructions, when executed by a processor of an electronic device, cause the electronic device to perform the redundant method as described in the second aspect.
[0023] In this embodiment, by deploying subsystems at two separate sites, with each site's object control system containing two redundant systems, a four-system redundancy architecture is constructed in a different location. This ensures that even if a single site fails entirely, the two systems at the other site can still take over control completely, reducing the risk of single-point failure. When the primary object control system experiences an internal failure, it performs a failover itself without the need for a computer interlocking system, achieving millisecond-level seamless switching and ensuring uninterrupted service. When the primary object control system experiences a complete failure, either the first or second computer interlocking system detects and executes a system-level failover, transferring control to the backup site. This fault type-based hierarchical processing avoids unnecessary remote switching delays while ensuring disaster recovery capabilities under extreme failures, making the overall fault recovery time controllable. Furthermore, an output interlocking mechanism is implemented by setting a controllable switch, ensuring that only one control system has output permissions at any given time, thereby improving system security. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 These are schematic diagrams of the structure of a rail transit signaling system based on a geographically distributed four-system redundancy architecture, provided in some embodiments of this application. Figure 2 These are schematic diagrams of the structure of a rail transit signaling system based on a geographically distributed four-system redundancy architecture, provided in some other embodiments of this application; Figure 3 This is a schematic diagram of the remote communication method between the first computer interlocking system and the second computer interlocking system provided in some embodiments of this application; Figure 4 This is a schematic diagram of a remote pulse transmission between a first computer interlocking system and a second computer interlocking system provided in some embodiments of this application; Figure 5 This is a flowchart illustrating the redundant methods provided in some embodiments of this application; Figure 6 These are schematic diagrams of the structure of electronic devices provided in some embodiments of this application. Detailed Implementation
[0026] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples.
[0027] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0028] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0029] OC (Object Controller): The object controller is a key device in the new generation of fully electronic computer interlocking systems. It replaces traditional relay circuits to achieve direct control and status acquisition of trackside equipment. As the direct control unit for trackside equipment, the reliability of the OC system directly affects the safe operation of critical equipment such as turnouts and signals.
[0030] CI (Computer Interlocking): The computer interlocking system is one of the core subsystems of the rail transit signaling system, used to ensure the safety of trains and shunting operations within stations.
[0031] RC (Route Controller): The route controller is an upgraded product of the computer interlocking system (CI). It is the traffic commander of the station interlocking system and is responsible for dynamically allocating train routes.
[0032] PSD (Platform Screen Doors): Platform screen doors.
[0033] ESB (Emergency Stop Button): Emergency stop button.
[0034] Primary: refers to the entity currently selected by the system to perform control or output functions.
[0035] Backup: refers to an entity that is synchronized with the primary entity but does not perform output.
[0036] Before providing a more detailed description of the embodiments of this application, the relevant technologies are introduced. As mentioned earlier, the rail transit signaling system suffers from severe architectural limitations. Both the CI / RC system and the OC system adopt a traditional single-site dual-system redundancy architecture. In the CI / RC system, the system can only be deployed at a single location, with primary / backup switching achieved through two boards in the same cabinet. Under this architecture, system redundancy relies entirely on local hardware; if a failure occurs at the deployment location, the entire control system faces the risk of complete failure. Similarly, the OC system also adopts a single-site dual-system mode, with the primary and backup OCs located at the same location, achieving limited hot standby switching through board-level redundancy. This architecture cannot cope with regional fault risks and has a serious single-point failure vulnerability. Therefore, the core problem of the relevant technologies lies in the lack of true multi-system collaboration capability across different locations.
[0037] Therefore, in order to improve the safety and disaster recovery capabilities of rail transit signaling systems, this application provides a rail transit signaling system and disaster recovery method based on a geographically distributed four-system redundancy architecture.
[0038] The following is a detailed description of the rail transit signaling system based on a geographically dispersed four-system redundancy architecture provided in the embodiments of this application, with reference to the accompanying drawings.
[0039] See Figure 1 This is a schematic diagram of a rail transit signaling system based on a geographically dispersed four-system redundancy architecture, provided in an embodiment of this application. Figure 1 As shown, the system includes a first subsystem and a second subsystem. The first subsystem is deployed at a first physical site, and the second subsystem is deployed at a second physical site, which is geographically separated from the first physical site. The first and second subsystems have the same or equivalent hardware and control capabilities.
[0040] In this embodiment, the first subsystem includes a first computer interlocking system and a first object control system, the first object control system including a first trackside control module and a second trackside control module; the second subsystem includes a second computer interlocking system and a second object control system, the second object control system including a third trackside control module and a fourth trackside control module.
[0041] In this embodiment, the first trackside control module, the second trackside control module, the third trackside control module, and the fourth trackside control module are four independent and mutually redundant trackside control modules. Thus, through a dual-location deployment and a dual-system design at each location, a geographically dispersed four-system redundant architecture is formed, achieving system-level redundancy.
[0042] In some embodiments of this application, each trackside control module includes a switch machine control unit, a data acquisition unit, and a processing unit. The switch machine control unit and the data acquisition unit can both be connected to the processing unit via the internal bus of the trackside control module.
[0043] The switch machine control unit is used to control the switch machine. The switch machine control unit includes switch machine boards. These boards can be installed in dedicated boxes within each trackside control module. Thus, in practical applications, for the same turnout, there are four controllable switch machine boards. Employing a one-master, three-backup board redundancy strategy ensures that a single board failure does not affect system functionality.
[0044] The acquisition unit includes a multi-channel acquisition node for parallel acquisition of the position information of the same turnout. Through the multi-channel acquisition node, multi-source acquisition data of the same turnout can be acquired. The acquisition unit can transmit the multi-source acquisition data to the processing unit through the internal bus.
[0045] In practical applications, multiple acquisition nodes in the acquisition unit can be redundantly configured within the control system using scattered plug-in boxes. These nodes acquire turnout position information by directly acquiring data from the automatic switch contacts. The automatic switch is the core component of the switch machine, primarily undertaking three key functions: indication, control, and safety protection. Compared to acquiring turnout position information through a single acquisition node, using multiple acquisition nodes in parallel ensures that the status of trackside equipment can still be successfully obtained even in the event of a single-point acquisition failure.
[0046] The processing unit is the mainboard of the controller, used to generate equipment status information based on multi-source acquired data transmitted by the acquisition unit, and to synchronize the equipment status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system. The equipment status information indicates the working status of trackside equipment, including turnouts. For example, the equipment status information may include one or more of the following: turnout position, health status, operation status, and fault code. The turnout position indicates the current location of the turnout; the health status indicates the self-test status of the switch machine board; the operation status indicates whether a turnout switching operation is being performed; and the fault code indicates the type of fault occurring in the control system.
[0047] In some embodiments of this application, the processing unit may employ a state fusion algorithm to generate equipment state information. Specifically, after receiving multi-source acquisition data transmitted by the acquisition unit, the processing unit may perform consistency verification on the multi-source acquisition data; if the multi-source acquisition data is consistent, it may generate reliable equipment state information based on the multi-source acquisition data and synchronize the equipment state information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
[0048] In some embodiments of this application, in order to improve the accuracy of consistency verification, before performing consistency verification on multi-source collected data, validity verification can be performed on the multi-source collected data to remove abnormal data.
[0049] In some embodiments of this application, in the event of inconsistencies in multi-source acquired data, in order to improve the accuracy of equipment status information synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system, the processing unit may initiate a delayed confirmation mechanism to perform the following processing within a second delay window: Receive data from multiple sources; Perform consistency verification on each received multi-source data collection; If the received multi-source acquisition data is consistent, the equipment status information generated based on the multi-source acquisition data will be synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system; Alternatively, if the received multi-source acquisition data is inconsistent, the device status information is generated based on the multi-source acquisition data using the majority voting principle. If consistent multi-source acquisition data is not received during the delay period, the final device status information is determined based on the device status information generated during the delay period, and the final device status information is synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
[0050] Multiple data acquisition nodes collect the position information of the same turnout. If a data acquisition node experiences an error due to hardware damage, communication interference, or contact oxidation, leading to inconsistencies in the multi-source data, this error usually accounts for a low percentage of the total acquisition count; for example, only one error may occur out of three nodes. By employing a majority voting principle, a few abnormal data points can be automatically filtered out, and the correct status can be output, avoiding misjudgments of the entire system due to a single node failure.
[0051] Through the above-mentioned delay processing mechanism, a real fault is only confirmed after the inconsistent state of multi-source data acquisition continues for the entire delay window, avoiding misjudgment of the state caused by transient anomalies such as communication jitter and contact glitch, and reducing unnecessary switching.
[0052] In some embodiments of this application, if consistent multi-source acquisition data is not received within the second delay window, the processing unit can determine the proportion of each type of device status information generated within the second delay window, determine the device status information with the highest proportion as the final device status information, and synchronize the final device status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system. Here, the proportion of each type of device status information within the second delay window refers to the ratio of the number of that type of device status information generated within the second delay window to the total number of device status information generated within the second delay window. For example, if a total of 5 device status information items are generated within the second delay window, and 2 of them are "faults," then the proportion of the "fault" device status information is 40%.
[0053] By using this method, the equipment information with the highest proportion is taken as the final equipment status information, which enables the status information of trackside equipment to be reported to the first computer interlocking system and the second computer interlocking system, and improves the accuracy of the reported equipment status information.
[0054] In some embodiments of this application, within the second delay window after the delay processing mechanism is started, the processing unit can maintain the original output state without interrupting normal control.
[0055] In some embodiments of this application, the length of the second delay window of the delay processing mechanism can be determined based on the time required for the processing unit to execute a complete processing chain. For example, it can be determined based on the time required for communication interruption, the processing time, and the execution time. Furthermore, the delay duration can be adjusted according to the on-site communication quality and equipment characteristics. This allows for setting different delay durations for different equipment models and communication conditions, thus providing engineering adaptability. For example, the delay duration of the delay processing mechanism can be set to 4-5 seconds.
[0056] In practical applications, one of the four trackside control modules (first, second, third, and fourth) is the primary trackside control module, while the other three are backup trackside control modules. The primary trackside control module performs trackside equipment control tasks. During operation, the primary trackside control module synchronizes its status and output to the three backup trackside control modules, enabling intelligent coordination and rapid fault switching among the four modules. For ease of distinction, the object control system including the primary trackside control module in the first and second object control systems will be referred to as the primary object control system, and the other object control system will be referred to as the backup object control system.
[0057] In some embodiments of this application, each object control system can monitor its included trackside control modules to determine whether each trackside control module has failed. If the primary object control system detects a failure in its included primary trackside control module (such as a single board or system failure), it can switch the primary trackside control module to an unavailable state and switch another included trackside control module to a new primary trackside control module. That is, when the primary object control system experiences an internal single-system failure, it can autonomously complete the primary / backup switch of its internal trackside control modules. In this way, the impact of the failure can be mitigated promptly through internal autonomous switching.
[0058] The first and second computer interlocking systems are communicatively connected to the first and second object control systems. They are used to switch a trackside control module from the backup object control system to a new primary trackside control module in the event of a complete failure of the primary object control system in either system. A complete failure refers to a failure that renders all trackside control modules in the primary object control system unusable. For example, a complete failure may include at least one of the following: internal bus dual-network failure, dual-system mainboard failure, external communication dual-network failure, or circuit breaker failure. Thus, when a complete failure occurs in the primary object control system due to internal reasons or a single point of failure such as a power outage at the site, continued operation can be achieved by switching to a trackside control module deployed at another site, improving system disaster recovery capability and reliability.
[0059] In some embodiments of this application, the first object control system and the second object control system are respectively connected to the first computer interlocking system and the second computer interlocking system through independent communication networks to ensure the reliability of the communication link.
[0060] In some embodiments of this application, the first computer interlocking system and the second computer interlocking system detect the communication status with the first object control system and the second object control system through periodic heartbeat detection, and receive fault codes sent by the first object control system and the second object control system to determine whether the first object control system or the second object control system has malfunctioned and the type of malfunction.
[0061] Based on the above description, the embodiments of this application automatically execute corresponding control system failover strategies based on fault types. The failover principles and fault classifications are shown in the table below: Through the aforementioned intelligent failover mechanism based on fault type, the remote object control system can quickly and safely complete the primary / backup switchover under various fault scenarios, thereby improving the system's disaster recovery capability.
[0062] In some embodiments of this application, if multiple trackside control modules output drive signals simultaneously, it may lead to turnout malfunction or equipment damage. Therefore, to improve system safety, ensuring that only one trackside control module in the first and second object control systems has output authority at any given time, each trackside control module can be equipped with a controllable switch. The controllable switch is used to control the output of its respective trackside control module to be on or off. The first and second computer interlocking systems, as the main control devices of the object control system, can be used to set the state of the controllable switches. Based on this, when the system starts, the controllable switch of the primary trackside control module can be set to the on state, and the controllable switches of other trackside control modules besides the primary trackside control module can be set to the off state. In the event of a switchover, the first or second computer interlocking system can set the controllable switch of the new primary trackside control module to the on state, and the controllable switches of other trackside control modules besides the new primary trackside control module to the off state. In this way, by implementing physical interlocking of the output circuit through hardware, only one switch machine board has the right to output at any given time, thereby improving output safety.
[0063] In some embodiments of this application, the controllable switch may be a relay.
[0064] In some embodiments of this application, to further enhance the output safety of the system, the output permissions of the trackside control module can be verified at the software level through the first computer interlocking system and the second computer interlocking system to prevent erroneous output. Specifically, the first computer interlocking system and the second computer interlocking system can receive the output of the trackside control module. When the output of the switch machine board in at least two trackside control modules is received simultaneously, it is determined that there is an erroneous output. At this time, in order to avoid misoperation, the switch machine can be kept in its original state.
[0065] In this embodiment, by deploying subsystems at two separate sites, with each site's object control system containing two redundant systems, a four-system redundancy architecture is constructed in a different location. This ensures that even if a single site fails entirely, the two systems at the other site can still take over control completely, reducing the risk of single-point failure. When the primary object control system experiences an internal failure, it performs a failover itself without the need for a computer interlocking system, achieving millisecond-level seamless switching and ensuring uninterrupted service. When the primary object control system experiences a complete failure, either the first or second computer interlocking system detects and executes a system-level failover, transferring control to the backup site. This fault type-based hierarchical processing avoids unnecessary remote switching delays while ensuring disaster recovery capabilities under extreme failures, making the overall fault recovery time controllable. Furthermore, an output interlocking mechanism is implemented by setting a controllable switch, ensuring that only one control system has output permissions at any given time, thereby improving system security.
[0066] In some embodiments of this application, when the primary trackside control system detects a fault in the primary trackside control module, for example, when the primary trackside control module experiences a switch machine board drive malfunction, the primary trackside control system automatically performs a primary / backup switchover within the control unit. This switches the primary trackside control module to an unavailable state and replaces its other trackside control module with a new primary trackside control module, thus enabling the backup switch machine board. After the switchover is complete, the primary trackside control system reports the switchover event to the first and second computer interlocking systems. This event indicates that the primary trackside control module has been switched. This allows the first and second computer interlocking systems to promptly know which trackside control module is currently being used for drive, thereby reducing the impact on the overall system operation. During the switchover process, output continuity is maintained, achieving seamless switching and preventing turnout malfunctions.
[0067] In some embodiments of this application, when a complete system failure occurs in the primary object control system, a system-level failover and off-site disaster recovery switch are performed by either the first or second computer interlocking system. Specifically, when the first or second computer interlocking system detects a complete system failure in the primary object control system, such as a dual-system failure or a communication terminal failure, a failover delay protection mechanism is activated to perform the following processing within a first delay window: Verify that the standby control system is in an available state; If it is determined that the standby object control system is in an available state, switch one of the trackside control modules in the standby object control system to the new primary trackside control module, and switch the standby object control system to the new primary object control system. Set the controllable switch of the new primary trackside control module to the ON state, and set the controllable switches of other trackside control modules other than the new primary trackside control module to the OFF state.
[0068] Through the above methods, during the delay period, either the first or second computer interlocking system verifies the availability of the backup control system, such as determining whether communication is normal, whether the mainboard is online, and whether the switch machine board is available. If the backup control system itself has a fault, the switchover is aborted and an alarm is triggered to avoid handing control over to an unavailable system, ensuring a high success rate for disaster recovery switching. After the switchover is completed, output interlocking is enabled to ensure that only the new primary trackside control module has output permissions at any given time, thereby reducing potential dual-primary conflicts at the switchover boundary and preventing damage to the turnout caused by simultaneous driving by two trackside control modules. After the switchover, the primary and backup status of the control system is updated and synchronized to other relevant systems. This ensures that subsequent fault detection, load balancing, and maintenance operations are based on the correct primary and backup roles, avoiding decision-making errors caused by state asynchrony.
[0069] In some embodiments of this application, the length of the first delay window may be determined based on at least one of the following: The duration of communication interruption between the object control system and the computer interlocking system, the duration of local fault handling in the object control system, the processing time of the switch machine board, the maximum duration of a single operation of the switch machine board, and the maximum duration of communication between the object control system and the switch machine board.
[0070] By employing the above-described configuration, when the first delay window is greater than or equal to the maximum single operation duration of the switch machine board, it can be ensured that if the turnout is switching when the reverse switch is initiated, the system will wait for it to complete its transition and lock before switching control. This prevents dangerous situations caused by power outages or switching midway through the turnout. When the first delay window is greater than or equal to the communication interruption duration between the object control system and the computer interlocking system, if communication is briefly interrupted due to load fluctuations or electromagnetic interference, the system will not immediately initiate a reverse switch but will wait for communication to resume. This avoids unnecessary disaster recovery actions caused by brief fluctuations and improves system stability. When the first delay window is greater than or equal to the local fault handling duration of the object control system, the object control system is allowed to attempt local self-recovery when an internal anomaly is detected. If recovery can be achieved within the handling duration, there is no need to switch to a backup station, reducing the frequency of system-level reverse switches and lowering the decision-making burden on the first and second computer interlocking systems. When the first delay window is greater than or equal to the processing time of the switch machine board and the communication time between the object control system and the switch machine board, it can be ensured that the current electronic processing can be completed before the switch machine board receives the switching command or stop command, so as to avoid sudden changes in the output signal and achieve smooth handover.
[0071] In some embodiments of this application, the length of the first delay window can be set to a value greater than or equal to the sum of the communication interruption execution duration between the object control system and the computer interlocking system, the local fault handling duration of the object control system, the processing duration of the switch machine board, the maximum duration of a single operation of the switch machine board, and the maximum communication duration between the object control system and the switch machine board.
[0072] For example, if the communication interruption execution time between the object control system and the computer interlocking system is 1.5 seconds, the local fault handling time of the object control system is 0.3 seconds, the processing time of the switch machine board is 0.15 seconds, the maximum single operation time of the switch machine board is 15 seconds, and the maximum communication time between the object control system and the switch machine board is 0.3 seconds, then the length of the first delay window can be set to be greater than or equal to the sum of the above durations. For example, the length of the first delay window can be set to 18 seconds.
[0073] The first delay window, set in the above manner, covers the complete time chain from the occurrence of a fault to a safe switchover, ensuring that the turnout does not switch prematurely while it is in operation, and avoiding indefinite waiting. For non-operational scenarios, the actual switchover time can be shortened according to actual needs.
[0074] In some embodiments of this application, the following security controls are implemented to further enhance system security: When communication with the switch machine board in the primary trackside control module is interrupted, the primary trackside control module can be automatically downgraded to standby mode and no longer participate in output driving. This reduces the risk of the board becoming stuck in primary mode due to communication failure and avoids output loss of control.
[0075] If the primary trackside control module malfunctions during turnout switching, the primary and backup control systems will be switched on only after the turnout switching action is completed. This improves the integrity of turnout switching and eliminates dangerous situations caused by power outages midway.
[0076] When a power supply abnormality is detected in the primary control system or primary trackside control module, such as a voltage drop, circuit breaker tripping, or power module failure, all output drive signals are immediately cut off, causing the switch machine to stop operating and remain in its current position or in a power-off locked state. This reduces uncontrolled output.
[0077] The following section explains the control system switching process using a specific scenario.
[0078] Scenario 1: Handling a partial fault in the primary object control system: Fault phenomenon: The controllable switch of the main railside control module in the main object control system cabinet has malfunctioned.
[0079] Processing flow: The primary object control system automatically switches to the backup railside control module in the same cabinet, thus completing disaster recovery seamlessly.
[0080] Impact assessment: Zero business interruption will be achieved, and system reliability will not be affected.
[0081] Scenario 2: Disaster recovery handling for a complete failure of the primary object control system: Fault symptom: A power outage at the site where the primary object control system is deployed causes the primary object control system to completely fail.
[0082] Processing procedure: When the first or second computer interlocking system detects a communication interruption in the primary object control system, it initiates a delayed switchover process to switch to the trackside control module in the backup object control system.
[0083] Switching result: Services resume within the first delay window, and the standby object control system takes over control.
[0084] This application, through the aforementioned technology, achieves a comprehensive upgrade of the object control system from traditional single-site dual-system redundancy to off-site four-system redundancy, providing a highly reliable and highly available execution-level control solution for rail transit signaling systems. In some embodiments of this application, one of the first physical site and the second physical site is the primary site, and the other is a backup site, also known as an off-site disaster recovery site. Which of the two physical sites serves as the primary site can be specified according to actual needs; this embodiment does not specifically limit this. When the rail transit signaling system is operating normally, the output of the primary site is selected as the final system output for executing signal control tasks. For example, if the first physical site is the primary site and the second physical site is the backup site, then when the rail transit signaling system is operating normally, the output of the first subsystem is selected as the final system output. The backup site is mainly used for temporary control when the primary site fails. Therefore, to reduce costs, the subsystem deployed at the backup site can only possess the main functions of the subsystem deployed at the primary site, thus reducing the cost of the subsystem deployed at the backup site. The main functions can be set based on actual needs. For example, the main functions may include, but are not limited to, route control, signal and turnout control, etc.
[0085] In some embodiments of this application, the two trackside control modules deployed at the primary site can be housed in two independent cabinets. For example, the trackside control modules deployed at the primary site include two independent cabinets, OC-1 and OC-2. The two trackside control modules deployed at the backup site can be housed in the same cabinet. For example, the backup site can be housed in an independent cabinet OC-S, which includes two trackside control modules, OC-S-1 and OC-S-2. This reduces the deployment cost of the backup site.
[0086] For example, such as Figure 2 As shown, the first physical station is the primary station, in which the two independent cabinets OC-1 and OC-2 can collect relevant information about signals, PSD, ESB and turnouts respectively, and control signals, PSD, ESB and turnouts. The second physical station serves as a backup station, in which the OC-S cabinet can collect relevant information about turnouts and control turnout machines.
[0087] In some embodiments of this application, the first computer interlocking system and the second computer interlocking system can also form a geographically distributed four-system redundancy architecture similar to the first object control system and the second object control system to improve system security. Specifically, the first computer interlocking system includes a first control system and a second control system; the second computer interlocking system includes a third control system and a fourth control system; wherein, one of the first, second, third, and fourth control systems is the primary control system, and the other three are backup control systems. The primary control system is used to synchronize status information and output signals to the backup control system so that the backup control system is in a hot standby state.
[0088] To address the challenge of deep collaboration among the first computer interlocking system, the second computer interlocking system, the first object control system, and the second object control system in a geographically dispersed four-system architecture, the rail transit signaling system also includes a unified intelligent arbitration platform. Through centralized decision-making and distributed execution, seamless collaboration and rapid fault recovery among the four systems are achieved.
[0089] An intelligent arbitration platform is used to monitor the health status of each control system. If a failure is detected in the primary control system, the primary control system is switched to an unavailable state, and the backup control system in the primary computer interlocking system is switched to the new primary control system. Alternatively, if a complete failure is detected in the primary computer interlocking system, one of the backup control systems in the backup computer interlocking system is switched to the new primary control system. The primary computer interlocking system is the computer interlocking system to which the primary control system belongs, and the backup computer interlocking systems are the computer interlocking systems in the first and second computer interlocking systems other than the primary computer interlocking system. In the above embodiments, by deploying the intelligent arbitration platform and multiple subsystems deployed in different locations, and setting up dual control systems in the subsystems, rapid detection and seamless switching of control systems in the event of communication anomalies or equipment failures are achieved. This improves operational security, fault tolerance, and business continuity, and can meet the disaster recovery requirements in scenarios with high security requirements.
[0090] In some embodiments of this application, the first computer interlocking system may be a CI system or an RC system. Similarly, the second computer interlocking system may also be a CI system or an RC system. This embodiment does not specifically limit this.
[0091] In some embodiments of this application, each control system may employ a two-out-of-two secure computing unit, wherein the computing unit may be a CPU processor.
[0092] For example, the first computer interlocking system includes two control systems, CI-A and CI-B. CI-A includes computing unit 1 and computing unit 2, forming a 2-out-of-2 safety computing unit. CI-B includes computing unit 3 and computing unit 4, forming another independent 2-out-of-2 safety computing unit. The second computer interlocking system includes two control systems, RC-A and RC-B. RC-A includes computing unit 5 and computing unit 6, forming a 2-out-of-2 safety computing unit. RC-B includes computing unit 7 and computing unit 8, forming yet another 2-out-of-2 safety computing unit.
[0093] See Figure 3 This is a schematic diagram illustrating the remote communication method between the first and second computer interlocking systems. The "communication control board" refers to the communication control board.
[0094] See Figure 4 This is a schematic diagram of the remote pulse transmission between the first computer interlocking system and the second computer interlocking system.
[0095] The intelligent arbitration platform serves as the central hub for intelligent arbitration and switching in the rail transit signaling system. It collects and compares in real time the output status, health, and network latency of four systems (such as CI-A, CI-B, RC-A, and RC-B) of the first and second computer interlocking systems, and controls the switching between primary and backup states.
[0096] In practical applications, the primary control system sends its outputs and system status to each backup control system via a dedicated communication link, ensuring that the backup control systems always maintain the same control state as the primary control system, thus achieving hot standby. For example, if the CI-A system in the first computer interlocking system is the primary control system, then the CI-A system sends its outputs and system status to the CI-B system, RC-A system, and RC-B system via dedicated communication links.
[0097] In some embodiments of this application, the intelligent arbitration platform, as the core decision-making unit for the collaboration of the four systems, can continuously collect comprehensive information such as the operating status, load, and network quality of each control system to assess the health status of each control system and thus determine whether a primary / backup switch is required for the control system.
[0098] In some embodiments of this application, to improve the accuracy of health status assessment, the intelligent arbitration platform can monitor multiple dimensions of indicators for each control system, such as communication status, processing performance, and resource utilization. Based on these multi-dimensional indicators, a weighted scoring algorithm is used to determine the health score of each control system. In this way, the health status is determined comprehensively by multiple dimensions of indicators, reducing misjudgments caused by using a single indicator.
[0099] In some embodiments of this application, the weight coefficients corresponding to each dimension in the weighted scoring algorithm can be dynamically adjusted according to the importance of the dimension. This allows for adaptation to different operational scenarios.
[0100] In some embodiments of this application, in order to optimize the switching effect, the intelligent arbitration platform can also predict the probability of the control system failing. By comparing the failure probability of the primary control system with a probability threshold, it can determine whether there is a failure risk in the primary control system. If a failure risk is determined, a primary-backup switch can be performed in advance before the primary control system fails, thereby achieving predictive maintenance and reducing emergency switching caused by sudden failures.
[0101] In some embodiments of this application, the intelligent arbitration platform, when determining the health status of the control system, may consider not only multiple current indicators such as communication status, processing performance, and resource utilization, but also the failure probability as a one-dimensional data point for weighted scoring calculation. This reduces the likelihood of switching to a control system with a high failure probability during primary / standby switching, thus minimizing the occurrence of secondary failures. In some embodiments of this application, the intelligent arbitration platform can identify potential failure risks of the control system in advance through historical data analysis and machine learning algorithms. For example, a fault identification model can be pre-trained based on historical operating data, and then the failure probability of the control system can be predicted using the fault identification model based on real-time acquired operating data.
[0102] In some embodiments of this application, the arbitration mechanism adopted by the intelligent arbitration platform may include the following: a. Routine Operation: When the intelligent arbitration platform determines that the primary computer interlocking system is functioning normally, the control switch selects the output of the primary computer interlocking system as the final system output. For example, if the first computer interlocking system is the primary computer interlocking system, then either the CI-A system or the CI-B system output can be selected as the final system output.
[0103] b. Switching within the primary system: If the primary control system in the primary computer interlocking system fails, the intelligent arbitration platform can redirect the system to another control system within the primary computer interlocking system as the final system output, achieving fault tolerance within the primary site. For example, if the first computer interlocking system is the primary computer interlocking system, and the CI-A system is the primary control system, then when the CI-A system fails, the CI-B system will be switched to the new primary control system, and the output of the CI-B system will be used as the final system output.
[0104] c. Inter-system switching: If the primary computer interlocking system experiences a complete failure, such as a site power outage, the intelligent arbitration platform can immediately control the switching switch to seamlessly transfer the system output to the healthy system of the backup computer interlocking system, thus achieving disaster recovery in a different location. For example, if the first computer interlocking system is the primary computer interlocking system, and a power outage occurs at the site where the first computer interlocking system is located, the intelligent arbitration platform can switch the RC-A or RC-B system in the second computer interlocking system to the new primary control system, and use the output of the new primary control system as the final system output. Optionally, when the intelligent arbitration platform can switch the control system in the backup computer interlocking system to the new primary control system, it can select the control system with the best health status from each control system in the backup computer interlocking system as the new primary control system. This can reduce the occurrence of secondary failures.
[0105] d. Dynamic Upgrade / Downgrade: The intelligent arbitration platform supports online maintenance and upgrades. After completing the primary / backup switchover, the platform can monitor whether the initially designated primary computer interlocking system has completed fault repair. If repaired, the platform can, according to its strategy, switch back to the primary computer interlocking system without disruption, restoring the initial deployment state. For example, after the initially designated primary computer interlocking system completes fault repair, the platform can comprehensively consider multiple optimization objectives such as system stability, switchover time, and business impact, selecting a time that will not interfere with system operations to switch back to the primary computer interlocking system. This reduces system downtime.
[0106] Through the above scheme, the computer interlocking system, with its dual-site deployment and dual-system design, constructs four independent and mutually redundant control systems, significantly increasing redundancy. The off-site deployment of the primary and backup systems, combined with the intelligent switching control of the intelligent arbitration platform, can withstand site-level disasters and improve business continuity.
[0107] In some embodiments of this application, under a geographically dispersed four-system architecture, at least one of the following methods can be used to improve data consistency among the systems: Method 1: Adopt an improved two-phase commit protocol to ensure consistency of data operations across systems.
[0108] Method 2: Establish a data version management mechanism to prevent data conflicts and loss through data version verification. For example, each trackside control module and each control system can maintain data version number information. Based on the data version number information, the received data is verified to reject duplicate data, thus preventing data conflicts and identifying whether data loss has occurred.
[0109] Method 3: Real-time data synchronization between the four systems is achieved through a high-speed private network.
[0110] In some embodiments of this application, the intelligent arbitration platform, in addition to performing primary / backup control system switching, can also be used to implement load balancing and resource scheduling. Specifically, the intelligent arbitration platform may employ intelligent resource scheduling and load balancing mechanisms: Dynamic load assessment: Real-time monitoring of system load, prediction of load trends, and dynamic assessment of system load; dynamic allocation of control tasks based on system load and capacity; dynamic adjustment of system resource allocation based on business needs, thereby achieving elastic scaling of resources.
[0111] In some embodiments of this application, at least one of the following basic designs can be adopted to improve the security of communication and data consistency between four remote systems: Design 1: Construct a multi-layered secure communication system. This includes, for example, transport layer encryption, two-way authentication, communication link redundancy, and link quality monitoring. Transport layer encryption can include end-to-end encryption of inter-system communication data using national cryptographic algorithms. Two-way authentication can be achieved by establishing a two-way authentication mechanism based on digital certificates. Communication link redundancy can include establishing multiple physically isolated communication links between each system. Link quality monitoring can include real-time monitoring of the quality of each link to achieve intelligent routing selection.
[0112] Design 2: Establish high-precision (e.g., nanosecond-level accuracy) system-wide clock synchronization. For example, a dual-mode GPS and BeiDou clock source can be used, combined with a high-precision clock server to achieve multi-source clock synchronization. Clock drift can also be predicted and compensated based on linear regression algorithms. Furthermore, all system events can be precisely time-stamped to ensure consistent event sequence.
[0113] Design 3: Employ an improved distributed consistency protocol to ensure strong data consistency across the four systems. For example, derivative algorithms such as Paxos can be used to synchronize four replicas, improving data consistency across multiple replicas. Alternatively, rapid fault detection based on heartbeat mechanisms and timeout checks can be employed to promptly identify system failures. Furthermore, the latest data status can be automatically synchronized after the failed system recovers.
[0114] Design 4: Employ a complete operation audit and security traceability mechanism. For example, it can record information such as the executor, time, and result of all critical operations, monitor security events in real time, issue timely alarms, and discover abnormal operation patterns based on behavioral analysis to issue timely warnings.
[0115] Design 5: Performance optimization for large-scale data synchronization. For example, lossless compression algorithms can be used to compress and transmit data, reducing network bandwidth consumption. Systems can synchronize only changed data, reducing the amount of data to be synchronized. Multiple operation requests can be processed in batches to improve processing efficiency.
[0116] In some embodiments of this application, considering that when a wide area network (WAN) experiences a split-brain (i.e., network outage but systems deployed at both sites are still available), relying solely on Paxos-like protocols may result in temporary unavailability due to election timeouts. Therefore, to address the split-brain and timing issues related to cross-regional data consistency, a persistent distributed arbitration storage, such as a high-speed SSD-based SAN or distributed KV storage, can be deployed at a location other than the first and second sites. This is referred to as the arbitration disk. Based on this, each control system can create a periodically renewed lease with a TTL (Time To Live) validity period on the arbitration disk before becoming the primary control system. When a network partition occurs, the lease mechanism on the arbitration disk ensures that at most one control system within the partition can hold a valid lease, thereby granting that control system output rights as the primary control system and reducing the simultaneous output of two or more control systems due to network partitions.
[0117] In addition, for primary control systems with existing leases, if the lease cannot be renewed within the validity period, it will automatically be downgraded to standby status and stop outputting.
[0118] Based on the rail transit signaling system with a geographically dispersed four-system redundancy architecture provided in the above embodiments, this application also provides a specific implementation of the disaster recovery method applied to the above-mentioned rail transit signaling system with a geographically dispersed four-system redundancy architecture. The disaster recovery method proposed in this application will be described below with reference to embodiments.
[0119] See Figure 5 The redundancy method provided in this application includes the following steps S510-S540.
[0120] S510. Monitor the health status of its own trackside control modules through the main object control system.
[0121] S520. In the event of a failure of the primary trackside control module in the primary object control system, the primary trackside control module is switched to an unavailable state, and another trackside control module in the primary object control system is switched to a new primary trackside control module.
[0122] S530. Monitor the health status of each object control system through the first computer interlocking system or the second computer interlocking system.
[0123] S540. In the event of a complete failure of the primary object control system, through the first or second computer interlocking system, one trackside control module in the backup object control system is switched to a new primary trackside control module, and the controllable switch of the new primary trackside control module is set to the ON state, while the controllable switches of other trackside control modules besides the new primary trackside control module are set to the OFF state.
[0124] Figure 6 A schematic diagram of the hardware structure of the electronic device provided in an embodiment of this application is shown.
[0125] Electronic device 600 may include processor 601 and memory 602 storing computer program instructions.
[0126] Specifically, the processor 601 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0127] Memory 602 may include a large-capacity memory for data or instructions. For example, and not limitingly, memory 602 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 602 may include removable or non-removable (or fixed) media. Where appropriate, memory 602 may be internal or external to electronic device 600. In a particular embodiment, memory 602 is a non-volatile solid-state memory. Memory 602 may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, generally, memory 602 includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it performs the operations described in any of the redundant methods in the above embodiments.
[0128] The processor 601 implements any of the redundancy methods described in the above embodiments by reading and executing computer program instructions stored in the memory 602.
[0129] In one example, electronic device 600 may further include communication interface 603 and bus 610. For example, Figure 6 As shown, the processor 601, memory 602, and communication interface 603 are connected through bus 610 and complete communication with each other.
[0130] The communication interface 603 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0131] Bus 610 includes hardware, software, or both, that couples components of electronic device 600 together. For example, and not limitingly, the bus may include Accelerated Graphics Port (AGP) or other graphics buses, Enhanced Industry Standard Architecture (EISA) buses, Front Side Bus (FSB), HyperTransport (HT) interconnects, Industry Standard Architecture (ISA) buses, Infinite Bandwidth Interconnects, Low Pin Count (LPC) buses, memory buses, Microchannel Architecture (MCA) buses, Peripheral Component Interconnect (PCI) buses, PCI-Express (PCI-X) buses, Serial Advanced Technology Attachment (SATA) buses, Video Electronics Standards Association Local (VLB) buses, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 610 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, any suitable bus or interconnect is contemplated herein.
[0132] Furthermore, in conjunction with the redundancy methods in the above embodiments, this application embodiment can provide a computer storage medium for implementation. This computer storage medium stores computer program instructions; when these computer program instructions are executed by a processor, they implement any of the redundancy methods in the above embodiments.
[0133] This application also provides a computer program product, including a computer program, which, when executed, implements any of the redundant methods described in the above embodiments.
[0134] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0135] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0136] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0137] The aspects of this disclosure have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by special-purpose hardware performing the specified functions or actions, or can be implemented by a combination of special-purpose hardware and computer instructions.
[0138] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A rail transit signaling system based on a geographically dispersed four-system redundancy architecture, characterized in that, include: The first subsystem, deployed at the first physical site, includes a first computer interlocking system and a first object control system. The first object control system includes a first trackside control module and a second trackside control module. The second subsystem is deployed at the second physical site, which is geographically separated from the first physical site. The second subsystem includes a second computer interlocking system and a second object control system. The second object control system includes a third trackside control module and a fourth trackside control module. One of the first trackside control module, the second trackside control module, the third trackside control module, and the fourth trackside control module is the main trackside control module, and the other three are backup trackside control modules. The main trackside control module is used to perform trackside equipment control tasks. Each trackside control module includes a controllable switch, which is used to control the output of the trackside control module to be turned on or off. The primary object control system is used to monitor the health status of its own trackside control modules. In the event of a failure of its primary trackside control module, the primary trackside control module is switched to an unavailable state, and another trackside control module it includes is switched to a new primary trackside control module. The primary object control system is the object control system that includes the primary trackside control module in the first object control system and the second object control system. The first and second computer interlocking systems are communicatively connected to the first and second object control systems, respectively, and are used to monitor the health status of each object control system. In the event of a complete failure of the primary object control system, one trackside control module in the backup object control system is switched to a new primary trackside control module, and the controllable switch of the new primary trackside control module is set to the ON state, while the controllable switches of other trackside control modules besides the new primary trackside control module are set to the OFF state. The backup object control system refers to the object control systems in the first and second object control systems other than the primary object control system.
2. The system according to claim 1, characterized in that, The first object control system and the second object control system are respectively connected to the first computer interlocking system and the second computer interlocking system through independent communication networks.
3. The system according to claim 1, characterized in that, The overall machine failure includes at least one of the following: Internal bus dual-network failure, dual-system motherboard failure, external communication dual-network failure, circuit breaker failure.
4. The system according to claim 3, characterized in that, The first computer interlocking system and the second computer interlocking system periodically detect the communication status with the first object control system and the second object control system, and receive fault codes sent by the first object control system and the second object control system, in order to determine whether the first object control system and the second object control system have malfunctioned and the type of malfunction.
5. The system according to claim 1, characterized in that, The primary trackside control system is further configured to report a switching event to the first computer interlocking system and the second computer interlocking system when the primary trackside control module is switched to an unavailable state and another trackside control module contained therein is switched to a new primary trackside control module. The switching event is used to indicate that the primary trackside control module has been switched.
6. The system according to claim 1, characterized in that, The first computer interlocking system or the second computer interlocking system is used to activate a reverse delay protection mechanism in the event of a complete failure in the primary object control system, so as to perform the following processing within a first delay window: Verify whether the backup object control system is in an available state; If it is determined that the backup object control system is in an available state, switch one of the trackside control modules in the backup object control system to a new primary trackside control module, and switch the backup object control system to a new primary object control system. Set the controllable switch of the new primary trackside control module to the ON state, and set the controllable switches of other trackside control modules besides the new primary trackside control module to the OFF state.
7. The system according to claim 6, characterized in that, The window length of the first delay window is determined based on at least one of the following: The duration of communication interruption between the object control system and the computer interlocking system, the duration of local fault handling in the object control system, the processing time of the switch machine board, the maximum duration of a single operation of the switch machine board, and the maximum duration of communication between the object control system and the switch machine board.
8. The system according to claim 1, characterized in that, Each trackside control module includes a switch machine control unit, a data acquisition unit, and a processing unit; The acquisition unit includes multiple acquisition nodes, which are used to acquire the position information of the same turnout in parallel to obtain multi-source acquisition data; The acquisition unit is also used to transmit the multi-source acquisition data to the processing unit; The processing unit is used to generate equipment status information based on the multi-source acquired data; and to synchronize the equipment status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system; wherein the equipment status information is used to indicate the working status of the trackside equipment.
9. The system according to claim 8, characterized in that, The processing unit is used to perform consistency verification on the multi-source acquired data; if the multi-source acquired data is consistent, it generates equipment status information based on the multi-source acquired data and synchronizes the equipment status information to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
10. The system according to claim 9, characterized in that, The processing unit is configured to initiate a delayed confirmation mechanism in the event of inconsistency in the multi-source collected data, and to perform the following processing within a second delay window: Receive data from multiple sources; Perform consistency verification on each received multi-source data collection; If the received multi-source acquisition data is consistent, the equipment status information generated based on the multi-source acquisition data will be synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system; Alternatively, if the received multi-source acquisition data is inconsistent, the device status information is generated based on the multi-source acquisition data using the majority voting principle. If consistent multi-source acquisition data is not received within the second delay window, the final device status information is determined based on the device status information generated during the delay period. The final device status information is then synchronized to the first computer interlocking system, the second computer interlocking system, and other object control systems in the rail transit signaling system.
11. The system according to claim 10, characterized in that, The processing unit is configured to determine the device status information with the highest proportion among the device status information generated within the second delay window as the final device status information if consistent multi-source acquisition data is not received within the second delay window.
12. The system according to claim 1, characterized in that, The first computer interlocking system includes a first control system and a second control system; The second computer interlocking system includes a third control system and a fourth control system; One of the first control system, the second control system, the third control system, and the fourth control system is the primary control system, and the other three are backup control systems. The rail transit signaling system also includes an intelligent arbitration platform; The primary control system is used to synchronize status information and output signals to the backup control system, so that the backup control system is in a hot standby state. The intelligent arbitration platform is used to monitor the health status of each control system; if a failure is determined in the primary control system, the primary control system is switched to an unavailable state, and the backup control system in the primary computer interlocking system is switched to a new primary control system; or, if a complete failure is determined in the primary computer interlocking system, one of the backup control systems in the backup computer interlocking system is switched to a new primary control system; wherein, the primary computer interlocking system is the computer interlocking system to which the primary control system belongs, and the backup computer interlocking system is the computer interlocking system other than the primary computer interlocking system in the first computer interlocking system and the second computer interlocking system.
13. The system according to claim 12, characterized in that, The intelligent arbitration platform is used to acquire multi-dimensional information of each control system and calculate the multi-dimensional information to determine the health status of the control system.
14. The system according to claim 12, characterized in that, The control system is also used to maintain data version number information, perform version verification on the received data based on the data version number information, reject data with duplicate versions, and identify whether data loss has occurred.
15. A disaster recovery method, characterized in that, The method, applied to any one of claims 1-14, for a rail transit signaling system based on a geographically dispersed four-system redundancy architecture, comprises: The primary object control system monitors the health status of its own trackside control modules. If the primary trackside control module in the primary object control system fails, the primary trackside control module is switched to an unavailable state, and another trackside control module in the primary object control system is switched to a new primary trackside control module. The health status of each object control system is monitored through the first or second computer interlocking system. In the event of a complete failure of the primary object control system, one trackside control module in the backup object control system is switched to a new primary trackside control module, and the controllable switch of the new primary trackside control module is set to the on state, while the controllable switches of other trackside control modules other than the new primary trackside control module are set to the off state.