Scheduling system, method and device for realizing automatic fault detection and intelligent emergency switching, processor and readable storage medium thereof

Through the two-site, three-center deployment plan and intelligent emergency switching technology, the problems of fault detection delay and data loss in traditional securities trading systems have been solved, low-latency, high-reliability fault detection and emergency switching have been achieved, ensuring the high availability and data integrity of the trading system.

CN120692190APending Publication Date: 2025-09-23GUOTAI JUNAN SECURITIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510738713.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

Traditional securities trading systems lack automatic fault detection and real-time monitoring, resulting in high risks of data loss and inconsistency during disaster recovery switching, frequent delayed detection and false alarms, and affecting business continuity and system recovery capabilities.

Method used

A two-site, three-center deployment scheme is adopted, and real-time synchronization of the main center, backup center, and disaster recovery center is achieved through multicast and unicast. The operation and maintenance management system automatically detects faults and triggers intelligent emergency switching. Utilizing fault detection algorithms and visual component status monitoring, component-level alarms and disaster recovery switching are implemented.

Benefits of technology

It achieves one-minute-level low-latency, high-reliability fault detection and emergency switching, ensuring transaction data integrity and service continuity, reducing the risk of data loss and inconsistency, and improving the system's automatic detection and recovery capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120692190A_ABST
    Figure CN120692190A_ABST
Patent Text Reader

Abstract

The invention relates to a dispatching system for realizing automatic fault detection and intelligent emergency switching, which comprises a main center, a standby center and a disaster center, and is characterized in that the main center and the standby center are intercommunicated through multicast and are synchronized in real time; each of the main center, the standby center or the disaster center is provided with a distributed core transaction platform, each distributed core transaction platform comprises an operation and maintenance management system, a business management system, a transaction engine, an offer service module, an access gateway and a database, and the business management system comprises a data adaptation module and a data persistence module. By adopting the scheduling system, method and device for realizing automatic seven-fault detection and intelligent emergency switching, the processor and the computer readable storage medium thereof, one-minute-level switching is realized, the characteristics of low delay and high reliability are realized, the capabilities of automatic detection, alarm, evaluation and quick recovery of the system when the system faces faults are improved, and the system has a wide application prospect. And the integrity of transaction data and the continuity of service are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of securities trading, and in particular to the field of automatic fault detection and disaster recovery switching scheduling of securities core trading systems. Specifically, it refers to a scheduling system, method, device, processor and computer-readable storage medium thereof for realizing automatic fault detection and intelligent emergency switching. Background Art

[0002] In order to ensure the high availability of distributed securities trading systems and to ensure the normal and smooth operation of transactions when trading components fail or even major accidents occur, the present invention aims to realize automatic fault detection and intelligent emergency switching of distributed securities trading systems. The system adopts a two-site three-center deployment solution, uses the operation and maintenance system monitoring subcomponents to collect abnormal data, and provides visual system alarm information and fault information, realizing system fault detection of the new generation core trading system, and initiating specific emergency switching processes according to different emergency levels.

[0003] Traditional securities trading systems lack automatic fault detection. There's no way to deploy monitoring agents at key nodes to collect real-time system status data and issue timely warnings. Disaster recovery switchovers also pose a higher risk of data loss, especially during the switchover process. Various factors (such as network latency and equipment failure) can cause data to be delayed and lost. The risk of data loss is particularly high in asynchronous replication, as the master node doesn't wait for confirmation from the slave node after writing data before deeming the operation successful. The risk of data inconsistency is also higher, potentially leading to business interruptions or data errors during the switchover process, severely impacting business continuity. Furthermore, automatic fault detection in disaster recovery systems often experiences significant latency, with a delay between the occurrence of a fault and its identification by the monitored system. This delay can impact the system's timely response and recovery capabilities, leading to missed and false positives. Fault detection units can generate false positives due to external interference or other factors, resulting in unnecessary system switches and wasted resources. Summary of the Invention

[0004] The purpose of the present invention is to overcome the shortcomings of the above-mentioned prior art and provide a scheduling system, method, device, processor and computer-readable storage medium thereof for realizing automatic fault detection and intelligent emergency switching with good integrity, good continuity and a wide range of applicability.

[0005] To achieve the above objectives, the present invention provides a dispatching system, method, device, processor, and computer-readable storage medium for automatic fault detection and intelligent emergency switching as follows:

[0006] The dispatching system for realizing automatic fault detection and intelligent emergency switching has the following main features: the system includes a main center, a backup center and a disaster center; the main center and the backup center communicate with each other via multicast and are synchronized in real time; the communication between the main center and the disaster center, and between the backup center and the disaster center, is unicast;

[0007] The main center, backup center or disaster center are all equipped with a distributed core trading platform, which includes multiple components, including an operation and maintenance management system, a business management system, a trading engine, a quotation service module, an access gateway and a database. The business management system includes a data adaptation module and a data persistence module.

[0008] The operation and maintenance management system, business management system, transaction engine, quotation service module, access gateway and database are all interconnected, their business data are communicated through a distributed message bus, and monitoring and management data streams are communicated and interacted through TCP.

[0009] Preferably, the operation and maintenance management system automatically detects faults, monitors component status in a visual manner, uses fault detection algorithms to analyze indicator data obtained from component logs, and triggers fault alarms of different levels based on abnormal indicators; the operation and maintenance management system presets warning methods and responds to different strategies based on different warning levels; the operation and maintenance management system switches according to different types of disaster recovery based on the alarm level triggered by the preset disaster recovery configuration.

[0010] Preferably, the operation and maintenance management system triggers an alarm and performs disaster recovery switching, specifically including the following steps:

[0011] Obtain disaster recovery switching information such as the component instance name and disaster recovery component instance name that need to be switched from the database, and generate a disaster recovery switching work order; switch according to the disaster recovery work order, and update the disaster recovery node status to switching in progress; switch at the data message bus level; and send component-level switching instructions through the disaster recovery node management system gateway.

[0012] Preferably, the data adaptation module is connected and interacted with the peripheral system and the back-end system, and is used to process the data of the peripheral system before the market opens and send it to the corresponding back-end component, and send the data to other trading systems for clearing after the market closes; the main center, backup center and disaster center all include data adaptation modules, and are all connected to the database of the main center.

[0013] Preferably, if the main center of the system fails and the backup center is normal, the attribute of the data adapter module of the backup center is changed from "backup" to "master", and the data adapter module of the backup center performs the task and sends the data file to the backup center and the disaster center;

[0014] If both the primary and backup centers fail, the attribute of the data adaptation module of the disaster center will be changed from "disaster" to "primary", and the data adaptation module of the disaster center will execute data entry and data exit, and the data adaptation module of the disaster center will send data files to the disaster center;

[0015] If the main center and the backup center are normal and the disaster center fails, the data will be put on the scene, the operation and maintenance management system will eliminate all abnormal servers in the disaster center, and the data adaptation module of the main center will not send files to the disaster center.

[0016] Preferably, the main center and the backup center both include a data persistence module and are connected to the database of the main center. The data persistence module of the main center synchronizes messages to the data persistence module of the backup center. When the system is not performing disaster recovery switching, the data persistence module of the disaster center only subscribes to messages. When the system is performing disaster recovery switching, the disaster recovery switching identifier of the data persistence module is set. When the disaster data persistence module obtains the disaster switching identifier and performs the storage operation.

[0017] Preferably, if the data persistence module of the main center of the system fails, the data persistence module of the backup center automatically switches and takes over the work according to the storage breakpoint; if the transaction engine does not perform disaster recovery switching, the disaster center query engine sends a message to the data persistence module of the disaster center, and at the same time synchronizes the full amount of transaction data through asynchronous synchronization between the databases of the main center and the disaster center.

[0018] Preferably, the operation and maintenance management system sends a disaster recovery instruction to the transaction engine by calling the interface provided by the management system gateway. After receiving the disaster recovery instruction, the transaction engine traverses all the quotation contract number generators in the partition and skips X numbers of the quotation contract numbers in the corresponding generator. The value of X is set by the global configuration value. The query engine of the disaster recovery node receives the same instruction and performs synchronous operation.

[0019] Preferably, the quotation service module receives a real-time disaster recovery switching instruction from the management system gateway, connects to the exchange gateway, queries the return breakpoint from the trading engine, and synchronizes the return with the exchange gateway. The quotation service module normally places orders and receives exchange returns.

[0020] Preferably, before the primary disaster switching is performed, the gateway switch state of the access gateway of the disaster center is closed and does not provide external services. When the disaster recovery switching is performed, the gateway switch is turned on and provides external services.

[0021] Preferably, both the main center and the disaster center include databases and the databases of the main center and the disaster center are independent; after the agent of the disaster center receives the SQL request, it synchronizes it to the main center. If a disaster recovery switch occurs, the agent connection properties of the database are manually modified to point the agent of the disaster center to the disaster cluster.

[0022] The scheduling method for realizing automatic fault detection and intelligent emergency switching based on the above system is characterized in that the method comprises the following steps:

[0023] (1) If both the primary and backup centers fail, set a data flow timeout for the message bus replication between the primary, backup, and disaster centers;

[0024] (2) After a timeout occurs, the disaster center's message bus writes the offline events of the primary and backup center clusters to the disaster center's domain server, and the operation and maintenance management system continuously polls the disaster center's domain server;

[0025] (3) If the operation and maintenance management system finds through the domain server that the primary center and backup center clusters are offline, the operation and maintenance management system initiates disaster recovery switching.

[0026] Preferably, the step (3) specifically includes the following steps:

[0027] (3.1) Check the status of the master instance;

[0028] (3.2) Stop the central node service;

[0029] (3.3) Carry out disaster reduction;

[0030] (3.4) Start the sampling component of the delay analysis tool;

[0031] (3.5) Enable the disaster recovery gateway.

[0032] Preferably, the step (3.1) specifically includes the following steps:

[0033] (3.1.1) Get all disaster recovery information from the disaster recovery information table;

[0034] (3.1.2) Monitor the status of the master instance, send the list of disaster instances to the corresponding IP address of the domain server, and return information on whether all corresponding master instances are offline;

[0035] (3.1.3) If all corresponding master instances are offline, the disaster shedding conditions are considered to have been met, and the disaster shedding information is inserted into the system work order table, and step (3.1.4) is continued; otherwise, the disaster shedding conditions are not met, and the step ends;

[0036] (3.1.4) Update the disaster recovery flags of the data persistence module, data push module, data recovery module, and market data module in the centralized configuration module table;

[0037] (3.1.5) Update the configuration of the disaster data adaptation module.

[0038] Preferably, the step (3.2) specifically includes the following steps:

[0039] (3.2.1) Based on the list of disaster instances, obtain the data center where the disaster instance is located and all computer rooms under the data center, and obtain all instances in non-disaster computer rooms as a list of instances that need to be stopped;

[0040] (3.2.2) Call the ASF script to stop the domain service instance;

[0041] (3.2.3) Verify the stop result. If it contains an instance that failed to execute, return the name of the instance that failed to execute to the front end.

[0042] Preferably, the step (3.3) specifically includes the following steps:

[0043] (3.3.1) Get all disaster recovery information from the disaster recovery information table;

[0044] (3.3.2) Obtain the disaster recovery log for the day from the system work order table, update the Redis node status to "switching in progress", and update the disaster recovery information status in the system work order table to "switching in progress";

[0045] (3.3.3) Switching is performed at the message bus level. According to the list of disaster instances, a switch request is sent and the message bus level switch is polled to see if it is complete.

[0046] (3.3.4) Based on the component names with disaster attributes obtained from the disaster recovery information table, send real-time disaster cutover instructions according to the partitions and obtain the instruction results;

[0047] (3.3.5) If the command result is obtained, the disaster recovery is successful; if an exception occurs during the process of obtaining the command result, the switch fails and the name of the component that failed to switch is returned.

[0048] Preferably, the step (3.4) specifically includes the following steps:

[0049] (3.4.1) Obtain the instance of the time delay sampling component of the disaster center according to the system type;

[0050] (3.4.2) Call the ASF script to pull up the delay sampling component of the disaster center.

[0051] Preferably, the step (3.5) specifically includes the following steps:

[0052] (3.5.1) Obtain the access gateway instance and data exchange system instance according to the system type;

[0053] (3.5.2) Send instructions and obtain instruction processing results.

[0054] Preferably, the method further comprises the following steps:

[0055] When the primary center fails, export a snapshot from any remaining domain server in the backup center, and rebuild the domain server in the backup center as a single instance based on the snapshot. Promote the backup center database to the primary database, and modify the database connection domain name of the backup center to point to the backup center database.

[0056] When the backup center fails, the system automatically completes the switch.

[0057] The main features of the dispatching device for realizing automatic fault detection and intelligent emergency switching are as follows:

[0058] a processor configured to execute computer-executable instructions;

[0059] The memory stores one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching are implemented.

[0060] The scheduling processor for realizing automatic fault detection and intelligent emergency switching has the main feature that the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching are realized.

[0061] The main feature of the computer-readable storage medium is that a computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching.

[0062] The scheduling system, method, device, processor and computer-readable storage medium thereof for realizing automatic fault detection and intelligent emergency switching of the present invention can achieve one-minute switching compared with the traditional transaction system's manual intervention fault emergency switching processing method, and has the characteristics of low latency and high reliability. It can improve the system's ability to automatically detect, alarm, evaluate and quickly recover when facing faults, ensure the integrity of transaction data and the continuity of services, and effectively solve the shortcomings of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 This is a two-site, three-center framework diagram of the core transaction system of the scheduling system that implements automatic fault detection and intelligent emergency switching in the present invention.

[0064] Figure 2 This is a flow chart of the switching disaster center of the dispatching system that realizes automatic fault detection and intelligent emergency switching of the present invention.

[0065] Figure 3 Schematic diagram of the disaster recovery switching steps of the scheduling system for realizing automatic fault detection and intelligent emergency switching of the present invention.

[0066] Figure 4 This is a schematic diagram of checking the status of the master instance in the disaster recovery switching step of the scheduling system that realizes automatic fault detection and intelligent emergency switching of the present invention.

[0067] Figure 5 This is a schematic diagram of stopping central node services in the disaster recovery switching step of the scheduling system for realizing automatic fault detection and intelligent emergency switching according to the present invention.

[0068] Figure 6 The diagram is a schematic diagram of disaster switching in the disaster recovery switching step of the scheduling system for realizing automatic fault detection and intelligent emergency switching of the present invention.

[0069] Figure 7 The present invention is a schematic diagram of clearing disaster recovery information in the disaster recovery switching step of the scheduling system for realizing automatic fault detection and intelligent emergency switching.

[0070] Figure 8 This is a schematic diagram of enabling a disaster recovery gateway in the disaster recovery switching step of the scheduling system for realizing automatic fault detection and intelligent emergency switching of the present invention.

[0071] Figure 9 This is a schematic diagram of starting a sampling component of a disaster delay analysis tool in the disaster recovery switching step of the scheduling system for realizing automatic fault detection and intelligent emergency switching of the present invention. DETAILED DESCRIPTION

[0072] In order to more clearly describe the technical content of the present invention, further description is given below in conjunction with specific embodiments.

[0073] The scheduling system for realizing automatic fault detection and intelligent emergency switching of the present invention includes a main center, a backup center and a disaster center. The main center and the backup center communicate with each other through multicast and are synchronized in real time. The main center and the disaster center, as well as the backup center and the disaster center, use unicast.

[0074] The main center, backup center or disaster center are all equipped with a distributed core trading platform, which includes multiple components, including an operation and maintenance management system, a business management system, a trading engine, a quotation service module, an access gateway and a database. The business management system includes a data adaptation module and a data persistence module.

[0075] The operation and maintenance management system, business management system, transaction engine, quotation service module, access gateway and database are all interconnected, their business data are communicated through a distributed message bus, and monitoring and management data streams are communicated and interacted through TCP.

[0076] As a preferred embodiment of the present invention, the operation and maintenance management system automatically detects faults, monitors them in a visual manner using component status, uses a fault detection algorithm to analyze indicator data obtained from component logs, and triggers fault alarms of different levels based on abnormal indicators; the operation and maintenance management system presets warning methods and responds to different strategies based on different warning levels; the operation and maintenance management system switches according to different types of disaster recovery based on the alarm level triggered by the preset disaster recovery configuration.

[0077] As a preferred embodiment of the present invention, the operation and maintenance management system triggers an alarm and performs disaster recovery switching, specifically including the following steps:

[0078] Obtain disaster recovery switching information such as the component instance name and disaster recovery component instance name that need to be switched from the database, and generate a disaster recovery switching work order; switch according to the disaster recovery work order, and update the disaster recovery node status to switching in progress; switch at the data message bus level; and send component-level switching instructions through the disaster recovery node management system gateway.

[0079] As a preferred embodiment of the present invention, the data adaptation module is connected and interacted with the peripheral system and the background system, and is used to process the data of the peripheral system before the market opens and send it to the corresponding background component, and send the data to other trading systems for clearing after the market closes; the main center, backup center and disaster center all include data adaptation modules, and are all connected to the database of the main center.

[0080] As a preferred embodiment of the present invention, if the main center of the system fails and the backup center is normal, the attribute of the data adapter module of the backup center is changed from "backup" to "master", and the data adapter module of the backup center performs the task and sends the data file to the backup center and the disaster center;

[0081] If both the primary and backup centers fail, the attribute of the data adaptation module of the disaster center will be changed from "disaster" to "primary", and the data adaptation module of the disaster center will execute data entry and data exit, and the data adaptation module of the disaster center will send data files to the disaster center;

[0082] If the main center and the backup center are normal and the disaster center fails, the data will be put on the scene, the operation and maintenance management system will eliminate all abnormal servers in the disaster center, and the data adaptation module of the main center will not send files to the disaster center.

[0083] As a preferred embodiment of the present invention, the main center and the backup center both include a data persistence module and are connected to the database of the main center. The data persistence module of the main center synchronizes messages to the data persistence module of the backup center. When the system is not performing disaster recovery switching, the data persistence module of the disaster center only subscribes to messages. When the system is performing disaster recovery switching, the disaster recovery switching identifier of the data persistence module is set. When the disaster data persistence module obtains the disaster switching identifier and performs the storage operation.

[0084] As a preferred embodiment of the present invention, if the data persistence module of the main center of the system fails, the data persistence module of the backup center will automatically switch and take over the work according to the storage breakpoint; if the transaction engine does not perform disaster recovery switching, the disaster center query engine will send a message to the data persistence module of the disaster center, and at the same time, the full amount of transaction data will be synchronized through asynchronous synchronization between the databases of the main center and the disaster center.

[0085] As a preferred embodiment of the present invention, the operation and maintenance management system sends a disaster recovery instruction to the transaction engine by calling the interface provided by the management system gateway. After receiving the disaster recovery instruction, the transaction engine traverses all the quotation contract number generators in the partition and skips X numbers of the quotation contract numbers in the corresponding generator. The value of X is set by the global configuration value. The query engine of the disaster recovery node receives the same instruction and performs synchronization operations.

[0086] As a preferred embodiment of the present invention, the quotation service module receives a real-time disaster recovery switching instruction issued by the management system gateway, connects to the exchange gateway, queries the report breakpoint from the trading engine, and synchronizes the report with the exchange gateway. The quotation service module normally places orders and receives exchange reports.

[0087] As a preferred embodiment of the present invention, before the main disaster switching is performed, the gateway switch state of the access gateway of the disaster center is closed and does not provide external services. When the disaster recovery switching is performed, the gateway switch is turned on and provides external services.

[0088] As a preferred embodiment of the present invention, the main center and the disaster center both include databases, and the databases of the main center and the disaster center are independent; after the agent of the disaster center receives the SQL request, it synchronizes it to the main center. If a disaster recovery switch occurs, the agent connection properties of the database are manually modified to point the agent of the disaster center to the disaster cluster.

[0089] The present invention implements a scheduling method for automatic fault detection and intelligent emergency switching based on the above-mentioned system, wherein the method comprises the following steps:

[0090] (1) If both the primary and backup centers fail, set a data flow timeout for the message bus replication between the primary, backup, and disaster centers;

[0091] (2) After a timeout occurs, the disaster center's message bus writes the offline events of the primary and backup center clusters to the disaster center's domain server, and the operation and maintenance management system continuously polls the disaster center's domain server;

[0092] (3) If the operation and maintenance management system finds through the domain server that the primary center and backup center clusters are offline, the operation and maintenance management system initiates disaster recovery switching.

[0093] As a preferred embodiment of the present invention, the step (3) specifically includes the following steps:

[0094] (3.1) Check the status of the master instance;

[0095] (3.2) Stop the central node service;

[0096] (3.3) Carry out disaster reduction;

[0097] (3.4) Start the sampling component of the delay analysis tool;

[0098] (3.5) Enable the disaster recovery gateway.

[0099] As a preferred embodiment of the present invention, the step (3.1) specifically includes the following steps:

[0100] (3.1.1) Get all disaster recovery information from the disaster recovery information table;

[0101] (3.1.2) Monitor the status of the master instance, send the list of disaster instances to the corresponding IP address of the domain server, and return information on whether all corresponding master instances are offline;

[0102] (3.1.3) If all corresponding master instances are offline, the disaster shedding conditions are considered to have been met, and the disaster shedding information is inserted into the system work order table, and step (3.1.4) is continued; otherwise, the disaster shedding conditions are not met, and the step ends;

[0103] (3.1.4) Update the disaster recovery flags of the data persistence module, data push module, data recovery module, and market data module in the centralized configuration module table;

[0104] (3.1.5) Update the configuration of the disaster data adaptation module.

[0105] As a preferred embodiment of the present invention, the step (3.2) specifically includes the following steps:

[0106] (3.2.1) Based on the list of disaster instances, obtain the data center where the disaster instance is located and all computer rooms under the data center, and obtain all instances in non-disaster computer rooms as a list of instances that need to be stopped;

[0107] (3.2.2) Call the ASF script to stop the domain service instance;

[0108] (3.2.3) Verify the stop result. If it contains an instance that failed, the name of the instance that failed to execute will be returned to the front end.

[0109] As a preferred embodiment of the present invention, the step (3.3) specifically includes the following steps:

[0110] (3.3.1) Get all disaster recovery information from the disaster recovery information table;

[0111] (3.3.2) Obtain the disaster recovery log for the day from the system work order table, update the Redis node status to "switching in progress", and update the disaster recovery information status in the system work order table to "switching in progress";

[0112] (3.3.3) Switching is performed at the message bus level. According to the list of disaster instances, a switch request is sent and the message bus level switch is polled to see if it is complete.

[0113] (3.3.4) Based on the component names with disaster attributes obtained from the disaster recovery information table, send real-time disaster cutover instructions according to the partitions and obtain the instruction results;

[0114] (3.3.5) If the command result is obtained, the disaster recovery is successful; if an exception occurs during the process of obtaining the command result, the switch fails and the name of the component that failed to switch is returned.

[0115] As a preferred embodiment of the present invention, the step (3.4) specifically includes the following steps:

[0116] (3.4.1) Obtain the instance of the time delay sampling component of the disaster center according to the system type;

[0117] (3.4.2) Call the ASF script to pull up the delay sampling component of the disaster center.

[0118] As a preferred embodiment of the present invention, the step (3.5) specifically includes the following steps:

[0119] (3.5.1) Obtain the access gateway instance and data exchange system instance according to the system type;

[0120] (3.5.2) Send instructions and obtain instruction processing results.

[0121] As a preferred embodiment of the present invention, the method further comprises the following steps:

[0122] When the primary center fails, export a snapshot from any remaining domain server in the backup center, and rebuild the domain server in the backup center as a single instance based on the snapshot. Promote the backup center database to the primary database, and modify the database connection domain name of the backup center to point to the backup center database.

[0123] When the backup center fails, the system automatically completes the switch.

[0124] The scheduling device for realizing automatic fault detection and intelligent emergency switching of the present invention comprises:

[0125] a processor configured to execute computer-executable instructions;

[0126] The memory stores one or more computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching are implemented.

[0127] The scheduling processor of the present invention realizes automatic fault detection and intelligent emergency switching, wherein the processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the various steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching are realized.

[0128] The computer-readable storage medium of the present invention stores a computer program thereon, and the computer program can be executed by a processor to implement the various steps of the above-mentioned scheduling method for realizing automatic fault detection and intelligent emergency switching.

[0129] In order to ensure the high availability of distributed securities trading systems and to ensure the normal and smooth operation of transactions when trading components fail or even major accidents occur, the present invention aims to realize automatic fault detection and intelligent emergency switching of distributed securities trading systems. The system adopts a two-site three-center deployment solution, uses the operation and maintenance system monitoring subcomponents to collect abnormal data, and provides visual system alarm information and fault information, realizing system fault detection of the new generation core trading system, and initiating specific emergency switching processes according to different emergency levels.

[0130] In this paper, the term "node" is a broader concept. Nodes are divided by business dimension, with each node responsible for a specific business. A node may contain dozens of machines, including several business components, such as the access gateway (AGW), core trading engine (TE), and offer (ORS). An instance refers to a specific component deployed on a machine, such as the master access gateway AGW_11 under the spot 11 master node.

[0131] The present invention discloses a method for automatically detecting faults in a distributed securities trading system and for scheduling intelligent emergency switching. The method includes the following contents:

[0132] Automatic fault detection and alarm: Monitoring agents are deployed at key nodes of the distributed trading system to collect system operation status data in real time. The monitoring center aggregates and analyzes the data from each monitoring agent in real time, and uses fault detection algorithms to automatically identify system faults. Once a fault is detected, an alarm is sent to the operation and maintenance personnel through a preset alarm method (pop-up window, SMS, email, etc.);

[0133] Fault level assessment: The operation and maintenance center automatically assesses the fault level based on the fault type, impact scope, and current system status according to the preset fault level assessment strategy. Different levels correspond to different emergency response strategies (fault recovery, hot standby switching, disaster recovery switching, etc.);

[0134] According to different emergency response strategies, a specific emergency switching process is initiated. During the switching process, master-slave data breakpoint synchronization, link switching instructions, component switching instructions, etc. are used to ensure data consistency and integrity, realize fault isolation, and achieve automatic switching of applications and links.

[0135] In a specific embodiment of the present invention, the distributed core trading platform is composed of components such as an access gateway (AGW), a trading engine (CTE), a query engine (CQE), a comprehensive quotation service (COS), an order service (ORS), a data exchange (DXS), a business management system (BSS), and an operation and maintenance management system (OSS).

[0136] The business management system consists of a data adaptation module (DA), a data persistence module (DP), a data push module (DK), a data recovery module (DR), etc.; the operation and maintenance management system consists of a centralized configuration module (CCS), a management system gateway (BOSGW), etc.

[0137] Component business data is communicated through the distributed message bus (AMI), which is the business bus; component monitoring and management data flows are communicated and interacted through TCP.

[0138] The transaction system completes the core transaction process based on the communication capabilities of the message bus and the business processing capabilities of each application.

[0139] The business and operation and maintenance management components constitute the business operation and maintenance management system (BOS), which is a front-end system that provides various operation and management functions for securities firms' operation management personnel, operation and maintenance personnel, etc. The remaining components are back-end components.

[0140] To achieve the above objectives, the present invention provides an automatic fault detection and disaster recovery switching scheduling method based on a new generation securities core trading system, and the components included include:

[0141] Preferably, the operation and maintenance management system adopts automatic fault detection. The system uses a visual component status method to monitor the system as a whole and each component, uses a fault detection algorithm to analyze the indicator data obtained from the component log, and triggers different levels of fault alarms based on abnormal indicators.

[0142] Preferably, the operation and maintenance management system presets warning methods, including alarm handling pages, warning pop-up windows, text messages, emails, etc., and returns warnings to operation and maintenance personnel, responding to different strategies according to different warning levels.

[0143] Preferably, the operation and maintenance management system has disaster recovery design and switching functions. According to the disaster recovery configuration preset in the system management and the triggered alarm level, the components that need to be switched and the corresponding component instances are switched according to different types of disaster recovery, including hot standby, disaster recovery and other disaster recovery types, to achieve node-level or data center-level disaster recovery switching.

[0144] Preferably, when the operation and maintenance management system triggers an alarm and initiates a disaster recovery switchover, it first retrieves the disaster recovery switchover information, such as the component instance name and the disaster recovery component instance name, from the database and generates a disaster recovery switchover work order. The switchover is then performed according to the work order, with the disaster recovery node status updated to "switching in progress." The switchover is then initiated at the data message bus level, and finally, a component-level switchover instruction is sent through the disaster recovery node management system gateway.

[0145] Preferably, the data recovery module, data push module, and data persistence component switching of the operation and maintenance management system all rely on the message bus. Real-time instructions or data switch switching can be used to achieve one-click switching of primary and standby components, resume transmission from breakpoints, and switch the data exchange system gateway to send component-level instructions to the background, send switching instructions and open a new thread to obtain the switching results, and finally update them to the database.

[0146] In order to realize disaster recovery design and switching functions, the present invention adopts a two-site three-center deployment scheme, specifically: one master and one backup instance are deployed in the main center, one backup instance is deployed in the backup center, and one disaster instance is deployed in the disaster recovery center. Multicast communication is carried out between the main center and the backup center, and the main and backup systems are synchronized in real time.

[0147] When the primary instance of a component in the primary or backup center experiences an anomaly or server failure, automatic failover occurs without manual intervention. The business system's data loss probability (RPO) is zero, and the system recovery time (RTO) is less than 10 seconds. Unicast communication is used between the primary and backup centers and the disaster center, and asynchronous data replication is used between the primary and disaster centers.

[0148] Accordingly, the disaster recovery design of the relevant modules of the operation and maintenance management system is as follows:

[0149] ①Data adaptation module disaster recovery

[0150] The data adaptation module interacts with peripheral and backend systems. Its primary functions are to process data from peripheral systems before the market opens and send it to the corresponding backend components (data onboarding). After the market closes, it sends the data to other trading systems for clearing (data offboarding). The module's disaster recovery design deploys a set of data adaptation module components in the primary, backup, and disaster recovery centers. Two instances are deployed in the primary center, one in the backup center, and one in the disaster recovery center.

[0151] Preferably, the system ensures that the main center, backup center and disaster center can all be connected to the peripheral systems related to the data adaptation module and can all obtain consistent on-site files.

[0152] Preferably, the main center, backup center, and disaster center of the system can independently complete the corresponding functions of the data adaptation module. Under normal circumstances, all data adaptation modules are connected to the main center database.

[0153] Preferably, the instance properties of the system deployed in different data centers are configured as primary, backup, and disaster according to the properties of the data center. In the absence of abnormalities, the functions of all data adaptation modules are completed by the data adaptation module of the main center, and the data adaptation module of the main center will send data files to the primary, backup, and disaster data centers.

[0154] As a preferred embodiment of the present invention, when a failure occurs in the main center, if the backup center is normal, the instance attribute of the data adaptation module of the backup center will be changed from "backup" to "master", and then the task will be executed in the backup operation and maintenance management system interface (at this time the operation and maintenance management system will prompt to eliminate all abnormal servers in the main center), and the corresponding task will be executed by the data adaptation module of the backup center, and the data adaptation module of the backup center will send data files to the backup and disaster center (the main center server will not send files because it is eliminated due to a failure).

[0155] As a preferred embodiment of the present invention, when both the main center and the backup center fail, the instance attribute of the data adaptation module of the disaster center is changed from "disaster" to "master", and then the on- and off-stage operations are executed in the disaster operation and maintenance management system interface (at this time the operation and maintenance management system will prompt to eliminate all abnormal servers in the main and backup centers), and the corresponding tasks will be executed by the data adaptation module of the disaster center, and the data adaptation module of the disaster center will send data files to the disaster center (the main and backup center servers will not send files if they are eliminated due to failure).

[0156] As a preferred embodiment of the present invention, when performing disaster shedding of the data adaptation module, it will be determined whether the instance is in the startup state. It is necessary to stop the disaster data adaptation module instance and then modify the instance attribute state, and add the operation of switching the disaster data adaptation module instance attributes to the automatic disaster shedding process.

[0157] As a preferred embodiment of the present invention, when the main center and the backup center are normal and a fault occurs in the disaster center, the on-site steps are executed at this time, and the operation and maintenance management system will prompt to eliminate all abnormal servers in the disaster center. The main center data adaptation module will not send files to the disaster data center, and there will be no impact on the transactions of the main and backup centers. A certain amount of high availability will be lost on the same day due to the inability to switch to the disaster recovery center.

[0158] As a preferred embodiment of the present invention, in the operation and maintenance management system, before executing data entry and exit, if a server failure occurs, the operation and maintenance management system interface will prompt the abnormal server to be removed. The interface can choose to automatically remove the abnormal server and then perform daily operation and maintenance. If the machine fails after the operation and maintenance process is started, the server may be stuck due to the lack of prior removal, which may lead to abnormal step timeout. In order to avoid the step timeout for a long time, a reasonable timeout parameter is set for each step. The operation and maintenance process steps of the operation and maintenance management system and the entry and exit steps of the data adaptation module are respectively set with a timeout mechanism, and the setting value is 1.5 to 2 times the normal execution time of the corresponding step.

[0159] ②Data persistence module disaster recovery

[0160] The data persistence module is a data persistence component. Its main center and backup center are deployed in cluster mode. The main center module synchronizes messages to the backup center module, and both are connected to the main database.

[0161] Ideally, when the primary center's data persistence module fails, the backup center's data persistence module can automatically switch to take over based on the data entry breakpoint. The disaster center's data persistence module has no cluster relationship with the primary and backup centers at the message bus level; its data comes from the disaster center's query engine.

[0162] Preferably, when the transaction system has not performed disaster recovery switching, the disaster center query engine will send a message to the disaster center data persistence module. At the same time, the transaction data will be synchronized in full through asynchronous synchronization between the main disaster database. Before the disaster recovery switch, the disaster data persistence module will not perform storage processing after receiving the message.

[0163] Preferably, a disaster recovery switching identifier of the data persistence module is added to the centralized configuration module table. When disaster recovery switching is not performed, the disaster center data persistence module only subscribes to messages and does not store data in the warehouse. When disaster recovery switching is performed, the disaster recovery switching identifier of the data persistence module is set. When the disaster data persistence module obtains the disaster switching identifier, the data storage operation is performed, and the breakpoint value is taken from the breakpoint value recorded by the main data persistence module to continue storing data in the warehouse.

[0164] ③Transaction engine disaster recovery

[0165] After the trading engine's disaster recovery instance switches to the primary instance, orders that were lost but already submitted to the exchange will have duplicate contract numbers with new orders, leading to duplicate submissions and inconsistent reporting. To address this issue, we implement a skipping method for contract number allocation during the disaster recovery switchover process.

[0166] Preferably, the operation and maintenance management system sends a disaster recovery instruction to the trading engine by calling an interface provided by the management system gateway. Upon receiving the disaster recovery instruction, the trading engine traverses all bid contract number generators in the partition and skips X bid contract numbers in the corresponding generators. The value of X is set using a global configuration value, defaulting to 500, and this variable supports real-time in-memory modification. The query engine on the disaster recovery node will receive the same instruction and perform synchronous operations.

[0167] Preferably, considering that the business operation and maintenance management system uses the customer contract number as one of the primary keys for order entry, it is also necessary to skip the customer contract number, and the logic is consistent with the quotation contract number skipping method.

[0168] ④ Disaster recovery of quotation service module

[0169] The disaster recovery instance of the Quotation Service Module will not be connected to the exchange. After receiving the real-time disaster recovery switchover instruction from the management system gateway, the Quotation Service Module will start connecting to the exchange gateway, query the return breakpoint from the trading engine, and synchronize the return with the exchange gateway (the Quotation Service Module will not search for the original order for the return during the disaster recovery switchover). After that, the Quotation Service Module will place orders and receive exchange returns normally.

[0170] Preferably, the main disaster offer configuration is consistent. The offer setting page adds the disaster center offer configuration, and the offer configuration of the main disaster center corresponds one to one, and the same offer has the same offer group configuration.

[0171] Preferably, the primary center's offer groups are compressed, maintaining the same number of active offer groups and retaining only one offer channel within each offer group. Disaster recovery offer configurations automatically generate offer group information, offer channels, and offer service routing rules, eliminating the need for manual configuration. Only key information such as the offer gateway IP and port number needs to be configured.

[0172] ⑤Access gateway disaster recovery

[0173] The access gateway is deployed in an active-active configuration, and the disaster center access gateway does not have disaster-resistant properties. Before a primary-disaster failover occurs, the access gateway in the disaster center is in a closed state and does not provide external services. When a disaster recovery failover occurs, the gateway is opened and provides external services.

[0174] Preferably, in high availability design, the access gateway is used as an independent gateway cluster in the disaster recovery center. The orders originally reported by the main access gateway cluster will be at risk of losing routing information after disaster recovery switching occurs. However, considering that in the data center-level fault switching scenario, both the access gateway cluster and the access middleware will undergo center-level fault switching, so a certain amount of historical returns and entrustment confirmation losses can be accepted.

[0175] ⑥Database disaster recovery

[0176] The database is deployed in cluster mode, so the application layer does not need to worry about the underlying database data synchronization issues.

[0177] Ideally, the proxy connections for the primary and disaster clusters are completely independent. Under normal circumstances, the disaster proxy will synchronize SQL requests received to the primary cluster. If a failover occurs, manually modify the database proxy connection properties to point the disaster proxy to the disaster cluster, and the disaster cluster will resume normal operation.

[0178] Preferably, the main instance connection strings of all application layers remain unchanged, and the disaster instance has a new connection string pointing to the disaster proxy connection of the database. When the disaster recovery is switched, the proxy layer is switched, and the application layer does not need to be changed or restarted.

[0179] Preferably, since the database is deployed in cluster mode, there is no recovery process involved. After the primary cluster is restored, the disaster cluster will automatically synchronize data to the primary cluster. Manually redirect the disaster cluster's proxy connection to the primary cluster, and the database will return to its original working state.

[0180] In the specific implementation of the present invention, the disaster recovery switching operation process and the disaster recovery switching process of the operation and maintenance management system will be described respectively when the main center fails, the backup center in the same city fails, and the main center and backup center fail.

[0181] As a preferred embodiment of the present invention, the domain server (DomainServer) assumes the role of the configuration center and arbitrator of the message bus and can monitor the corresponding data indicators. Since TCP unicast can be communicated between the main center and the backup center, and the distance between the main center and the backup center is relatively close, a domain server cluster is shared between the main center and the backup center. Three instances are deployed in the main center and two instances are deployed in the backup center. An independent cluster is deployed in the disaster center, and there is only one instance in the cluster. The initial configuration files of the above two domain server clusters are exactly the same, and each works independently.

[0182] As a preferred embodiment of the present invention, when the main center fails, the domain server of the backup center suspends work because it cannot reach the majority. At this time, the background business components related to the message bus cannot complete the failover. It is necessary to export a snapshot from any remaining domain server in the backup center, and rebuild the domain server in the backup center as a single instance based on the snapshot. Run the script to promote the backup center database to the main database, and modify the database connection domain name of the backup center management system component to point to the backup center database. After that, the background business components related to the message bus find that the domain server is available, and the background components continue to complete the main-backup switch.

[0183] As a preferred embodiment of the present invention, when the same-city backup center fails, the system automatically completes the switch without the need for operation and maintenance intervention. However, the background service components based on the message bus may be suspended for 10 seconds due to the arbitration and switching mechanism of the message bus.

[0184] As a preferred embodiment of the present invention, when both the main center and the backup center fail, the data flow timeout time of the message bus replication message between the main center, the backup center and the disaster center is set to 10 seconds. After the timeout occurs, the message bus of the disaster center will write the main center and backup center cluster offline events to the domain server of the disaster center.

[0185] As a preferred embodiment of the present invention, the disaster recovery operation and maintenance management system continuously polls the domain server of the disaster center every 6 seconds. When the operation and maintenance management system finds through the domain server that the main center and backup center clusters are offline, it immediately reminds the operation and maintenance to start the disaster recovery switching process.

[0186] As a preferred embodiment of the present invention, the proxy connection properties of the database are manually modified to point the disaster proxy to the disaster cluster, and the disaster cluster starts to work normally.

[0187] As a preferred embodiment of the present invention, click Start Disaster Recovery Switching in the operation and maintenance management system page, and the operation and maintenance management system writes a message bus disaster recovery switching instruction to the domain server of the disaster center to complete the disaster recovery switching at the background component message bus level.

[0188] As a preferred embodiment of the present invention, the operation and maintenance management system issues a disaster recovery switching instruction to the backend component, and performs disaster recovery switching at the backend component business level in the order of trading engine, quotation service / comprehensive quotation service, and access gateway, that is, the quotation service / comprehensive quotation service is connected to the exchange quotation for report synchronization, and the access gateway is allowed to accept client connections.

[0189] As a preferred embodiment of the present invention, when the primary center and the backup center are unavailable and the system background components are not started, the operation process of disaster recovery switching is as follows:

[0190] (1) Start the redis sentinel in the disaster center through the script command and write the disaster recovery switch flag to redis;

[0191] (2) The database is routed to the backup database in the southern computer room for access;

[0192] (3) Re-operate the data adaptation module at the disaster center;

[0193] (4) Start the background components of the disaster center through the operation and maintenance management system page;

[0194] (5) Click on the page to actively start disaster recovery switching (the domain server will not have a prompt indicating that the primary center and backup center are offline). The operation and maintenance management system writes the disaster recovery switching instruction of the message bus to the domain server of the disaster center, completing the disaster recovery switching at the background component message bus level.

[0195] (6) The operation and maintenance management system issues a disaster recovery switch command to the backend components, and performs disaster recovery switch at the backend component business level in the order of trading engine, quotation service / comprehensive quotation service, and access gateway. That is, the quotation service / comprehensive quotation service connects to the exchange quotation for report synchronization, and the access gateway accepts client connections.

[0196] As a preferred embodiment of the present invention, the main process of disaster recovery in the operation and maintenance management system includes checking the status of the main instance, stopping the central node service, performing disaster recovery, starting the sampling component of the delay analysis tool, and opening the disaster recovery gateway.

[0197] The main process of checking the status of the master instance is as follows:

[0198] (1) Obtain all disaster recovery information from the disaster recovery information table, including the disaster instance list and domain server address;

[0199] (2) Monitor the status of the master instance, send the list of disaster instances to the corresponding IP address of the domain server, and return information on whether all corresponding master instances are offline;

[0200] (3) If all corresponding master instances are offline, the disaster shedding condition is considered to have been met, and the disaster shedding information is inserted into the system work order table; otherwise, the disaster shedding condition is not met;

[0201] (4) Update the disaster cut-off flags of the data persistence module, data push module, data recovery module, and market data module in the centralized configuration module table, and update the parameter value field of the disaster instance parameter name field of the four modules, disaster.switch.status, to 1 in the database configuration table;

[0202] (5) The configuration of the disaster data adaptation module is updated. The usage type field of the data adaptation module instance whose operating system type field is kunpeng in the instance deployment information table is updated to 1 (primary).

[0203] The process of stopping the central node service is mainly as follows:

[0204] (1) Based on the list of disaster instances, obtain the data center where the disaster instance is located and all computer rooms under the data center. Then obtain all instances in non-disaster computer rooms as the list of instances that need to be stopped;

[0205] (2) Call the ASF script to asynchronously stop all instances in the instance list except the domain service, and finally stop the domain service instance;

[0206] (3) Verify the stop result. If there is a failure, return the name of the failed instance to the front end.

[0207] The disaster recovery process is as follows:

[0208] (1) Obtain all disaster recovery information from the disaster recovery information table, including the list of disaster instances, domain server addresses, and component names with disaster attributes;

[0209] (2) Obtain the disaster recovery log for the day from the system work order table, update the redis node status to switching, and update the status of the disaster recovery information in the system work order table to switching;

[0210] (3) Switching is performed at the message bus level. According to the list of disaster instances, a switching request is sent and the message bus level switching is polled to see whether it is completed.

[0211] (4) Based on the component names with disaster attributes obtained from the disaster recovery information table, a real-time disaster cutover command is sent according to the partition and the command results are obtained. When the command results are obtained, it indicates that the disaster cutover is successful. If an exception occurs during the process of obtaining the command results, the default switchover fails and the name of the component that failed to switch is returned.

[0212] The main process of starting the sampling component of the latency analysis tool is as follows:

[0213] (1) Obtain the instance of the time delay sampling component of the disaster center according to the system type;

[0214] (2) Call the ASF script to start the delay sampling component of the disaster center (before disaster shedding, the disaster delay sampling component will not be started on the main operation and maintenance management system).

[0215] The process of enabling the disaster recovery gateway is as follows:

[0216] (1) Obtain the access gateway instance and data exchange system instance according to the system type;

[0217] (2) Send instructions and obtain instruction processing results.

[0218] The automatic fault detection and alarm of the operation and maintenance management system can monitor system components and the global situation, and can comprehensively monitor system components and overall operating status. The system has established different warning levels and provides component status in the form of graphics and list pop-ups, which can help operation and maintenance personnel to find problems in a timely and rapid manner, and respond to different strategies according to different warning levels, thereby greatly helping the operation and maintenance team to quickly and accurately identify problems and flexibly take corresponding measures based on different warning levels.

[0219] Under normal working conditions, only the transaction business nodes of the main center provide market transaction services, and the business nodes of the backup center and disaster center do not provide services to the outside world.

[0220] When disaster recovery switching occurs, the new generation of core transaction system can achieve one-minute switching: when the main and backup data center components or multiple points fail, automatic switching can be achieved without manual intervention, the recovery time target is less than 10 seconds, and the data loss is 0; the recovery time target for switching from the main center failure to the backup center is less than 1 minute, and the data loss is 0; the disaster recovery system supports data center-level fault switching and node-level fault switching. Through automatic detection, manual triggering, and automatic switching, the tolerable data loss is less than 10 seconds, the recovery time target is less than 5 minutes, and the switching time is less than 1 minute.

[0221] The system implements breakpoint-resume transmission. After a disaster recovery switch occurs, the data is re-pulled and reported based on the breakpoint of the main synchronization. For example, the data persistence module of the backup data center can automatically switch to take over the work based on the storage breakpoint. The breakpoint value is taken from the breakpoint value recorded by the main data persistence module to continue to store data, ensuring that data is not duplicated or lost.

[0222] The specific application scenario of the technical solution of the present invention is: an ultra-large-scale distributed securities trading system consisting of three centers in two locations and dozens of modules, hundreds of servers and thousands of instances. The scale is large, and the technical solution is relatively difficult, which is not at the same level as the difficulty of the existing solutions in this field.

[0223] The disaster recovery solution in this technical solution meets the disaster recovery requirements of complex, high-frequency, and massively complex systems. Based on a distributed database system with three centers across two locations, this architecture can already support 40 million customers and hundreds of millions of transaction volumes, demonstrating a high data carrying capacity. This distributed trading system, capable of processing tens of thousands of transactions per second, provides real-time intraday alerts and disaster recovery, supporting both on-site active / standby failover after node-level failures and remote failover after sudden disasters.

[0224] The technical solution of the present invention uses specific key indicators to carry out disaster recovery design for the transaction links and all modules of the entire transaction system, ensuring that in the event of a failure in any link or even a failure of the entire link, disaster recovery switching can be used to ensure that data is not duplicated or lost. It is very complete and has high reliability.

[0225] The technical solution of the present invention can achieve independent operation and maintenance of a single data center after switching, and the transaction core can seamlessly access external systems or customers; after the main center is restored, the system can also be quickly restored.

[0226] Based on the full-module disaster recovery design and AMI message bus, the present invention can achieve minute-level disaster recovery and data synchronization, ensuring that data is not duplicated or lost.

[0227] The technical solution of the present invention adopts a detailed distributed system and distributed database architecture, describes in detail the early warning and switching process of large-scale distributed systems and the disaster recovery plan of each module of the system, and clearly gives key indicators such as data switching experiments, loss rate, and recovery delay, so as to achieve rapid early warning and disaster recovery for sudden disasters.

[0228] The specific implementation scheme of this embodiment can be found in the relevant descriptions in the above embodiments and will not be repeated here.

[0229] It can be understood that the same or similar parts of the above embodiments can be referenced to each other, and the contents not described in detail in some embodiments can refer to the same or similar contents in other embodiments.

[0230] It should be noted that, in the description of the present invention, the terms "first", "second", etc. are used for descriptive purposes only and should not be understood as indicating or implying relative importance. In addition, in the description of the present invention, unless otherwise specified, the meaning of "plurality" is at least two.

[0231] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0232] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution device. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0233] Those skilled in the art will understand that all or part of the steps in the method for implementing the above-mentioned embodiment can be completed by instructing related hardware through a program, and the corresponding program can be stored in a computer-readable storage medium. When the program is executed, it includes one of the steps of the method embodiment or a combination thereof.

[0234] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0235] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.

[0236] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0237] The scheduling system, method, device, processor and computer-readable storage medium thereof for realizing automatic fault detection and intelligent emergency switching of the present invention can achieve one-minute switching compared with the traditional transaction system's manual intervention fault emergency switching processing method, and has the characteristics of low latency and high reliability. It can improve the system's ability to automatically detect, alarm, evaluate and quickly recover when facing faults, ensure the integrity of transaction data and the continuity of services, and effectively solve the shortcomings of traditional methods.

[0238] In this specification, the present invention has been described with reference to specific embodiments thereof. However, it will be apparent that various modifications and variations may be made without departing from the spirit and scope of the present invention. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A dispatching system for realizing automatic fault detection and intelligent emergency switching, characterized in that: The system includes a main center, a backup center and a disaster center. The main center and the backup center communicate with each other through multicast, and the main center and the backup center are synchronized in real time. The communication between the main center and the disaster center and between the backup center and the disaster center is unicast. The main center, backup center or disaster center are all equipped with a distributed core trading platform. The distributed core trading platform includes multiple components, including an operation and maintenance management system, a business management system, a trading engine, a quotation service module, an access gateway and a database. The business management system includes a data adaptation module and a data persistence module. The operation and maintenance management system, business management system, trading engine, quotation service module, access gateway and database are all interconnected, and their business data are communicated through a distributed message bus, and monitoring and management data streams are communicated and interacted through TCP.

2. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The operation and maintenance management system automatically detects faults, monitors component status in a visual manner, uses fault detection algorithms to analyze indicator data obtained from component logs, and triggers fault alarms of different levels based on abnormal indicators; the operation and maintenance management system presets warning methods and responds to different strategies based on different warning levels; the operation and maintenance management system switches according to the alarm level triggered by the preset disaster recovery configuration and the different types of disaster recovery.

3. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The operation and maintenance management system triggers an alarm and performs disaster recovery switching, specifically including the following steps: Obtain disaster recovery switching information such as the component instance name and disaster recovery component instance name that need to be switched from the database, and generate a disaster recovery switching work order; switch according to the disaster recovery work order, and update the disaster recovery node status to switching in progress; switch at the data message bus level; and send component-level switching instructions through the disaster recovery node management system gateway.

4. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The data adaptation module is connected and interacted with the peripheral system and the back-end system, and is used to process the data of the peripheral system before the market opens and send it to the corresponding back-end components, and send the data to other trading systems for clearing after the market closes; the main center, backup center and disaster center all include data adaptation modules, and are all connected to the database of the main center.

5. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 4, characterized in that: If the main center of the system fails and the backup center is normal, the attribute of the data adapter module of the backup center is changed from "backup" to "master", and the data adapter module of the backup center performs the task and sends the data file to the backup center and the disaster center; If both the primary and backup centers fail, the attribute of the data adaptation module in the disaster center is changed from "disaster" to "primary". The data adaptation module in the disaster center performs data entry and data exit, and sends data files to the disaster center. If the main center and the backup center are normal and the disaster center fails, the data will be put on the scene, the operation and maintenance management system will eliminate all abnormal servers in the disaster center, and the data adaptation module of the main center will not send files to the disaster center.

6. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The main center and the backup center both include data persistence modules and are connected to the database of the main center. The data persistence module of the main center synchronizes messages to the data persistence module of the backup center. When the system is not performing disaster recovery switching, the data persistence module of the disaster center only subscribes to messages. When the system is performing disaster recovery switching, the disaster recovery switching identifier of the data persistence module is set. When the disaster data persistence module obtains the disaster switching identifier and performs the storage operation.

7. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 6, characterized in that: If the data persistence module of the main center of the system fails, the data persistence module of the backup center will automatically switch and take over the work according to the storage breakpoint; if the transaction engine does not perform disaster recovery switching, the disaster center query engine will send a message to the data persistence module of the disaster center, and at the same time, the full amount of transaction data will be synchronized through asynchronous synchronization between the databases of the main center and the disaster center.

8. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The operation and maintenance management system sends a disaster recovery instruction to the transaction engine by calling the interface provided by the management system gateway. After receiving the disaster recovery instruction, the transaction engine traverses all the offer contract number generators in the partition and skips X numbers of the offer contract numbers in the corresponding generator. The value of X is set by the global configuration value. The query engine of the disaster recovery node receives the same instruction and performs synchronization operations.

9. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: The quotation service module receives the real-time disaster recovery switching instruction issued by the management system gateway, connects to the exchange gateway, queries the return breakpoint from the trading engine, and synchronizes the return with the exchange gateway. The quotation service module normally places orders and receives exchange returns.

10. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: Before the main disaster switching is performed in the system, the gateway switch state of the access gateway of the disaster center is closed and does not provide external services. When the disaster recovery switching is performed, the gateway switch is turned on and provides external services.

11. The dispatching system for realizing automatic fault detection and intelligent emergency switching according to claim 1, characterized in that: Both the main center and the disaster center include databases, and the databases of the main center and the disaster center are independent; after the agent of the disaster center receives the SQL request, it synchronizes it to the main center. If a disaster recovery switch occurs, the agent connection properties of the database are manually modified to point the agent of the disaster center to the disaster cluster.

12. A scheduling method for realizing automatic fault detection and intelligent emergency switching based on the system of claim 1, characterized in that: The method comprises the following steps: (1) If both the primary and backup centers fail, set a data flow timeout for the message bus replication between the primary, backup, and disaster centers; (2) After a timeout occurs, the disaster center's message bus writes the offline events of the primary and backup center clusters to the disaster center's domain server, and the operation and maintenance management system continuously polls the disaster center's domain server; (3) If the operation and maintenance management system finds through the domain server that the primary center and backup center clusters are offline, the operation and maintenance management system initiates disaster recovery switching.

13. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 12, characterized in that: The step (3) specifically includes the following steps: (3.1) Check the status of the master instance; (3.2) Stop the central node service; (3.3) Carry out disaster reduction; (3.4) Start the sampling component of the delay analysis tool; (3.5) Enable the disaster recovery gateway.

14. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 13, characterized in that: The step (3.1) specifically includes the following steps: (3.1.1) Get all disaster recovery information from the disaster recovery information table; (3.1.2) Monitor the status of the master instance, send the list of disaster instances to the corresponding IP address of the domain server, and return information on whether all corresponding master instances are offline; (3.1.3) If all corresponding master instances are offline, the disaster shedding conditions are considered to have been met, and the disaster shedding information is inserted into the system work order table, and step (3.1.4) is continued; otherwise, the disaster shedding conditions are not met, and the step ends; (3.1.4) Update the disaster recovery flags of the data persistence module, data push module, data recovery module, and market data module in the centralized configuration module table; (3.1.5) Update the configuration of the disaster data adaptation module.

15. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 13, characterized in that: The step (3.2) specifically includes the following steps: (3.2.1) Based on the list of disaster instances, obtain the data center where the disaster instance is located and all computer rooms under the data center, and obtain all instances in non-disaster computer rooms as a list of instances that need to be stopped; (3.2.2) Call the ASF script to stop the domain service instance; (3.2.3) Verify the stop result. If it contains an instance that failed, the name of the instance that failed to execute will be returned to the front end.

16. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 13, characterized in that: The step (3.3) specifically includes the following steps: (3.3.1) Get all disaster recovery information from the disaster recovery information table; (3.3.2) Obtain the disaster recovery log for the day from the system work order table, update the Redis node status to "switching in progress", and update the disaster recovery information status in the system work order table to "switching in progress"; (3.3.3) Switching is performed at the message bus level. According to the list of disaster instances, a switch request is sent and the message bus level switch is polled to see if it is complete. (3.3.4) Based on the component names with disaster attributes obtained from the disaster recovery information table, send real-time disaster cutover instructions according to the partitions and obtain the instruction results; (3.3.5) If the command result is obtained, the disaster recovery is successful; if an exception occurs during the process of obtaining the command result, the switch fails and the name of the component that failed to switch is returned.

17. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 13, characterized in that: The step (3.4) specifically includes the following steps: (3.4.1) Obtain the instance of the time delay sampling component of the disaster center according to the system type; (3.4.2) Call the ASF script to pull up the delay sampling component of the disaster center.

18. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 13, characterized in that: The step (3.5) specifically includes the following steps: (3.5.1) Obtain the access gateway instance and data exchange system instance according to the system type; (3.5.2) Send instructions and obtain instruction processing results.

19. The scheduling method for realizing automatic fault detection and intelligent emergency switching according to claim 12, characterized in that: The method further comprises the following steps: When the primary center fails, export a snapshot from any remaining domain server in the backup center, and rebuild the domain server in the backup center as a single instance based on the snapshot. Promote the backup center database to the primary database, and modify the database connection domain name of the backup center to point to the backup center database. When the backup center fails, the system automatically completes the switch.

20. A dispatching device for realizing automatic fault detection and intelligent emergency switching, characterized in that: The device comprises: a processor configured to execute computer-executable instructions; A memory storing one or more computer-executable instructions, wherein when the computer-executable instructions are executed by the processor, the steps of the scheduling method for realizing automatic fault detection and intelligent emergency switching as described in any one of claims 12 to 19 are implemented.

21. A dispatch processor that implements automatic fault detection and intelligent emergency switching, characterized in that: The processor is configured to execute computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the scheduling method for realizing automatic fault detection and intelligent emergency switching described in any one of claims 12 to 19 are implemented.

22. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the various steps of the scheduling method for realizing automatic fault detection and intelligent emergency switching as described in any one of claims 12 to 19.