Fault recovery method, apparatus and device for security transaction system, and medium

Through real-time monitoring and switching of trading nodes, combined with snapshot services and message queue data recovery mechanism, the problem of lack of universality and high resource consumption in the existing technology is solved, and the rapid failure recovery of securities trading systems and high availability deployment with low resource consumption is achieved.

CN119988096APending Publication Date: 2025-05-13JIANGSU SECURITIES
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510106288.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing securities trading system failure recovery solutions lack universality and consume a lot of resources, especially in the fast trading environment, it is difficult to meet the requirements of high availability and low resource consumption.

Method used

By monitoring the operating status of transaction nodes in real time, restarting or switching to standby nodes, and using snapshot services and message queues to achieve rapid data recovery and update, supporting a multi-master and one-standard deployment method to reduce server resource requirements.

Benefits of technology

It realizes rapid failure recovery of trading nodes, reduces data recovery complexity, saves server resources, and can meet the requirements of RPO=0 and RTO<10S in an extremely fast trading environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988096A_ABST
    Figure CN119988096A_ABST
Patent Text Reader

Abstract

The invention discloses a failure recovery method, device and equipment of a security transaction system and a medium, and belongs to the technical field of application system failure recovery. The method comprises the steps that when it is monitored that a transaction node cannot run, the transaction node is restarted, or a standby node is started to become a new transaction node; setting the newly started transaction node as a fault state; acquiring operation data of the transaction node by using the snapshot service and writing the operation data into a message queue; the transaction node pulls the operation data from the message queue, obtains return data during the abnormal period of the transaction node from a downstream node of the transaction node, and updates the data to the memory database; and setting the transaction node to be in a normal service state. The fault recovery method is high in speed and universality, is not influenced by development languages, development frameworks, business logics and the like, can ensure the data consistency of fault recovery, can support a plurality of transaction systems by one set of snapshot service, and saves server resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of application system fault recovery, and relates to a fault recovery method, device, equipment and medium for a securities trading system. Background Art

[0002] For securities trading systems, hardware or software failures will cause transactions to stop, so it is necessary to ensure high availability of system services in the event of hardware or software failures. A common solution is to switch to a backup node when the primary node fails. After the switch, the backup node becomes the primary node and continues to provide trading services.

[0003] The design of the master-slave solution usually includes the following three types: 1) Real-time persistence of messages through streaming files, and then restoring memory to achieve fault recovery by replaying messages. This method writes the upstream and downstream messages received by the trading system into the record file in sequence. When the trading system needs to be restored, the record file is replayed to recalculate and restore the memory. 2) The solution of memory synchronization between the master and the standby nodes to achieve fault recovery. In this method, the master node processes orders and updates the memory, and synchronizes the memory changes to the standby node in real time, which is generally an asynchronous method. When switching between the master and the standby, it is necessary to align the master and standby memory data before switching, and then switch the upstream and downstream connections. 3) The solution of simultaneous calculation between the master and the standby nodes to achieve fault recovery. In this solution, the master and standby nodes receive messages at the same time, process orders and update memory at the same time. Among them, there are generally two ways for the coordination of the master and standby nodes. One is that the standby node does not send messages to the outside, and only the master node sends messages to the outside, to ensure that the standby node does not interfere with the transaction. The other is that the master and standby nodes send messages at the same time, but use a common agent to sort and deduplicate the messages sent to ensure data correctness. When the master and standby nodes switch, message alignment is required to ensure data consistency before and after the switch. However, all three solutions have defects. Their implementation methods are bound to specific trading systems. They do not solve the problem of trading system fault recovery in a unified and standardized service manner and are not universal. They are all deployed in a one-master-one-backup manner, requiring more machine resources, especially for ultra-fast trading using physical machines and deployed in exchange computer rooms, where resources are even more scarce. Summary of the invention

[0004] The purpose of the present invention is to overcome the deficiencies in the prior art and to provide a fault recovery method, device, equipment and medium for a securities trading system, which can enable a faulty node to quickly resume operation.

[0005] To achieve the above object, the present invention is implemented by adopting the following technical solutions:

[0006] In a first aspect, the present invention provides a fault recovery method for a securities trading system, comprising:

[0007] Real-time monitoring of the operating status of one or more trading nodes in the securities trading system;

[0008] When it is detected that a transaction node cannot run, restart the transaction node, or start the standby node to make the standby node a new transaction node;

[0009] Set the newly started transaction node to a faulty state;

[0010] Using the snapshot service to obtain the operating data of the transaction node and write it into the message queue;

[0011] The transaction node pulls the operation data from the message queue, obtains the feedback data during the abnormal period of the transaction node from its downstream node, and updates the data to the memory database;

[0012] The transaction node is set to a normal service state.

[0013] Furthermore, if a standby node is started, when the inoperable transaction node is restored, it is set as a new standby node.

[0014] Furthermore, the operation data of the trading node includes the accounts processed by the trading node, as well as the accounts' unfulfilled orders, funds and position data.

[0015] Further, the snapshot service is in a multi-active form, including: multiple snapshot service nodes in a multi-active architecture form a snapshot service group, a snapshot service group can serve multiple transaction nodes at the same time, and the data of a transaction node is stored in a snapshot service group;

[0016] The snapshot service node manages multiple memory databases.

[0017] Furthermore, the memory database is provided with a plurality of data tables, and the data tables include a primary key field and a data body field;

[0018] One of the memory databases serves one or more transaction nodes, and different data table names are used to distinguish data stored in different transaction nodes in the same memory database.

[0019] Further, updating the data to the memory database includes: the transaction node pushes the data to the message queue, and its corresponding snapshot service node obtains the data from the message queue and stores the data in the corresponding data table in the memory database;

[0020] The snapshot service node identifies the data by the primary key of the data in the message queue.

[0021] Furthermore, the message queue is Kafka or Pulsar.

[0022] In a second aspect, the present invention further provides a fault recovery device for a securities trading system, the device comprising:

[0023] An operation status monitoring module is used to monitor the operation status of one or more trading nodes in the securities trading system in real time;

[0024] The transaction node startup module is used to restart the transaction node or start the standby node to make the standby node a new transaction node when it is detected that the transaction node cannot run;

[0025] A fault status setting module, used to set a newly started transaction node to a fault status;

[0026] An operation data acquisition module, used to acquire the operation data of the transaction node by using a snapshot service and write the data into a message queue;

[0027] A memory database update module is used for the transaction node to pull the operation data from the message queue, and obtain the feedback data of the transaction node during the abnormal period from its downstream node, and update the data to the memory database;

[0028] The normal service state setting module is used to set the transaction node to a normal service state.

[0029] In a third aspect, the present invention further provides a computer device, comprising:

[0030] Memory for storing computer programs;

[0031] A processor is used to execute the computer program to implement the steps of the above-mentioned fault recovery method of the securities trading system.

[0032] In a fourth aspect, the present invention further provides a computer-readable storage medium having a computer program stored thereon, the program being executed by a processor to implement the steps of the above-mentioned securities trading system failure recovery method.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] The fault recovery method of the securities trading system provided by the present invention is that multiple trading nodes only need to connect with the snapshot service through an interface to implement the logic of pulling data, which is not affected by the development language, development framework, business logic, etc., and greatly reduces the complexity of recovering the status data of the trading node; the present invention is deployed in the form of multiple masters and one backup, and a set of snapshot services can support multiple trading systems, which greatly saves server resources, especially in the field of ultra-fast trading, where physical machines are mainstream and deployed in the exchange hosting room; the fault recovery method provided by the present invention pulls data at a speed of 40,000 TPS. When the funds, positions, and in-transit order data that need to be pulled reach 300,000, data push can be completed within 10 seconds, and the system can be quickly restored, which can meet the requirements of RPO=0 and RTO<10S of the securities trading system. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 A schematic diagram of a flow chart of a fault recovery method for a securities trading system provided by an embodiment of the present invention;

[0036] Figure 2 It is a structural diagram of data interaction in a securities trading system in an embodiment of the present invention;

[0037] Figure 3 A schematic diagram of the structure of the snapshot service in an embodiment of the present invention;

[0038] Figure 4 A schematic diagram of a process for restoring transaction node data in an embodiment of the present invention;

[0039] Figure 5 A schematic diagram of the structure of a fault recovery device for a securities trading system provided by an embodiment of the present invention;

[0040] Figure 6 An internal structure diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] The technical solution of the present invention is described in detail below through the accompanying drawings and specific embodiments. The same reference numerals in the accompanying drawings indicate the same or similar components or parts. It should be understood by those skilled in the art that these drawings are not necessarily drawn to scale. The embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. In the absence of conflict, the embodiments of the present application and the technical features in the embodiments can be combined with each other.

[0042] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.

[0043] Embodiment 1:

[0044] like Figures 1 to 4 As shown, an embodiment of the present invention provides a fault recovery method for a securities trading system. Figure 1 The flowchart is a schematic diagram of the fault recovery method of the securities trading system. This flowchart only shows the logical sequence of the method described in this embodiment. Under the premise of no conflict, in other possible embodiments of the present invention, different methods can be used. Figure 1 The steps shown or described are accomplished in the order shown.

[0045] The fault recovery method of the securities trading system provided in this embodiment can be applied to a terminal, and can be executed by a fault recovery device of the securities trading system, which can be implemented by software and / or hardware, and can be integrated in a terminal.

[0046] See also Figure 1 The method of the embodiment of the present invention specifically comprises steps 1 to 6, wherein:

[0047] Step 1: Monitor the operating status of one or more trading nodes in the securities trading system in real time.

[0048] The monitoring service in the securities trading system will periodically call the inspection interface, and judge whether there are any abnormalities in the nodes of the securities trading system that cause the nodes to stop running based on the information and time returned by the interface. In addition, the upstream node of the trading node will send the order entrustment to the trading node. Whether the trading node returns the response to the upstream node's order entrustment in a timely manner can also be used to judge whether the trading node is running normally.

[0049] When the trading node is in normal operation, the data such as orders, transactions, funds, positions, etc. will be pushed to the snapshot service through the message queue in real time, allowing the snapshot service to update the memory data in real time. The data status of the snapshot service will always be consistent with the current status of the trading system.

[0050] Step 2: When it is detected that the transaction node cannot run, restart the transaction node, or start the standby node to make the standby node a new transaction node.

[0051] When a transaction node is detected to have an abnormality and stops running, depending on the specific situation, you can choose to restart the transaction node and restore its data so that it can operate normally, or replace the abnormal transaction node with a standby node, making the standby node a new transaction node and restoring its data.

[0052] If the inoperable transaction node is replaced by a standby node, the inoperable transaction node will be set as a new standby node after it can operate normally. Because the standby node of the present invention does not need to run services in real time, multiple transaction nodes can share one standby node.

[0053] The present invention does not need to set up a backup node for each transaction node, but deploys servers in the form of multiple masters and one backup, which greatly saves server resources.

[0054] Step 3: Set the newly started transaction node to a faulty state.

[0055] After the inoperable trading node or the replaced new trading node is started, it will be immediately set to a faulty state and will no longer accept order data from its upstream nodes and feedback data from its downstream nodes.

[0056] Step 4: Use the snapshot service to obtain the operating data of the transaction node and write it into the message queue.

[0057] Step 5: The transaction node pulls the operation data from the message queue, obtains the feedback data during the transaction node abnormality period from its downstream node, and updates the data to the memory database.

[0058] like Figure 2 As shown, under normal conditions, the securities trading system of the present invention will push data changes to its message queue (MQ) in real time, and the snapshot service obtains data from the message queue and stores it in the memory database.

[0059] If the snapshot service stops working due to an exception, data is pulled again from the message queue to restore the snapshot service. Therefore, the message queue of the present invention needs to have a message persistence function, and Kafka or Pulsar can be used as the message queue.

[0060] The snapshot service is in the form of Multi-site Active, which is a high-availability distributed system architecture. Its core feature is that it provides services to the outside world simultaneously in multiple data centers, and the service status of each data center is the same, and it can respond to multiple user requests in real time.

[0061] In the present invention, the multi-active data center is a snapshot service node. Specifically, Figure 3As shown, multiple snapshot service nodes in a multi-active architecture form a snapshot service group. A snapshot service group can serve multiple transaction nodes at the same time. The data of a transaction node is stored in a snapshot service group.

[0062] A snapshot service node can have one or more memory databases, each of which has multiple data tables. The data tables store data through primary key fields and data body fields. A memory database can serve different transaction nodes, and different data table names are used to distinguish data stored in different transaction nodes in the same memory database.

[0063] The snapshot service node only uses the primary key to identify the data. When the snapshot service node obtains data from the message queue and stores it in the memory database, if the primary key does not exist in the data table corresponding to the transaction node, the data is directly inserted. Otherwise, the new data in the message queue is used to overwrite the old data corresponding to the primary key in the data table. In addition, the snapshot service node does not perform data calculations to ensure the consistency of the data stored in the memory database and the transaction node data. The data content of the transaction node is stored as an aggregate field, and the snapshot service node does not care about the content in the field. Therefore, the universality of the fault recovery method of the present invention is guaranteed, and a set of snapshot services can support multiple different transaction systems.

[0064] After the transaction node receives orders from the upstream node and the downstream node performs feedback data processing, it sends the updated data to the corresponding snapshot service node through the message queue. After the snapshot service node receives the data, it updates the memory data in real time. After the transaction node sends the data to the message queue, the high availability of the data is guaranteed by the distributed mechanism of the message queue.

[0065] Figure 4 The interactive process of each node when the present invention restores the transaction node data is shown, including: 1) starting the transaction node and putting it in a faulty state; 2) requesting the snapshot service to pull the operating data; 3) the snapshot service writes the operating data to the message queue, and the transaction node pulls the operating data from the message queue; 4) the transaction node restores the operating data to the memory; 5) the transaction node completes the transaction data generated during the transaction node abnormality from the downstream node; 6) the transaction node updates the memory and returns to the normal service state.

[0066] The operation data of the trading node includes the account being processed by the trading node, as well as information such as the account's unfulfilled orders, funds and position data.

[0067] Because the executed orders and their transaction records have no impact on system recovery and continued trading, the present invention only uses the snapshot service to pull the operating data of the transaction node and writes it into the message queue corresponding to the transaction node, which greatly reduces the amount of data pulled from the message queue by the transaction node and reduces the time required to restore the transaction node. In addition, multiple transaction nodes only need to connect with the snapshot service through the interface and implement the logic of pulling data, which can be directly applied to the trading, risk control, strategy and other systems of the securities trading system without being affected by the development language, development framework, business logic, etc.

[0068] Step 6: Set the transaction node to normal service status.

[0069] Through online testing and verification, the recovery method of the present invention can pull data at a speed of 40,000 TPS. When the funds, positions, and in-transit order data to be pulled reach 300,000, data pulling can be completed within 10 seconds, and the system can be quickly restored.

[0070] Embodiment 2:

[0071] Based on the same inventive concept as in Example 1, the embodiment of the present invention also provides a fault recovery device for a securities trading system for implementing the fault recovery method for the securities trading system. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiment of the fault recovery device for a securities trading system provided below can refer to the limitations of the fault recovery method for a securities trading system above, and will not be repeated here.

[0072] like Figure 5 As shown, an embodiment of the present invention provides a fault recovery device for a securities trading system, comprising:

[0073] An operation status monitoring module is used to monitor the operation status of one or more trading nodes in the securities trading system in real time;

[0074] The transaction node startup module is used to restart the transaction node or start the standby node to make the standby node a new transaction node when it is detected that the transaction node cannot run;

[0075] A fault status setting module, used to set a newly started transaction node to a fault status;

[0076] An operation data acquisition module, used to acquire the operation data of the transaction node by using a snapshot service and write the data into a message queue;

[0077] A memory database update module is used for the transaction node to pull the operation data from the message queue, and obtain the feedback data of the transaction node during the abnormal period from its downstream node, and update the data to the memory database;

[0078] The normal service state setting module is used to set the transaction node to a normal service state.

[0079] Embodiment 3:

[0080] The embodiment of the present invention further provides a computer device, which may be a server, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, the fault recovery method of the securities trading system in the aforementioned embodiment is implemented.

[0081] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0082] Embodiment 4:

[0083] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the following method are implemented:

[0084] Real-time monitoring of the operating status of one or more trading nodes in the securities trading system;

[0085] When it is detected that a transaction node cannot run, restart the transaction node, or start the standby node to make the standby node a new transaction node;

[0086] Set the newly started transaction node to a faulty state;

[0087] Using the snapshot service to obtain the operating data of the transaction node and write it into the message queue;

[0088] The transaction node pulls the operation data from the message queue, obtains the feedback data during the abnormal period of the transaction node from its downstream node, and updates the data to the memory database;

[0089] The transaction node is set to a normal service state.

[0090] It should be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products, and therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0091] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0092] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0093] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0094] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present invention and the claims, which all fall within the protection of the present invention.

Claims

1. A fault recovery method for a securities trading system, characterized in that: include: Real-time monitoring of the operating status of one or more trading nodes in the securities trading system; When it is detected that a transaction node cannot run, restart the transaction node, or start the standby node to make the standby node a new transaction node; Set the newly started transaction node to a faulty state; Using the snapshot service to obtain the operating data of the transaction node and write it into the message queue; The transaction node pulls the operation data from the message queue, obtains the feedback data during the abnormal period of the transaction node from its downstream node, and updates the data to the memory database; The transaction node is set to a normal service state.

2. The fault recovery method of the securities trading system according to claim 1, characterized in that: If a standby node is started, it will be set as the new standby node after the inoperable transaction node is restored.

3. The fault recovery method of the securities trading system according to claim 1, characterized in that: The operation data of the trading node includes the accounts processed by the trading node, as well as the accounts' unfulfilled orders, funds and position data.

4. The fault recovery method of the securities trading system according to claim 1, characterized in that: The snapshot service is in a multi-active form, including: multiple snapshot service nodes in a multi-active architecture form a snapshot service group, a snapshot service group can serve multiple transaction nodes at the same time, and the data of a transaction node is stored in a snapshot service group; The snapshot service node manages multiple memory databases.

5. The fault recovery method of the securities trading system according to claim 4, characterized in that: The memory database is provided with a plurality of data tables, wherein the data tables include a primary key field and a data body field; One of the memory databases serves one or more transaction nodes, and different data table names are used to distinguish data stored in different transaction nodes in the same memory database.

6. The fault recovery method of the securities trading system according to claim 5, characterized in that: Updating the data to the memory database includes: the transaction node pushes the data to the message queue, and its corresponding snapshot service node obtains the data from the message queue and stores the data in the corresponding data table in the memory database; The snapshot service node identifies the data by the primary key of the data in the message queue.

7. The fault recovery method of the securities trading system according to claim 6, characterized in that: The message queue is Kafka or Pulsar.

8. A fault recovery device for a securities trading system, characterized in that: include: An operation status monitoring module is used to monitor the operation status of one or more trading nodes in the securities trading system in real time; The transaction node startup module is used to restart the transaction node or start the standby node to make the standby node a new transaction node when it is detected that the transaction node cannot run; A fault status setting module, used to set a newly started transaction node to a fault status; An operation data acquisition module, used to acquire the operation data of the transaction node by using a snapshot service and write the data into a message queue; A memory database update module is used for the transaction node to pull the operation data from the message queue, and obtain the feedback data of the transaction node during the abnormal period from its downstream node, and update the data to the memory database; The normal service state setting module is used to set the transaction node to a normal service state.

9. A computer device, characterized in that: include: Memory for storing computer programs; A processor, configured to execute the computer program to implement the steps of the fault recovery method for a securities trading system according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the fault recovery method of the securities trading system described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Message processing system and method based on queuing machine

    CN121029446A