A method and device for ensuring data consistency in a multi-active architecture
By using redo logs and failure logging mechanisms in the distributed search engine ES proxy subsystem of the financial industry, the problem of data loss during downtime of ES clusters is solved, and data consistency and real-time under the multi-active architecture are achieved.
Patent Information
- Application Number
- CN202110997544.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-08-27
AI Technical Summary
In the financial industry, distributed search engine ES will become unavailable when encountering power outages in the computer room, and ES does not support data synchronization between different clusters, making it difficult to ensure data consistency under a multi-active architecture, and data downtime may result in data loss after writing and waiting for disk flushing.
By introducing redo logs and failure log mechanisms into the ES proxy subsystem, data written to the ES cluster are synchronized, and when the write failure information is received, the data in the redo log is played back to the failure log to obtain the update exception data. When the recovery data task is started, the data in the failure log is written to the ES cluster again.
Ensure that data is not lost when the ES cluster is down, data consistency is achieved under a multi-active architecture, avoiding the situation where historical abnormal data is not written to the ES cluster, and improving the real-time and processing efficiency of data.
Smart Images

Figure CN113722398B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of financial technology (Fintech), and in particular, to a method and device for ensuring data consistency in a multi-active architecture. Background Art
[0002] With the development of computer technology, more and more technologies are applied in the financial field, and traditional finance is gradually transforming into financial technology. However, due to the security and real-time requirements of the financial industry, higher requirements are also put forward for technology.
[0003] Currently, in the financial industry, a distributed search engine (elasticsearch, ES) is used, and based on ES, a large amount of data can be searched, analyzed, and explored. However, in actual use, if the entire computer room loses power, ES will still become unavailable. Therefore, in actual use, ES needs to be deployed in a multi-active architecture to improve system reliability. However, since ES itself does not support data synchronization between different clusters, if a multi-active architecture deployment method is adopted, other solutions need to be adopted to ensure that the data of at least two ES clusters is the same.
[0004] Although, in related technologies, when implementing data consistency of ES clusters, a solution for data synchronization based on applications is provided. However, in this solution, since the order of data writing is to write to memory first and then to the corresponding disk, when the data is written to the ES cluster and the write success is feedback, if the machine crashes while waiting for the data to be flushed from memory to disk, the data has the risk of loss. Summary of the Invention
[0005] The present invention provides a method and device for ensuring data consistency in a multi-active architecture, which are used to solve the problem of easy data loss in the prior art.
[0006] In a first aspect, the present invention provides a method for ensuring data consistency in a multi-active architecture, which is applied to a distributed search engine ES proxy subsystem, and includes: pulling first data to be written from a distributed log system, and synchronously writing the first data to be written into each ES cluster; when receiving a write failure message sent by any one of the ES clusters, writing the write success data recorded in the current redo log back to the failure log to obtain updated abnormal data in the failure log; wherein, the redo log includes the write success data in each ES cluster retained within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated abnormal data is data that has not been successfully written to any ES cluster; if, starting from the start time of the recovery data task, when a preset start duration is reached, the updated abnormal data in the failure log is synchronously written into each ES cluster.
[0007] In the above method, when the ES proxy subsystem determines that any ES cluster has an exception, that is, when it receives the write failure information sent by any one of the ES clusters in each ES cluster, it will replay the successfully written data recorded in the current redo log to the failure log to obtain the updated exception data in the failure log. When the data recovery task is started and reaches the preset startup duration, the updated exception data in the failure log will be written into each ES cluster again, thus ensuring that the data will not be lost. Specifically, since the current redo log records the data successfully written in each ES cluster reserved within a preset duration, and this preset duration is greater than the disk flushing duration corresponding to the ES cluster, the data can still be retained when the ES cluster crashes, and this data is written into each ES cluster again. Therefore, the purpose of data not being lost when the ES cluster crashes can be achieved.
[0008] Optionally, after pulling the first data to be written from the distributed log system, the method further includes: determining whether there is historical exception data in the failure log; when it is determined that there is historical exception data in the failure log, writing the first data to be written into the failure log, and stopping writing the first data to be written into each ES cluster.
[0009] Based on the above method, the situation where the historical exception data is not written into each ES cluster while the first data to be written is written into each ES cluster can be avoided, that is, the timeliness of the data written into each ES cluster is ensured as much as possible.
[0010] Optionally, after writing the first data to be written into the failure log, the method further includes: before pulling the second data to be written from the distributed log system, determining the pull waiting duration. When it is determined that the pull waiting duration is reached since the moment of pulling the first data to be written, pulling the second data to be written, and sending the second data to be written to the local database; after a first preset duration, obtaining the second data to be written from the local database, and synchronously writing the second data to be written into each ES cluster; where the second data to be written is the content different from that pulled after pulling the first data to be written from the distributed log system.
[0011] In the above method, when there is exception data, the consumption rate of pulling messages in the distributed system can be reduced, so that the state of data delay can be avoided, and the real-time performance of the data is improved.
[0012] Optionally, determine the pull waiting duration, including: determining the pull order of the second data to be written; wherein, the pull order is determined according to the pull time for the data pulled from the distributed log system after the historical exception data already exists in the failure log; determining the previous pull waiting duration, and taking N times of the previous pull waiting duration as the pull waiting duration; the previous pull waiting duration is: the pull waiting duration corresponding to the data pulled in the previous pull order of the pull order of the second data to be written; wherein, when the previous pull order is the first time, then determine that the previous pull waiting duration is the processing duration for the historical exception data, and N is a positive integer greater than 2.
[0013] In the above method, the pull waiting duration gradually increases with the pull order of the data to be written. That is, when the pull order of the second data to be pulled is relatively later, more data is written to the failure log, and the waiting duration for subsequent data pulling is also longer, so that there is sufficient time to process the data in the failure log and improve the processing efficiency.
[0014] Optionally, when receiving the write failure information sent by any one of the ES clusters in the respective ES clusters, replay the data recorded in the current redo log to the failure log, including: determining whether the updated exception data recorded in the failure log exists in the data recorded in the current redo log; when the updated exception data recorded in the failure log does not exist in the data recorded in the current redo log, insert the data recorded in the redo log into the failure log; when the updated exception data recorded in the failure log exists in the data recorded in the current redo log, determine the timestamp of the updated exception data recorded in the failure log; determine whether the timestamp is greater than the timestamp of the data recorded in the redo log; when it is determined that the timestamp is not greater than the timestamp of the data recorded in the redo log, replay the data recorded in the current redo log to the failure log.
[0015] Based on the above method, the data recorded in the redo log stored in the local database can be replayed to the failure log, thereby reducing the data storage amount in the local database, and reducing the steps for the ES proxy subsystem to obtain exception data, and improving the processing efficiency of the system.
[0016] Optionally, when the preset startup duration is reached since the start time of the data recovery task, the updated abnormal data in the failure log is synchronously written into each of the ES clusters, including: setting a one-by-one lock grabbing mechanism for each data recovery task in the data recovery project; when the lock grabbing for any one of the data recovery tasks is successful, creating a child thread corresponding to any one of the data recovery tasks, and based on the child thread, performing: when the preset startup duration is reached since the start time of the data recovery task, synchronously writing the updated abnormal data in the failure log into each of the ES clusters; when each of the ES clusters feeds back a write success message, releasing the lock corresponding to the data recovery task.
[0017] Based on the above method, the sequentiality of the write operation can be ensured as much as possible when writing abnormal data into the ES clusters, thereby ensuring the consistency when writing data into each of the ES clusters.
[0018] Optionally, after synchronously writing the updated abnormal data in the failure log into each of the ES clusters, the method further includes: feeding back an offset to the distributed log system; after completing the feedback of the offset to the distributed log system and determining that a pre-subscribed topic has been adjusted, determining not to pull data from the distributed log system.
[0019] In the above method, after completing the feedback of the offset to the distributed log system, when it is determined that a pre-subscribed topic has been adjusted, it is determined not to pull data from the distributed log system, so that the data synchronization from the business system to the ES proxy subsystem can be completed quickly and efficiently.
[0020] In a second aspect, the present invention provides a device for ensuring data consistency in a multi-active architecture, which is applied to a distributed search engine ES proxy subsystem, including: a processing unit, configured to pull first data to be written from a distributed log system and synchronously write the first data to be written into each of the ES clusters; an obtaining unit, configured to, when receiving a write failure message sent by any one of the ES clusters in each of the ES clusters, replay the write success data recorded in the current redo log into the failure log to obtain the updated abnormal data in the failure log; wherein, the redo log includes the write success data in each of the ES clusters retained within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated abnormal data is the data that has not been successfully written into any one of the ES clusters; a recovery writing unit, configured to, when the preset startup duration is reached since the start time of the data recovery task, synchronously write the updated abnormal data in the failure log into each of the ES clusters.
[0021] Optionally, after pulling the first data to be written from the distributed logging system, the processing unit is further configured to: determine whether historical exception data already exists in the failure log; when it is determined that the historical exception data already exists in the failure log, write the first data to be written into the failure log, and stop writing the first data to be written into each ES cluster.
[0022] Optionally, after writing the first data to be written into the failure log, the processing unit is further configured to: determine a pull waiting duration from the distributed logging system before pulling the second data to be written; when it is determined that the time since pulling the first data to be written reaches the pull waiting duration, pull the second data to be written and send the second data to be written to the local database; after a first preset duration, obtain the second data to be written from the local database, and synchronously write the second data to be written into each ES cluster; wherein, the second data to be written is: data with different content pulled from the distributed logging system after pulling the first data to be written.
[0023] Optionally, the processing unit is specifically configured to: determine the pull order of the second data to be written; wherein, the pull order is determined according to the pull time for the data pulled from the distributed logging system after the historical exception data already exists in the failure log; determine the previous pull waiting duration, and use N times the previous pull waiting duration as the pull waiting duration; the previous pull waiting duration is: the pull waiting duration corresponding to the data pulled in the previous pull order of the second data to be written; wherein, when the previous pull order is the first time, determine the previous pull waiting duration as the processing duration for the historical exception data, and N is a positive integer greater than 2.
[0024] Optionally, the obtaining unit is specifically configured to: determine whether the updated exception data recorded in the failure log exists in the data recorded in the current redo log; when the updated exception data recorded in the failure log does not exist in the data recorded in the current redo log, insert the data recorded in the redo log into the failure log; when the updated exception data recorded in the failure log exists in the data recorded in the current redo log, determine the timestamp of the updated exception data recorded in the failure log; determine whether the timestamp is greater than the timestamp of the data recorded in the redo log; when it is determined that the timestamp is not greater than the timestamp of the data recorded in the redo log, replay the data recorded in the current redo log into the failure log.
[0025] Optionally, the recovery writing unit is specifically configured to: set a one-by-one lock grabbing mechanism for each recovery data task in the recovery data items; when the lock grabbing for any one of the recovery data tasks is successful, establish a sub-thread corresponding to any one of the recovery data tasks, and based on the sub-thread, execute: starting from the start time of the recovery data task, when a preset start duration is reached, synchronously write the update exception data in the failure log to each of the ES clusters; when each of the ES clusters feedbacks a write success message, release the lock corresponding to the recovery data task.
[0026] Optionally, after synchronously writing the update exception data in the failure log to each of the ES clusters, the processing unit is further configured to: feedback an offset to the distributed log system; after completing the feedback of the offset to the distributed log system and determining that a pre-subscribed topic has been adjusted, determine not to pull data from the distributed log system.
[0027] For the beneficial effects of the second aspect and each optional device of the second aspect, reference may be made to the beneficial effects of the first aspect and each optional method of the first aspect, which will not be elaborated here.
[0028] In a third aspect, the present invention provides a computer device, including a program or instruction, which when executed, is used to execute the first aspect and each optional method of the first aspect.
[0029] In a fourth aspect, the present invention provides a storage medium, including a program or instruction, which when executed, is used to execute the first aspect and each optional method of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments.
[0031] Figure 1 It is a schematic diagram of an optional application scenario provided by an embodiment of the present invention;
[0032] Figure 2 It is a schematic diagram of the step flow of a method for ensuring data consistency in a multi-active architecture provided by an embodiment of the present invention;
[0033] Figure 3 It is another schematic diagram of the step flow of a method for ensuring data consistency in a multi-active architecture provided by an embodiment of the present invention;
[0034] Figure 4 It is a schematic diagram of the structure of a device for ensuring data consistency in a multi-active architecture provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations on the technical solution of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.
[0036] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such images can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.
[0037] To facilitate understanding of the technical solution provided by the embodiments of the present invention, some key terms used in the embodiments of the present invention are explained here first:
[0038] 1. Distributed indexing engine (ElasticSearch, ES): It is a distributed, highly scalable, and highly real-time search and data analysis engine. It can easily endow a large amount of data with the capabilities of search, analysis, and exploration. In addition, it can be used in Java,.NET (C#), PHP, Python, Apache Groovy, Ruby, and many other languages.
[0039] 2. Kafka: It is a distributed, partitioned, multi-replica, multi-subscriber distributed log system coordinated by ZooKeeper. This distributed log system can be responsible for transferring data from one application to another. The application only needs to focus on the data and does not need to care about how the data is transferred between two or more applications. And the distributed message passing is based on a reliable message queue, and messages are asynchronously passed between the client application and the message system.
[0040] Specifically, the message is persisted to a topic, and the consumer can subscribe to one or more topics. The consumer can consume all the data in the topic. The same piece of data can be consumed by multiple consumers, and the data will not be deleted immediately after being consumed. And the data in the topic can be divided into one or more partitions, that is to say, each topic has at least one partition.
[0041] 3. Exception Log: Used to record the data that fails to be written into the ES cluster.
[0042] 4. Redo Log: Used to record the data that is successfully written into the ES cluster in the current batch.
[0043] 5. Trans log: When it is confirmed that the data has been successfully written into the ES cluster, the data will be recorded in the translog. If the ES cluster crashes, the data will be restored from the Translog when restarting to ensure data is not lost. However, there is a certain time interval from when the translog is written into memory to when it is flushed to disk.
[0044] The design concept of the embodiments of the present invention is briefly introduced below:
[0045] Currently, in the solutions provided by related technologies, when implementing data consistency in the ES cluster, a solution for data synchronization based on applications is provided. However, in this solution, since the data is written in the order of first written into memory and then written into the corresponding disk, when the data is written into the ES cluster and the feedback indicates successful writing, if the system crashes while waiting for the data to be flushed from memory to disk, there is a risk of data loss.
[0046] In view of this, the embodiments of the present invention provide a method for ensuring data consistency in a multi-active architecture. In this method, when synchronously writing the information to be written into each ES cluster, the failure log and redo log mechanisms are utilized in the synchronization logic. When the ES proxy subsystem receives the write failure information sent by any ES cluster, it will replay the data in the redo log of the data successfully written into each ES cluster retained within a preset duration into the failure log. When the preset startup duration is reached since the start time of the data recovery task, the updated exception data in the failure log will be synchronously written into each of the ES clusters, that is, the data in the redo log will be written into the ES again, thereby ensuring that the data will definitely not be lost.
[0047] After introducing the design concept of the embodiments of the present invention, some simple introductions are made below to the application scenarios applicable to the technical solution for ensuring data consistency in the multi-active architecture of the embodiments of the present invention. It should be noted that the application scenarios described in the embodiments of the present invention are for more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those of ordinary skill in the art know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.
[0048] Refer to Figure 1As shown in the figure, it is a schematic diagram of an application scenario in an embodiment of the present invention. Specifically, in this scenario diagram, it includes a business system 101, a distributed logging system 102, an ES proxy subsystem 103, and an ES cluster 104. Among them, ES clusters 104-1, 104-2, ……, 104-n can be used by different users.
[0049] Specifically, the business system 101 can convert requests to be written into the ES cluster based on the application into messages and send them to the distributed logging system 102, which is, for example, a kafka distributed logging system. Then the ES proxy subsystem 103 pulls messages from the distributed logging system 102 and synchronously writes the pulled messages into each ES cluster 104 associated with the ES proxy subsystem 103. Specifically, when the ES proxy subsystem 103 receives the write success information feedback from the ES cluster 104, it will write the successfully written data into the redo log, and the redo log will save the data within a preset duration, which is, for example, 10 seconds; when the ES proxy subsystem 103 receives the write failure information feedback from the ES cluster 104, it will write the failed written data into the failure log, and at the same time replay the data in the redo log into the failure log, and then rewrite the data in the failure log into the ES cluster 104. Specifically, how to rewrite the data in the failure log into the ES cluster 104 will be described in detail later.
[0050] In an embodiment of the present invention, the ES cluster 104 can be a server cluster or a distributed system composed of multiple physical servers, or can also be a server cluster or a distributed system composed of cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, but is not limited thereto.
[0051] Among them, between the business system 101, the distributed logging system 102, the ES proxy subsystem 103, and the ES cluster 104, and between each ES cluster 104, they can be directly or indirectly communicatively connected through one or more networks 105. The network 105 can be a wired network or a wireless network. For example, the wireless network can be a mobile cellular network or can be a Wireless-Fidelity (WIFI) network. Of course, it can also be other possible networks, and the embodiments of the present invention do not limit this.
[0052] To further illustrate the solution of the method for ensuring data consistency in the multi-active architecture provided by the embodiments of the present invention, the following will be described in detail in conjunction with the accompanying drawings and specific implementation manners. Although the embodiments of the present invention provide the method operation steps as shown in the following embodiments or drawings, in the method, more or fewer operation steps may be included based on routine or non-creative labor. In the steps where there is no necessary causal relationship logically, the execution order of these steps is not limited to the execution order provided by the embodiments of the present invention. When the method is actually processed or executed by the device, it can be executed in the method order shown in the embodiments or drawings or executed in parallel (such as in an application environment of a parallel processor or multi-threaded processing).
[0053] The following will be described in conjunction with Figure 2 the following method flow chart to illustrate the method for ensuring data consistency in the multi-active architecture in the embodiments of the present invention. The following will introduce the method flow of the embodiments of the present invention.
[0054] Step 201: Pull the first data to be written from the distributed log system and synchronously write the first data to be written into each ES cluster.
[0055] In the embodiments of the present invention, when the ES proxy subsystem is started, it can read the configuration information of the distributed log system, and then subscribe to the corresponding topic, so as to pull the data to be written corresponding to the topic. Further, the ES proxy subsystem can determine whether there is historical abnormal data in the failure log. When it is determined that there is historical abnormal data in the failure log, write the first data to be written into the failure log and stop writing the first data to be written into each ES cluster.
[0056] Specifically, when the ES proxy subsystem determines that there is historical abnormal data in the failure log, it can determine that the data to be written pulled before pulling the first data to be written has not been successfully synchronously written into each ES cluster. In order to ensure the timeliness of the data written into each ES cluster, the ES proxy subsystem can write the first data to be written into the failure log and stop writing the first data to be written into each ES cluster.
[0057] In an embodiment of the present invention, after the ES proxy subsystem writes the first data to be written into the failure log, the ES proxy subsystem determines a pull waiting duration from the distributed log system before pulling the second data to be written. When the time since determining to pull the first data to be written reaches the pull waiting duration, the second data to be written is pulled and sent to the local database. In this way, it is possible to avoid writing the information to be written that is behind in the pull order to the ES cluster in advance, resulting in incorrect timing of the data written to the ES cluster. Specifically, the pull duration is determined based on the processing cycle of historical exception data and a preset policy.
[0058] In an embodiment of the present invention, when the ES proxy subsystem determines that historical exception data already exists in the failure log, it determines the pull order of the second data to be written; wherein, the pull order is determined according to the pull time for the data pulled from the distributed log system after historical exception data already exists in the failure log. Then, the previous pull waiting duration can be determined, and N times the previous pull waiting duration is used as the pull waiting duration; the previous pull waiting duration is: the pull waiting duration corresponding to the data pulled in the previous time of the pull order of the second data to be written; wherein, when the previous pull order is the first time, it is determined that the previous pull waiting duration is the processing duration for the historical exception data, and N is a positive integer greater than 2.
[0059] For example, assume N is 3. If the pull order of the second data to be written is 3, and the pull waiting duration corresponding to the data with a pull order of 2, that is, the previous pull waiting duration, is two minutes, then the waiting duration for the second data to be written can be determined to be six minutes.
[0060] For example, assume N is 2 and the processing duration for the historical exception data is one minute. Before pulling the data a to be written, first check whether there is historical exception data in the failure log. If there is historical exception data in the failure log, that is, it is determined that the pull order of the data a to be written is 1, that is, the pull order is the first time, then it can be determined that the pull waiting duration for the data a to be written is the processing duration for the historical exception data, which is one minute. It can be seen that after pulling the data a to be written, it is possible to wait for one minute before pulling new data.
[0061] And so on. When it is determined that the pull order of the data d to be written is 2, then it can be determined that the data pulled in the previous time of the pull order of 2, that is, the data a to be written. Then, correspondingly, it can be determined that the pull waiting duration is twice that of one minute, which is two minutes. When it is determined that the pull order of the data f to be written is 3, then it can be determined that the data pulled in the previous time of the pull order of 3 is the data d to be written. Then, correspondingly, it can be determined that the pull waiting duration is twice that of two minutes, which is four minutes.
[0062] Specifically, when there is no historical exception data in the failure log, there is no need to wait, and new data to be written can be directly pulled. When the ES proxy subsystem next pulls data to be written from the distributed log system, if it is determined that there is historical exception data in the failure log before pulling the data, the pulling order of the data is determined to be 1, and the corresponding pulling waiting duration is one minute, that is, new data is pulled again one minute after pulling this data.
[0063] In this way, when there is a backlog of historical exception data, the pulling speed of pulling data from the distributed system can be postponed. After the backlog of historical exception data in the failure log is processed, the data pulled from the distributed system can be directly written into each ES cluster.
[0064] In an embodiment of the present invention, after a first preset duration, the second data to be written is obtained from the local database, and the second data to be written is synchronously written into each ES cluster; wherein, the second data to be written is: data with different content pulled from the distributed log system after pulling the first data to be written.
[0065] In an embodiment of the present invention, when the ES proxy subsystem determines that there is no historical exception data in the failure log, the first data to be written can be synchronously written into each ES cluster. When the write success feedback from the ES cluster is received, one copy of the data successfully written into the ES cluster is written into the redo log.
[0066] Step 202: When the write failure information sent by any one of the ES clusters in each ES cluster is received, the write success data recorded in the current redo log is replayed into the failure log to obtain the updated exception data in the failure log; wherein, the redo log includes the data successfully written in each ES cluster retained within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated exception data is the data that has not been successfully written into any ES cluster.
[0067] In an embodiment of the present invention, when the ES proxy subsystem receives the write failure information sent by any one of the ES clusters in each ES cluster, it can be determined that a failure has occurred in any one of the ES clusters. This failure is, for example, that the ES cluster crashes or the area where the ES cluster is located loses power and cannot provide power for the ES cluster, then it can be determined that the data synchronously written into each ES cluster this time is written fails.
[0068] In an embodiment of the present invention, when the ES proxy subsystem receives the write failure information sent by any one of the ES clusters in each ES cluster, the write success data recorded in the current redo log is replayed into the failure log to obtain the updated exception data in the failure log. Specifically, the updated exception data in the failure log can be obtained by, but not limited to, the following steps:
[0069] Step a: Determine whether the updated exception data recorded in the failure log exists in the data recorded in the current redo log;
[0070] Specifically, the ES proxy subsystem can compare the current index and primary key data in the updated exception data with the current index and primary key data in the data recorded in the current redo log, so as to determine whether the updated exception data recorded in the failure log exists in the data recorded in the current redo log.
[0071] Step b: When the updated exception data recorded in the failure log does not exist in the data recorded in the current redo log, insert the data recorded in the redo log into the failure log;
[0072] Step c: When the updated exception data recorded in the failure log exists in the data recorded in the current redo log, determine the timestamp of the updated exception data recorded in the failure log
[0073] Step e: Determine whether the timestamp is greater than the timestamp of the data recorded in the redo log;
[0074] Step f: When it is determined that the timestamp is not greater than the timestamp of the data recorded in the redo log, replay the data recorded in the current redo log into the failure log.
[0075] It can be seen that in the embodiment of the present invention, when the timestamp of the data in the failure log is greater than the timestamp of the data recorded in the redo log, it is determined that the data in the failure log does not need to be updated; when it is determined that the timestamp of the data in the failure log is not greater than the timestamp of the data recorded in the redo log, the data recorded in the current redo log is replayed into the failure log, that is, the data in the failure log is updated, so as to ensure that the data recorded in the failure log is the latest data that has not been successfully written into each ES cluster as much as possible.
[0076] Step 203: If the preset startup duration is reached since the startup moment of the recovery data task, synchronously write the updated exception data in the failure log into each ES cluster.
[0077] In the embodiment of the present invention, the ES subsystem can set a one-by-one lock grabbing mechanism for each recovery data task in the recovery data project; when the lock grabbing for any recovery data task is successful, a child thread corresponding to any recovery data task is established, and based on the child thread, the following is executed: If the preset startup duration is reached since the startup moment of the recovery data task, synchronously write the updated exception data in the failure log into each ES cluster; when each ES cluster feedbacks the write success information, release the lock corresponding to the recovery data task.
[0078] In an embodiment of the present invention, based on the lock grabbing mechanism, the recovery tasks corresponding to each piece of abnormal data can be started one by one to ensure the timeliness of the data rewritten into each ES cluster.
[0079] Next, a specific example is used to illustrate the method for ensuring data consistency in the multi-active architecture provided in the embodiments of the present invention.
[0080] Refer to Figure 3 As shown, it is another implementation flowchart of the method for ensuring data consistency in the multi-active architecture in the embodiments of the present invention.
[0081] Step 301: The ES proxy subsystem sends a request message for subscribing to topics to the distributed log system, so that the distributed log system obtains the corresponding messages from the business system and stores them in the corresponding storage area.
[0082] In an embodiment of the present application, the ES proxy subsystem is a subscriber for subscribing to messages from the distributed log system, the business system is a publisher for publishing messages to the distributed log system, and the distributed log system can allocate the data carried in the write request issued by the business system based on the application to the corresponding Partition according to the hash value corresponding to the primary key. Specifically, in order to strictly ensure the consumption order of messages, the number of partitions can be set to 1, that is, one topic corresponds to one partition. Further, the ES proxy subsystem can obtain the corresponding messages from the partition corresponding to the distributed log system.
[0083] Step 302: The ES proxy subsystem synchronously initializes the ES synchronizers corresponding to each ES cluster, determines the first quantity of the first data to be written to be pulled according to the number of the corresponding ES clusters, and pulls the first quantity of the first data to be written from the distributed log system.
[0084] In an embodiment of the present application, the ES proxy subsystem can synchronize the ES synchronizers corresponding to each ES cluster, so as to ensure that the initial states of the initialized ES synchronizers are the same as much as possible, providing a good implementation basis for subsequent data synchronization.
[0085] Step 303: Each ES synchronizer corresponding to the ES proxy subsystem determines whether there is historical abnormal data in the failure log; if there is no historical abnormal data in the failure log, step 304 is executed; if there is historical abnormal data in the failure log, step 307 is executed.
[0086] In an embodiment of the present application, the manner in which each ES synchronizer corresponding to the ES proxy subsystem determines whether there is historical abnormal data in the failure log can be referred to the execution of step 203, which will not be elaborated here.
[0087] Step 304: Each synchronizer corresponding to the ES proxy subsystem pulls a first quantity of first data to be written from the distributed log system and synchronously writes them into each ES cluster respectively.
[0088] In the embodiment of the present application, when each synchronizer corresponding to the ES proxy subsystem determines that there is no historical abnormal data in the failure log, it pulls a first quantity of first data to be written from the distributed log system and synchronously writes them into each ES cluster respectively.
[0089] Step 305: When each synchronizer corresponding to the ES proxy subsystem receives a write failure message sent by the corresponding ES cluster, it executes Step 307; when it receives a write success message sent by the corresponding ES cluster, it executes Step 306.
[0090] Step 306: The ES proxy subsystem writes the data successfully written into the corresponding ES cluster into the redo log.
[0091] In the embodiment of the present application, the ES proxy subsystem is also provided with a redo log deletion task, and this redo log deletion task can delete the data newly written into the redo log when the time interval from the moment when the data newly written into the redo log is determined to the current moment reaches a preset duration. Specifically, this preset duration is greater than the disk flushing duration corresponding to the ES cluster.
[0092] For example, if the disk flushing duration corresponding to the ES cluster is 5 seconds, the preset duration can be set to 10 seconds. In this way, even if the newly written data has not been written to the disk corresponding to the ES cluster, the latest written data is still saved in the redo log, providing an implementation basis for subsequent data recovery.
[0093] Step 307: The ES proxy subsystem replays the write success data recorded in the current redo log into the failure log to obtain the updated abnormal data in the failure log.
[0094] In the embodiment of the present application, when each synchronizer corresponding to the ES proxy subsystem receives a write failure message sent by the corresponding ES cluster, it can execute the foregoing Steps a - f, so as to replay the write success data recorded in the current redo log into the failure log and obtain the updated abnormal data in the failure log.
[0095] Step 308: When the ES proxy subsystem successfully locks a recovery data task and establishes a child thread corresponding to any recovery data task, it determines to recover the updated abnormal data.
[0096] Step 309: The ES proxy subsystem executes based on a child thread: Starting from the startup moment of the data recovery task, when the preset startup duration is reached, the updated abnormal data in the failure log is synchronously written to each ES cluster;
[0097] Step 310: When the ES proxy subsystem receives the write success information feedback from each ES cluster, the data successfully written to each ES cluster is written to the redo log, and the lock corresponding to the data recovery task is released; when receiving the write failure information sent by any ES cluster, step 307 is executed.
[0098] Step 311: When each synchronizer corresponding to the ES proxy subsystem determines that the synchronization is successful, feedback information is sent to the distributed log system.
[0099] Step 312: The ES proxy subsystem determines whether the subscribed topics to the distributed log system are adjusted. If adjusted, the subscription to the topics from the distributed log system is cancelled.
[0100] It can be seen that in the embodiment of the present application, based on the distributed log system, the message queue is used to asynchronously write the data to be written. The business system does not need to pay attention to the specific writing logic, reducing the cost of the business system accessing the ES proxy subsystem.
[0101] Moreover, considering that even when the ES proxy subsystem receives the write success information, the data written to the ES cluster may not be written to the disk. If a power failure occurs at this time, a certain amount of data will still be lost. After the ES cluster resumes operation, it can send the write failure information to the ES proxy subsystem. Thus, the ES proxy subsystem will replay the data successfully written to the redo log last time to the failure log. When the data recovery task starts and the preset startup duration is reached, the data in the redo log, that is, the abnormal data in the failure log, can be written to the ES cluster again to ensure that the data will not be lost.
[0102] Refer to Figure 4As shown in the figure, the present invention provides a device for ensuring data consistency in a multi-live architecture, which is applied to a distributed search engine ES proxy subsystem, including: a processing unit 401, configured to pull the first data to be written from a distributed log system, and synchronously write the first data to be written into each ES cluster; an obtaining unit 402, configured to, when receiving a write failure message sent by any one of the ES clusters in each ES cluster, replay the successfully written data recorded in the current redo log into the failure log to obtain the updated abnormal data in the failure log; wherein, the redo log includes the successfully written data in each ES cluster retained within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated abnormal data is the data that has not been successfully written into any one of the ES clusters; a recovery write unit 403, configured to, if the preset startup duration is reached since the startup moment of the recovery data task, synchronously write the updated abnormal data in the failure log into each ES cluster.
[0103] Optionally, after pulling the first data to be written from the distributed log system, the processing unit 401 is further configured to: determine whether there is historical abnormal data in the failure log; when it is determined that there is historical abnormal data in the failure log, write the first data to be written into the failure log, and stop writing the first data to be written into each ES cluster.
[0104] Optionally, after writing the first data to be written into the failure log, the processing unit 401 is further configured to: determine a pull waiting duration from the distributed log system before pulling the second data to be written; when it is determined that the pull waiting duration is reached since the moment of pulling the first data to be written, pull the second data to be written, and send the second data to be written to a local database; after a first preset duration, obtain the second data to be written from the local database, and synchronously write the second data to be written into each ES cluster; wherein, the second data to be written is the data with different content pulled from the distributed log system after pulling the first data to be written.
[0105] Optionally, the processing unit 401 is specifically configured to: determine the pulling order of the second data to be written, where the pulling order is determined according to the pulling time of the data pulled from the distributed log system after the historical exception data already exists in the failure log; determine the previous pulling waiting duration, and use N times of the previous pulling waiting duration as the pulling waiting duration, where the previous pulling waiting duration is the pulling waiting duration corresponding to the data pulled in the previous time of the pulling order of the second data to be written; and when the previous pulling order is the first time, determine the previous pulling waiting duration as the processing duration of the historical exception data, and N is a positive integer greater than 2.
[0106] Optionally, the obtaining unit 402 is specifically configured to: determine whether the updated exception data recorded in the failure log exists in the data recorded in the current redo log; when the updated exception data recorded in the failure log does not exist in the data recorded in the current redo log, insert the data recorded in the redo log into the failure log; when the updated exception data recorded in the failure log exists in the data recorded in the current redo log, determine the timestamp of the updated exception data recorded in the failure log; determine whether the timestamp is greater than the timestamp of the data recorded in the redo log; and when it is determined that the timestamp is not greater than the timestamp of the data recorded in the redo log, replay the data recorded in the current redo log to the failure log.
[0107] Optionally, the recovery writing unit 403 is specifically configured to: set a one-by-one lock grabbing mechanism for each recovery data task in the data item to be recovered; when the lock grabbing for any one of the recovery data tasks is successful, establish a sub-thread corresponding to any one of the recovery data tasks, and based on the sub-thread, execute: starting from the start time of the recovery data task, when a preset start duration is reached, synchronously write the updated exception data in the failure log to each of the ES clusters; and when each of the ES clusters feedbacks a successful writing message, release the lock corresponding to the recovery data task.
[0108] Optionally, after synchronously writing the updated exception data in the failure log to each of the ES clusters, the processing unit 401 is further configured to: feedback an offset to the distributed log system; and after completing the feedback of the offset to the distributed log system and determining that the pre-subscribed topic is adjusted, determine not to pull data from the distributed log system.
[0109] An embodiment of the present invention provides a computer device, including a program or an instruction, which when executed, is used to execute a method for ensuring data consistency in a multi-active architecture provided by an embodiment of the present invention and any optional method.
[0110] An embodiment of the present invention provides a storage medium, including a program or instruction, which, when executed, is used to execute a method for ensuring data consistency in a multi-active architecture provided by an embodiment of the present invention and any optional method.
[0111] Finally, it should be noted that those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.
[0112] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0113] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0114] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A method for ensuring data consistency in a multi-active architecture, characterized in that, Applied to the ES proxy subsystem of a distributed search engine, including: Pull the first data to be written from the distributed log system and synchronously write the first data to be written into each ES cluster; When receiving the write failure information sent by any one of the ES clusters, replay the successfully written data recorded in the current redo log to the failure log to obtain the updated abnormal data in the failure log; wherein, the redo log includes the successfully written data in each ES cluster reserved within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated abnormal data is the data that has not been successfully written into any one of the ES clusters; If, starting from the start time of the recovery data task, when the preset start duration is reached, synchronously write the updated abnormal data in the failure log into each ES cluster.
2. The method according to claim 1, wherein After pulling the first data to be written from the distributed log system, the method further includes: Determine whether there is historical abnormal data in the failure log; When it is determined that there is historical abnormal data in the failure log, write the first data to be written into the failure log and stop writing the first data to be written into each ES cluster.
3. The method according to claim 2, wherein After writing the first data to be written into the failure log, the method further includes: Before pulling the second data to be written from the distributed log system, determine the pull waiting duration. When it is determined that the pull waiting duration is reached since the time of pulling the first data to be written, pull the second data to be written and send the second data to be written to the local database; After the first preset duration, obtain the second data to be written from the local database and synchronously write the second data to be written into each ES cluster; wherein, the second data to be written is the data with different content pulled from the distributed log system after pulling the first data to be written.
4. The method according to claim 3, characterized in that, Determining the pull waiting duration includes: Determine the pull order of the second data to be written; wherein, the pull order is determined according to the pull time for the data pulled from the distributed log system after there is historical abnormal data in the failure log; Determine the previous pull waiting duration and use N times the previous pull waiting duration as the pull waiting duration; the previous pull waiting duration is the pull waiting duration corresponding to the data pulled in the previous pull order of the second data to be written; Wherein, when the previous pull order is the first time, determine the previous pull waiting duration as the processing duration for the historical abnormal data, and N is a positive integer greater than 2.
5. The method according to claim 1, wherein When receiving the write failure information sent by any one of the ES clusters, replaying the data recorded in the current redo log to the failure log includes: Judge whether the updated abnormal data recorded in the failure log exists in the data recorded in the current redo log; When the updated exception data recorded in the failure log does not exist in the data recorded in the current redo log, insert the data recorded in the redo log into the failure log; When the updated exception data recorded in the failure log exists in the data recorded in the current redo log, determine the timestamp of the updated exception data recorded in the failure log; Judge whether the timestamp is greater than the timestamp of the data recorded in the redo log; When it is determined that the timestamp is not greater than the timestamp of the data recorded in the redo log, replay the data recorded in the current redo log to the failure log.
6. The method according to claim 1, characterized in that, If, since the start time of the data recovery task, a preset start duration is reached, synchronously write the updated exception data in the failure log to each ES cluster, including: Set a one-by-one lock grabbing mechanism for each data recovery task in the data recovery project; When the lock grabbing for any one of the data recovery tasks is successful, establish a child thread corresponding to any one of the data recovery tasks, and based on the child thread, execute: If, since the start time of the data recovery task, a preset start duration is reached, synchronously write the updated exception data in the failure log to each ES cluster; When each ES cluster feeds back a write success message, release the lock corresponding to the data recovery task.
7. The method according to claim 1, characterized in that, After synchronously writing the updated exception data in the failure log to each ES cluster, the method further includes: Feed back the offset to the distributed log system; After completing the feedback of the offset to the distributed log system and determining that the pre-subscribed topic is adjusted, determine not to pull data from the distributed log system.
8. An apparatus for ensuring data consistency in a multi-active architecture, characterized in that Applied to a distributed search engine ES proxy subsystem, including: A processing unit, configured to pull the first data to be written from the distributed log system and synchronously write the first data to be written to each ES cluster; An obtaining unit, configured to, when receiving a write failure message sent by any one of the ES clusters in each ES cluster, replay the successfully written data recorded in the current redo log to the failure log and obtain the updated exception data in the failure log; wherein, the redo log includes the successfully written data in each ES cluster reserved within a preset duration; the preset duration is greater than the disk flushing duration corresponding to the ES cluster; the updated exception data is the data that has not been successfully written to any ES cluster; A recovery writing unit, configured to, if, since the start time of the data recovery task, a preset start duration is reached, synchronously write the updated exception data in the failure log to each ES cluster.
9. A computer device, characterized in that, Includes a program or instruction, and when the program or instruction is executed, the method described in any one of claims 1 to 7 is executed.
10. A storage medium, characterized in that, Includes a program or instruction, and when the program or instruction is executed, the method described in any one of claims 1 to 7 is executed.
Citation Information
Patent Citations
Database recovery method and device, storage medium and database system
CN110955556A
Pre-written log record sorting system in database cluster
CN112131318A