Multi-cluster data search method, device and computer readable storage medium

CN116204120BActive Publication Date: 2026-09-25中国邮政储蓄银行股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211698853.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-28
Publication Date
2026-09-25
Estimated Expiration
2042-12-28

AI Technical Summary

Technical Problem

[0010]本申请的主要目的在于提供一种多集群的数据搜索方法、数据搜索装置与计算机可读存储介质,以至少解决现有技术中难以较为可靠地对集群间的数据进行读写的问题

Benefits of technology

[0021]应用本申请的技术方案,所述的多集群的数据搜索方法中,基于目标SDK和预定写入策略,将从数据源中读取的第一目标数据,写入至目标ES集群中;基于预定读取策略和目标SDK,从目标ES集群中读取第二目标数据,且将读取的第二目标数据发送至目标UI界面,以将第二目标数据显示在目标UI界面。本申请的数据搜索方法中,开发了跨集群统一读写的服务组件,即目标SDK,同时设置了较为可靠的数据写入策略和数据读取策略,即预定写入策略和预定读取策略。基于目标SDK和预定写入策略,实现了较为可靠地将第一目标数据写入目标ES集群,以及基于目标SDK和预定读取策略,实现了较为可靠地对第二目标数据进行读取,这样保证了对数据的写入和读取均较为可靠以及数据一致性较高,保证了集群部署的复杂性向应用透明,从而解决了现有技术中难以较为可靠地对集群间的数据进行读写的问题。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116204120B_ABST
    Figure CN116204120B_ABST
Patent Text Reader

Abstract

The application provides a multi-cluster data search method, device and computer readable storage medium. The data search method comprises the following steps: reading first target data from a data source, and writing the first target data into a target ES cluster based on a target SDK and a predetermined writing strategy, the target ES cluster comprising at least one of a first ES cluster and a second ES cluster, and the predetermined writing strategy comprising at least one of simultaneous writing, fault-tolerant writing and specified writing; reading second target data from the target ES cluster based on a predetermined reading strategy and the target SDK, and sending the second target data to a target UI interface, the predetermined reading strategy comprising at least one of specified reading and default reading. The application ensures that the writing and reading of data are relatively reliable and the data consistency is high, ensures that the complexity of cluster deployment is transparent to the application, and thus solves the problem that it is difficult to reliably read and write data between clusters in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and more specifically, to a multi-cluster data search method, a data search device, and a computer-readable storage medium. Background Technology

[0002] In scenarios involving detailed content queries, the new core system, compared to the old one, requires more detailed information, a longer time span, and simpler query steps. To ensure high availability of the search service, the new core system employs a two-site, three-center deployment. The two centers within the same city utilize a dual-active strategy, while the off-site center serves as disaster recovery storage. This multi-center deployment not only meets cluster-level disaster recovery requirements but also guarantees real-time read / write performance and data integrity for massive amounts of search content.

[0003] Most existing highly available big data search systems are deployed in two forms:

[0004] 1) Single cluster multi-zone deployment: This involves deploying a single cluster across data centers, with machines distributed across at least two locations. This deployment method only involves one cluster for data writing and querying, making data processing relatively simple and avoiding the need for rewriting and upgrading based on open-source clients.

[0005] 2) Multi-cluster, multi-zone deployment: This involves deploying multiple independent clusters across multiple data centers. This deployment method utilizes Elasticsearch's cross-cluster replication and cross-cluster search capabilities to perform data searches across multiple independent clusters. However, this method requires configuring local clusters (leaders) and remote clusters (followers) and does not involve rewriting or upgrading open-source clients.

[0006] The above deployment method also has the following problems:

[0007] 1) Single-cluster multi-zone deployment: To avoid split-brain phenomena, there must be a distinction between large and small centers. If the large center with many master-eligible nodes goes down, Elasticsearch becomes unavailable. The cluster can only survive if a small center goes down. This requires high network performance between centers. If there are network problems between multiple centers that prevent them from communicating, incorrect master election behavior may occur.

[0008] 2) Multi-cluster, multi-zone deployment: The local cluster connects to the remote cluster via TCP (Transmission Control Protocol). If the entire local cluster fails, it will be unable to connect to the remote cluster, resulting in data loss during writes and unavailability for data reads.

[0009] Therefore, there is an urgent need for a method that can read and write data between clusters in a relatively flexible and reliable manner while ensuring the high availability of the search system. Summary of the Invention

[0010] The main objective of this application is to provide a multi-cluster data search method, data search device, and computer-readable storage medium, so as to at least solve the problem in the prior art that it is difficult to reliably read and write data between clusters.

[0011] To achieve the above objectives, according to one aspect of this application, a multi-cluster data search method is provided, comprising: reading first target data from a data source, and writing the first target data into a target ES cluster based on a target SDK and a predetermined write strategy, wherein the target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster, and the predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write, wherein the first ES cluster is a cluster deployed in a first data center, the second ES cluster is a cluster deployed in a second data center, and the target SDK is a service component having at least write and read functions; reading second target data from the target ES cluster based on a predetermined read strategy and the target SDK, and sending the second target data to a target UI interface, wherein the predetermined read strategy includes at least one of the following: specified read and default read.

[0012] Optionally, after writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: if the predetermined write strategy is simultaneous writing and the first target data is successfully written into both the first ES cluster and the second ES cluster, sending a response result to the UI interface, wherein the response result indicates that the first target data was successfully written; if the predetermined write strategy is fault-tolerant writing and the first target data is successfully written into either the first ES cluster or the second ES cluster, sending the response result to the UI interface; if the predetermined write strategy is specified writing and the first target data is successfully written into a specified ES cluster, sending the response result to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

[0013] Optionally, after writing the first target data into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: if the first target data is not successfully written into the target ES cluster based on the target SDK and the predetermined write strategy, storing the first target data into the target database; and writing the first target data in the target database into the target ES cluster again based on the target SDK and the predetermined write strategy.

[0014] Optionally, after writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write policy, the data search method further includes: generating an alarm log if the first target data is not successfully written back into the target ES cluster based on the target SDK and the predetermined write policy; and sending the alarm log to the operation and maintenance system so that operation and maintenance personnel can locate and analyze the problem based on the alarm log.

[0015] Optionally, reading second target data from the target ES cluster based on a predetermined reading strategy and the target SDK includes: reading the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; or reading the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0016] Optionally, when the predetermined read strategy is the specified read, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined read strategy is the specified read and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; when the predetermined read strategy is the default read, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined read strategy is the default read and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, reading the second target data from the backup ES cluster.

[0017] Optionally, the data search method further includes: based on the target SDK, periodically probing multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and saving the first cluster list and the second cluster list to a local cache.

[0018] Optionally, reading the first target data from the data source includes: using a DTS component to read the data source to obtain predetermined data; and using the DTS component to perform format conversion on the predetermined data to obtain the first target data.

[0019] According to another aspect of this application, a multi-cluster data search apparatus is provided, comprising: a first execution unit, configured to read first target data from a data source and write the first target data into a target ES cluster based on a target SDK and a predetermined write strategy, wherein the target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster, and the predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write, wherein the first ES cluster is a cluster deployed in a first data center, the second ES cluster is a cluster deployed in a second data center, and the target SDK is a service component having at least write and read functions; and a second execution unit, configured to read second target data from the target ES cluster based on a predetermined read strategy and the target SDK, and send the second target data to a target UI interface, wherein the predetermined read strategy includes at least one of the following: specified read and default read.

[0020] According to another aspect of this application, a computer-readable storage medium is provided, the computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform any of the multi-cluster data search methods.

[0021] Applying the technical solution of this application, the multi-cluster data search method, based on the target SDK and a predetermined write strategy, writes the first target data read from the data source into the target ES cluster; based on the predetermined read strategy and the target SDK, it reads the second target data from the target ES cluster and sends the read second target data to the target UI interface to display the second target data on the target UI interface. In the data search method of this application, a cross-cluster unified read / write service component, namely the target SDK, is developed, and relatively reliable data write and read strategies, namely predetermined write and read strategies, are set. Based on the target SDK and the predetermined write strategy, relatively reliable writing of the first target data into the target ES cluster and relatively reliable reading of the second target data are achieved. This ensures that both data writing and reading are relatively reliable and that data consistency is high, ensuring that the complexity of cluster deployment is transparent to the application, thereby solving the problem in the prior art of reliably reading and writing data between clusters. Attached Figure Description

[0022] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0023] Figure 1 A hardware structure block diagram of a mobile terminal for performing a multi-cluster data search method is shown in an embodiment of this application;

[0024] Figure 2 A flowchart illustrating a multi-cluster data search method provided by an embodiment of this application is shown.

[0025] Figure 3 A structural block diagram of a target SDK provided by an embodiment of this application is shown;

[0026] Figure 4 A flowchart illustrating the writing of first target data to a target Elasticsearch cluster according to an embodiment of this application is shown;

[0027] Figure 5 A flowchart illustrating a data writing process provided by an embodiment of this application is shown;

[0028] Figure 6 A flowchart illustrating an anomaly during the writing of first target data to a target Elasticsearch cluster, as provided in an embodiment of this application, is shown.

[0029] Figure 7 A flowchart illustrating a timed liveness detection method provided in an embodiment of this application is shown;

[0030] Figure 8 A schematic diagram of the structure of a multi-cluster data search device provided in an embodiment of this application is shown. Detailed Implementation

[0031] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0032] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0033] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0034] For ease of description, the following explains some of the nouns or terms used in the embodiments of this application:

[0035] The new core system refers to the bank's next-generation core system, which employs a distributed and microservice architecture to deploy core systems such as deposit and remittance systems. The deposit core system for individual customers provides transaction details for withdrawals, inquiries, and transfers, allowing customers to search in real time.

[0036] Elasticsearch (ES) is an open-source, distributed, scalable, real-time search and data analysis engine, and also a document-oriented, non-relational database.

[0037] Flink: A streaming job processing framework and distributed processing engine for stateful computation on unbounded and bounded data streams.

[0038] As described in the background section, it is difficult to reliably read and write data between clusters in the prior art. To solve the problem of the difficulty in reliably reading and writing data between clusters in the prior art, embodiments of this application provide a multi-cluster data search method, a data search device, and a computer-readable storage medium.

[0039] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0040] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a multi-cluster data search method according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0041] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the device information display method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned networks may include wireless networks provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0042] This embodiment provides a multi-cluster data search method that runs on a mobile terminal, computer terminal, or similar computing device. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0043] Figure 2 This is a flowchart of a multi-cluster data search method according to an embodiment of this application. For example... Figure 2 As shown, this data search method includes the following steps:

[0044] Step S201: Read the first target data from the data source, and write the first target data into the target ES cluster based on the target SDK and the predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least write and read functions.

[0045] Specifically, such as Figure 3As shown, the target SDK integrates native clients for ES and Flink at the underlying level (i.e., respectively). Figure 3 The Elasticsearch-rest-connector and Flink-connector-elasticsearch shown above, on this basis, specifically implement the multi-datacenter read / write implementation layer (i.e. Figure 3 The Mdp-search-es-impl, Mdp-search-core-impl, and Mdp-search-es-flink shown) and the interface layer (i.e. Figure 3 The project includes Mdp-search-api, Mdp-search-core-api, and Mdp-search-autoconfig, and provides a quick-integration starter sub-project (i.e., ...) for upper-layer applications. Figure 3 (As shown in Mdp-search-starter and Mdp-flink-starter).

[0046] Specifically, the target SDK provides flexible dual-write functionality and highly fault-tolerant cross-datacenter search functionality, and integrates functions such as multi-datacenter liveness detection and monitoring alerts.

[0047] Specifically, the first Elasticsearch cluster mentioned above is a cluster deployed in the first data center, and the second Elasticsearch cluster mentioned above is a cluster deployed in the second data center. Within each data center's single cluster, nodes are divided into Master Nodes and Data Nodes.

[0048] Step S202: Based on the predetermined reading strategy and the target SDK, read the second target data from the target ES cluster and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0049] In the aforementioned multi-cluster data search method, based on the target SDK and a predetermined write strategy, first target data read from the data source is written to the target ES cluster; based on the predetermined read strategy and the target SDK, second target data is read from the target ES cluster and sent to the target UI interface to display the second target data. In the data search method of this application, a cross-cluster unified read / write service component, namely the target SDK, is developed, and relatively reliable data write and read strategies, namely predetermined write and read strategies, are set. Based on the target SDK and the predetermined write strategy, relatively reliable writing of the first target data to the target ES cluster and relatively reliable reading of the second target data are achieved. This ensures that both data writing and reading are relatively reliable and that data consistency is high, ensuring that the complexity of cluster deployment is transparent to the application, thereby solving the problem of reliably reading and writing data between clusters in existing technologies.

[0050] In one specific embodiment, such as Figure 4 As shown, a first data center and a second data center are set up within the same city (i.e., "dual-center within the same city" in the "two-location three-center" architecture). When a data write request is received in either data center, the DTS component reads data from the data source in the corresponding data center to obtain the first target data, and writes the first target data to the Kafka cluster. The Flink job consumes and parses the first target data in the Kafka cluster. Afterwards, the Flink job calls the target SDK to write the parsed first target data to the first Elasticsearch (ES) cluster and the second ES cluster (where, Figure 4 The dashed line represents remote writing (while the implementation is local writing). If the first target data is not successfully written to the first ES cluster and the second ES cluster, the exception handling process is initiated. This can largely guarantee data integrity and data consistency across multiple data centers.

[0051] In another specific embodiment, such as Figure 5 As shown, data writing involves steps such as data verification, data format conversion, common processing, data splitting, and data sinking. Specifically, a Flink job consumes data from a Kafka cluster to obtain the first target data. Then, the first target data undergoes data verification, data format conversion, common processing, data splitting, and data sinking to write the first target data into the index database (first ES cluster and second ES cluster).

[0052] In the specific implementation process, in order to further provide a more flexible dual-write function and further ensure the reliability of data writing, after step S201, the data search method can also be implemented through the following steps: when the predetermined writing strategy is simultaneous writing and the first target data is successfully written to the first ES cluster and the second ES cluster, a response result is sent to the UI interface, and the response result is a result indicating that the first target data was successfully written; when the predetermined writing strategy is fault-tolerant writing and the first target data is successfully written to either the first ES cluster or the second ES cluster, the response result is sent to the UI interface; when the predetermined writing strategy is specified writing and the first target data is successfully written to a specified ES cluster, the response result is sent to the UI interface, and the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster. In this scheme, as shown in Table 1, when the predefined write strategy is simultaneous write, the write is successful only if the first target data is successfully written to both the first ES cluster and the second ES cluster simultaneously; when the predefined write strategy is fault-tolerant write, the write is successful if either the first ES cluster or the second ES cluster is successfully written; when the predefined write strategy is specified write, the write is successful if the specified ES cluster is successfully written. This further ensures that the write strategy for the first target data is more flexible and reliable.

[0053] Table 1

[0054]

[0055] In the specific implementation process, after step S201 above, the data search method can also be implemented through the following steps: if the first target data is not successfully written to the target ES cluster based on the target SDK and the predetermined writing strategy, the first target data is stored in the target database; based on the target SDK and the predetermined writing strategy, the first target data in the target database is written to the target ES cluster again. Specifically, if the first target data is not successfully written to the target ES cluster (i.e., the first ES cluster and the second ES cluster) based on the target SDK and the predetermined writing strategy, the first target data is stored in the target database to wait for another attempt to write to the target ES cluster. This further ensures that the first target data can be written to the target ES cluster more accurately and reliably.

[0056] In practical applications, such as Figure 6As shown, after reading the first target data from the data source using the DTS component, the first target data is written to the MQ cluster (specifically, it can be written to the Kafka cluster). The Flink Job consumes the first target data from the Kafka cluster and parses it. The parsed first target data is then written to the first Elasticsearch (ES) cluster and the second ES cluster, respectively. If the first target data cannot be successfully written to the first and second ES clusters, the exception handling process begins. The specific steps of the exception handling process are: storing the first target data that failed to be written to the first and second ES clusters into the target database, and enabling timed compensation logic. Specifically, the timed compensation logic is: calling the target SDK again to query the target database, parse the data, and write it. If calling the target SDK again and the pre-defined write strategy fails to successfully write the first target data to the first and second ES clusters, the exception handling process resumes, generating an alarm log and sending it to the operations and maintenance platform so that operations and maintenance personnel can locate and analyze the problem based on the alarm log.

[0057] In some embodiments, in order to read the second target data more flexibly, step S202 can be implemented by the following steps: when the predetermined reading strategy is the specified reading, the second target data is read from a specified ES cluster, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; when the predetermined reading strategy is the default reading, the second target data is read from a default ES cluster, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0058] In practical applications, when the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: if the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, then reading the second target data from a backup ES cluster, where the backup ES cluster is an ES cluster deployed in a backup data center; if the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: if the predetermined reading strategy is the default reading and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, then reading the second target data from the backup ES cluster. In this embodiment, if the second target data cannot be read from the specified ES cluster or the default ES cluster for a predetermined number of consecutive times, then the second target data can be read from the backup ES cluster. This further ensures the high fault tolerance of the data search method of this application and further achieves highly available search performance.

[0059] In the specific implementation process, the above data search method also includes: based on the above target SDK, performing periodic liveness detection on multiple ES clusters in the above first data center and the above second data center to obtain a first cluster list and a second cluster list, wherein the above first cluster list is a list of available ES clusters and the above second cluster list is a list of unavailable ES clusters; saving the above first cluster list and the above second cluster list to a local cache, thereby further ensuring the high availability of the search service (i.e., writing data and reading data).

[0060] In one specific embodiment of this application, such as Figure 7 As shown, when the application starts, it loads the local configuration and simultaneously starts an Apoll0 listener to monitor for changes. If a configuration change is detected, the Apollo configuration is loaded, with Apollo configuration having higher priority than the local configuration. Simultaneously, based on the target SDK, multiple Elasticsearch clusters in the first and second data centers are periodically probed for availability, resulting in lists of the first and second clusters. These lists are stored locally in a cache (i.e., a cached availability list). The program periodically probes each data center again, promptly removing unavailable clusters from the first cluster list and adding them to the second cluster list; conversely, once an unavailable cluster becomes available, it is added back to the first cluster list.

[0061] In some embodiments, reading first target data from a data source includes: using a DTS component to read the data source to obtain predetermined data; using the DTS component to perform format conversion on the predetermined data to obtain the first target data. This ensures that the data format of the first target data meets the data writing format of the first ES cluster and the second ES cluster, further ensuring that the writing of the first target data is relatively accurate and reliable.

[0062] This application also provides a multi-cluster data search device. It should be noted that the multi-cluster data search device of this application can be used to execute the multi-cluster data search method provided in this application. This device is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0063] The following describes the multi-cluster data search device provided in the embodiments of this application.

[0064] Figure 8 This is a schematic diagram of the structure of a multi-cluster data search device according to an embodiment of this application. Figure 8 As shown, the data search device includes:

[0065] The first execution unit 10 is used to read the first target data from the data source and write the first target data into the target ES cluster based on the target SDK and the predetermined writing strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined writing strategy includes at least one of the following: simultaneous writing, fault-tolerant writing, and specified writing. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least writing and reading functions.

[0066] The second execution unit 20 is used to read second target data from the target ES cluster based on a predetermined reading strategy and the target SDK, and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0067] In the aforementioned multi-cluster data search device, the first execution unit is used to write first target data read from the data source to the target ES cluster based on the target SDK and a predetermined write strategy; the second execution unit is used to read second target data from the target ES cluster based on the predetermined read strategy and the target SDK, and send the read second target data to the target UI interface to display the second target data on the target UI interface. In the data search device of this application, a cross-cluster unified read / write service component, namely the target SDK, is developed, and relatively reliable data write and read strategies, namely predetermined write and read strategies, are set. Based on the target SDK and the predetermined write strategy, relatively reliable writing of the first target data to the target ES cluster and relatively reliable reading of the second target data are achieved. This ensures that both data writing and reading are relatively reliable and that data consistency is high, ensuring that the complexity of cluster deployment is transparent to the application, thereby solving the problem of reliably reading and writing data between clusters in the prior art.

[0068] In one specific embodiment, such as Figure 4 As shown, a first data center and a second data center are set up within the same city (i.e., "dual-center within the same city" in the "two-location three-center" architecture). When a data write request is received in either data center, the DTS component reads data from the data source in the corresponding data center to obtain the first target data, and writes the first target data to the Kafka cluster. The Flink job consumes and parses the first target data in the Kafka cluster. Afterwards, the Flink job calls the target SDK to write the parsed first target data to the first Elasticsearch (ES) cluster and the second ES cluster (where, Figure 4 The dashed line represents remote writing (while the implementation is local writing). If the first target data is not successfully written to the first ES cluster and the second ES cluster, the exception handling process is initiated. This can largely guarantee data integrity and data consistency across multiple data centers.

[0069] In another specific embodiment, such as Figure 5 As shown, data writing involves steps such as data verification, data format conversion, common processing, data splitting, and data sinking. Specifically, a Flink job consumes data from a Kafka cluster to obtain the first target data. Then, the first target data undergoes data verification, data format conversion, common processing, data splitting, and data sinking to write the first target data into the index database (first ES cluster and second ES cluster).

[0070] In the specific implementation process, in order to further provide a more flexible dual-write function and further ensure the reliability of data writing, the data search device also includes a first sending unit, a second sending unit, and a third sending unit. The first sending unit is used to send a response result to the UI interface when the predetermined writing strategy is simultaneous writing and the first target data is successfully written to both the first ES cluster and the second ES cluster. The response result indicates that the first target data was successfully written. The second sending unit is used to send the response result to the UI interface when the predetermined writing strategy is fault-tolerant writing and the first target data is successfully written to either the first ES cluster or the second ES cluster. The third sending unit is used to send the response result to the UI interface when the predetermined writing strategy is specified writing and the first target data is successfully written to a specified ES cluster. The specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster. In this scheme, as shown in Table 1, when the predefined write strategy is simultaneous write, the write is successful only if the first target data is successfully written to both the first ES cluster and the second ES cluster simultaneously; when the predefined write strategy is fault-tolerant write, the write is successful if either the first ES cluster or the second ES cluster is successfully written; when the predefined write strategy is specified write, the write is successful if the specified ES cluster is successfully written. This further ensures that the write strategy for the first target data is more flexible and reliable.

[0071] In the specific implementation process, after step S201, the data search device further includes a storage unit and a third execution unit. The storage unit is used to store the first target data into a target database if, based on the target SDK and the predetermined write strategy, the first target data is not successfully written into the target ES cluster. The third execution unit is used to write the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy. Specifically, if, based on the target SDK and the predetermined write strategy, the first target data is not successfully written into the target ES cluster (i.e., the first ES cluster and the second ES cluster), the first target data is stored in the target database to await another attempt to write it into the target ES cluster. This further ensures that the first target data can be written into the target ES cluster more accurately and reliably.

[0072] In practical applications, such as Figure 6As shown, after reading the first target data from the data source using the DTS component, the first target data is written to the MQ cluster (specifically, it can be written to the Kafka cluster). The Flink Job consumes the first target data from the Kafka cluster and parses it. The parsed first target data is then written to the first Elasticsearch (ES) cluster and the second ES cluster, respectively. If the first target data cannot be successfully written to the first and second ES clusters, the exception handling process begins. The specific steps of the exception handling process are: storing the first target data that failed to be written to the first and second ES clusters into the target database, and enabling timed compensation logic. Specifically, the timed compensation logic is: calling the target SDK again to query the target database, parse the data, and write it. If calling the target SDK again and the pre-defined write strategy fails to successfully write the first target data to the first and second ES clusters, the exception handling process resumes, generating an alarm log and sending it to the operations and maintenance platform so that operations and maintenance personnel can locate and analyze the problem based on the alarm log.

[0073] In some embodiments, to further facilitate more flexible reading of the second target data, the second execution unit includes a first execution module and a second execution module. The first execution module is used to read the second target data from a specified ES cluster when the predetermined reading strategy is specified reading. The specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster. The second execution module is used to read the second target data from a default ES cluster when the predetermined reading strategy is default reading. The default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0074] In practical applications, the aforementioned data search device further includes a fourth execution unit, configured to read the second target data from a designated ES cluster when the predetermined reading strategy is the specified reading, and then read the second target data from a backup ES cluster when the predetermined reading strategy is the specified reading and the second target data fails to be read from the designated ES cluster for a predetermined number of consecutive reads. The backup ES cluster is an ES cluster deployed in a backup data center. The aforementioned data search method further includes a fifth execution unit, configured to read the second target data from a default ES cluster when the predetermined reading strategy is the default reading, and then read the second target data from the backup ES cluster when the predetermined reading strategy is the default reading and the second target data fails to be read from the default ES cluster for a predetermined number of consecutive reads. In this embodiment, if the second target data fails to be read from either the designated ES cluster or the default ES cluster for a predetermined number of consecutive reads, the second target data can be read from the backup ES cluster. This further ensures the high fault tolerance of the data search method of this application and further achieves highly available search performance.

[0075] In the specific implementation process, the data search device further includes a liveness detection unit, which is used to periodically detect the liveness of multiple ES clusters in the first data center and the second data center based on the target SDK, and obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; the first cluster list and the second cluster list are saved to the local cache, thereby further ensuring the high availability of the search service (i.e., writing data and reading data).

[0076] In one specific embodiment of this application, such as Figure 7 As shown, when the application starts, it loads the local configuration and simultaneously starts an Apoll0 listener to monitor for changes. If a configuration change is detected, the Apollo configuration is loaded, with Apollo configuration having higher priority than the local configuration. Simultaneously, based on the target SDK, multiple Elasticsearch clusters in the first and second data centers are periodically probed for availability, resulting in lists of the first and second clusters. These lists are stored locally in a cache (i.e., a cached availability list). The program periodically probes each data center again, promptly removing unavailable clusters from the first cluster list and adding them to the second cluster list; conversely, once an unavailable cluster becomes available, it is added back to the first cluster list.

[0077] In some embodiments, the first execution unit includes a reading module and a conversion module. The reading module is used to read the data source using a DTS component to obtain predetermined data. The conversion module is used to convert the predetermined data using the DTS component to obtain the first target data. This ensures that the data format of the first target data meets the data writing format of the first ES cluster and the second ES cluster, further ensuring that the writing of the first target data is relatively accurate and reliable.

[0078] The aforementioned multi-cluster data search device includes a processor and a memory. The first execution unit and the second execution unit, etc., are all stored as program units in the memory. The processor executes these program units stored in the memory to achieve the corresponding functions. All of the above modules are located in the same processor; alternatively, the modules may be located in different processors in any combination.

[0079] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can address the problem of reliably reading and writing data between clusters, a problem that is difficult to solve in existing technologies.

[0080] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0081] This invention provides a computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the multi-cluster data search method.

[0082] Specifically, multi-cluster data search methods include:

[0083] Step S201: Read the first target data from the data source, and write the first target data into the target ES cluster based on the target SDK and the predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least write and read functions.

[0084] Specifically, such as Figure 3 As shown, the target SDK integrates native clients for ES and Flink at the underlying level (i.e., respectively). Figure 3The Elasticsearch-rest-connector and Flink-connector-elasticsearch shown above, on this basis, specifically implement the multi-datacenter read / write implementation layer (i.e. Figure 3 The Mdp-search-es-impl, Mdp-search-core-impl, and Mdp-search-es-flink shown) and the interface layer (i.e. Figure 3 The project includes Mdp-search-api, Mdp-search-core-api, and Mdp-search-autoconfig, and provides a quick-integration starter sub-project (i.e., ...) for upper-layer applications. Figure 3 (As shown in Mdp-search-starter and Mdp-flink-starter).

[0085] Specifically, the target SDK provides flexible dual-write functionality and highly fault-tolerant cross-data center search functionality, and integrates functions such as multi-data center liveness detection and monitoring alerts.

[0086] Specifically, the first Elasticsearch cluster mentioned above is a cluster deployed in the first data center, and the second Elasticsearch cluster mentioned above is a cluster deployed in the second data center. Within each data center's single cluster, nodes are divided into Master Nodes and Data Nodes.

[0087] Step S202: Based on the predetermined reading strategy and the target SDK, read the second target data from the target ES cluster and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0088] Optionally, after writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: if the predetermined write strategy is simultaneous writing and the first target data is successfully written into both the first ES cluster and the second ES cluster, sending a response result to the UI interface, wherein the response result indicates that the first target data was successfully written; if the predetermined write strategy is fault-tolerant writing and the first target data is successfully written into either the first ES cluster or the second ES cluster, sending the response result to the UI interface; if the predetermined write strategy is specified writing and the first target data is successfully written into a specified ES cluster, sending the response result to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

[0089] Optionally, after writing the first target data into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: if the first target data is not successfully written into the target ES cluster based on the target SDK and the predetermined write strategy, storing the first target data into the target database; and writing the first target data from the target database into the target ES cluster again based on the target SDK and the predetermined write strategy.

[0090] Optionally, after writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: generating an alarm log if the first target data is not successfully written back into the target ES cluster based on the target SDK and the predetermined write strategy; and sending the alarm log to the operation and maintenance system so that operation and maintenance personnel can locate and analyze the problem based on the alarm log.

[0091] Optionally, reading the second target data from the target ES cluster based on the predetermined reading strategy and the target SDK includes: reading the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; or reading the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0092] Optionally, when the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; when the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined reading strategy is the default reading and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, reading the second target data from the backup ES cluster.

[0093] Optionally, the above data search method further includes: based on the target SDK, periodically probing multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and saving the first cluster list and the second cluster list to a local cache.

[0094] Optionally, reading the first target data from the data source includes: using a DTS component to read the data source and obtain predetermined data; using the DTS component to perform format conversion on the predetermined data and obtain the first target data.

[0095] This invention provides a processor for running a program, wherein the program executes the multi-cluster data search method.

[0096] Specifically, multi-cluster data search methods include:

[0097] Step S201: Read the first target data from the data source, and write the first target data into the target ES cluster based on the target SDK and the predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least write and read functions.

[0098] Specifically, such as Figure 3 As shown, the target SDK integrates native clients for ES and Flink at the underlying level (i.e., respectively). Figure 3 The Elasticsearch-rest-connector and Flink-connector-elasticsearch shown above, on this basis, specifically implement the multi-datacenter read / write implementation layer (i.e. Figure 3 The Mdp-search-es-impl, Mdp-search-core-impl, and Mdp-search-es-flink shown) and the interface layer (i.e. Figure 3 The project includes Mdp-search-api, Mdp-search-core-api, and Mdp-search-autoconfig, and provides a quick-integration starter sub-project (i.e., ...) for upper-layer applications. Figure 3(As shown in Mdp-search-starter and Mdp-flink-starter).

[0099] Specifically, the target SDK provides flexible dual-write functionality and highly fault-tolerant cross-datacenter search functionality, and integrates functions such as multi-datacenter liveness detection and monitoring alerts.

[0100] Specifically, the first Elasticsearch cluster mentioned above is a cluster deployed in the first data center, and the second Elasticsearch cluster mentioned above is a cluster deployed in the second data center. Within each data center's single cluster, nodes are divided into Master Nodes and Data Nodes.

[0101] Step S202: Based on the predetermined reading strategy and the target SDK, read the second target data from the target ES cluster and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0102] Optionally, after writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: if the predetermined write strategy is simultaneous writing and the first target data is successfully written into both the first ES cluster and the second ES cluster, sending a response result to the UI interface, wherein the response result indicates that the first target data was successfully written; if the predetermined write strategy is fault-tolerant writing and the first target data is successfully written into either the first ES cluster or the second ES cluster, sending the response result to the UI interface; if the predetermined write strategy is specified writing and the first target data is successfully written into a specified ES cluster, sending the response result to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

[0103] Optionally, after writing the first target data into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: if the first target data is not successfully written into the target ES cluster based on the target SDK and the predetermined write strategy, storing the first target data into the target database; and writing the first target data from the target database into the target ES cluster again based on the target SDK and the predetermined write strategy.

[0104] Optionally, after writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: generating an alarm log if the first target data is not successfully written back into the target ES cluster based on the target SDK and the predetermined write strategy; and sending the alarm log to the operation and maintenance system so that operation and maintenance personnel can locate and analyze the problem based on the alarm log.

[0105] Optionally, reading the second target data from the target ES cluster based on the predetermined reading strategy and the target SDK includes: reading the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; or reading the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0106] Optionally, when the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; when the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined reading strategy is the default reading and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, reading the second target data from the backup ES cluster.

[0107] Optionally, the above data search method further includes: based on the target SDK, periodically probing multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and saving the first cluster list and the second cluster list to a local cache.

[0108] Optionally, reading the first target data from the data source includes: using a DTS component to read the data source and obtain predetermined data; using the DTS component to perform format conversion on the predetermined data and obtain the first target data.

[0109] This invention provides a device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs at least the following steps:

[0110] Step S201: Read the first target data from the data source, and write the first target data into the target ES cluster based on the target SDK and the predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least write and read functions.

[0111] Step S202: Based on the predetermined reading strategy and the target SDK, read the second target data from the target ES cluster and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0112] The devices mentioned in this article can be servers, PCs, tablets, mobile phones, etc.

[0113] Optionally, after writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: if the predetermined write strategy is simultaneous writing and the first target data is successfully written into both the first ES cluster and the second ES cluster, sending a response result to the UI interface, wherein the response result indicates that the first target data was successfully written; if the predetermined write strategy is fault-tolerant writing and the first target data is successfully written into either the first ES cluster or the second ES cluster, sending the response result to the UI interface; if the predetermined write strategy is specified writing and the first target data is successfully written into a specified ES cluster, sending the response result to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

[0114] Optionally, after writing the first target data into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: if the first target data is not successfully written into the target ES cluster based on the target SDK and the predetermined write strategy, storing the first target data into the target database; and writing the first target data from the target database into the target ES cluster again based on the target SDK and the predetermined write strategy.

[0115] Optionally, after writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: generating an alarm log if the first target data is not successfully written back into the target ES cluster based on the target SDK and the predetermined write strategy; and sending the alarm log to the operation and maintenance system so that operation and maintenance personnel can locate and analyze the problem based on the alarm log.

[0116] Optionally, reading the second target data from the target ES cluster based on the predetermined reading strategy and the target SDK includes: reading the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; or reading the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0117] Optionally, when the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; when the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined reading strategy is the default reading and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, reading the second target data from the backup ES cluster.

[0118] Optionally, the above data search method further includes: based on the target SDK, periodically probing multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and saving the first cluster list and the second cluster list to a local cache.

[0119] Optionally, reading the first target data from the data source includes: using a DTS component to read the data source and obtain predetermined data; using the DTS component to perform format conversion on the predetermined data and obtain the first target data.

[0120] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program having at least the following method steps:

[0121] Step S201: Read the first target data from the data source, and write the first target data into the target ES cluster based on the target SDK and the predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in the first data center, the second ES cluster is a cluster deployed in the second data center, and the target SDK is a service component with at least write and read functions.

[0122] Step S202: Based on the predetermined reading strategy and the target SDK, read the second target data from the target ES cluster and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading.

[0123] Optionally, after writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: if the predetermined write strategy is simultaneous writing and the first target data is successfully written into both the first ES cluster and the second ES cluster, sending a response result to the UI interface, wherein the response result indicates that the first target data was successfully written; if the predetermined write strategy is fault-tolerant writing and the first target data is successfully written into either the first ES cluster or the second ES cluster, sending the response result to the UI interface; if the predetermined write strategy is specified writing and the first target data is successfully written into a specified ES cluster, sending the response result to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

[0124] Optionally, after writing the first target data into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: if the first target data is not successfully written into the target ES cluster based on the target SDK and the predetermined write strategy, storing the first target data into the target database; and writing the first target data from the target database into the target ES cluster again based on the target SDK and the predetermined write strategy.

[0125] Optionally, after writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: generating an alarm log if the first target data is not successfully written back into the target ES cluster based on the target SDK and the predetermined write strategy; and sending the alarm log to the operation and maintenance system so that operation and maintenance personnel can locate and analyze the problem based on the alarm log.

[0126] Optionally, reading the second target data from the target ES cluster based on the predetermined reading strategy and the target SDK includes: reading the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading, wherein the specified ES cluster includes at least one of the following: the first ES cluster and the second ES cluster; or reading the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the default ES cluster includes at least one of the following: the first ES cluster and the second ES cluster.

[0127] Optionally, when the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; when the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined reading strategy is the default reading and the second target data cannot be read from the default ES cluster for a predetermined number of consecutive times, reading the second target data from the backup ES cluster.

[0128] Optionally, the above data search method further includes: based on the target SDK, periodically probing multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and saving the first cluster list and the second cluster list to a local cache. Optionally, reading the first target data from the data source includes: using the DTS component to read the data source to obtain predetermined data; and using the DTS component to perform format conversion on the predetermined data to obtain the first target data.

[0129] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0130] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0131] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0132] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0134] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0135] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0136] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0137] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0138] As can be seen from the above description, the embodiments of this application achieve the following technical effects:

[0139] 1) In the multi-cluster data search method of this application, based on the target SDK and a predetermined write strategy, first target data read from the data source is written to the target ES cluster; based on the predetermined read strategy and the target SDK, second target data is read from the target ES cluster and sent to the target UI interface to display the second target data. The data search method of this application develops a cross-cluster unified read / write service component, namely the target SDK, and sets relatively reliable data write and read strategies, namely predetermined write and read strategies. Based on the target SDK and the predetermined write strategy, relatively reliable writing of the first target data to the target ES cluster and relatively reliable reading of the second target data are achieved. This ensures that both data writing and reading are relatively reliable and that data consistency is high, ensuring that the complexity of cluster deployment is transparent to the application, thereby solving the problem of reliably reading and writing data between clusters in existing technologies.

[0140] 2) In the multi-cluster data search device of this application, the first execution unit is used to write the first target data read from the data source to the target ES cluster based on the target SDK and a predetermined write strategy; the second execution unit is used to read the second target data from the target ES cluster based on the predetermined read strategy and the target SDK, and send the read second target data to the target UI interface to display the second target data on the target UI interface. The data search device of this application develops a cross-cluster unified read / write service component, namely the target SDK, and sets relatively reliable data write and read strategies, namely the predetermined write and read strategies. Based on the target SDK and the predetermined write strategy, relatively reliable writing of the first target data to the target ES cluster and relatively reliable reading of the second target data are achieved. This ensures that both data writing and reading are relatively reliable and that data consistency is high, ensuring that the complexity of cluster deployment is transparent to the application, thereby solving the problem in the prior art of reliably reading and writing data between clusters.

[0141] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A multi-cluster data search method, characterized in that, include: Read first target data from the data source and write the first target data into the target ES cluster based on the target SDK and a predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in a first data center, and the second ES cluster is a cluster deployed in a second data center. The target SDK is a service component with at least write and read functions. Based on a predetermined reading strategy and the target SDK, second target data is read from the target ES cluster and sent to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading. When the predetermined write strategy is simultaneous write, the first target data is successfully written to both the first ES cluster and the second ES cluster simultaneously, and the write is successful; when the predetermined write strategy is fault-tolerant write, the write is successful if either the first ES cluster or the second ES cluster is successfully written; when the predetermined write strategy is specified write, the write is successful if the specified ES cluster is successfully written, and the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster. Based on the target SDK, periodic liveness detection is performed on multiple ES clusters in the first data center and the second data center to obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters. Save the first cluster list and the second cluster list to the local cache; When the predetermined read strategy is the specified read, the second target data is read from the specified ES cluster; when the predetermined read strategy is the default read, the second target data is read from the default ES cluster, wherein the specified ES cluster and the default ES cluster each include at least one of the following: the first ES cluster and the second ES cluster; When the predetermined reading strategy is the specified reading, after reading the second target data from the specified ES cluster, the data search method further includes: when the predetermined reading strategy is the specified reading and the second target data cannot be read from the specified ES cluster for a predetermined number of consecutive times, reading the second target data from a backup ES cluster, wherein the backup ES cluster is an ES cluster deployed in a backup data center; When the predetermined reading strategy is the default reading, after reading the second target data from the default ES cluster, the data search method further includes: when the predetermined reading strategy is the default reading and the second target data is not read from the default ES cluster for the predetermined number of consecutive times, reading the second target data from the backup ES cluster.

2. The data search method according to claim 1, characterized in that, After writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: If the predetermined write strategy is simultaneous write, and the first target data is successfully written to the first ES cluster and the second ES cluster, a response result is sent to the UI interface, and the response result is a result indicating that the first target data was successfully written. If the predetermined write strategy is the fault-tolerant write, and the first target data is successfully written to either the first ES cluster or the second ES cluster, the response result is sent to the UI interface. When the predetermined write strategy is the specified write and the first target data is successfully written to the specified ES cluster, the response result is sent to the UI interface, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster.

3. The data search method according to claim 1, characterized in that, After writing the first target data into the target ES cluster based on the target SDK and a predetermined write strategy, the data search method further includes: If the first target data is not successfully written to the target ES cluster based on the target SDK and the predetermined write strategy, the first target data will be stored in the target database. Based on the target SDK and the predetermined write policy, the first target data in the target database is written to the target ES cluster again.

4. The data search method according to claim 3, characterized in that, After writing the first target data from the target database back into the target ES cluster based on the target SDK and the predetermined write strategy, the data search method further includes: If, based on the target SDK and the predetermined write policy, the first target data is not successfully written to the target ES cluster again, an alarm log is generated. The alarm logs are sent to the operation and maintenance system so that operation and maintenance personnel can locate and analyze problems based on the alarm logs.

5. The data search method according to any one of claims 1 to 4, characterized in that, Read the first target data from the data source, including: The DTS component is used to read the data source to obtain the predetermined data; The DTS component is used to convert the format of the predetermined data to obtain the first target data.

6. A multi-cluster data search device, characterized in that, include: The first execution unit is configured to read first target data from a data source and write the first target data into a target ES cluster based on a target SDK and a predetermined write strategy. The target ES cluster includes at least one of the following: a first ES cluster and a second ES cluster. The predetermined write strategy includes at least one of the following: simultaneous write, fault-tolerant write, and specified write. The first ES cluster is a cluster deployed in a first data center, and the second ES cluster is a cluster deployed in a second data center. The target SDK is a service component with at least write and read functions. The second execution unit is configured to read second target data from the target ES cluster based on a predetermined reading strategy and the target SDK, and send the second target data to the target UI interface. The predetermined reading strategy includes at least one of the following: specified reading and default reading. The multi-cluster data search device is further configured to: when the predetermined write strategy is simultaneous write, if the first target data is successfully written to both the first ES cluster and the second ES cluster, then the write is successful; when the predetermined write strategy is fault-tolerant write, if either the first ES cluster or the second ES cluster is successfully written, then the write is successful; when the predetermined write strategy is specified write, if the specified ES cluster is successfully written, then the write is successful, wherein the specified ES cluster is at least one of the following: the first ES cluster and the second ES cluster. The liveness detection unit is used to periodically detect the liveness of multiple ES clusters in the first data center and the second data center based on the target SDK, and obtain a first cluster list and a second cluster list, wherein the first cluster list is a list of available ES clusters and the second cluster list is a list of unavailable ES clusters; and save the first cluster list and the second cluster list to a local cache. The multi-cluster data search device is further configured to read the second target data from a specified ES cluster when the predetermined reading strategy is the specified reading; and to read the second target data from a default ES cluster when the predetermined reading strategy is the default reading, wherein the specified ES cluster and the default ES cluster each include at least one of the following: the first ES cluster and the second ES cluster; The fourth execution unit is configured to read the second target data from the designated ES cluster when the predetermined reading strategy is the designated reading, and then read the second target data from the backup ES cluster when the predetermined reading strategy is the designated reading and the second target data is not read from the designated ES cluster for a predetermined number of consecutive times. The backup ES cluster is an ES cluster deployed in a backup data center. The fifth execution unit is configured to read the second target data from the backup ES cluster after reading the second target data from the default ES cluster when the predetermined reading strategy is the default reading and the second target data has not been read from the default ES cluster for the predetermined number of consecutive times.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the multi-cluster data search method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • File processing system and method

    CN109960687A

  • Unified management method for multi-data center dual-stack container cloud platform

    CN113778613A