Data synchronization method and device, equipment and medium

By updating the cluster identifier in the distributed database and extracting data from the replication source, the problem of cross-system and cross-tenant backup is solved, enabling flexible data synchronization and backup.

CN120929535APending Publication Date: 2025-11-11JINZHUAN INFORMATION TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511095151.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Traditional distributed database disaster recovery solutions have strict system and tenant limitations, making it impossible to achieve cross-system and cross-tenant disaster recovery synchronization.

Method used

By obtaining the identifier of the source cluster and updating it to the identifier of the target cluster, data and identifier are obtained separately. Backup data content is extracted from the replication source of the source cluster and synchronized to the target cluster, supporting cross-system and cross-tenant backup.

Benefits of technology

It enables cross-system and cross-tenant data backup, improving the flexibility and accuracy of backup and breaking through the system and tenant limitations of traditional solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120929535A_ABST
    Figure CN120929535A_ABST
Patent Text Reader

Abstract

The invention discloses a data synchronization method and device, equipment and a medium. The method comprises the following steps: acquiring a first identifier of a source cluster and a second identifier of a to-be-synchronized target cluster; acquiring cluster metadata of the source cluster, and updating a first identifier in the cluster metadata into the second identifier; obtaining a first copy source of the source cluster, and extracting backup data content from the first copy source; and synchronizing the content of the source cluster to the target cluster according to the cluster metadata and the backup data content. According to the embodiment of the invention, the flexibility of the backup system can be improved, and cross-system disaster recovery synchronization is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of database technology, and in particular to a data synchronization method, apparatus, device, and medium. Background Technology

[0002] With the acceleration of digital transformation, distributed databases face core challenges such as heterogeneous environment collaboration, cross-tenant isolation, and multi-level disaster recovery.

[0003] Traditional distributed database disaster recovery solutions have strict system and tenant limitations, only support one-to-one replication mode, and require the source and destination tenant numbers to be consistent, making it impossible to achieve cross-system and cross-tenant disaster recovery synchronization. Summary of the Invention

[0004] This invention provides a data synchronization method, apparatus, device, and medium that can increase the flexibility of backup systems and enable cross-system disaster recovery synchronization.

[0005] According to one aspect of the present invention, an embodiment of the present invention provides a data synchronization method, the method comprising:

[0006] Obtain the first identifier of the source cluster and the second identifier of the target cluster to be synchronized;

[0007] Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier;

[0008] Obtain the first replication source of the source cluster, and extract backup data content from the first replication source;

[0009] Based on the cluster metadata and the backup data, the content of the source cluster is synchronized to the target cluster.

[0010] According to another aspect of the present invention, embodiments of the present invention also provide a data synchronization device, the device comprising:

[0011] The cluster identifier acquisition module is used to acquire the first identifier of the source cluster and the second identifier of the target cluster to be synchronized.

[0012] The cluster identifier update module is used to obtain the cluster metadata of the source cluster and update the first identifier in the cluster metadata to the second identifier;

[0013] The replication source data extraction module is used to obtain the first replication source of the source cluster and extract backup data content from the first replication source;

[0014] The data synchronization module is used to synchronize the content of the source cluster to the target cluster based on the cluster metadata and the backup data content.

[0015] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0016] At least one processor; and

[0017] A memory that is communicatively connected to at least one processor; wherein,

[0018] The memory stores a computer program that can be executed by at least one processor, such that the at least one processor is able to perform the data synchronization method of any embodiment of the present invention.

[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for causing a processor to execute and implement the data synchronization method of any embodiment of the present invention.

[0020] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the data synchronization method described in any embodiment of the present invention.

[0021] The technical solution of this invention updates the source cluster identifier in the metadata of the source cluster to the target cluster identifier, and extracts backup content data from the replication source of the source cluster. It can separate the data and the identifier for acquisition, and can obtain the data content to be backed up from any replication source in the source cluster. It can separate the identifier and the data content for backup, thereby realizing cross-system backup. It solves the problem that cross-system and cross-tenant backup is not possible in the prior art, and can flexibly select the system for backup, improving the flexibility of backup.

[0022] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart of a data synchronization method provided according to an embodiment of the present invention;

[0025] Figure 2 This is a flowchart of a data synchronization method according to an embodiment of the present invention;

[0026] Figure 3 This is a flowchart of a method for synchronizing updated metadata according to an embodiment of the present invention;

[0027] Figure 4 This is a flowchart of a data synchronization method provided according to an embodiment of the present invention;

[0028] Figure 5 This is a flowchart of a data priority configuration method provided by an embodiment of the present invention;

[0029] Figure 6 This is a structural diagram of a data synchronization device provided according to an embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] The acquisition, storage, and application of driving trajectory points and other related technologies in the technical solutions of this invention comply with relevant laws and regulations and do not violate public order and good morals.

[0034] Figure 1This is a flowchart illustrating a data synchronization method provided in an embodiment of the present invention. This embodiment is applicable to situations where data backup is performed on a cluster within a database. The method can be executed by a data synchronization device, which can be implemented in hardware and / or software.

[0035] See Figure 1 The data synchronization methods shown include:

[0036] S101. Obtain the first identifier of the source cluster and the second identifier of the target cluster to be synchronized.

[0037] In this context, the source cluster refers to the original cluster that needs to be backed up. The target cluster refers to the cluster obtained after backing up the data. The source database includes at least one cluster, and among all the clusters included in the source database, the cluster that needs to be backed up for disaster recovery can be the source cluster. Identifiers are used to distinguish different clusters. The first identifier can be the cluster identifier of the source cluster. The second identifier can be the cluster identifier of the target cluster. The first identifier can be obtained by querying the cluster identifier of the source cluster. The second identifier can be randomly generated, or a cluster identifier different from all the cluster identifiers of the backed-up clusters of the source cluster can be generated as the second identifier. In some embodiments, the cluster identifiers of the backup clusters of the same source cluster can be cumulatively named according to a preset rule, thus obtaining a new cluster identifier for the backup cluster based on the cluster identifiers of the historical backup clusters according to the preset rule. For example, the existing cluster identifiers of the backup clusters of the source cluster are 1, 2, and 3. The second identifier of the target cluster can be 4.

[0038] In an optional embodiment, the source cluster belongs to the source database, which includes at least one source cluster, and different source clusters are used to store data of different business types.

[0039] Different clusters within the source database contain different types of data. For example, the source database can be a database for a device system. The source database may include clusters A, B, and C. Cluster A contains device attributes, cluster B contains device operational data, and cluster C contains device permission information. Clusters B and C are source clusters, while cluster A is not. A source database can include at least one source cluster. The business types of the data stored in each source cluster differ.

[0040] In some embodiments, the cluster containing the key data in the source database cluster can be identified as the source cluster.

[0041] Dividing the source database into at least one cluster and performing independent backup processing for each source cluster allows for processing only a portion of the data in the source database, precise determination of the backup data range, and flexible adjustment of the backup data range.

[0042] It is evident that by configuring source clusters in the source database and performing data synchronization backups for each source cluster, the flexibility and accuracy of data backups can be improved.

[0043] S102. Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier.

[0044] Cluster metadata can refer to data describing the data within a cluster. The source cluster's cluster metadata can also refer to data describing the data within the source cluster. Replacing all the content of the first identifier in the source cluster's cluster metadata with the second identifier results in updated cluster metadata equivalent to the target cluster's cluster metadata.

[0045] S103. Obtain the first replication source of the source cluster, and extract backup data content from the first replication source.

[0046] Here, the first replication source can refer to the cluster that needs to acquire the data content of the source cluster. The first replication source is the cluster that has historically backed up the source cluster. The backup data content can refer to the data content in the first replication source. The backup data content is identical to the data content in the source cluster.

[0047] S104. Synchronize the content of the source cluster to the target cluster based on the cluster metadata and the backup data content.

[0048] Specifically, cluster metadata is used as the metadata for the backup data content. The backup data content is stored in the target cluster's storage space, and the cluster metadata is used as the metadata for the data content in the target cluster. This achieves synchronization between the source cluster and the target cluster. The cluster metadata can refer to the cluster metadata after updating the second identifier.

[0049] In one example, a financial system deploys a distributed database with two clusters: Cluster A (ID=1) and Cluster B (ID=2). Cross-system disaster recovery from Cluster A to Cluster B needs to be implemented. Cross-cluster ID synchronization: When Cluster A (ID=1) synchronizes data to Cluster B (ID=2), the cluster ID in the synchronized data is changed from 1 to 2 using the metadata conversion mechanism of the MDS (Metadata Section). This MDS is responsible for storing and managing disaster recovery synchronization metadata, including cluster mapping relationships, priority configurations, distributed database identifiers, etc., to achieve metadata conversion and synchronization.

[0050] When synchronizing metadata, the cluster ID of the source distributed database is converted to the cluster ID of the destination distributed database. This includes the conversion of metadata such as data dictionary, indexes, and account information to ensure the consistency of metadata in the destination distributed database.

[0051] This allows for the decoupling of cluster IDs between the source and destination distributed databases, enabling data synchronization between different cluster IDs through a metadata conversion mechanism. It also allows a source distributed database to synchronize data with multiple destination distributed databases using different cluster IDs. Furthermore, a tenant mapping mechanism is established to overcome tenant number consistency limitations, enabling cross-tenant data synchronization and routing, and supporting cross-system disaster recovery for heterogeneous databases (such as distributed databases).

[0052] The technical solution of this invention updates the source cluster identifier in the metadata of the source cluster to the target cluster identifier, and extracts backup content data from the replication source of the source cluster. It can separate the data and the identifier for acquisition, and can obtain the data content to be backed up from any replication source in the source cluster. It can separate the identifier and the data content for backup, thereby realizing cross-system backup. It solves the problem that cross-system and cross-tenant backup is not possible in the prior art, and can flexibly select the system for backup, improving the flexibility of backup.

[0053] In an optional embodiment, the data synchronization method further includes: when a change in data is detected in the source cluster, obtaining a second replication source of the source cluster; extracting updated data content from the second replication source; obtaining cluster metadata of the source cluster, updating a first identifier in the cluster metadata to a second identifier; and sending the cluster metadata and the updated data content to the target cluster so that the target cluster can synchronize the changes.

[0054] In this context, "data change" can refer to changes in the data content within the source cluster. For example, changes may occur in the storage location, structure, or attribute values ​​of the data. These changes can include additions, deletions, and modifications. The second replication source can be a backup source that backs up the target cluster for updated data content. The second replication source is independent of the first replication source and can be the same as or different from the first. For example, if the first replication source fails, the second replication source may be different from the first. Alternatively, the second replication source may be the same as the first. "Updated data content" can refer to changes in the data content within the meta-cluster.

[0055] In some embodiments, the content of the first identifier in the cluster metadata of the source cluster is completely replaced with the second identifier, and the updated cluster metadata is equivalent to the cluster metadata of the target cluster. The cluster metadata is used as the metadata for the updated data content. The updated data content is stored in the storage space of the target cluster, and the cluster metadata is used as the metadata for the data content in the target cluster, thus synchronizing the source cluster to the target cluster. Here, the cluster metadata can refer to the cluster metadata after updating the second identifier.

[0056] In one example, such as Figure 2 As shown, the data synchronization method also includes

[0057] S201, Data in the source cluster has changed.

[0058] S202, Capture data changes.

[0059] Data changes in the source cluster can be captured through the Data Replication and Synchronization Processor (DRSP) engine.

[0060] S203. Obtain the first identifier of the source cluster and the second identifier of the target cluster.

[0061] The first identifier of the source cluster and the second identifier of the target cluster can be obtained through MDS. The source cluster, target cluster, and replication source are clusters of GoldenDB (Database), a financial-grade transactional distributed database.

[0062] S204, Metadata Transformation: The first identifier in the cluster metadata is updated to the second identifier.

[0063] S205, Related information about the conversion cluster identifier.

[0064] Data processing is performed through GTM (Global Transaction Manager), handling cluster identity-related information in addition to the second identity update. This information includes, for example, data fields, indexes, and account information.

[0065] S206. Select the first replication source based on priority.

[0066] The first replication source is selected from the alternative replication sources in the source cluster according to priority.

[0067] S207. Synchronize the contents of the source cluster to the target cluster.

[0068] S208, Data synchronized by the target cluster application.

[0069] Users can query data from the source cluster by accessing the target cluster. It should be noted that the target cluster cannot add, delete, or modify synchronized data.

[0070] As can be seen, by implementing a backup of the source cluster in the target cluster and then synchronizing and updating the data that changes in the source cluster in real time, real-time synchronization can be achieved, thereby improving the accuracy of backup and disaster recovery.

[0071] In an optional embodiment, the synchronization of cluster metadata changes in the source cluster can be achieved by directly obtaining the changed cluster metadata from the source cluster and updating it in the target cluster. The data synchronization method further includes: when a change in cluster metadata is detected in the source cluster, obtaining the updated metadata of the source cluster; updating the first identifier in the updated metadata to the second identifier; and sending the updated metadata to the target cluster, so that the target cluster performs the synchronization change.

[0072] In one example, such as Figure 3 As shown, the methods for synchronizing cluster metadata include:

[0073] S301. The source cluster metadata has changed.

[0074] S302, Capture metadata changes.

[0075] S303, Obtain the first identifier of the source cluster and the second identifier of the target cluster.

[0076] S304, Query cluster mapping relationship.

[0077] S305, Metadata Conversion: Update the first identifier in the changed updated metadata to the second identifier.

[0078] S306. Synchronize the updated metadata to the target cluster.

[0079] S307, Update metadata for target cluster applications.

[0080] Figure 4 This is a flowchart illustrating a data synchronization method provided in an embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention further specifies obtaining the first replication source of the source cluster as follows: obtaining at least one candidate replication source of the source cluster; and selecting a first replication source from among the candidate replication sources according to their priority.

[0081] It should be noted that for parts not described in detail in the embodiments of the present invention, please refer to the descriptions in other embodiments.

[0082] See Figure 4 The data synchronization methods shown include:

[0083] S401. Obtain the first identifier of the source cluster and the second identifier of the target cluster to be synchronized.

[0084] S402. Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier.

[0085] S403. Obtain at least one alternative replication source from the source cluster.

[0086] In this context, a backup replication source can refer to a cluster that stores the data content of the source cluster. A source cluster can have at least one backup replication source. A backup replication source can be the source cluster itself. If the source cluster has not been backed up, it is used as a backup replication source. If the source cluster has been backed up, both the backed-up cluster and the source cluster itself are used as backup replication sources.

[0087] S404. Select a first replication source from the candidate replication sources according to their priority.

[0088] Priority is used to filter replication sources. The priority of a candidate replication source can be determined based on its relative position to the target cluster, its performance, operational status, and load. The candidate replication source with the highest priority can be selected as the first replication source. If there are multiple candidate replication sources with the highest priority, one can be randomly selected as the first replication source.

[0089] S405. Extract backup data content from the first copy source.

[0090] S406. Synchronize the content of the source cluster to the target cluster based on the cluster metadata and the backup data content.

[0091] This invention optimizes disaster recovery resource allocation and improves data synchronization efficiency by selecting a first replication source from at least one candidate replication source in the source cluster.

[0092] In an optional embodiment, selecting a first replication source based on the priority of each of the candidate replication sources includes: selecting priority levels one by one according to the order of priority levels; for each priority level, filtering out the candidate replication source with the highest priority among the candidate replication sources at the same priority level; and determining the candidate replication source with the highest priority at the last level as the first replication source.

[0093] The priority can include at least one priority level. The highest-priority candidate replication source in each priority level is selected, and the highest-priority candidate replication source among all the final priority levels is used as the first replication source. The priority levels are used to further subdivide priorities, achieving more granular priority filtering.

[0094] The order of priority levels varies. The order can be configured to prioritize or defer different priority levels. For example, you can prioritize based on the order of priority levels, filtering out the highest priority candidate replication sources one by one, and then, within the next priority level, continue filtering from the highest priority candidate replication sources selected in the previous priority level to find the highest priority candidate replication source.

[0095] In some embodiments, the priority levels are arranged in chronological order, including level A, level B, and level C. From the candidate replication sources, at least one candidate replication source with the highest priority at level A is selected, resulting in *a* candidate replication sources. From these *a* candidate replication sources, at least one candidate replication source with the highest priority at level B is selected, resulting in one candidate replication source. Since there is one remaining candidate replication source, level C does not need further priority selection. This remaining candidate replication source is directly used as the first replication source.

[0096] As can be seen, by configuring at least one priority level and selecting the highest priority candidate replication source for each priority level to obtain the first replication source, fine-grained filtering of replication sources can be achieved. Furthermore, replication sources can be selected more accurately to meet different needs and flexibly adapt to complex network environments.

[0097] In an optional embodiment, the priority hierarchy includes: data center level, team level, and address level.

[0098] In this context, "data center" can refer to the space where database servers are deployed. "Data center level" can refer to a priority level based on physical distance. "Team level" can refer to a priority level based on database nodes. "Address level" can refer to a priority level based on network addresses. Here, "network address" can refer to an Internet Protocol address (IP address).

[0099] In some embodiments, priority fields (such as priority_type) can be added to the metadata, supporting priority fields at the data center level, team level, and address level.

[0100] For data center tiers, alternative replication sources belonging to the same data center as the target cluster have higher priority, and alternative replication sources in data centers with high availability have higher priority.

[0101] At the team level, core alternative replication sources have higher priority. Core alternative replication sources can refer to alternative replication sources for core or key business operations.

[0102] At the address level, candidate replication sources belonging to the critical address range have higher priority, and the critical address range can be specified by the user. By specifying the IP range, flexible adaptation to complex network environments can be achieved.

[0103] In one example, such as Figure 5 As shown, methods for priority configuration and dynamic selection may include:

[0104] S501. Initialize DRSP configuration.

[0105] DRSP can configure the priority of each alternative replication source and generate priority configuration information for each alternative replication source.

[0106] S502, Database node loading priority configuration information.

[0107] The database node loads the priority configuration information for each alternative replication source.

[0108] S503, Data synchronization is triggered, and the priority level of the alternative replication source is obtained.

[0109] A data node hierarchy can also be added before S504. For the data node hierarchy, the replication source of the primary data node is selected first based on the data node identifier.

[0110] S504. For data center level, select the highest priority replication source based on the data center identifier.

[0111] S505. For team-level applications, select the highest priority replication source based on the team identifier.

[0112] S506. For the address level, filter the highest priority replication source according to the address identifier to obtain the first replication source.

[0113] S507. Synchronize the backup data from the first replication source to the target cluster.

[0114] As can be seen, by configuring priority levels including data center level, team level and address level, multi-dimensional priority configuration can be achieved, the replication source can be dynamically adjusted, the disaster recovery resource allocation can be optimized, the data synchronization efficiency can be improved, and the intelligence and adaptability of the disaster recovery system can be enhanced.

[0115] In an optional embodiment, after selecting a first replication source based on the priority of each of the alternative replication sources, the method further includes: when the selected first replication source fails, updating the selected first replication source based on the priority of the remaining alternative replication sources.

[0116] Among these steps, the faulty replication source is excluded, and from the remaining alternative replication sources, the highest priority alternative replication source is reselected as the first replication source.

[0117] In some embodiments, for the data center level, cluster B nodes in the same data center as cluster A are preferentially selected as the replication source, and when a node in the same data center fails, the replication source is automatically switched to a node in another data center.

[0118] It is evident that by updating the selection of the first replication source based on the priority of the remaining alternative replication sources when the first replication source fails, and obtaining the data backup content from the normally functioning alternative replication sources, the accuracy and completeness of the data are improved.

[0119] In some embodiments, the system implementing the data synchronization method may include: a metadata node, a disaster recovery architecture engine, a system core database, and a network topology. The metadata node is responsible for storing and managing disaster recovery synchronization metadata, including cluster mapping relationships, priority configurations, and distributed database identifiers, enabling metadata conversion and synchronization. The disaster recovery architecture engine is deployed in each distributed database cluster to handle data synchronization tasks, dynamically selecting replication sources based on priority and performing cross-cluster ID data conversion. The system core database (relational database) stores disaster recovery synchronization configuration information through a metadata disaster recovery architecture information table, supporting triplet primary keys to achieve multi-synchronization source management. The network topology can support cascading or star topologies of source distributed databases and multiple destination distributed databases. Each cluster shares network information through big data nodes, and the metadata node coordinates data synchronization based on priority and cluster mapping relationships.

[0120] Scripted deployment (such as disaster recovery architecture scripts) can automate the writing of priority configurations and distributed database identifiers to the system's core database. This supports dynamic modification of synchronization priorities via operational commands, improving operational efficiency and reducing costs. Automated script deployment and dynamic configuration interfaces reduce manual intervention, enhancing the maintainability and reliability of the disaster recovery system. Disaster recovery architecture information is loaded into the metadata process, synchronization status is monitored in real time, and replication sources are dynamically adjusted based on priority, optimizing disaster recovery resource allocation.

[0121] For cascading network topologies: cluster A synchronizes to cluster B, then further synchronizes to cluster C, and so on, forming a cascading scenario. Cross-cluster data routing is achieved through a cluster mapping table. Simultaneously, one-to-many data synchronization is also supported for star network topologies: cluster A acts as the core of a star structure, simultaneously synchronizing data to branches such as cluster B and cluster C. Cross-system disaster recovery synchronization can be achieved, realizing cluster ID independence in distributed database disaster recovery synchronization, breaking the hard constraint that "cluster IDs must be the same," supporting data synchronization between heterogeneous clusters, achieving cluster decoupling, and solving the cluster identifier consistency problem during cross-cluster synchronization through a cluster mapping table and metadata transformation engine. This enables cross-system disaster recovery for heterogeneous databases, improving the flexibility and reliability of the disaster recovery system, reducing operational costs, and meeting the complex disaster recovery needs of large-scale distributed systems.

[0122] Figure 6 This is a schematic diagram of a data synchronization device provided in an embodiment of the present invention. The embodiment of the present invention is applicable to situations where data synchronization is performed in response to a user's request. The device can execute a data synchronization method and can be implemented in hardware and / or software.

[0123] See Figure 6 The data synchronization device shown includes:

[0124] Cluster identifier acquisition module 601 is used to acquire the first identifier of the source cluster and the second identifier of the target cluster to be synchronized;

[0125] The cluster identifier update module 602 is used to obtain the cluster metadata of the source cluster and update the first identifier in the cluster metadata to the second identifier;

[0126] The replication source data extraction module 603 is used to obtain the first replication source of the source cluster and extract backup data content from the first replication source;

[0127] The data synchronization module 604 is used to synchronize the content of the source cluster to the target cluster based on the cluster metadata and the backup data content.

[0128] The technical solution of this invention updates the source cluster identifier in the metadata of the source cluster to the target cluster identifier, and extracts backup content data from the replication source of the source cluster. It can separate the data and the identifier for acquisition, and can obtain the data content to be backed up from any replication source in the source cluster. It can separate the identifier and the data content for backup, thereby realizing cross-system backup. It solves the problem that cross-system and cross-tenant backup is not possible in the prior art, and can flexibly select the system for backup, improving the flexibility of backup.

[0129] Optionally, the source data extraction module can be copied, specifically for:

[0130] Obtain at least one alternative replication source from the source cluster;

[0131] Select a first replication source from the candidate replication sources according to their priority.

[0132] Optionally, the source data extraction module can be copied, specifically for:

[0133] Select priority levels one by one according to their order of priority.

[0134] For each priority level, among the candidate replication sources at the same priority level, the highest priority candidate replication source is selected;

[0135] The candidate replication source with the highest priority at the last level is determined as the first replication source.

[0136] Optionally, the priority levels include: data center level, team level, and address level.

[0137] Optionally, the data synchronization device also includes:

[0138] The replication source update module is used to update the selected first replication source according to the priority of each of the candidate replication sources after selecting a first replication source based on the priority of each candidate replication source. When the selected first replication source fails, the module updates the selected first replication source according to the priority of the remaining candidate replication sources.

[0139] Optionally, the data synchronization device also includes:

[0140] The data update and backup module is used for:

[0141] When a change in data is detected in the source cluster, a second replication source of the source cluster is obtained;

[0142] Extract updated data from the second replication source;

[0143] Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier;

[0144] The cluster metadata and the updated data content are sent to the target cluster so that the target cluster can synchronize the changes.

[0145] Optionally, the source cluster belongs to the source database, and the source database includes at least one source cluster, with different source clusters used to store data of different business types.

[0146] The data synchronization device provided in the embodiments of the present invention can execute the data synchronization method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the data synchronization method.

[0147] Figure 7 A schematic diagram of the structure of an electronic device 700 that can be used to implement an embodiment of the present invention is shown.

[0148] like Figure 7 As shown, the electronic device 700 includes at least one processor 701 and a memory, such as a read-only memory (ROM) 702 and a random access memory (RAM) 703, communicatively connected to the at least one processor 701. The memory stores computer programs executable by the at least one processor. The processor 701 can perform various appropriate actions and processes based on the computer program stored in the ROM 702 or loaded into the RAM 703 from storage unit 708. The RAM 703 can also store various programs and data required for the operation of the electronic device 700. The processor 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0149] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, such as keyboard, mouse, etc.; output unit 707, such as various types of displays, speakers, etc.; storage unit 708, such as disk, optical disk, etc.; and communication unit 709, such as network card, modem, wireless transceiver, etc. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0150] Processor 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 701 performs the various methods and processes described above, such as data synchronization methods.

[0151] In some embodiments, the data synchronization method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 700 via ROM 702 and / or communication unit 709. When the computer program is loaded into RAM 703 and executed by processor 701, one or more steps of the data synchronization method described above may be performed. Alternatively, in other embodiments, processor 701 may be configured to perform the data synchronization method by any other suitable means (e.g., by means of firmware).

[0152] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0153] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0154] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0155] To provide interaction with a user, the systems and techniques described herein can be implemented on an operational detection device. This electronic device includes: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0156] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0157] A computing system can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system. It addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability.

[0158] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0159] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data synchronization method, characterized in that, The method includes: Obtain the first identifier of the source cluster and the second identifier of the target cluster to be synchronized; Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier; Obtain the first replication source of the source cluster, and extract backup data content from the first replication source; Based on the cluster metadata and the backup data, the content of the source cluster is synchronized to the target cluster.

2. The method according to claim 1, characterized in that, The step of obtaining the first replication source of the source cluster includes: Obtain at least one alternative replication source from the source cluster; Select a first replication source from the candidate replication sources according to their priority.

3. The method according to claim 2, characterized in that, The step of selecting a first replication source from the candidate replication sources according to their priorities includes: Select priority levels one by one according to their order of priority. For each priority level, among the candidate replication sources at the same priority level, the highest priority candidate replication source is selected; The candidate replication source with the highest priority at the last level is determined as the first replication source.

4. The method according to claim 3, characterized in that, The priority levels include: data center level, team level, and address level.

5. The method according to claim 2, characterized in that, After selecting a first replication source from the candidate replication sources according to their priorities, the process further includes: When the selected primary replication source fails, the selected primary replication source is updated according to the priority of the remaining alternative replication sources.

6. The method according to claim 1, characterized in that, Also includes: When a change in data is detected in the source cluster, a second replication source of the source cluster is obtained; Extract the updated data from the second replication source; Obtain the cluster metadata of the source cluster, and update the first identifier in the cluster metadata to the second identifier; The cluster metadata and the updated data content are sent to the target cluster so that the target cluster can synchronize the changes.

7. The method according to claim 1, characterized in that, The source cluster belongs to the source database, and the source database includes at least one source cluster. Different source clusters are used to store data of different business types.

8. A data synchronization device, characterized in that, The device includes: The cluster identifier acquisition module is used to acquire the first identifier of the source cluster and the second identifier of the target cluster to be synchronized. The cluster identifier update module is used to obtain the cluster metadata of the source cluster and update the first identifier in the cluster metadata to the second identifier; The replication source data extraction module is used to obtain the first replication source of the source cluster and extract backup data content from the first replication source; The data synchronization module is used to synchronize the content of the source cluster to the target cluster based on the cluster metadata and the backup data content.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the data synchronization method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the data synchronization method of any one of claims 1-7.

Citation Information

Patent Citations

  • Copy exception recovery method and device based on storage cluster, and computer equipment

    CN108647118A

  • File reading method and system, metadata server and user equipment

    CN110022338A

  • Data information synchronization method and device, electronic equipment and medium

    CN111581285A

  • Data backup method and device, equipment and storage medium

    CN113064766A

  • Data migration method and device, server and storage medium

    CN114024956A