Data synchronization method, device, electronic device and storage medium

By performing data unit granularity management and point-to-point data transmission between the master and slave database nodes in the data center, the problem of inefficient remote data synchronization is solved, efficient and accurate data synchronization is achieved, and resource waste and data conflicts are reduced.

CN119336841BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411366508.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2025-09-23
Estimated Expiration
2044-09-27

AI Technical Summary

Technical Problem

When synchronizing data between multiple data centers, existing technologies have difficulty in efficiently handling the needs of remote data synchronization, especially when the distance between data centers is long, resulting in low synchronization efficiency and waste of resources.

Method used

By managing data units at a granular level between the master and slave database nodes in the data center, we screen out globally accessible data units and utilize peer-to-peer data transfer tools for remote data synchronization, reducing the amount of synchronized data and improving synchronization efficiency. Furthermore, an encryption gateway handles data conversion between different database types, ensuring data security and synchronization accuracy.

Benefits of technology

It achieves efficient data synchronization between remote data centers, reduces the amount of synchronized data, improves synchronization efficiency, avoids resource waste and data key value conflicts, and ensures the service quality of online business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336841B_ABST
    Figure CN119336841B_ABST
Patent Text Reader

Abstract

The present application discloses a data synchronization method, device, electronic device, and storage medium, and relates to the computer field, and in particular to artificial intelligence fields such as cloud computing and big data. The specific implementation scheme is as follows: in response to a data update of a first data unit in a first master database node of a first data center, synchronizing the first updated data and first attribute information of the first updated data in the first data unit to a first slave database node of the first data center; determining the access type corresponding to the first data unit; in response to the access type corresponding to the first data unit being global access, synchronizing the first updated data and first attribute information to a second master database node of a second data center through the first slave database node; wherein the distance between the first data center and the second data center is greater than a first threshold; synchronizing the first updated data and first attribute information to a second slave database node of the second data center through the second master database node.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, in particular to artificial intelligence fields such as cloud computing and big data, and specifically to a data synchronization method, device, electronic device and storage medium. Background Art

[0002] With the development of database cluster technology, data centers may be set up in multiple cities. Each data center can be responsible for business services in the corresponding geographical area, and there is a need for data synchronization between different data centers. Summary of the Invention

[0003] The present application provides a data synchronization method, device, electronic device and storage medium.

[0004] According to one aspect of the present application, a data synchronization method is provided, comprising:

[0005] In response to a data update occurring in a first data unit in a first master database node in a first data center, synchronizing first updated data in the first data unit and first attribute information of the first updated data to a first slave database node in the first data center;

[0006] determining an access type corresponding to the first data unit;

[0007] In response to the access type corresponding to the first data unit being global access, synchronizing the first updated data and the first attribute information to a second master database node in the second data center through the first slave database node; wherein a distance between the first data center and the second data center is greater than a first threshold;

[0008] The first updated data and the first attribute information are synchronized to the second slave database node of the second data center through the second master database node.

[0009] According to another aspect of the present application, a data synchronization device is provided, comprising:

[0010] a first synchronization module, configured to synchronize, in response to a data update occurring in a first data unit in a first master database node in a first data center, first updated data and first attribute information of the first updated data in the first data unit to a first slave database node in the first data center;

[0011] A first determining module, configured to determine an access type corresponding to the first data unit;

[0012] a second synchronization module, configured to synchronize the first updated data and the first attribute information to a second master database node in the second data center via the first slave database node in response to the access type corresponding to the first data unit being global access; wherein the distance between the first data center and the second data center is greater than a first threshold;

[0013] A third synchronization module is used to synchronize the first updated data and the first attribute information to the second slave database node of the second data center through the second master database node.

[0014] According to another aspect of the present application, an electronic device is provided, including:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method described in the above embodiment.

[0018] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method according to the above embodiment.

[0019] According to another aspect of the present application, a computer program product is provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.

[0020] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.

[0022] Figure 1 A flowchart of a data synchronization method provided in one embodiment of the present application;

[0023] Figure 2 A flowchart of a data synchronization method provided in another embodiment of the present application;

[0024] Figure 3 A flowchart of a data synchronization method provided in another embodiment of the present application;

[0025] Figure 4 A flowchart of a data synchronization method provided in another embodiment of the present application;

[0026] Figure 5 A schematic diagram of a data synchronization process provided in an embodiment of the present application;

[0027] Figure 6 A schematic diagram of the structure of a data synchronization device provided in one embodiment of the present application;

[0028] Figure 7 It is a block diagram of an electronic device used to implement the data synchronization method of an embodiment of the present application. DETAILED DESCRIPTION

[0029] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] The following describes the data synchronization method, device, electronic device, and storage medium according to the embodiments of the present application with reference to the accompanying drawings.

[0031] Figure 1 A flowchart of a data synchronization method provided in accordance with an embodiment of the present application.

[0032] The data synchronization method of the embodiment of the present application can be executed by the data synchronization device of the embodiment of the present application, and the device can be configured in an electronic device.

[0033] Among them, the electronic device can be any device with computing capabilities, such as a personal computer, mobile terminal, server, etc. The mobile terminal can be, for example, a vehicle-mounted device, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.

[0034] Exemplarily, the data synchronization method of the embodiment of the present application can be executed by a data synchronization system, which can include a first data center, a second data center, etc., or can be executed by a data management system for managing data center synchronization, etc.

[0035] like Figure 1 As shown, the data synchronization method includes:

[0036] Step 101: In response to a data update of a first data unit in a first master database node of a first data center, first updated data and first attribute information of the first updated data in the first data unit are synchronized to a first slave database node of the first data center.

[0037] In this application, multiple data centers can be deployed for one or more applications, and these data centers can provide business services for these applications. The first data center can be any data center of the multiple deployed data centers.

[0038] The first data center may include a first master database node and at least one first slave database node. The first slave database node may be responsible for processing all write operations, such as insert, update, and delete, and may also manage transaction logs to record all data changes and coordinate and manage data synchronization of the first slave database node.

[0039] In addition, the first database center can support read-write separation. For example, the application can send write operations to the first master database node and read operations to the first slave database node as needed to optimize database performance.

[0040] In this application, the database on the first master database node may be pre-unitized to obtain at least one data unit. A data unit may be a data table or a row of data in a database table, and this application does not limit the granularity of the data unit.

[0041] For example, you can first determine the databases that each application can access, and then divide the data of the same access type into a data unit for the databases that the application can access. For example, you can use a locally accessed data table as a data unit and a globally accessed data table as a data unit.

[0042] Among them, local access can refer to data in a data unit being directly accessed by a local application or service, or can refer to data in a data unit being accessed by a local business; global access can refer to data in a data unit being accessed by a global business.

[0043] For example, local services may include local virtual machine information management, local database instance information management, local server information management, etc., and global services may include virtual machine creation rules, billing rules and policy management, global account management, etc.

[0044] If the first data unit in the first master database node is updated, the first updated data in the first data unit and the first updated information of the first updated data can be synchronized to the first slave data node, thereby realizing data synchronization between the master and slave database nodes in the first data center.

[0045] Among them, the data update of the first data unit may be caused by a write operation on the first data unit, where the write operation may be, for example, an insert operation, an update operation, a delete operation, etc. For example, if a row of data in the first data unit is deleted, then the first updated data is the row of data.

[0046] The first attribute information may include but is not limited to the write operation corresponding to the first update data, the source information of the first update data, the source database to which the first update data belongs, the master database node information to which the first update data belongs, the operation source of the first update data, etc.

[0047] The source information of the first update data may be used to indicate to which master database node or database the first update data is initially written. For example, the source information may include a database node representation, a database identifier, and the like.

[0048] The source database to which the first updated data belongs may refer to the database into which the first updated data is initially written.

[0049] The primary database node information to which the first updated data belongs refers to the node information of the primary database node to which the first updated data is initially written, such as a node identifier.

[0050] Among them, the operation source of the first updated data can be a synchronization task or a business service. Based on the operation source of the first updated data, it can be determined whether the first data unit is written due to the synchronization task or due to the provision of business services.

[0051] It should be noted that the first data unit may be one or more, and there is no limitation on this.

[0052] Step 102: Determine the access type corresponding to the first data unit.

[0053] In this application, each data unit has attribute information, wherein the attribute information of the data unit may include but is not limited to the access type of the data unit, information about the database to which the data unit belongs, information about the database node where the database to which the data unit belongs is located, etc.

[0054] Exemplarily, the access type corresponding to the first data unit may be determined based on the attribute information of the first data unit. The access type corresponding to the first data unit may be global access or local access.

[0055] Step 103 : In response to the access type corresponding to the first data unit being global access, the first updated data and the first attribute information are synchronized to the second master database node of the second data center through the first slave database node.

[0056] The second data center and the first data center may provide business services for the same application. In addition, the distance between the second data center and the first data center may be greater than a first threshold, which means that the second data center and the first data center are considered to be in different locations.

[0057] The second data center may include a second master database node and at least one second slave database node. The second slave database node may be responsible for processing all write operations, such as insert, update, and delete, and may also manage transaction logs to record all data changes and coordinate and manage data synchronization of the second slave database nodes.

[0058] In this application, if the access type corresponding to the first data unit is global access, it can be considered that the first data unit contains global data, then the first slave database node can synchronize the first updated data and the first attribute information to the second master database node of the second data center.

[0059] For example, a peer-to-peer data transmission tool can be used to synchronize data between remote data centers. For example, the peer-to-peer data transmission tool can subscribe to the transaction log of the first slave database node. When a data unit in the first slave database node is updated, a synchronization task is triggered. The transaction log can be parsed to obtain the first updated data and the first attribute information, and the first updated data and the first attribute information are synchronized to the second master database node.

[0060] In this application, the updated data that needs to be synchronized is screened out through the access type of the data unit, and then synchronized to the main database node of the remote data center, thereby reducing the amount of synchronized data and improving the efficiency of remote data synchronization.

[0061] In this application, the first data center and the second data center can provide services at the same time, and the remote data center is not a cold standby solution, so it will not cause waste of resource costs.

[0062] In addition, the databases in the first data center and the second data center can be metadata databases, and the metadata databases in the first data center and the second data center can be of different types. For example, the metadata database in data center 1 can be of MySQL type, and the metadata database in data center 2 can be of PostgreSQL type.

[0063] It should be noted that there can be one or more second data centers, and there is no limitation on this.

[0064] Step 104: Synchronize the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node.

[0065] In this application, the second master database node has a data update due to the first updated data and the first attribute information, and the first updated data and the first data information can be synchronized to the second slave database node, so that the data of the second database center is consistent, and data synchronization between the first data center and the second data center is realized.

[0066] In an embodiment of the present application, if a data unit in the master database node of the first data center is updated, it can be synchronized to the slave database node of the first data center to achieve data consistency in the first data center. If the access type of the first data unit is global access, the first updated data and its attribute information are synchronized to the second master database node of the second data center in a different location, and then synchronized to the slave database node of the second data center by the second master database node. In this way, the data that needs to be synchronized between the remote data centers is filtered out based on the access type of the data unit, the amount of synchronized data is reduced, and thus the data synchronization efficiency between the remote data centers is improved. Data synchronization is performed at the data unit granularity, which can improve data synchronization efficiency. Moreover, synchronization from the slave database node of the first data center to the master database node of the second data center can reduce the pressure on the master database node of the first data center.

[0067] Figure 2 A flowchart of a data synchronization method provided in another embodiment of the present application.

[0068] like Figure 2 As shown, the data synchronization method includes:

[0069] Step 201: In response to a data update of a first data unit in a first master database node of a first data center, first updated data and first attribute information of the first updated data in the first data unit are synchronized to a first slave database node of the first data center.

[0070] In the present application, step 201 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0071] For example, if there are multiple first slave database nodes, after receiving the first updated data and first attribute information synchronized by the first master database node and performing synchronized updates, the multiple first slave database nodes can send synchronization confirmation messages to the first master database node. When the first master database node receives a synchronization confirmation message for the first updated data from any first slave database node, it can determine that the data in the first data center is consistent. This eliminates the need to wait for synchronization confirmation messages from all first slave database nodes in the first data center before determining that the data in the first data center is consistent. This significantly improves data synchronization performance and prevents the master data node from waiting for extended periods of time.

[0072] Exemplarily, the first data center can be set up in multiple geographical areas, and the distance between the multiple geographical areas is less than the second threshold. Multiple first slave database nodes can be set up in each geographical area, and the first master database node can be set up in one of the geographical areas. For any geographical area, if the first master database node receives a synchronization confirmation message for the first updated data sent by any first slave database node in the geographical area, it is determined that the data in the geographical area is consistent.

[0073] For example, the first data center has two computer rooms, which are located in different areas of the same city. The first master database node is set in one of the computer rooms. If the first master database node receives a synchronization confirmation message sent by any first slave database node in a certain computer room, it can be determined that the data in the computer room is consistent.

[0074] Therefore, there is no need to wait for the synchronization confirmation messages sent by all the first slave database nodes in the geographical area to determine the consistency of the data in the geographical area. This can greatly improve the performance of data synchronization and avoid the master data node waiting for too long.

[0075] Step 202: Determine the access type corresponding to the first data unit.

[0076] In the present application, step 202 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0077] Step 203 : In response to the access type corresponding to the first data unit being global access, determine whether the data source of the first update data is the first master database node based on the first source information in the first attribute information.

[0078] In the present application, the first attribute information may include the first source information of the first updated data. If the access type corresponding to the first data unit is global access, the data of the first data unit can be considered as globally accessed data. Then, based on the first source information, it is further determined whether the data source of the first updated data is the first master database node, that is, whether the first updated data is written into the first data unit by the first master database node due to providing business services.

[0079] Exemplarily, the first source information may include a database node identifier, and the database node identified by the database node identifier is the data source of the first updated data. If the database node identifier is the same as the identifier of the first main database node, it can be determined that the data source of the first updated data is the first main database node. If the database node identifier is different from the identifier of the first main database node, it can be determined that the data source of the first updated data is not the first main database node, and the first updated data is written to the first data unit due to a synchronization task.

[0080] Step 204 : In response to the data source of the first update data being the first master database node, the first update data and the first attribute information are synchronized to the second master database node via the first slave database node.

[0081] In this application, if the access type corresponding to the first data unit is global access, and the data source of the first updated data is the first master database node, the first slave database node synchronizes the first updated data and the first attribute information to the second master database node.

[0082] For example, there may be multiple first slave database nodes. The first slave database node responsible for providing offline business services can synchronize the first updated data and first attribute information to the second master database node. This can reduce the pressure on the first slave database node responsible for providing online business services, avoid affecting online business services, and ensure the quality of online business services.

[0083] Step 205: Synchronize the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node.

[0084] In the present application, step 205 can be implemented in any of the embodiments of the present application, so it will not be described in detail here.

[0085] In an embodiment of the present application, when the access type corresponding to the first data unit is global access, it is further determined based on the source information of the first updated data whether the data source of the first updated data is the first master database node. If the data source is the first master database node, the first updated data and its attribute information are synchronized to the master database node of the remote data center, which can further improve the accuracy of the data to be synchronized between the remote data centers, reduce the amount of synchronized data, and improve synchronization efficiency.

[0086] Figure 3 A flowchart of a data synchronization method provided in another embodiment of the present application.

[0087] like Figure 3 As shown, the data synchronization method includes:

[0088] Step 301: In response to a data update of a first data unit in a first master database node of a first data center, first updated data and first attribute information of the first updated data in the first data unit are synchronized to a first slave database node of the first data center.

[0089] In this application, step 301 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0090] Step 302: Determine the access type corresponding to the first data unit.

[0091] In this application, step 302 can be implemented in any of the embodiments of this application, so it will not be described here in detail.

[0092] Step 303: In response to the access type corresponding to the first data unit being global access, obtain load information of each first slave database node.

[0093] In the present application, there may be multiple first data units and multiple first slave database nodes. If the access type corresponding to the first data unit is global access, the load information of each first slave database node may be obtained. The load information may include the number of tasks being processed and to be processed.

[0094] Step 304: Determine the number of synchronization tasks corresponding to each first slave database node according to the load information.

[0095] In this application, the load of the first slave database can be determined based on the load information of each first slave database node, and the number of synchronization tasks that each first slave database node can undertake can be determined based on the load, the total number of synchronization tasks, the number of first slave database nodes, etc.

[0096] If the access type for multiple first data units is global access, the total number of synchronization tasks is the same as the number of data update records for the multiple first data units. For example, if there are 100 data units with data updates, 80 of which have global access access and 1000 data update records, the total number of synchronization tasks is 1000.

[0097] For example, if first slave database nodes with a load less than a preset load are selected based on the load of each first slave database node, the number of synchronization tasks corresponding to each of the selected first slave database nodes is determined based on the number of selected first slave database nodes and the total number of synchronization tasks. The greater the load of a first slave database node, the fewer synchronization tasks it can handle.

[0098] It is understandable that the number of synchronization tasks corresponding to some first slave database nodes in the first data center may be zero.

[0099] Step 305 : Determine, from the plurality of first data units, a first data unit corresponding to each first slave database node according to the number of synchronization tasks.

[0100] In the present application, first update data equal to the number of synchronization tasks can be selected from multiple first data units of global type and allocated to the first slave database nodes, so that each first slave database node can be allocated a corresponding number of synchronization tasks.

[0101] Step 306: Synchronize the first updated data and the first attribute information corresponding to each first slave database node to the second master database node through each first slave database node.

[0102] In this application, for the first slave database node assigned the synchronization task, the first slave database node synchronizes the first updated data and the first attribute information of the first updated data to the second master database node, thereby realizing the parallel synchronization tasks between remote data centers and improving data synchronization efficiency.

[0103] Step 307: Synchronize the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node.

[0104] In this application, step 307 can be implemented in any of the embodiments of this application, so it will not be described here in detail.

[0105] In an embodiment of the present application, in the case where there are multiple first data units with data changes and multiple first slave database nodes, a corresponding number of synchronization tasks can be assigned to the first slave database nodes based on the load information of each first slave database node, and the multiple first slave database nodes can synchronize in parallel to the master database node of the remote data center, thereby reducing delays and improving the data synchronization efficiency between remote data centers.

[0106] Figure 4 A flowchart of a data synchronization method provided in another embodiment of the present application.

[0107] like Figure 4 As shown, the data synchronization method further includes:

[0108] Step 401 : In response to a data update of a second data unit in a second master database node, second updated data and second attribute information of the second updated data in the second data unit are synchronized to a second slave database node.

[0109] The second data unit may be one or more. The explanation of the second attribute information can refer to the first attribute information and will not be repeated here.

[0110] In this application, the method by which the second master database point synchronizes the second updated data and the second attribute information to the second slave database node is the same as the method by which the first master database point synchronizes the first updated data and the first attribute information to the second slave database node, so they will not be repeated here.

[0111] Step 402: Determine the access type corresponding to the second data unit.

[0112] In this application, you can refer to the above explanation of determining the access type corresponding to the first data unit, so it will not be repeated here.

[0113] Step 403 : In response to the access type corresponding to the second data unit being global access, the second updated data and the second attribute information are synchronized to the first master database node through the second slave database node.

[0114] In the present application, the second attribute information may include the second source information of the second updated data. If the access type corresponding to the second data unit is global access, the data of the second data unit can be considered as globally accessed data. Then, based on the second source information, it is further determined whether the data source of the second updated data is the second main database node, that is, whether the second updated data is written into the second data unit by the second main database node due to providing business services.

[0115] Exemplarily, the second source information may include a database node identifier, and the database node identified by the database node identifier is the data source of the second updated data. If the database node identifier is the same as the identifier of the first main database node, it can be determined that the data source of the second updated data is the second main database node. If the database node identifier is different from the identifier of the second main database node, it can be determined that the data source of the second updated data is not the second main database node, and the second updated data is written to the second data unit due to the synchronization task.

[0116] If the access type corresponding to the second data unit is global access, and the data source of the second updated data is the second master database node, the second slave database node synchronizes the second updated data and the second attribute information to the first master database node.

[0117] Step 404: synchronize the second updated data and the second attribute information to the first slave database node through the first master database node.

[0118] In this application, please refer to the above explanation of synchronizing the first updated data and the first attribute information to the second slave database node through the second master database node, which will not be repeated here.

[0119] In an embodiment of the present application, when the first data unit of the first master database node of the first data center changes and the second data unit is globally accessed, the first updated data and the first attribute information are first synchronized to the first slave database node, and then synchronized to the second master database node of the second data center by the first slave database node. Similarly, when the second master database node of the second data center updates data and the second data unit is globally accessed, the second updated data and the second attribute information can be first synchronized to the second slave database node, and then synchronized to the first master database node of the first data center by the second slave database node, thereby realizing bidirectional data synchronization of globally accessed data between remote data centers.

[0120] Moreover, during bidirectional data synchronization, it is possible to determine whether the data source of the updated data is the master database node of the slave database node that initiated the data synchronization based on the source information of the updated data. If so, it is synchronized to the remote data center, thereby avoiding bidirectional data loops, reducing the amount of synchronized data, saving resources, avoiding data key conflicts, and improving the accuracy of bidirectional data synchronization.

[0121] In one embodiment of the present application, based on the above embodiment, encryption gateways can be deployed respectively for the first data center and the second data center. The encryption gateway has a heterogeneous protocol conversion function. Taking the first data center as an example, the encryption gateway can first determine whether the first database type corresponding to the data to be written is consistent with the second database type of the second data center. If the first database type and the second database type are inconsistent, it means that the data to be written cannot be directly stored in the first data center. The data to be written can be converted into data of the first database type through the encryption gateway, and then encrypted to obtain first update data, and the first update data is written into the first data unit.

[0122] Exemplarily, after receiving an access request from a client, the encryption gateway can convert the data to be written from the first database type into data of the first database type and encrypt it before writing it into the main database node of the first database center if the type of the database from which the data is to be written is different from the database type of the first data center.

[0123] For example, if the data to be synchronized is received from the second data center, the data to be synchronized may be decrypted first to obtain the data to be written, and then the above judgment may be performed through the encryption gateway.

[0124] Therefore, the data to be written is converted into a type suitable for storage in the data center through the encryption gateway, thereby improving the generalization capability of the data center.

[0125] In addition, the encryption gateway can be integrated into the data center. The deployment of the encryption gateway is loosely coupled with the data center, which can improve the data center's ability to provide services.

[0126] In addition, the encryption algorithm used by the encryption gateway can be stored separately from the data center and seamlessly connected to ensure data security.

[0127] In one embodiment of the present application, based on the above embodiment, multiple slave database nodes in a data center can be categorized into slave database nodes that provide online business services and slave database nodes that provide offline business services, thereby separating online and offline business services. For example, online business may primarily involve primary key index query services, while offline business may primarily involve statistical reporting and analytical query services.

[0128] Taking the first data center as an example, if the first data center receives a database access request for a target business, it can first determine the business type of the target business. If the business type of the target business is an online business, the database access request can be sent to the first slave database node that provides online business services. If the business type of the target business is an offline business, the database access request can be sent to the first slave database node that provides offline business services.

[0129] For example, the service type of the target service may be determined based on the attribute information of the target service, such as the real-time requirement.

[0130] Therefore, by assigning online services and offline services to different slave database nodes to provide services, online services are separated from offline services, which can avoid affecting online services and ensure the service quality of online services and offline services.

[0131] In order to facilitate understanding of the data synchronization method of this application, Figure 5 Provide explanation. Figure 5 A schematic diagram of a data synchronization process provided in an embodiment of the present application.

[0132] like Figure 5 As shown, there are two data centers, Data Center 1 and Data Center 2, and the distance between the two data centers is greater than a first threshold. Data Center 1 is located in Area 1 and Area 2 of the same city. This means that Data Center 1 has two computer rooms, one in Area 1 and the other in Area 2.

[0133] If data centers are set up in multiple regions (for the above-mentioned geographical regions), the slave data nodes in each region can be grouped together. For example, the slave database nodes in region 1 in data center 1 are divided into group 1, and the slave database nodes in region 2 in data center 1 are divided into group 2. Taking region 2 as an example, if master1 receives a synchronization confirmation message from any slave database node in region 2, that is, receives a synchronization confirmation message from any slave database node in group 2, it can be determined that the data in region 2 is consistent.

[0134] If data is updated in the data unit of the master database node master1 in data center 1, it can be synchronized to the slave database nodes slave1, slave2, ..., slave8. If the access type corresponding to the data unit where the updated data is located is global access, and the data source of the updated data is master1, then it can be synchronized to the master database node master2 in data center 2 by the slave database node slave6.

[0135] Similarly, if a data unit in the master database node master2 in data center 2 is updated, it can be synchronized to the slave database nodes slave9, slave10, slave11, and slave12. If the access type corresponding to the data unit where the updated data is located is global access and the data source of the updated data is master2, then the slave database node slave12 can synchronize it to the master database node master1 in data center 1. This allows for bidirectional data synchronization based on the data source of the updated data, avoiding bidirectional data loops and data key conflicts.

[0136] When synchronizing data between remote data centers, point-to-point data transmission tools can be used to perform high-performance point-to-point data transmission, thereby improving data synchronization performance.

[0137] In addition, the slave database nodes in the data center can be divided into two categories: one for providing services for online business and the other for providing services for offline business. For example, Figure 5 In data center 1, online business services are provided from database nodes slave1, slave2, slave3, and slave4, and offline business services are provided from database nodes slave5, slave6, slave7, and slave8. In data center 2, online business services are provided from database nodes slave9 and slave10, and offline business services are provided from database nodes slave11 and slave12.

[0138] In addition, the services provided by the database can also be divided into local and global services. An online service can be a local service or a global service, and an offline service can be a local service or a global service. Among them, the access type of the data unit accessed by the global service is global access.

[0139] In order to implement the above embodiment, the embodiment of the present application also proposes a data synchronization device. Figure 6 A schematic diagram of the structure of a data synchronization device provided in one embodiment of the present application.

[0140] like Figure 6 As shown, the data synchronization device 600 includes:

[0141] A first synchronization module 610 is configured to synchronize first updated data and first attribute information of the first updated data in a first data unit in a first master database node in a first data center to a first slave database node in the first data center in response to a data update in the first data unit;

[0142] A first determining module 620, configured to determine an access type corresponding to the first data unit;

[0143] a second synchronization module 630 configured to synchronize the first updated data and the first attribute information to a second master database node in the second data center via the first slave database node in response to the access type corresponding to the first data unit being global access; wherein the distance between the first data center and the second data center is greater than a first threshold;

[0144] The third synchronization module 640 is configured to synchronize the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node.

[0145] Optionally, the first attribute information includes first source information of the first update data, and the second synchronization module 630 is configured to:

[0146] determining, based on the first source information, whether a data source of the first update data is the first master database node;

[0147] In response to the data source of the first update data being the first master database node, the first update data and the first attribute information are synchronized to the second master database node through the first slave database node.

[0148] Optionally, there are multiple first data units, multiple first slave database nodes, and the second synchronization module 630 is configured to:

[0149] Obtaining load information of each of the first slave database nodes;

[0150] Determine the number of synchronization tasks corresponding to each first slave database node according to the load information;

[0151] Determining, from the plurality of first data units, first updated data corresponding to each of the first slave database nodes according to the number of synchronization tasks;

[0152] The first updated data and the first attribute information corresponding to each first slave database node are synchronized to the second master database node through each first slave database node.

[0153] Optionally, the device further comprises:

[0154] a fourth synchronization module, configured to synchronize second updated data and second attribute information of the second updated data in the second data unit to the second slave database node in response to data update of the second data unit in the second master database node;

[0155] a second determining module, configured to determine an access type corresponding to the second data unit;

[0156] a fifth synchronization module, configured to synchronize the second updated data and the second attribute information to the first master database node through the second slave database node in response to the access type corresponding to the second data unit being global access;

[0157] A sixth synchronization module is used to synchronize the second updated data and the second attribute information to the first slave database node through the first master database node.

[0158] Optionally, the second attribute information includes second source information of the second update data, and the fifth synchronization module is configured to:

[0159] determining, according to the second source information, whether a data source of the second update data is the second master database node;

[0160] In response to the data source of the second update data being the second master database node, the second update data and the second attribute information are synchronized to the first master database node through the second slave database node.

[0161] Optionally, the device further comprises:

[0162] a third determination module, configured to determine, through an encryption gateway, whether the first database type corresponding to the data to be written is consistent with the second database type of the first data center;

[0163] A writing module is used to convert and encrypt the data to be written through the encryption gateway in response to the inconsistency between the first database type and the second database type, obtain the first update data, and write the first update data into the first data unit.

[0164] Optionally, the first data center is set in multiple geographical areas, the distance between the multiple geographical areas is less than a second threshold, and the geographical areas are provided with multiple first slave database nodes. The apparatus further includes:

[0165] The fourth determination module is configured to determine, for any geographical area, that data in the any geographical area is consistent in response to the first master database node receiving a synchronization confirmation message for the first updated data sent by any first slave database node in the any geographical area.

[0166] Optionally, the device further comprises:

[0167] a fifth determining module, configured to determine a business type of the target business in response to receiving a database access request of the target business;

[0168] A sending module, configured to send the database access request to a first slave database node providing online business services in response to the business type being an online business;

[0169] The sending module is further configured to send the database access request to a first slave database node providing offline business services in response to the business type being an offline business.

[0170] It should be noted that the explanation of the aforementioned data synchronization method embodiment is also applicable to the data synchronization device of this embodiment, so it will not be repeated here.

[0171] In an embodiment of the present application, if a data unit in the master database node of the first data center is updated, it can be synchronized to the slave database node of the first data center to achieve data consistency in the first data center. If the access type of the first data unit is global access, the first updated data and its attribute information are synchronized to the second master database node of the second data center in a different location, and then synchronized to the slave database node of the second data center by the second master database node. In this way, the data that needs to be synchronized between the remote data centers is filtered out based on the access type of the data unit, the amount of synchronized data is reduced, and thus the data synchronization efficiency between the remote data centers is improved. Data synchronization is performed at the data unit granularity, which can improve data synchronization efficiency. Moreover, synchronization from the slave database node of the first data center to the master database node of the second data center can reduce the pressure on the master database node of the first data center.

[0172] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0173] Figure 7 A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0174] like Figure 7As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 702 or a computer program loaded from a storage unit 708 into a RAM (Random Access Memory) 703. Various programs and data required for the operation of the device 700 can also be stored in the RAM 703. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An I / O (Input / Output) interface 705 is also connected to the bus 704.

[0175] Various components in device 700 are connected to I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, optical disk, etc.; and a communication unit 709, such as a network card, modem, wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0176] The computing unit 701 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Unit), various specialized AI (Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above, such as the data synchronization method. For example, in some embodiments, the data synchronization method can be implemented as a computer software program that is tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the data synchronization method described above can be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to execute the data synchronization method in any other appropriate manner (eg, by means of firmware).

[0177] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System on Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0178] The program code for implementing the methods of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow charts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0179] In the context of the present application, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, RAM, ROM, EPROM (Electrically Programmable Read-Only-Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0180] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0181] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.

[0182] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The client-server relationship is established by computer programs running on the respective computers and establishing a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and poor scalability of traditional physical hosts and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0183] According to an embodiment of the present application, the present application further provides a computer program product, which, when an instruction processor in the computer program product is executed, executes the data synchronization method proposed in the above embodiment of the present application.

[0184] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this application can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this application can be achieved. This is not a limitation herein.

[0185] The above specific embodiments do not constitute a limitation on the scope of protection of this application. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application shall be included within the scope of protection of this application.

Claims

1. A data synchronization method, comprising: In response to a data update occurring in a first data unit in a first master database node in a first data center, synchronizing first updated data in the first data unit and first attribute information of the first updated data to a first slave database node in the first data center; determining an access type corresponding to the first data unit; In response to the access type corresponding to the first data unit being global access, synchronizing the first updated data and the first attribute information to a second master database node in a second data center through the first slave database node; wherein a distance between the first data center and the second data center is greater than a first threshold; Synchronizing the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node; There are multiple first data units and multiple first slave database nodes, and synchronizing the first updated data and the first attribute information to the second master database node of the second data center through the first slave database node includes: Obtaining load information of each of the first slave database nodes; Determine the number of synchronization tasks corresponding to each first slave database node according to the load information; Determining, from the plurality of first data units, first updated data corresponding to each of the first slave database nodes according to the number of synchronization tasks; The first updated data and the first attribute information corresponding to each first slave database node are synchronized to the second master database node through each first slave database node.

2. The method according to claim 1, wherein The first attribute information includes first source information of the first updated data, and synchronizing the first updated data and the first attribute information to the second master database node of the second data center through the first slave database node includes: determining, based on the first source information, whether a data source of the first update data is the first master database node; In response to the data source of the first update data being the first master database node, the first update data and the first attribute information are synchronized to the second master database node through the first slave database node.

3. The method of claim 1 , further comprising: In response to a data update of the second data unit in the second master database node, synchronizing the second updated data in the second data unit and the second attribute information of the second updated data to the second slave database node; determining an access type corresponding to the second data unit; In response to the access type corresponding to the second data unit being global access, synchronizing the second updated data and the second attribute information to the first master database node through the second slave database node; The second updated data and the second attribute information are synchronized to the first slave database node through the first master database node.

4. The method according to claim 3, wherein: The second attribute information includes second source information of the second updated data, and synchronizing the second updated data and the second attribute information to the first master database node through the second slave database node further includes: determining, according to the second source information, whether a data source of the second update data is the second master database node; In response to the data source of the second update data being the second master database node, the second update data and the second attribute information are synchronized to the first master database node through the second slave database node.

5. The method of claim 1 , further comprising: Determining, through the encryption gateway, whether the first database type corresponding to the data to be written is consistent with the second database type of the first data center; In response to the first database type being inconsistent with the second database type, the data to be written is converted and encrypted by the encryption gateway to obtain the first update data, and the first update data is written into the first data unit.

6. The method of claim 1, wherein: The first data center is set in multiple geographical areas, the distance between the multiple geographical areas is less than a second threshold, and the geographical areas are provided with multiple first slave database nodes. The method further includes: For any geographical area, in response to the first master database node receiving a synchronization confirmation message for the first updated data sent by any first slave database node in the any geographical area, it is determined that the data in the any geographical area is consistent.

7. The method of claim 1 , further comprising: In response to receiving a database access request for a target business, determining a business type of the target business; In response to the business type being an online business, sending the database access request to a first slave database node that provides online business services; In response to the service type being an offline service, the database access request is sent to a first slave database node that provides offline service.

8. A data synchronization device, comprising: a first synchronization module, configured to synchronize, in response to a data update occurring in a first data unit in a first master database node in a first data center, first updated data and first attribute information of the first updated data in the first data unit to a first slave database node in the first data center; A first determining module, configured to determine an access type corresponding to the first data unit; a second synchronization module, configured to synchronize the first updated data and the first attribute information to a second master database node in a second data center via the first slave database node in response to the access type corresponding to the first data unit being global access; wherein the distance between the first data center and the second data center is greater than a first threshold; a third synchronization module, configured to synchronize the first updated data and the first attribute information to the second slave database node in the second data center through the second master database node; There are multiple first data units, multiple first slave database nodes, and the second synchronization module is used to: Obtaining load information of each of the first slave database nodes; Determine the number of synchronization tasks corresponding to each first slave database node according to the load information; Determining, from the plurality of first data units, first updated data corresponding to each of the first slave database nodes according to the number of synchronization tasks; The first updated data and the first attribute information corresponding to each first slave database node are synchronized to the second master database node through each first slave database node.

9. The device according to claim 8, wherein The first attribute information includes first source information of the first update data, and the second synchronization module is used to: determining, based on the first source information, whether a data source of the first update data is the first master database node; In response to the data source of the first update data being the first master database node, the first update data and the first attribute information are synchronized to the second master database node through the first slave database node.

10. The apparatus of claim 8, further comprising: a fourth synchronization module, configured to synchronize second updated data and second attribute information of the second updated data in the second data unit to the second slave database node in response to data update of the second data unit in the second master database node; a second determining module, configured to determine an access type corresponding to the second data unit; a fifth synchronization module, configured to synchronize the second updated data and the second attribute information to the first master database node through the second slave database node in response to the access type corresponding to the second data unit being global access; A sixth synchronization module is used to synchronize the second updated data and the second attribute information to the first slave database node through the first master database node.

11. The device according to claim 10, wherein The second attribute information includes second source information of the second update data, and the fifth synchronization module is configured to: determining, according to the second source information, whether a data source of the second update data is the second master database node; In response to the data source of the second update data being the second master database node, the second update data and the second attribute information are synchronized to the first master database node through the second slave database node.

12. The apparatus of claim 8, further comprising: a third determination module, configured to determine, through an encryption gateway, whether the first database type corresponding to the data to be written is consistent with the second database type of the first data center; A writing module is used to convert and encrypt the data to be written through the encryption gateway in response to the inconsistency between the first database type and the second database type, obtain the first update data, and write the first update data into the first data unit.

13. The apparatus of claim 8, wherein: The first data center is set in multiple geographical areas, the distance between the multiple geographical areas is less than a second threshold, and the geographical areas are provided with multiple first slave database nodes. The apparatus further includes: The fourth determination module is configured to determine, for any geographical area, that data in the any geographical area is consistent in response to the first master database node receiving a synchronization confirmation message for the first updated data sent by any first slave database node in the any geographical area.

14. The apparatus of claim 8, further comprising: a fifth determining module, configured to determine a business type of the target business in response to receiving a database access request of the target business; A sending module, configured to send the database access request to a first slave database node providing online business services in response to the business type being an online business; The sending module is further configured to send the database access request to a first slave database node providing offline business services in response to the business type being an offline business.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data synchronization method and device, computer equipment and computer readable storage medium

    CN112883119A

  • Data processing method and system

    CN113656496A