Data Synchronization Method, Device, Terminal, and Storage Medium

By performing tree-level library division and offline bypass processing on the subscription data set, a tree-shaped synchronization system is established, which solves the problem of insufficient flexibility in data synchronization in the existing technology, and realizes efficient data synchronization between a large number of databases and target objects.

CN117235087BActive Publication Date: 2025-07-18BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311337656.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-16
Publication Date
2025-07-18
Estimated Expiration
2043-10-16

AI Technical Summary

Technical Problem

In the prior art, data synchronization can only realize real-time synchronization between source databases and a small number of databases, and cannot realize real-time synchronization between source databases and a large number of databases, resulting in poor flexibility in data synchronization.

Method used

By hierarchically dividing the subscription data sets sent by multiple databases, a tree synchronization system is established. Each upstream node can correspond to multiple downstream nodes. The data stored in the leaf node instance supports target object subscription, and high fan-out synchronization is achieved through offline bypass and distributed file systems.

Benefits of technology

It realizes data synchronization between a large number of databases and a large number of target objects, improves the flexibility and throughput of data synchronization, and can maintain stability and efficient transmission in high fan-out scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117235087B_ABST
    Figure CN117235087B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a data synchronization method, apparatus, terminal, and storage medium, which relate to the computer field, and particularly to the database field. The specific implementation solution is as follows: receiving a subscription data set sent by a database set within a duration threshold; performing tree-level database partitioning on the subscription data set to obtain a tree-shaped synchronization system, wherein an instance is set for each node in the tree-shaped synchronization system, and the instance in the root node receives and stores the data in the subscription data set, and the data stored in any upstream instance supports subscription by multiple downstream instances adjacent to any upstream instance; synchronizing the data stored in the leaf node instance set to multiple target objects, wherein the leaf node instance set includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system. The present disclosure adopting the above solution can improve the flexibility of data synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a data synchronization method, an apparatus, a terminal, and a storage medium. Background Art

[0002] Data synchronization is used to achieve real-time synchronization from a data source to a destination. Its basic components are a data source, a data terminal, and data transmission. In related technologies, only real-time synchronization between a source database and a small number of databases can be achieved, and the flexibility of data synchronization is poor. Summary of the Invention

[0003] The present disclosure provides a data synchronization method, an apparatus, a terminal, and a storage medium, and the main purpose is to improve the flexibility of data synchronization.

[0004] According to an aspect of the present disclosure, a data synchronization method is provided, including:

[0005] Receiving a set of subscribed data sent by a set of databases within a duration threshold;

[0006] Performing tree-level database partitioning on the set of subscribed data to obtain a tree-shaped synchronization system, where an instance is set for each node in the tree-shaped synchronization system, the instance in the root node receives and stores the data in the set of subscribed data, and the data stored in any upstream instance supports subscription by multiple downstream instances adjacent to the any upstream instance;

[0007] Synchronizing the data stored in the set of leaf node instances to multiple target objects, where the set of leaf node instances includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system.

[0008] According to another aspect of the present disclosure, a data synchronization apparatus is provided, including:

[0009] A data receiving unit, configured to receive a set of subscribed data sent by a set of databases within a duration threshold;

[0010] A data database partitioning unit, configured to perform tree-level database partitioning on the set of subscribed data to obtain a tree-shaped synchronization system, where an instance is set for each node in the tree-shaped synchronization system, the instance in the root node receives and stores the data in the set of subscribed data, and the data stored in any upstream instance supports subscription by multiple downstream instances adjacent to the any upstream instance;

[0011] A data publishing unit, configured to synchronize the data stored in the set of leaf node instances to multiple target objects, where the set of leaf node instances includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system.

[0012] According to another aspect of the present disclosure, a terminal is provided, including:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of the foregoing aspects.

[0016] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method according to any one of the foregoing aspects.

[0017] According to another aspect of the present disclosure, a computer program product is provided, including a computer program which, when executed by a processor, implements the method according to any one of the foregoing aspects.

[0018] In one or more embodiments of the present disclosure, by performing tree-level sub-database partitioning on the subscription data sets sent by multiple databases, each upstream node can correspond to multiple downstream nodes, and the data stored in the leaf node instances can all support target object subscriptions. That is to say, the tree-shaped synchronization system obtained through tree-level sub-database partitioning can fan out a large number of instances for target object synchronization. Therefore, data synchronization between a large number of databases and a large number of target objects can be achieved, and the flexibility of data synchronization is high.

[0019] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0021] Figure 1 A background schematic diagram showing a data synchronization method provided by an embodiment of the present disclosure;

[0022] Figure 2 A flowchart showing the first data synchronization method provided by an embodiment of the present disclosure;

[0023] Figure 3 A flowchart showing the second data synchronization method provided by an embodiment of the present disclosure;

[0024] Figure 4Schematic diagram of an architecture of a data synchronization method provided by an embodiment of the present disclosure;

[0025] Figure 5 Schematic diagram of a structure of a data synchronization device provided by an embodiment of the present disclosure;

[0026] Figure 6 Block diagram of a terminal for implementing the data synchronization method of an embodiment of the present disclosure. Detailed implementation manners

[0027] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0028] It should be noted that when real-time synchronization is performed between a source database and a small number of databases, it can be implemented through a source change capture module, a message queue module, and a terminal synchronization module. Among them, due to the limitation of the single-partition subscription: publication throughput ratio of the message queue module, the cloud product can only achieve a single-partition subscription: publication throughput ratio of 1:2, that is, a single source database can only be synchronously and timely with two databases.

[0029] According to some embodiments, Figure 1 Schematic background diagram of a data synchronization method provided by an embodiment of the present disclosure. As Figure 1 shown, the essence of computational advertising is to match user needs and customer businesses to achieve the best balance among user experience, customer conversion, and platform efficiency. From the implementation level, in the field of advertising placement, the data synchronization method is applied, and the data synchronization system corresponding to the data synchronization method is composed of a placement system and a retrieval system. A customer expresses the product or service to be promoted in the placement system, and the advertising data placed by the customer will be synchronously and timely to the retrieval system. When a user searches for a product / information stream product and triggers a retrieval request, the request is sent to the user product system and also sent to the (commercial) retrieval system. The retrieval system completes the matching of the current traffic x placement and returns the advertising result, and the advertising result and the natural result are returned simultaneously and presented on the user interface.

[0030] In some embodiments, as the traffic system continues to grow and in order to achieve better matching, the retrieval complexity also increases. For example, the number of retrieval instances can be as high as tens of thousands and can continue to rise. This means that for customer placements, it is necessary to fan out from the database of the placement system to tens of thousands of instances. However, in the prior art, only real-time synchronization between the source database and a small number of databases can be achieved, and real-time synchronization between the source database and a large number of databases (for example, more than ten thousand databases) cannot be achieved.

[0031] The present application will be described in detail below with reference to specific embodiments.

[0032] In the first embodiment, as Figure 2 shown, Figure 2 FIG. 1 shows a schematic flow chart of a first data synchronization method provided by an embodiment of the present disclosure. This method can be implemented depending on a computer program and can run on a device for data synchronization. The computer program can be integrated in an application or run as an independent tool class application.

[0033] Among them, the data synchronization device can be a terminal with data synchronization function. The terminal includes but is not limited to: wearable devices, handheld devices, personal computers, tablet computers, vehicle-mounted devices, smart phones, computing devices or other processing devices connected to a wireless modem, etc. In different networks, the terminal can be called by different names. For example: user equipment, access terminal, user unit, user station, mobile station, mobile unit, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), fifth-generation mobile communication technology (5G) network, fourth-generation mobile communication technology (4G) network, third-generation mobile communication technology (3G) network or a terminal in a future evolved network, etc.

[0034] Specifically, the data synchronization method includes:

[0035] S101, receiving a set of subscription data sent by a set of databases within a duration threshold;

[0036] According to some embodiments, the duration threshold refers to the duration used by the terminal to perform one data synchronization. This duration threshold does not specifically refer to a certain fixed threshold. For example, the duration threshold can be adjusted according to the actual application scenario.

[0037] In some embodiments, a database set refers to a set formed by aggregating multiple databases that need to be synchronized, and each database corresponds one-to-one to the subscription data in the subscription data set. A database (DB) is an organic collection of a large amount of sharable data organized in a certain structure and stored in the terminal for a long time.

[0038] According to some embodiments, subscription data refers to the data that needs to be synchronized. For example, in the field of advertising placement, the subscription data may be advertising data.

[0039] It is easy to understand that when the terminal performs data synchronization, the terminal can receive the subscription data set sent by the database set within the duration threshold.

[0040] S102. Perform tree-level database partitioning on the subscription data set to obtain a tree-shaped synchronization system;

[0041] According to some embodiments, tree-level database partitioning means performing multi-level database partitioning on the subscription data set starting from the root node. Among them, the root node, also called the tree root, refers to the ancestor of all nodes except itself in the same tree. The root node has no parent node, and a parent node can obtain multiple child nodes through database partitioning. Therefore, after performing multi-level database partitioning in sequence, multiple leaf nodes can be obtained. The number of leaf nodes can be, for example, tens of thousands. A leaf node refers to a node without child nodes.

[0042] According to some embodiments, a tree-shaped synchronization system refers to a system that can perform data synchronization obtained after performing tree-level database partitioning on the subscription data set. An instance can be set for each node in the tree-shaped synchronization system. The instance in the root node can receive and store all the data in the subscription data set, and the data stored in any upstream instance can support multiple downstream instances adjacent to the any upstream instance to subscribe.

[0043] In some embodiments, an instance refers to a database program in a running state when the terminal executes the data synchronization method, and some memory spaces allocated for these database programs. An instance is located in the memory and only exists when the database is in a running state. The instance is responsible for implementing various functions such as providing network connections and reading and writing data files.

[0044] It should be noted that publish / subscribe is a message paradigm. The publisher of the data does not directly send the data to a specific subscriber, but publishes the data through a data channel so that the subscriber subscribing to the data can obtain it.

[0045] In some embodiments, any instance except the root node can subscribe to an upstream instance, and the instance can also publish the data stored inside it for multiple downstream instances to subscribe.

[0046] It is easy to understand that when the terminal obtains the subscription data set, the terminal can perform tree-level sub-database partitioning on the subscription data set to obtain a tree-shaped synchronization system.

[0047] S103, synchronize the data stored in the leaf node instance set to multiple target objects.

[0048] According to some embodiments, the leaf node instance set includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system.

[0049] In some embodiments, the target object refers to an object that needs to be synchronized with data. For example, in the field of advertising placement, the target object can be a retrieval system.

[0050] In some embodiments, when synchronizing the data stored in the leaf node instance set to multiple target objects, each leaf node instance can publish the data it stores for the target object to subscribe, and each target object can obtain the data stored in different leaf node instances by subscribing to different leaf node instances.

[0051] It is easy to understand that when the terminal obtains the tree-shaped synchronization system, the terminal can synchronize the data stored in the leaf node instance set to multiple target objects.

[0052] In summary, the method provided by the embodiments of the present disclosure performs tree-level sub-database partitioning on the subscription data sets sent by multiple databases. Each upstream node can correspond to multiple downstream nodes, and the data stored in the leaf node instances can all support subscription by the target objects. That is to say, the tree-shaped synchronization system obtained through tree-level sub-database partitioning can fan out a large number of instances for the target objects to synchronize. Therefore, data synchronization between a large number of databases and a large number of target objects can be achieved, and the flexibility of data synchronization is high. In addition, in the high-fan-out scenario, the throughput of data synchronization is also relatively high.

[0053] Please refer to Figure 3 , Figure 3 which shows a schematic flowchart of the second data synchronization method provided by the embodiments of the present disclosure. Specifically,

[0054] S201, receive the subscription data set sent by the database set within the duration threshold;

[0055] It should be noted that when receiving the subscription data set sent by the database set to the subscription data set, for example, the subscription data sent by the database node of any database in the database set can be received.

[0056] For example, in the field of advertising placement, if the advertising data of M users is received within the duration threshold, the subscription data sent by the database corresponding to each of these M users can be received.

[0057] S202. Determine the number of levels for sub - databases in the tree - level sub - database division of the subscription data set.

[0058] According to some embodiments, when determining the number of levels for sub - databases in the tree - level sub - database division of the subscription data set, the traffic data corresponding to the instance can be determined, and the fan - out quota corresponding to the instance can be determined according to the traffic data; in the case where the total fan - out quota is less than the number of databases in the database set, the number of levels is increased until the number of databases in the database set is not greater than the total fan - out quota. Therefore, the accuracy of determining the number of levels can be improved.

[0059] In some embodiments, the traffic data is used to indicate the amount of data that the current instance can provide within a certain duration. The traffic data itself is variable. In the case where the traffic data is less than the single - sub - database throughput data, the smaller the traffic data, the larger the fan - out quota of a single instance can be.

[0060] In some embodiments, the single - sub - database throughput data refers to the single - sub - database throughput that an instance can support. The single - sub - database throughput data is input by the user corresponding to the database set. For example, in the field of advertising placement, the single - sub - database throughput data can be determined by the advertising data of the target object.

[0061] According to some embodiments, the total fan - out quota corresponding to the tree - shaped synchronization system is determined by the number of levels and the fan - out quota.

[0062] For example, when each instance can support N subscriptions, if the tree - shaped synchronization system has X levels, then at the leaf nodes, N*X fan - outs can be supported, that is, the total fan - out quota is N*X.

[0063] It should be noted that the total throughput processed by a single instance can be determined according to the fan - out quota and the single - sub - database throughput. Specifically, the total throughput processed by a single instance is the product of the fan - out quota and the single - sub - database throughput, which can vividly reflect the improvement of transmission capacity and scalability, as well as the best data synchronization scheme for improving the customer budget and user traffic.

[0064] According to some embodiments, when determining the number of levels for sub - databases in the tree - level sub - database division of the subscription data set, the single - sub - database throughput data input for the instance can also be obtained; in the case where the number of databases in the database set is equal to the total fan - out quota and the single - sub - database throughput data is higher than the single - sub - database throughput data threshold, the number of levels is increased. Therefore, the single - fan - out throughput can be effectively reduced, the stability of data synchronization can be improved, and the throughput of data synchronization in the fan - out scenario can be improved.

[0065] According to some embodiments, the technical implementation information corresponding to the instance cluster can be obtained; the single-instance resource specification information can be obtained; and the single-shard throughput data threshold can be determined according to the technical implementation information and the single-instance resource specification information. Therefore, the accuracy of determining the single-shard throughput data threshold can be improved.

[0066] In some embodiments, the instance cluster includes all instances and has the function of lifecycle management for these instances at the same time. The functions of this lifecycle management include but are not limited to instance migration, restart, keep-alive, etc.

[0067] In some embodiments, the technical implementation information is used to indicate the cluster information adopted when performing lifecycle management on any instance. For example, the technical implementation information can indicate that the k8s cluster performs lifecycle management on instance A and the yarn cluster performs lifecycle management on instance B.

[0068] In some embodiments, the single-instance resource specification information is used to indicate the resource allocation information corresponding to the resource specification of any instance, and specifically used to indicate how much cpu / mem / disk can be allocated to any instance. For example, if instance A is cpu-intensive, then more cpu will be allocated in its resource specification, and it can be allocated 8cpu + 10Gmem + 500G hdd disk; if instance B is io-intensive, then more disk and mem will be allocated in its resource specification, and it can be allocated 1cpu + 20G mem + 1T ssd disk.

[0069] It should be noted that by adjusting the allocation parameters in the resource allocation information, the single-instance resource specification information can be adjusted. For example, 10Gmem can be adjusted to 20Gmem.

[0070] According to some embodiments, as the technical implementation of the cluster continues to improve, by adjusting the single-instance resource specification information, the single-shard throughput data threshold can be adjusted to make the fan-out utilization rate have sufficient redundancy. That is to say, the throughput of the entire tree-shaped synchronization system is sufficient at this time. That is to say, the throughput of data synchronization in the fan-out scenario is high.

[0071] S203, according to the number of levels, perform tree-level sharding on the subscribed data set through the sharding key to obtain a tree-shaped synchronization system;

[0072] According to some embodiments, the sharding key is also called the partitioning key and is used to determine which shard the data will be distributed to.

[0073] In some embodiments, taking a relational database as an example, the sharding key can be one or more column fields.

[0074] It should be noted that each instance in other nodes except the root node is responsible for the data of a sub-database. The data of the sub-database can be all the data stored in the upstream instance corresponding to the instance, or can be part of the data stored in the upstream instance corresponding to the instance.

[0075] According to some embodiments, when performing tree-level sub-database partitioning on the subscribed data set through the sub-database key, the data stored in the upstream instance can be divided into at least one data set; according to the sub-database key configured in the downstream instance, the downstream instance is controlled to subscribe to the data set corresponding to the sub-database key.

[0076] In some embodiments, the data sets correspond one-to-one with the sub-database keys configured in the downstream instances. That is to say, each instance except the root node can, according to the sub-database key configured by itself, screen the data corresponding to the sub-database key configured by itself from its upstream instance, and pull the data corresponding to the sub-database key configured by itself to the local of the instance. The data pulled to the local of the instance can also continue to be pulled by other downstream instances.

[0077] It should be noted that each data set published by an instance carries a sub-database identification (ID). According to the value of the sharding key, the sub-database ID where a certain data set is distributed can be calculated.

[0078] According to some embodiments, when controlling a downstream instance to subscribe to a data set corresponding to a sub-database key according to the sub-database key configured in the downstream instance, the sub-database ID corresponding to the sub-database key configured in the downstream instance can be determined; the downstream instance is controlled to subscribe to the data set corresponding to the sub-database ID.

[0079] In some embodiments, the sub-database IDs correspond one-to-one with the data sets.

[0080] For example, the data stored in instance A is divided into 8 data sets. Instance B only needs to obtain the data of the first data set among these 8 data sets. Therefore, the sub-database key configured in instance B is set to id.mod.8_1, and instance B can obtain the data with a sub-database ID of %8 == 1 in instance A.

[0081] It should be noted that the values of the sub-database keys configured in different instances can be the same. For example, multiple downstream instances can all subscribe to all the data stored in the same upstream instance.

[0082] S204, construct an offline bypass at any level of the tree-shaped synchronization system, and upload the data stored in the target instance at any level to the distributed file system through the offline bypass;

[0083] It should be noted that when there is surplus utilization rate in the upstream fan-out but the downstream traffic grows, the fan-out can be increased by expanding instances. However, for the stability of overall data synchronization, the throughput of a single fan-out needs to be strictly restricted. However, it is difficult to achieve rapid expansion under limited throughput. Therefore, it is necessary to build an offline bypass Uploader to pull historical data with high throughput to achieve rapid expansion.

[0084] According to some embodiments, an instance can also be set in the offline bypass. This instance can subscribe to the data stored in the target instance and upload it to an offline distributed file system (DFS) by file.

[0085] In some embodiments, the physical storage resources managed by the file system in the DFS are not necessarily directly connected to the local node, but are connected to the node through a computer network.

[0086] In some embodiments, the target instance refers to any instance in the layer corresponding to the offline bypass.

[0087] S205, publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance;

[0088] For example, when the layer corresponding to the offline bypass is the layer where the leaf nodes are located, the data stored in the distributed file system can be published to the target object. When the layer corresponding to the offline bypass is not the layer where the leaf nodes are located, the data stored in the distributed file system can be published to the downstream instance adjacent to the target instance.

[0089] According to some embodiments, when publishing the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance, the data stored in the distributed file system can be published to the target object or the downstream instance adjacent to the target instance through peer-to-peer technology.

[0090] In some embodiments, peer-to-peer (P2P) technology, also known as peer-to-peer internetworking technology, relies on the computing power and bandwidth of participants in the network rather than concentrating dependencies on a few servers. Therefore, it can improve the efficiency and stability of data transmission.

[0091] It should be noted that the number of data published from the distributed file system to the target object is not limited either, and can be multiple or one.

[0092] S206, synchronize the data stored in the leaf node instance set to multiple target objects.

[0093] According to some embodiments, the data stored in the leaf node instance set can be synchronized to the databases corresponding to each of the multiple target objects.

[0094] Taking a scenario as an example, Figure 4 The schematic architecture diagram of a data synchronization method provided by an embodiment of the present disclosure is shown. As Figure 4 shown, this data synchronization method is applied to the field of advertising placement. The subscription data set is stored in the DB. The tree-level database partitioning is performed on the subscription data set, and a total of X levels of database partitioning are performed. Each instance is set to support 25 subscriptions. Then, there are X * 25 leaf node instances. That is to say, this tree-shaped synchronization system can support X * 25 fanouts, that is, it can perform data synchronization between the target object and at most X * 25 databases.

[0095] In some embodiments, as Figure 4 shown, an offline bypass is set at the first level. The instances in this offline bypass can subscribe to the data published by the instances in the root node and upload the subscribed data to the DFS. Then, the instances in the second level can subscribe to the data stored in the DFS through P2P technology.

[0096] It should be noted that the method provided by the embodiments of the present disclosure can also be applied to search scenarios, in-app information flow scenarios, alliance scenarios, etc. As the volume and daily operation volume corresponding to data synchronization gradually increase, the present disclosure can ensure that the data stored in the subscription data set is synchronously diffused in real time to a large number (e.g., tens of thousands) of cross-regional in-memory database nodes, and the accuracy, throughput, and stability of data synchronization are high.

[0097] In summary, for the method provided in the embodiments of the present disclosure, first, by receiving the subscription data set sent by the database set within the duration threshold; determining the number of levels for sub-database when performing tree-level sub-database on the subscription data set; and performing tree-level sub-database on the subscription data set according to the number of levels through the sub-database key to obtain a tree-shaped synchronization system. Therefore, the sub-database efficiency and effect during tree-level sub-database can be improved, and when using this tree-shaped synchronization system for data synchronization, the throughput and flexibility of data synchronization are high. Next, an offline bypass is constructed at any level of the tree-shaped synchronization system, and the data stored in the target instance at any level is uploaded to the distributed file system through the offline bypass; the data stored in the distributed file system is published to the target object or the downstream instance adjacent to the target instance; therefore, high-throughput pulling of historical data can be supported, and rapid expansion can be achieved. Finally, the data stored in the leaf node instance set is synchronized to multiple target objects; therefore, tree-level sub-database is performed on the subscription data set sent by multiple databases, each upstream node can correspond to multiple downstream nodes, and the data stored in the leaf node instances can all support target object subscription, that is to say, the tree-shaped synchronization system obtained through tree-level sub-database can fan out a large number of instances for target object synchronization. Therefore, data synchronization between a large number of databases and a large number of target objects can be achieved, and the flexibility of data synchronization is high. In addition, in a high-fan-out scenario, the throughput of data synchronization is also relatively high.

[0098] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0099] The following are the embodiments of the present disclosure device, which can be used to execute the embodiments of the present disclosure method. For details not disclosed in the embodiments of the present disclosure device, please refer to the embodiments of the present disclosure method.

[0100] Please refer to Figure 5 , which shows a schematic structural diagram of a data synchronization device provided by an exemplary embodiment of the present disclosure. This data synchronization device can be implemented as all or part of the device through software, hardware, or a combination of both. The data synchronization device 500 includes a data receiving unit 501, a data sub-database unit 502, and a data publishing unit 503, where:

[0101] The data receiving unit 501 is configured to receive the subscription data set sent by the database set within the duration threshold;

[0102] The data sub-library unit 502 is used to perform tree-level sub-library division on the subscription data set to obtain a tree-shaped synchronization system. Among them, an instance is set for each node in the tree-shaped synchronization system. The instance in the root node receives and stores the data in the subscription data set, and the data stored in any upstream instance supports multiple downstream instances adjacent to any upstream instance to subscribe;

[0103] The data publishing unit 503 is used to synchronize the data stored in the leaf node instance set to multiple target objects. Among them, the leaf node instance set includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system.

[0104] According to some embodiments, when the data sub-library unit 502 is used to perform tree-level sub-library division on the subscription data set to obtain a tree-shaped synchronization system, it is specifically used for:

[0105] Determine the number of levels of sub-library division when performing tree-level sub-library division on the subscription data set;

[0106] According to the number of levels, perform tree-level sub-library division on the subscription data set through the sub-library key to obtain a tree-shaped synchronization system.

[0107] According to some embodiments, when the data sub-library unit 502 is used to perform tree-level sub-library division on the subscription data set through the sub-library key, it is specifically used for:

[0108] Divide the data stored in the upstream instance into at least one data set, and the data set corresponds one-to-one with the sub-library key configured in the downstream instance;

[0109] According to the sub-library key configured in the downstream instance, control the downstream instance to subscribe to the data set corresponding to the sub-library key.

[0110] According to some embodiments, when the data sub-library unit 502 is used to control the downstream instance to subscribe to the data set corresponding to the sub-library key according to the sub-library key configured in the downstream instance, it is specifically used for:

[0111] Determine the sub-library identifier corresponding to the sub-library key configured in the downstream instance;

[0112] Control the downstream instance to subscribe to the data set corresponding to the sub-library identifier, where the sub-library identifier corresponds one-to-one with the data set.

[0113] According to some embodiments, when the data sub-library unit 502 is used to determine the number of levels of sub-library division when performing tree-level sub-library division on the subscription data set, it is specifically used for:

[0114] Determine the traffic data corresponding to the instance, and determine the fan-out quota corresponding to the instance according to the traffic data, where the fan-out quota indicates the number of subscriptions supported by any instance;

[0115] In the case where the total fan-out quota is less than the number of databases in the database set, increase the number of levels until the number of databases in the database set is not greater than the total fan-out quota, where the total fan-out quota corresponding to the tree synchronization system is determined by the number of levels and the fan-out quota.

[0116] According to some embodiments, when the data database partitioning unit 502 is used to determine the number of levels of database partitioning during tree-level database partitioning of the subscription data set, it is specifically used for:

[0117] Obtain the single database throughput data input for the instance;

[0118] In the case where the number of databases in the database set is equal to the total fan-out quota and the single database throughput data is higher than the single database throughput data threshold, increase the number of levels.

[0119] According to some embodiments, the data database partitioning unit 502 is further used for:

[0120] Obtain the technical implementation information corresponding to the instance cluster, where the technical implementation information is used to indicate the cluster information adopted when performing life cycle management on any instance;

[0121] Obtain the single instance resource specification information, where the single instance resource specification information is used to indicate the resource allocation information corresponding to the resource specification of any instance;

[0122] Determine the single database throughput data threshold according to the technical implementation information and the single instance resource specification information.

[0123] According to some embodiments, the data synchronization device 500 further includes an offline expansion unit 504, and the offline expansion unit 504 is used for:

[0124] Build an offline bypass at any level of the tree synchronization system, and upload the data stored in the target instance at any level to the distributed file system through the offline bypass;

[0125] Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance.

[0126] According to some embodiments, when the offline expansion unit 504 is used to publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance, it is specifically used for:

[0127] Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance through the peer-to-peer technology.

[0128] It should be noted that: such as Figure 5As shown, the modules that must be included in the data synchronization device 500 are indicated by solid-line boxes, such as the data receiving unit 501, the data sub-library unit 502, and the data publishing unit 503; the modules that may or may not be included in the data synchronization device 500 are indicated by dashed-line boxes, such as the offline expansion unit 504.

[0129] It should be noted that when the data synchronization device provided in the above embodiments executes the data synchronization method, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data synchronization device provided in the above embodiments and the embodiments of the data synchronization method belong to the same concept. The implementation process is shown in the method embodiments and will not be elaborated here.

[0130] The serial numbers of the above embodiments of the present disclosure are only for description and do not represent the advantages or disadvantages of the embodiments.

[0131] In summary, the device provided in the embodiments of the present disclosure performs tree-level sub-library partitioning on the subscription data sets sent by multiple databases. Each upstream node can correspond to multiple downstream nodes, and the data stored in the leaf node instances can all support target object subscriptions. That is to say, the tree-shaped synchronization system obtained through tree-level sub-library partitioning can have a high fan-out of a large number of instances for target objects to synchronize. Therefore, data synchronization between a large number of databases and a large number of target objects can be achieved, and the flexibility of data synchronization is high.

[0132] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.

[0133] According to the embodiments of the present disclosure, the present disclosure also provides a terminal, a readable storage medium, and a computer program product.

[0134] Figure 6 The schematic block diagram of an example terminal 600 that can be used to implement the embodiments of the present disclosure is shown. The terminal is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The terminal can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0135] As Figure 6As shown, the terminal 600 includes a computing unit 601, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 602 or computer programs loaded from a storage unit 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the device 600 can also be stored. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0136] Multiple components in the device 600 are connected to the I / O interface 605, including: an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disc, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0137] The computing unit 601 can be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the computing unit 601 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 executes the various methods and processes described above, such as the data synchronization method. For example, in some embodiments, the data synchronization method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the data synchronization method described above can be executed. Alternatively, in other embodiments, the computing unit 601 can be configured to execute the data synchronization method in any other appropriate way (e.g., by means of firmware).

[0138] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0139] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0140] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0141] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0142] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain network.

[0143] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0144] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.

[0145] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data synchronization method, comprising: Receiving a set of subscription data sent by a database set within a duration threshold; Determining the number of levels for hierarchical database partitioning of the set of subscription data; Performing hierarchical database partitioning on the set of subscription data by a database partitioning key according to the number of levels to obtain a tree-shaped synchronization system, wherein an instance is set for each node in the tree-shaped synchronization system, the instance in the root node receives and stores the data in the set of subscription data, and the data stored in any upstream instance supports subscription by multiple downstream instances adjacent to the any upstream instance; Synchronizing the data stored in the set of leaf node instances to multiple target objects, wherein the set of leaf node instances includes multiple leaf node instances, and the leaf node instances are the instances set in the leaf nodes of the tree-shaped synchronization system; The determining the number of levels for hierarchical database partitioning of the set of subscription data includes: In the case where the total fan-out quota is less than the number of databases in the database set, increasing the number of levels until the number of databases in the database set is not greater than the total fan-out quota, wherein the total fan-out quota corresponding to the tree-shaped synchronization system is determined by the number of levels and the fan-out quota of a single instance, and the fan-out quota of a single instance indicates the number of subscriptions supported by the single instance.

2. The method according to claim 1, wherein The performing hierarchical database partitioning on the set of subscription data by a database partitioning key includes: Dividing the data stored in the upstream instance into at least one data set, and the data sets correspond one-to-one to the database partitioning keys configured in the downstream instances; Controlling the downstream instances to subscribe to the data sets corresponding to the database partitioning keys according to the database partitioning keys configured in the downstream instances.

3. The method according to claim 2, wherein, The controlling the downstream instances to subscribe to the data sets corresponding to the database partitioning keys according to the database partitioning keys configured in the downstream instances includes: Determining the database partition identifier corresponding to the database partitioning key configured in the downstream instance; Controlling the downstream instances to subscribe to the data sets corresponding to the database partition identifier, wherein the database partition identifier corresponds one-to-one to the data sets.

4. The method according to claim 1, wherein, The method further includes: Determining the traffic data corresponding to the instance, and determining the fan-out quota corresponding to the instance according to the traffic data, wherein the traffic data is used to indicate the amount of data provided by the instance within a certain duration.

5. The method according to claim 1, wherein, The determining the number of levels for hierarchical database partitioning of the set of subscription data includes: Obtaining the single database partition throughput data input for the instance; In the case where the number of databases in the database set is equal to the total fan-out quota and the single database partition throughput data is higher than the single database partition throughput data threshold, increasing the number of levels.

6. The method according to claim 5, wherein The method further includes: Obtaining the technical implementation information corresponding to the instance cluster, wherein the technical implementation information is used to indicate the cluster information adopted when performing lifecycle management on any of the instances; Obtaining the single instance resource specification information, wherein the single instance resource specification information is used to indicate the resource allocation information corresponding to the resource specification of any of the instances Determining the single database partition throughput data threshold according to the technical implementation information and the single instance resource specification information.

7. The method according to claim 1, wherein, The method further includes: Construct an offline bypass at any level of the tree-shaped synchronization system, and upload the data stored in the target instance at the any level to the distributed file system through the offline bypass; Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance.

8. The method according to claim 7, wherein The publishing the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance includes: Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance through peer-to-peer technology.

9. A data synchronization device, comprising: A data receiving unit, configured to receive a subscription data set sent by a database set within a duration threshold; A data sub-database unit, configured to determine the number of levels for tree-level sub-database of the subscription data set; perform tree-level sub-database on the subscription data set through a sub-database key according to the number of levels, to obtain a tree-shaped synchronization system, wherein an instance is set for each node in the tree-shaped synchronization system, the instance in the root node receives and stores the data in the subscription data set, and the data stored in any upstream instance supports multiple downstream instances adjacent to the any upstream instance to subscribe; A data publishing unit, configured to synchronize the data stored in the leaf node instance set to multiple target objects, wherein the leaf node instance set includes multiple leaf node instances, and the leaf node instance is an instance set in the leaf node of the tree-shaped synchronization system; When the data sub-database unit is configured to determine the number of levels for tree-level sub-database of the subscription data set, specifically: In the case where the total fan-out quota is less than the number of databases in the database set, increase the number of levels until the number of databases in the database set is not greater than the total fan-out quota, wherein the total fan-out quota corresponding to the tree-shaped synchronization system is determined by the number of levels and the fan-out quota of a single instance, and the fan-out quota of the single instance indicates the number of subscriptions supported by the single instance.

10. The device according to claim 9, wherein, When the data sub-database unit is configured to perform tree-level sub-database on the subscription data set through a sub-database key, specifically: Divide the data stored in the upstream instance into at least one data set, and the data sets correspond one-to-one to the sub-database keys configured in the downstream instances; According to the sub-database keys configured in the downstream instances, control the downstream instances to subscribe to the data sets corresponding to the sub-database keys.

11. The device according to claim 10, wherein, When the data sub-database unit is configured to control the downstream instances to subscribe to the data sets corresponding to the sub-database keys according to the sub-database keys configured in the downstream instances, specifically: Determine the sub-database identifier corresponding to the sub-database key configured in the downstream instance; Control the downstream instances to subscribe to the data sets corresponding to the sub-database identifier, wherein the sub-database identifier corresponds one-to-one to the data set.

12. The device according to claim 9, wherein, The data sub-database unit is further specifically configured to: Determine the traffic data corresponding to the instance, and determine the fan-out quota corresponding to the instance according to the traffic data, wherein the traffic data is used to indicate the amount of data provided by the instance within a certain duration.

13. The device according to claim 9, wherein, When the data sub-library unit is used to determine the number of levels of sub-libraries during the tree-level sub-library division of the subscription data set, it is specifically used for: Obtain the single sub-library throughput data input for the instance; Increase the number of levels when the number of databases in the database set is equal to the total fan-out quota and the single sub-library throughput data is higher than the single sub-library throughput data threshold.

14. The apparatus according to claim 13, wherein, The data sub-library unit is further used for: Obtain the technical implementation information corresponding to the instance cluster, where the technical implementation information is used to indicate the cluster information adopted when performing life cycle management on any one of the instances; Obtain the single-instance resource specification information, where the single-instance resource specification information is used to indicate the resource allocation information corresponding to the resource specification of any one of the instances Determine the single sub-library throughput data threshold according to the technical implementation information and the single-instance resource specification information.

15. The apparatus according to claim 9, wherein, The device further includes an offline expansion unit for: Build an offline bypass at any level of the tree-shaped synchronization system, and upload the data stored in the target instance at the any level to the distributed file system through the offline bypass; Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance.

16. The device according to claim 15, wherein When the offline expansion unit is used to publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance, it is specifically used for: Publish the data stored in the distributed file system to the target object or the downstream instance adjacent to the target instance through peer-to-peer technology.

17. A terminal, comprising: At least one processor; And A memory communicatively connected to the at least one processor; characterized in that The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

19. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1-8 when executed by a processor.

Citation Information

Patent Citations

  • Data transmission method, device and equipment and readable storage medium

    CN113254233A