A method for realizing cross-database type data synchronization based on configuration parameters
Through the method based on configuration parameters, data synchronization and automatic backup recommendations across database types are achieved, which solves the problem of inability to adjust and occupy server resources after backup in the existing technology, and improves backup efficiency and accuracy.
Patent Information
- Application Number
- CN202311765345.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-12-20
AI Technical Summary
The existing method of synchronizing data across database types cannot achieve differential adjustments after backup, occupying server resources affects system use, and real-time backup cannot be achieved, and automatic backup efficiency is low.
Using a configuration parameter-based method, select the corresponding connection driver and parameter configuration by determining the database type, create data synchronization logic, use message queues to backup asynchronously, and complete automatic backup recommendations through synchronous prediction algorithms.
Real-time backup of original business is realized, the workload of enterprise software applications is reduced, the efficiency and accuracy of data synchronization is improved, and efficient and automatic data synchronization is achieved.
Smart Images

Figure CN117786005B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of database data synchronization and provides an automated backup tool for enterprise management software users, thereby greatly improving the office efficiency of enterprise personnel. Specifically, the present invention relates to a method for realizing cross-database type data synchronization based on configuration parameters. Background Art
[0002] Synchronizing data across database types means synchronizing data from one database management system to another database management system, where the two database management systems may use different database types and architectures. This synchronization can be one-way, where only the data from one side is synchronized to the other, or it can be two-way, where the data on both sides are updated synchronously with each other. Synchronizing data across database types requires considering differences in data formats, data types, indexes, and query languages, and requires the use of specialized tools or application programming to implement. This technology is very common in cross-platform applications and enterprise-level applications because it allows data sharing and collaboration between different databases.
[0003] When applying enterprise management software, data backup is an essential part. Many enterprises need to modify the business data of different subsidiaries and departments after backup. Conventional data backup directly backs up the entire database, which cannot achieve differential adjustment. In addition, the backup often occupies server resources, affecting system use. In addition, conventional backup is not real-time.
[0004] For example, Chinese patent 201310719511.4 discloses a method for synchronizing data across database types based on configuration parameters, which includes configuration parameters, reading database data and writing it into memory, reading memory data and synchronizing the database, and is used for synchronizing database data of different data types. However, the above-mentioned method for synchronizing data across database types still has the following shortcomings in specific applications: if the business data of different subsidiaries and departments needs to be modified after backup, the conventional full backup method cannot achieve difference adjustment. Secondly, server resources will be occupied during backup, affecting the normal use of the system. In addition, conventional backup cannot achieve the effect of real-time backup. At the same time, the existing method for synchronizing data does not optimize the automatic synchronization between different databases, resulting in low efficiency of automatic backup.
[0005] Currently, no effective solution has been proposed for the problems in the related technologies. Summary of the invention
[0006] In view of the problems in the related art, the present invention proposes a method for synchronizing data across database types based on configuration parameters to overcome the above technical problems existing in the existing related art.
[0007] To this end, the specific technical solution adopted by the present invention is as follows:
[0008] A method for synchronizing data across database types based on configuration parameters, the method comprising the following steps:
[0009] S1. Determine several databases, and select corresponding database connection drivers and parameter configurations according to different database types, create data synchronization logic, and design system basic tables;
[0010] S2, backing up the database asynchronously based on the message queue and by the backup library execution program;
[0011] S3. Set the synchronization database prediction task and complete the automatic backup recommendation of the database based on the synchronization prediction algorithm.
[0012] Furthermore, the creation of data synchronization logic includes the following steps:
[0013] Read the data in the backup source database, convert the data type and format, and determine the data to be written into the backup destination database;
[0014] Obtain the characteristics and limitations of the backup destination database, and map the data types of the backup source database and the backup destination database;
[0015] Convert the grammatical rules between different databases, convert the data in the backup source database into the data format of the backup destination database, and perform data verification and calibration;
[0016] Set up error handling and exception handling mechanisms in data synchronization tasks to ensure the reliability and stability of data synchronization;
[0017] Test and optimize data synchronization tasks to ensure the correctness and efficiency of data synchronization.
[0018] Furthermore, the design of the system basic table includes the following steps:
[0019] Set the backup source database to have a unique primary key field;
[0020] The backup destination database must have a unique primary key, last modified time, and backup source database ID fields.
[0021] Furthermore, the method of asynchronously backing up the database based on the message queue and by executing a program in the backup library includes the following steps:
[0022] If a database structured query statement is executed, the database structured query statement is sent to a message queue, and the backup library execution program asynchronously receives the database structured query statement from the message queue, and executes the database structured query statement on the backup destination database at the same time;
[0023] Before executing the database structured query statement, the backup database execution program first verifies the last modification time of the data corresponding to the database structured query statement in the backup destination database;
[0024] If the last modification time is the same as the last modification time in the backup source database, the backup database execution program will directly execute the database structured query statement;
[0025] If the last modification time is the same as the last modification time in the backup source database but there is a difference, the backup database execution program performs other operations;
[0026] The other operations include choosing to ignore the database structured query statement and overwriting the data in the backup destination database.
[0027] Furthermore, the setting of the synchronization database prediction task and completing the automatic backup recommendation of the database based on the synchronization prediction algorithm includes the following steps:
[0028] Set the prediction task for the synchronization database and determine the start time of the prediction task;
[0029] Complete automatic database backup recommendations through synchronization prediction algorithms;
[0030] Among them, when setting the synchronization database prediction task and determining the start time of the prediction task, the time points include once, every day and every week. If the start time point of the prediction task is cancelled, the prediction task will no longer be executed, and the prediction task after expiration will no longer be executed.
[0031] Furthermore, the automatic backup recommendation of the database using the synchronization prediction algorithm includes the following steps:
[0032] Obtain a scoring data matrix of the backup destination database to the backup source database, where the number of backup source databases is n and the number of backup destination databases is m;
[0033] Cluster the data in all backup source databases using a clustering algorithm to obtain several data blocks;
[0034] The score matrix of the synchronization task of the backup destination database to the backup source database is calculated, and the score similarity of the synchronization task of two backup destination databases to one backup source database is calculated by cosine similarity, and the synchronization task includes several data blocks;
[0035] Calculate the preference of two backup destination databases for data blocks of synchronization tasks in one backup source database;
[0036] Multiply the score similarity by the preference to get the similarity value of the synchronization task;
[0037] If the similarity value is greater than or equal to a preset threshold, the two backup destination databases are marked as mutually recommended, and if one of the synchronization tasks completes data synchronization to one backup destination database, the data of the synchronization task is automatically recommended to the other backup destination database and marked.
[0038] Furthermore, the calculation of the score matrix of the synchronization task of the backup destination database to the backup source database includes the following steps:
[0039] Obtain the synchronization tasks for data synchronization of each backup source database and the scoring data of each synchronization task of the backup destination database, and calculate the scoring matrix of the synchronization tasks. The number of synchronization tasks for each backup source database is k;
[0040] Multiply the scoring matrix of the synchronization task by the scoring data matrix to obtain the score matrix of the backup destination database for each synchronization task in the backup source database:
[0041]
[0042] In the formula, m is the number of backup destination databases;
[0043] k is the number of synchronization tasks for the backup source database.
[0044] Furthermore, the calculation of the preference of two backup destination databases for data blocks of synchronization tasks in one backup source database includes the following steps:
[0045] Obtain the number of times the backup destination database accesses a data block in the backup source database and the number of synchronizations of the data block;
[0046] Divide the number of synchronizations by the number of accesses to obtain the interest value of the backup destination database for the data block;
[0047] Calculate the preference between the backup destination databases for a data block in the backup source database:
[0048]
[0049] Where, L ua The interest value of database u in data block a for backup purposes;
[0050] Lva is the interest value of the backup destination database v in data block a;
[0051] N is a non-zero natural number.
[0052] Furthermore, clustering the data in all backup source databases by a clustering algorithm to obtain a number of data blocks includes the following steps:
[0053] Randomize the data in any backup source database to obtain several data sets in different orders;
[0054] Cluster the data sets in different orders and obtain several clusters;
[0055] Draw a histogram of the cluster, with the horizontal axis being the attribute items of the data and the vertical axis being the frequency of occurrence of the attribute items;
[0056] The cluster with the largest benefit value is selected as the optimal cluster for this round, and the benefit value calculation formula is:
[0057]
[0058] In the formula, L is the number of clusters, C k is the kth cluster;
[0059] S is the area of the histogram, and W is the width of the histogram;
[0060] r is the rejection factor;
[0061] The optimal clustering is input into the data set of the next iteration, and the corresponding data set is randomly shuffled. At the same time, the cluster with the largest benefit value is selected as the optimal clustering of this round until the benefit value remains unchanged and each cluster is a data block.
[0062] Furthermore, the randomization comprises the following steps:
[0063] For a data set of X records, randomly select a record and swap it with the last record;
[0064] In the data set that has not been randomized, a record is selected and exchanged with the last record. After all records are exchanged, the randomization is completed.
[0065] The beneficial effects of the present invention are:
[0066] (1) The method of realizing cross-database type data synchronization based on configuration parameters of the present invention can realize real-time backup of original business. When the backup content needs to be adjusted and the original data cannot be modified, the content of the backup library can be modified without affecting the original business data. The workload on enterprise software application is greatly reduced.
[0067] (2) The present invention uses a synchronization prediction algorithm to complete automatic database backup recommendation, and can automatically recommend data of a synchronization task to another backup target database with similarity, thereby achieving accurate data recommendation, enabling different departments to obtain the same valuable data, and completing efficient and automatic data synchronization. In addition, in the process of clustering the data in all backup source databases through a clustering algorithm, the present invention uses an improved CLOPE algorithm and randomly disrupts the data, thereby improving the quality of data block clustering.
[0068] (3) The present invention is universal. The background development language can use any mainstream development language, such as Java, Python, C, etc.; the message queue can use message processing tools such as Redis and RabbitMQ; the database supports any relational database such as MySql and PostgrelSQL. Data synchronization is real-time, and the backup source is backed up without feeling, and the impact on the execution efficiency of the system can be ignored. For the difference statement, it can be screened, filtered and processed (ignored or overwritten) the difference content. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0070] Figure 1 The present invention is a flowchart of a method for synchronizing data across database types based on configuration parameters according to an embodiment of the present invention. DETAILED DESCRIPTION
[0071] To further illustrate each embodiment, the present invention provides drawings, which are part of the disclosure of the present invention and are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these contents, ordinary technicians in the field should be able to understand other possible implementations and advantages of the present invention. The components in the figures are not drawn to scale, and similar component symbols are generally used to represent similar components.
[0072] According to an embodiment of the present invention, a method for synchronizing data across database types based on configuration parameters is provided, which includes the table structure design of database business documents, the message queue application method, and the verification mechanism design of data synchronization. This method has a wide range of application needs in the field of enterprise management software applications.
[0073] The present invention is further described with reference to the accompanying drawings and specific embodiments. Figure 1As shown, according to a method for implementing cross-database type data synchronization based on configuration parameters according to an embodiment of the present invention, the method includes the following steps:
[0074] S1. Determine several databases, and select corresponding database connection drivers and parameter configurations according to different database types. Create data synchronization logic and design system basic tables.
[0075] In one embodiment, the creating data synchronization logic comprises the following steps:
[0076] Read the data in the backup source database, convert the data type and format, and determine the data to be written into the backup destination database;
[0077] Obtain the characteristics and limitations of the backup destination database, and map the data types of the backup source database and the backup destination database;
[0078] Convert the grammatical rules between different databases, convert the data in the backup source database into the data format of the backup destination database, and perform data verification and calibration;
[0079] Set up error handling and exception handling mechanisms in data synchronization tasks to ensure the reliability and stability of data synchronization; for example, you can set up retry mechanisms, exception capture, and log operations to deal with possible exceptions.
[0080] Test and optimize data synchronization tasks to ensure the correctness and efficiency of data synchronization. For example, you can compare the data in the source database and the target database to check whether the data synchronization is successful. At the same time, you can optimize the performance of the data synchronization logic to improve the speed and efficiency of data synchronization.
[0081] In one embodiment, the designing of the system basic table includes the following steps:
[0082] Set the backup source database to have a unique primary key field;
[0083] The backup destination database must have a unique primary key, last modified time, and backup source database ID fields.
[0084] S2. The database is backed up asynchronously based on the message queue and executed by the backup library.
[0085] In one embodiment, the method of asynchronously backing up the database based on the message queue and by executing a backup library program includes the following steps:
[0086] If a database structured query statement (SQL) is executed, the database structured query statement is sent to a message queue, and the backup library execution program asynchronously receives the database structured query statement from the message queue, and executes the database structured query statement on the backup destination database at the same time;
[0087] Before executing the database structured query statement, the backup database execution program first verifies the last modification time of the data corresponding to the database structured query statement in the backup destination database;
[0088] If the last modification time is the same as the last modification time in the backup source database, the backup database execution program will directly execute the database structured query statement;
[0089] If the last modification time is the same as the last modification time in the backup source database but there is a difference, the backup database execution program performs other operations;
[0090] The other operations include choosing to ignore the database structured query statement and overwriting the data in the backup target database or performing other processing.
[0091] A message queue is a communication mechanism for delivering messages between applications. It consists of three parts: a message producer (Producer), a message queue (Queue), and a message consumer (Consumer).
[0092] The message producer sends the message to the message queue, which temporarily stores the message and sends it to the message consumer. The message consumer reads the message from the message queue and processes it. The message queue provides a way of asynchronous communication. The producer and consumer do not have to be online at the same time. They can send and receive messages at any time. Message queues usually have the following characteristics: reliability, asynchrony, decoupling, and buffering.
[0093] The main features of the present invention are:
[0094] The application objects of the present invention are software implementation and application personnel.
[0095] The present invention can support different development languages (Java / Python / C) and different databases (MySQL / PostgreSQL).
[0096] The structure of all database tables is shown in Table 1:
[0097] Table 1 All table structures of the database
[0098] Field Name type Remark ID Integer Unique primary key LAST_UPDATE_TIME Datetime Data last modified time SYNC_ID Integer Backup source ID
[0099] Synchronous initiation is executed when the system interacts with the database, and the interactive SQL is sent to the message queue.
[0100] The synchronous execution program obtains the execution SQL through the message queue and executes the SQL in real time in the backup library. The specific code and comments are as follows (Python):
[0101]
[0102]
[0103]
[0104] S3. Set the synchronization database prediction task and complete the automatic backup recommendation of the database based on the synchronization prediction algorithm.
[0105] In one embodiment, the step of setting a synchronization database prediction task and completing automatic database backup recommendation based on a synchronization prediction algorithm includes the following steps:
[0106] Set the prediction task for the synchronization database and determine the start time of the prediction task;
[0107] Complete automatic database backup recommendations through synchronization prediction algorithms;
[0108] Among them, when setting the synchronization database prediction task and determining the start time of the prediction task, the time points include once, every day and every week. If the start time point of the prediction task is cancelled, the prediction task will no longer be executed, and the prediction task after expiration will no longer be executed.
[0109] In one embodiment, the automatic backup recommendation of the database using the synchronization prediction algorithm includes the following steps:
[0110] Obtain a scoring data matrix of the backup target database to the backup source database, and the scoring data is set to be one to five points, and the number of the backup source databases is n, and the number of the backup target databases is m;
[0111] Cluster the data in all backup source databases using a clustering algorithm to obtain several data blocks;
[0112] The score matrix of the synchronization task of the backup destination database to the backup source database is calculated, and the score similarity of the synchronization task of two backup destination databases to one backup source database is calculated by cosine similarity, and the synchronization task includes several data blocks;
[0113] Calculate the preference of two backup destination databases for data blocks of synchronization tasks in one backup source database;
[0114] Multiply the score similarity by the preference to get the similarity value of the synchronization task;
[0115] If the similarity value is greater than or equal to a preset threshold, the two backup destination databases are marked as mutually recommended, and if one of the synchronization tasks completes data synchronization to one backup destination database, the data of the synchronization task is automatically recommended to the other backup destination database and marked.
[0116] In one embodiment, the step of calculating the score matrix of the synchronization task of the backup destination database to the backup source database comprises the following steps:
[0117] Get the synchronization tasks for data synchronization of each backup source database and the scoring data of each synchronization task of the backup destination database (scored by professionals), and calculate the scoring matrix of the synchronization tasks. The number of synchronization tasks for each backup source database is k;
[0118] Multiply the scoring matrix of the synchronization task by the scoring data matrix to obtain the score matrix of the backup destination database for each synchronization task in the backup source database:
[0119]
[0120] In the formula, m is the number of backup destination databases;
[0121] k is the number of synchronization tasks for the backup source database.
[0122] In one embodiment, the step of calculating the preference of two backup destination databases for data blocks of synchronization tasks in one backup source database comprises the following steps:
[0123] Obtain the number of times the backup destination database accesses a data block in the backup source database and the number of synchronizations of the data block;
[0124] Divide the number of synchronizations by the number of accesses to obtain the interest value of the backup destination database for the data block;
[0125] Calculate the preference between the backup destination databases for a data block in the backup source database:
[0126]
[0127] Where, L ua The interest value of database u in data block a for backup purposes;
[0128] Lva is the interest value of the backup destination database v in data block a;
[0129] N is a non-zero natural number.
[0130] In one embodiment, clustering the data in all backup source databases by using a clustering algorithm to obtain a number of data blocks includes the following steps:
[0131] Randomize the data in any backup source database to obtain several data sets in different orders;
[0132] Cluster the data sets in different orders and obtain several clusters;
[0133] Draw a histogram of the cluster, with the horizontal axis being the attribute items of the data and the vertical axis being the frequency of occurrence of the attribute items;
[0134] The cluster with the largest benefit value is selected as the optimal cluster for this round, and the benefit value calculation formula is:
[0135]
[0136] In the formula, L is the number of clusters, C k is the kth cluster;
[0137] S is the area of the histogram, and W is the width of the histogram;
[0138] r is the exclusion factor, the larger r is, the more clusters there are;
[0139] The optimal clustering is input into the data set of the next iteration, and the corresponding data set is randomly shuffled. At the same time, the cluster with the largest benefit value is selected as the optimal clustering of this round until the benefit value remains unchanged and each cluster is a data block.
[0140] The property items include:
[0141] Numeric attributes: including real numbers, integers, decimals, etc. Categorical attributes: including discrete values, nominal values, etc. Temporal attributes: including date, time, timestamp, etc. Text attributes: including natural language text, code, etc. Geographical location attributes: including longitude, latitude, etc. Image, video, etc.
[0142] In one embodiment, the randomization comprises the following steps:
[0143] For a data set of X records, randomly select a record and swap it with the last record;
[0144] In the data set that has not been randomized, a record is selected and exchanged with the last record. After all records are exchanged, the randomization is completed.
[0145] In summary, a method for realizing cross-database type data synchronization based on configuration parameters of the present invention can realize real-time backup of original business. When the backup content needs to be adjusted and the original data cannot be modified, the content of the backup library can be modified without affecting the original business data. The workload on enterprise software applications is greatly reduced. The present invention completes the automatic backup recommendation of the database through the synchronization prediction algorithm, and can automatically recommend the data of a synchronization task to another backup destination database with similarity, realizes accurate data recommendation, enables different departments to obtain the same valuable data, and completes efficient and automatic data synchronization. In the process of clustering the data in all backup source databases through the clustering algorithm, the present invention adopts the improved CLOPE algorithm, and randomly disrupts the data, thereby improving the quality of data block clustering. The present invention has versatility, and the background development language can use any mainstream development language, such as java, python, C, etc.; the message queue can use message processing tools such as Redis and RabbitMQ; the database supports any relational database such as MySql, PostgrelSQL, etc. Data synchronization is real-time, and the backup source is realized without sense of backup, and the impact on the system execution efficiency can be ignored. For difference statements, you can filter, filter and process (ignore or overwrite) the difference content.
[0146] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for synchronizing data across database types based on configuration parameters. It is characterized in that The method comprises the following steps: S1. Determine several databases, and select corresponding database connection drivers and parameter configurations according to different database types, create data synchronization logic, and design system basic tables; S2, backing up the database asynchronously based on the message queue and by the backup library execution program; S3. Set the synchronization database prediction task and complete the automatic backup recommendation of the database based on the synchronization prediction algorithm; Among them, the automatic backup recommendations for the database completed through the synchronization prediction algorithm include: The score matrix of the synchronization task of the backup destination database to the backup source database is calculated, and the score similarity of the synchronization task of two backup destination databases to one backup source database is calculated by cosine similarity, and the synchronization task includes several data blocks; Calculate the preference of two backup destination databases for the data blocks of the synchronization task in one backup source database; Multiply the score similarity by the preference to get the similarity value of the synchronization task; If the similarity value is greater than or equal to a preset threshold, the two backup destination databases are marked as mutually recommended, and if one of the synchronization tasks completes data synchronization to one backup destination database, the data of the synchronization task is automatically recommended to the other backup destination database and marked.
2. A method for synchronizing data across database types based on configuration parameters according to claim 1, It is characterized in that The creation of data synchronization logic includes the following steps: Read the data in the backup source database, convert the data type and format, and determine the data to be written into the backup destination database; Obtain the characteristics and limitations of the backup destination database, and map the data types of the backup source database and the backup destination database; Convert the grammatical rules between different databases, convert the data in the backup source database into the data format of the backup destination database, and perform data verification and calibration; Set up error handling and exception handling mechanisms in data synchronization tasks to ensure the reliability and stability of data synchronization; Test and optimize data synchronization tasks to ensure the correctness and efficiency of data synchronization.
3. A method for synchronizing data across database types based on configuration parameters according to claim 1, It is characterized in that The design of the system basic table includes the following steps: Set the backup source database to have a unique primary key field; The backup destination database must have a unique primary key, last modified time, and backup source database ID fields.
4. A method for synchronizing data across database types based on configuration parameters according to claim 1, It is characterized in that The method of asynchronously backing up the database based on the message queue and by executing a backup library program comprises the following steps: If a database structured query statement is executed, the database structured query statement is sent to a message queue, and the backup library execution program asynchronously receives the database structured query statement from the message queue, and executes the database structured query statement on the backup destination database at the same time; Before executing the database structured query statement, the backup database execution program first verifies the last modification time of the data corresponding to the database structured query statement in the backup destination database; If the last modification time is the same as the last modification time in the backup source database, the backup database execution program will directly execute the database structured query statement; If the last modification time is the same as the last modification time in the backup source database but there is a difference, the backup database execution program performs other operations; The other operations include choosing to ignore the database structured query statement and overwriting the data in the backup destination database.
5. A method for synchronizing data across database types based on configuration parameters according to claim 1, It is characterized in that The method further includes the following steps before completing the automatic backup recommendation of the database by the synchronous prediction algorithm: Set the prediction task for the synchronization database and determine the start time of the prediction task; Among them, when setting the synchronization database prediction task and determining the start time of the prediction task, the time points include once, every day and every week. If the start time point of the prediction task is cancelled, the prediction task will no longer be executed, and the prediction task after expiration will no longer be executed.
6. A method for synchronizing data across database types based on configuration parameters according to claim 5, It is characterized in that Before calculating the score matrix of the synchronization task of the backup destination database to the backup source database, the following steps are also included: Get the score data matrix of the backup destination database to the backup source database. The number of backup source databases is n , the number of backup destination databases is m ; Cluster the data in all backup source databases using a clustering algorithm to obtain several data blocks; The step of calculating the score matrix of the synchronization task of the backup destination database to the backup source database comprises the following steps: Get the synchronization tasks for data synchronization of each backup source database and the scoring data of each synchronization task of the backup destination database, and calculate the scoring matrix of the synchronization tasks. The number of synchronization tasks for each backup source database is k ; Multiply the scoring matrix of the synchronization task by the scoring data matrix to obtain the score matrix of the backup destination database for each synchronization task in the backup source database: In the formula, m The number of databases for backup purposes; k The number of synchronization tasks for backing up the source database; The calculation to obtain the preference of two backup destination databases for data blocks of synchronization tasks in one backup source database comprises the following steps: Obtain the number of times the backup destination database accesses a data block in the backup source database and the number of synchronizations of the data block; Divide the number of synchronizations by the number of accesses to obtain the interest value of the backup destination database for the data block; Calculate the preference between the backup destination databases for a data block in the backup source database: In the formula, L ua Database for backup purposes u For data blocks a Interest value; LV Database for backup purposes v For data blocks a Interest value; N is a non-zero natural number.
7. A method for synchronizing data across database types based on configuration parameters according to claim 6, It is characterized in that The method of clustering the data in all backup source databases by using a clustering algorithm to obtain a number of data blocks includes the following steps: Randomize the data in any backup source database to obtain several data sets in different orders; Cluster the data sets in different orders and obtain several clusters; Draw a histogram of the cluster, with the horizontal axis being the attribute items of the data and the vertical axis being the frequency of occurrence of the attribute items; The cluster with the largest benefit value is selected as the optimal cluster for this round, and the benefit value calculation formula is: In the formula, L is the number of clusters, C k For the k clusters; S is the area of the histogram, W is the width of the histogram; r is the rejection factor; The optimal clustering is input into the data set of the next iteration, and the corresponding data set is randomly shuffled. At the same time, the cluster with the largest benefit value is selected as the optimal clustering of this round until the benefit value remains unchanged and each cluster is a data block.
8. A method for synchronizing data across database types based on configuration parameters according to claim 7, It is characterized in that The randomization includes the following steps: For a data set of X records, randomly select a record and swap it with the last record; In the data set that has not been randomized, a record is selected and exchanged with the last record. After all records are exchanged, the randomization is completed.
Citation Information
Patent Citations
A method for synchronizing data across database types based on configuration parameters
CN103699638B
Block storage adaptive backup system based on cloud environment
CN114020539A
Data synchronization method based on kettle and database logs
CN114036119A