Data synchronization system and method, data backup method and electronic equipment
By introducing synchronization middleware and reserved plug-in interfaces into the data synchronization system, the problem of data synchronization complexity between heterogeneous data sources is solved, and automated data synchronization is achieved, reducing development costs and resource consumption.
Patent Information
- Application Number
- CN202311617909.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-29
- Publication Date
- 2025-05-30
AI Technical Summary
Different business systems have different database structures, which leads to data synchronization between heterogeneous data sources that require dependence on specific tools or development, which increases the work burden of developers.
It provides a data synchronization system, including source-side database, destination-side database and synchronization middleware. The synchronization middleware includes scheduling module and working module. It adapts to any database through the reserved plug-in interface to achieve automated data synchronization.
Reduces the dependence of the data synchronization process between databases on specific tools, removes the limitations of the data synchronization process on the database data structure, reduces the work burden of developers, and improves resource utilization.
Smart Images

Figure CN120067204A_ABST
Abstract
Description
Background Art
[0002] A large amount of data is generated during the production and manufacturing process. The data of different business systems is stored in different databases. With the continuous expansion of business requirements, it is necessary to synchronize data between the databases corresponding to different business systems, or to reduce the database pressure of the production system, it is necessary to perform regular backups on the data. However, the structures of the databases of different business systems are different, and data synchronization between heterogeneous data sources requires relying on specific synchronization tools or corresponding development, which greatly increases the workload of developers.
[0003] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0004] The purpose of the present disclosure is to provide a data synchronization system and method, a data backup method, and an electronic device.
[0005] Other features and advantages of the present disclosure will become apparent through the following detailed description, or will be partially learned through the practice of the present disclosure.
[0006] According to a first aspect of the present disclosure, a data synchronization system is provided, including: at least one source database, at least one destination database, and a synchronization middleware, where the synchronization middleware is respectively connected to at least one source database and at least one destination database;
[0007] The synchronization middleware includes: a scheduling module and a working module. The scheduling module is configured to, according to the configured target task and the execution conditions of the target task, when the current execution conditions are met, schedule the working module to execute the target task based on the configuration class of the target task, and the configuration class includes the working configuration information of the target task;
[0008] The working module is configured to, according to the working configuration information, start corresponding working threads to perform data synchronization of the data to be synchronized associated with the target task between at least one source database and the corresponding at least one destination database; the at least one source database and the at least one destination database are respectively configured through the reserved plug-in interfaces of the synchronization middleware.
[0009] Optionally, the scheduling module includes a trigger, and the trigger is configured to configure the execution conditions; the scheduling module is configured to bind the trigger of the target task and the corresponding configuration class;
[0010] The configuration class scheduling based on the target task schedules the work module to execute the target task, including: when the execution condition is currently met, triggering the work module to execute the target task based on the configuration class through the trigger.
[0011] Optionally, the work configuration information includes the primary key sharding information of the table data, and the work module includes: a task splitting unit, configured to split the target task into multiple subtasks according to the sharding strategy of at least one source database or the primary key sharding information of the table data, and each subtask is used to perform data synchronization of different parts of the source data.
[0012] Optionally, the work configuration information includes the channel concurrency number, and the work module includes:
[0013] a subtask management unit, configured to form multiple subtasks into a task group according to the channel concurrency number, and manage the execution process of each subtask based on the task group, where the channel concurrency number is the same as the concurrent thread number of the task group.
[0014] Optionally, the work module corresponding to the subtask includes:
[0015] a data collector, configured to collect corresponding source data from the source database and write the source data into a data converter;
[0016] a data converter, configured to perform data conversion processing on the source data to obtain corresponding target data;
[0017] a data writer, configured to read the target data from the data converter and send the target data to the corresponding destination database.
[0018] Optionally, the work configuration information includes a flow control mode, and the data converter is configured to control the speed of the data synchronization according to the flow control mode.
[0019] Optionally, the data converter is used to:
[0020] determine the resource occupancy corresponding to the current data flow rate;
[0021] in response to the resource occupancy exceeding a preset threshold, trigger the current working thread to enter a sleep state, where the sleep time of the current working thread is determined based on the current data flow rate corresponding to the flow control mode and the preset flow rate.
[0022] Optionally, the working configuration information includes a data writing mode, and the data converter is configured to trigger and execute a corresponding target execution statement according to the data writing mode, where the target execution statement is used to instruct the destination database to delete historical data corresponding to the data to be written under the processing conditions corresponding to the data writing mode.
[0023] Optionally, the working module further includes:
[0024] A data monitoring unit, configured to collect at least one of processing performance data, synchronization failure data, and processing progress data during the execution of the target task in real time, and record and display the data collected in real time in a log.
[0025] Optionally, the data monitoring unit is further configured to monitor dirty data generated during the data synchronization process; the dirty data is error data generated during the data synchronization process; in response to the data proportion of the dirty data being greater than a preset threshold, an alarm message is generated.
[0026] According to a second aspect of the present disclosure, a data backup method is provided, including:
[0027] Obtain production data of each production link;
[0028] In response to currently meeting a preset data backup condition, send the data to be backed up in the production database to the target database through a synchronization middleware to complete the production data backup; the synchronization middleware includes: a scheduling module and a working module, where the scheduling module is configured to, when the data backup condition is reached currently, schedule the working module based on a configuration class to back up the data to be backed up to the target database; the production database and the target database are respectively configured through a reserved plug-in interface of the synchronization middleware.
[0029] Optionally, the method further includes:
[0030] Determine the number of backup data records backed up to the target database;
[0031] Perform consistency verification on the current data backup according to the number of backup data records and the number of data records to be backed up;
[0032] In response to the verification failing, re-perform data backup on the restored data, where the restored data is obtained by processing the data to be backed up according to the failure reason, and the failure reason is determined according to a preset rule and the consistency verification result.
[0033] According to a third aspect of the present disclosure, there is provided a data synchronization method applied to a synchronization middleware, including: in response to reaching an execution condition configured for a target task, collecting source data corresponding to the target task from at least one source database associated with the target task;
[0034] Performing data conversion processing on the source data according to the data storage attributes of at least one destination database associated with the target task to obtain corresponding target data;
[0035] Sending the corresponding target data to at least one destination database to complete data synchronization.
[0036] Wherein, the at least one source database and the at least one destination database are respectively configured through reserved plug-in interfaces of the synchronization middleware.
[0037] Optionally, the data storage attribute includes a data field attribute. Performing data conversion processing on the source data includes: matching the field attributes of the source data according to the data field attributes of at least one of the destination databases to obtain the target data for writing to the corresponding destination database.
[0038] Optionally, performing data conversion processing on the source data further includes: performing at least one of desensitization processing, completion processing, and data filtering on the source data.
[0039] Optionally, after completing data synchronization, the method further includes:
[0040] Performing consistency verification on the synchronization data of at least one destination database associated with the target task;
[0041] In response to the verification failing, triggering the re-execution of the target task.
[0042] According to a fourth aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the program is executed by a processor, it implements the data synchronization method in the above-mentioned embodiment.
[0043] According to a fifth aspect of the present disclosure, there is provided an electronic device, including: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the data synchronization method in the above-mentioned embodiment.
[0044] It can be seen from the above technical solutions that the data synchronization system and method in the exemplary embodiments of the present disclosure at least have the following advantages and positive effects:
[0045] The data synchronization system of the present disclosure, on the one hand, can configure the corresponding databases of both parties for data synchronization through the reserved plug-in interface of the synchronization middleware, which can adapt to any database, greatly reducing the dependence on specific tools during the data synchronization process between databases and lifting the restrictions on the data structure of the databases during the data synchronization process. On the other hand, when a new database needs to be connected, only the adaptation development of the new database for the reserved plug-in interface needs to be carried out, which can reduce the workload of developers and save resources. On the other hand, the synchronization middleware includes a scheduling module and a working module. The scheduling module can, according to the configured target task and the execution conditions of the target task, when the current execution conditions are met, schedule the working module to execute the target task based on the configuration class of the target task; the working module can, according to the working configuration information, start a working thread to perform data synchronization between at least one source database and the corresponding at least one destination database. In this way, users can realize the automatic data synchronization process of the data to be synchronized through the target task execution conditions and the configuration class, without personnel monitoring and manual triggering, saving human resources, and can realize the rational allocation of network resources and hardware resources through the arrangement of the execution conditions of different data synchronization processes, improving resource utilization.
[0046] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.
[0048] Figure 1 FIG. 1 schematically shows one of the schematic diagrams of the data synchronization system according to an embodiment of the present disclosure.
[0049] Figure 2 FIG. 2 schematically shows another schematic diagram of the data synchronization system according to an embodiment of the present disclosure.
[0050] Figure 3 FIG. 3 schematically shows a schematic diagram of the working module of a subtask according to an embodiment of the present disclosure.
[0051] Figure 4 FIG. 4 schematically shows a third schematic diagram of the data synchronization system according to an embodiment of the present disclosure.
[0052] Figure 5 FIG. 5 schematically shows the data synchronization execution process of the synchronization middleware according to an embodiment of the present disclosure.
[0053] Figure 6 A schematic diagram showing a data synchronization method according to an embodiment of the present disclosure is schematically illustrated.
[0054] Figure 7 A schematic diagram showing a data backup method according to an embodiment of the present disclosure is schematically illustrated.
[0055] Figure 8 A schematic diagram of modules of an electronic device according to an embodiment of the present disclosure is schematically illustrated. Detailed implementation manners
[0056] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.
[0057] In addition, the described features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that the technical solutions of the present disclosure can be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. can be used. In other cases, well-known methods, devices, implementations, or operations are not shown or described in detail to avoid obscuring aspects of the present disclosure.
[0058] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0059] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the contents and operations / steps, nor do they necessarily have to be executed in the order described. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.
[0060] The present disclosure is mainly directed to scenarios where data synchronization or data backup is required between different databases. For example, in the semiconductor manufacturing field, discrete industrial manufacturing field, or warehouse item management field, etc., data synchronization and data backup are required. For example, for the Advanced Planning and Scheduling (APS) system in the industrial manufacturing field, it needs to collect sales demand data, process route data, inventory, equipment, etc. data from systems such as CIM (Computer Integrated Manufacturing) and ERP (Enterprise Resource Planning), and perform logical processing on the collected raw data to generate scheduling data, sales forecasts, material consumption, etc. data. And during the industrial production process, a large amount of important data is generated every day, which will exert a greater pressure on the data storage end, and data backup needs to be performed regularly to reduce the pressure on the production library storage end and ensure data security. The application of the above APS system is only an exemplary introduction to the application scenario of the present disclosure, and the present disclosure can also be applied to other scenarios where data backup or data synchronization is required. The present disclosure places no restrictions on the application scenario.
[0061] In an embodiment of the present disclosure, a data synchronization system is proposed, as Figure 1 shown. The data synchronization system at least includes: at least one source database 110 (such as databases A1 to An in the figure), at least one destination database 120 (such as databases B1 to Bm in the figure), and a synchronization middleware 130. The synchronization middleware 130 is respectively connected to at least one source database 110 and at least one destination database 120;
[0062] The synchronization middleware 130 includes: a scheduling module 131 and a working module 132. The scheduling module 131 is used to, according to the configured target task and the execution conditions of the target task, when the current execution conditions are met, schedule the working module to execute the target task based on the configuration class of the target task. The configuration class includes the working configuration information of the target task;
[0063] The working module is used to, according to the working configuration information, start corresponding working threads to perform data synchronization of the data to be synchronized associated with the target task between at least one source database and the corresponding at least one destination database; the at least one source database and the at least one destination database are respectively configured through the reserved plug-in interfaces of the synchronization middleware.
[0064] The above Figure 1The data synchronization system of the embodiment can configure the corresponding databases on both sides of data synchronization through the reserved plug-in interface of the synchronization middleware, and can adapt to any database, greatly reducing the dependence on specific tools during the data synchronization process between databases and removing the restriction on the data structure of the database during the data synchronization process. On the other hand, when a new database needs to be connected, only the adaptability development of the new database for the reserved plug-in interface is required, which can reduce the workload of developers and save resources. On the further hand, the synchronization middleware includes a scheduling module and a working module. The scheduling module can, according to the configured target task and the execution conditions of the target task, when the current execution conditions are met, schedule the working module to execute the target task based on the configuration class of the target task; the working module can, according to the working configuration information, start a working thread to perform data synchronization between at least one source database and the corresponding at least one destination database. In this way, the user can realize the automatic data synchronization process of the data to be synchronized through the target task, execution conditions and configuration class, without personnel monitoring and manual triggering, saving human resources, and can realize the rational allocation of network resources and hardware resources through the arrangement of the execution conditions of different data synchronization processes, improving resource utilization rate.
[0065] To make the technical solution of the present disclosure clearer, the following will explain each part of the data synchronization system.
[0066] As Figure 2 shown, the data synchronization system of the present disclosure includes at least one source database 220 and at least one destination database 230 connected through a synchronization middleware 210. A star structure can be formed between each source database (such as source databases 1 to 4 in the figure) and each destination database (such as destination databases 1 to 4 in the figure) through the synchronization middleware. The database with any data structure can be configured as the source and destination of data synchronization through the reserved plug-in interface of the synchronization middleware, so that data synchronization can be realized between databases with various different data structures, reducing the dependence on specific tools (such as DBLink) during the data synchronization process and improving the flexibility of the data synchronization process.
[0067] In this exemplary embodiment, the source database and the destination database can be various heterogeneous data sources. For example, the source database and the destination database can be relational databases (such as MySQL, Oracle, SQL Server, etc.), HDFS (Hadoop Distributed File System), ODPS (Open Data Processing Service), HBase, etc. This example is not limited thereto. The synchronization middleware can be an intermediate component for communicating with the source database and the destination database, and can be arranged in an electronic device (such as a computer, a server) capable of communicating with the source database and the destination database.
[0068] The synchronization middleware 210 can include: a scheduling module and a working module. The scheduling module is used to schedule the working module to execute the target task when the current execution conditions of the configured target task and the target task are met.
[0069] In this exemplary embodiment, the target task can be a data synchronization task, a data update task, a data backup task, etc. The target task can be a synchronization task for one or more table data, and can include the source database and the source data information to be synchronized (such as the source data table name, table identifier, etc.), the destination database. The execution conditions can include the execution time and execution frequency, etc. When the current execution time or execution frequency is reached, the working module is automatically scheduled to execute the target task. Various configurations can be made for the target task on the configuration page or configuration table of the target task, and relevant configurations of the target task can also be made in the configuration class (such as JobDetail) of the target task. The scheduling module is responsible for the progress and schedule management of each task. For example, the scheduling module can use Quartz. Quartz is an open-source task scheduling management system that can perform task timing, regular, or calendar scheduling; it has good compatibility and can be compatible with the smallest standard Java program to the largest e-commerce system. For example, relevant configurations of the target task can be made through JobDetail of Quartz, and detailed information (relevant configurations of the target task) of JobDetails can be obtained through the GetJobDetails method during the task execution process.
[0070] For the working module, it is used to start a working thread to perform data synchronization between at least one source database and the corresponding at least one destination database according to the working configuration information.
[0071] In this exemplary embodiment, the working module is used to complete the data synchronization work of the source data corresponding to the target task, and one working module (Job module) can correspond to one task. Exemplarily, the scheduling module includes a trigger, and the trigger is used to configure the execution conditions (such as execution time or execution frequency, etc.); the scheduling module is used to bind the trigger of the target task and the corresponding configuration class; by binding the trigger and the corresponding configuration class, when the current execution condition corresponding to the trigger is reached (i.e., when the trigger is triggered), the working module executes the target task based on the configuration class (i.e., executes the job corresponding to the configuration class). In this example, through the binding of the configuration class and the trigger, when the trigger is triggered (the execution condition is reached), the working module can automatically execute the thread (job) corresponding to the bound configuration class, realizing the automatic and accurate execution of the synchronization task and avoiding manual monitoring and triggering.
[0072] In this example, the scheduling module can trigger the execution of the working module through the trigger, and the working module can start the working thread to execute the corresponding task. The working module is the central management node in the execution process of the target task and is used to manage each data processing process (such as data cleaning, task splitting, task grouping, data conversion, etc.) in the execution process of the target task. The working configuration information refers to the configuration information for the execution process of the target task, which can include the configuration information for the data cleaning process, the configuration information for the task splitting process, the configuration information for the task grouping process, and the configuration information for the data conversion process, etc.
[0073] Exemplarily, the working configuration information includes the primary key sharding information of the table data, and the working module includes: a task splitting unit, which is used to split the target task into multiple subtasks according to the splitting strategy of at least one source database or the primary key sharding information of the table data, and each subtask is used to execute the data synchronization of different parts of the source data.
[0074] In this exemplary embodiment, the primary key sharding information of the corresponding table data can be configured for the target task, that is, the primary key sharding information of the table data to be synchronized. Through the primary key sharding information, the rapid positioning of the table data can be realized, and the target task can be split based on the primary key sharding information to form multiple subtasks. Each subtask is used to execute a part of the data synchronization process of the source data. In this example, through task splitting, multiple subtasks can be executed concurrently to improve the synchronization efficiency. This example can also split tasks based on the splitting strategy of the source database to realize the rapid positioning of the source data and avoid the cross-sharding execution of the same subtask.
[0075] Exemplarily, the working configuration information includes the channel concurrency number, and the working module further includes: a subtask management unit, which is used to form multiple subtasks into a task group according to the channel concurrency number and manage the execution process of each subtask based on the task group. The channel concurrency number is the same as the concurrent thread number of the task group.
[0076] In this exemplary embodiment, for the convenience of managing multiple subtasks, the working module further includes a subtask management unit. The subtask management unit is used to form a task group from multiple subtasks according to the configured channel concurrency number, and manage the execution process of each subtask in the group based on the task group. The channel concurrency number is equal to the concurrent thread number of the group.
[0077] For example, after splitting multiple subtasks (Tasks), the working module can call the Scheduler module to recombine the split Tasks according to the configured channel concurrency number and assemble them into a TaskGroup (task group). For example, for a table with 1 million data records, if the primary key sharding is configured to be 10, then 100,000 records are divided into one Task, resulting in 10 tasks. If the channel concurrency number is configured to be 5, then 5 tasks are assembled into one TaskGroup, and a total of 2 TaskGroups are assembled. Each TaskGroup is responsible for running all the allocated Tasks concurrently. That is, each TaskGroup can wait for all the subtasks in the group to run to completion and consider the task execution finished. Through the TaskGroup, it is convenient to manage the subtasks in groups.
[0078] In some embodiments, as Figure 3 shown, the working module corresponding to the subtask can be designed to include:
[0079] A data collector 310, which is used to collect the corresponding source data from the source database and write the source data into the data converter;
[0080] A data converter 320, which is used to perform data conversion processing on the source data to obtain the corresponding target data;
[0081] A data writer 330, which is used to read the target data from the data converter and send the target data to the corresponding destination database.
[0082] In this exemplary embodiment, each subtask can be completed by a corresponding data collector (Reader), data converter (Translator), and data writer (Writer). The data collector can have a first reserved plug-in interface for configuring at least one source database corresponding to the target task. The data writer can have a second reserved plug-in interface for configuring at least one destination database corresponding to the target task. By performing corresponding configuration development of the source database and destination database in the data collector and data writer, data synchronization can be achieved between the configured source database and destination database. After the database is configured, the data collector can collect the source data corresponding to the current subtask from the source data storage DB and write the source data to the data converter corresponding to the current subtask. The data converter is used to connect the data collector and the data writer and can serve as a data transmission channel between the data collector and the data writer. It can manage a Channel object internally, and the Channel object can manage a MemoryChannel instance. The data collector and the data writer can achieve data writing from the data collector to the MemoryChannel and data reading from the MemoryChannel by holding references to the same MemoryChannel instance, and then write the data to the target DB of the destination database, thereby realizing data transmission between the source and the destination. The data converter can perform data conversion processing on the source data according to the target task to obtain target data that can be used to write to the destination data. For example, if the source database includes attribute information of A and B, and the destination database includes attribute information of A, the data converter can delete the corresponding B attribute and retain the A attribute to obtain target data that only retains the A attribute information. The data converter can also perform other processing such as transformation and update, and insertion of the attribute information of the table data. This example does not limit this.
[0083] In the above Figure 3 In the technical solution of the above embodiment, through the design of the data collector, data converter, and data writer for the subtask, a transmission data stream corresponding to the subtask is formed between the source and the destination, realizing the conversion and transmission process of the table data from large to small, facilitating the control and management of the synchronization process at the subtask granularity, and improving the controllability of the synchronization process.
[0084] In some embodiments, the working configuration information includes a flow control mode, and the data converter is used to control the speed of data synchronization according to the flow control mode.
[0085] In this exemplary embodiment, based on Figure 3The sub-task design shown can perform flow control on the synchronization process. The flow control mode can be configured in the configuration table of the target task. The flow control mode can include a first flow control mode based on channels (concurrency), a second flow control mode based on record flow, and a third flow control mode based on byte flow. Corresponding flow control thresholds can be configured for each flow control mode. When the data stream written by the data converter exceeds the corresponding flow control threshold, the data stream of the synchronization process is reduced or controlled to ensure the normal progress of other processing processes and avoid resource preemption problems. In this example, the flow control thresholds corresponding to each flow control mode can be the same or different, and this example does not limit this.
[0086] Exemplarily, the data flow rate can be controlled through the following steps: Determine the resource occupancy corresponding to the current data flow rate; in response to the resource occupancy exceeding a preset threshold, trigger the current working thread to enter the sleep state.
[0087] In this exemplary embodiment, a mapping relationship can be established between the network resource occupancy rate (such as bandwidth occupancy or computing resource occupancy) and the data flow rate. Look up the resource occupancy corresponding to the current data flow rate in the mapping relationship. When the resource occupancy exceeds the corresponding threshold, trigger the working thread corresponding to the current working module to enter the sleep state. The sleep state refers to the state where data processing is not performed. The sleep time for the current working thread to enter the sleep state can be determined based on the current data flow rate corresponding to the flow control mode and the preset flow rate. For example, sleep time = current byteSpeed or recordSpeed × interval time / corresponding threshold - interval time. The interval time is determined according to the data flow rate unit. For example, if the data flow rate unit is bps, the interval time is 1s.
[0088] In some embodiments, the work configuration information includes the data writing mode, and the data converter is used to trigger the execution of corresponding target execution statements according to the data writing mode.
[0089] In this exemplary embodiment, during the data synchronization process, the destination database may already have historical data corresponding to some of the data to be synchronized (such as in the data update process). To avoid duplicate data in the destination database from occupying storage resources, the data that already exists in the destination database can be deleted. The data writing mode can include a pre-writing mode and a post-writing mode. The pre-writing mode means deleting the corresponding historical data before writing the data, and the post-processing mode means deleting the corresponding historical data after writing the data. The target execution statement is used to instruct the destination database to delete the historical data corresponding to the data to be written under the processing conditions corresponding to the data writing mode. For example, the target execution statement corresponding to the pre-writing mode can be "preSql": ["delete from table"], and the execution result is that before each table is written to the destination database, the corresponding table name (table) will be deleted first. The target execution statement corresponding to the post-writing mode can be to delete the corresponding table name after the data to be written is written to the destination table.
[0090] In some embodiments, the working module further includes: a data monitoring unit, configured to collect at least one of the processing performance data, synchronization failure data, and processing progress data during the execution of the target task in real time, and record and display the data collected in real time.
[0091] In this exemplary embodiment, various types of data during the execution of the target task can also be collected in real time by the data monitoring unit, and the data can be directly returned or returned to the user interface for display after analysis and processing, so that the user can timely understand the task running situation and improve the user experience. The processing performance data can include process CPU, memory situation (such as memory occupancy rate, garbage collection GC situation), virtual machine situation, data traffic, transmission speed, etc. The synchronization failure data refers to the data that goes wrong during the process, which can include error data, error reasons, error times, etc. The processing progress data can include information such as the amount of processed data or the remaining amount of data, processing status, and execution progress. Exemplarily, relevant information during operation can be stored in Communication, reader, writer, error, and virtual machine information can be collected in real time, and relevant information can also be printed in the operation log.
[0092] Exemplarily, the data monitoring unit is further configured to monitor the dirty data generated during the data synchronization process; and generate an alarm message in response to the data proportion of the dirty data being greater than a preset threshold.
[0093] In this exemplary embodiment, the dirty data is the error data generated during the data synchronization process. For example, when exporting data from Excel to Oracle, if one column is of numeric type in Excel while the corresponding column in Oracle is of alphabetic type, this will cause a data type conversion error and form dirty data. The synchronization quality can be evaluated by monitoring the dirty data generated during the synchronization process; alternatively, when the number or proportion of the dirty data is greater than the corresponding preset threshold, an alarm message can be generated to alert the user that the synchronization has failed, facilitating the timely discovery and handling of abnormal situations during the synchronization process.
[0094] In some embodiments, the architecture of the synchronization middleware of the present disclosure is as Figure 4 shown. The synchronization middleware may include a scheduler 410, a trigger 420, and a work module 430. One work module may correspond to the synchronization process of the data of one table. The scheduler 410 can schedule the trigger 420 to trigger the work module 430 to execute tasks. Execution conditions can be configured in the trigger, and corresponding execution conditions can be configured for each target task. The work module can manage multiple subtasks, and each subtask can perform the data synchronization process corresponding to the subtask through a data collector 431, a data converter 432 (channel), and a data writer 433. Multiple subtasks can form a task group to facilitate the management of the subtasks. Through the above Figure 4 synchronization middleware, the data synchronization process between various heterogeneous data sources can be realized. When a new data source is added, only corresponding configuration needs to be performed at the reserved plug-in interface for the new data source, greatly reducing the workload of developers.
[0095] For example, the scheduler can use Quartz, as Figure 5As shown, multiple classes can be defined to perform tasks. The job control class is used to start the program or automatically restart all jobs. For example, the job service Service can be invoked. The job service Service is used to obtain configuration information from the database, generate the job detail class JobDetail and the trigger. For example, the job configuration information can be obtained through the GetJobDetails method, and the execution conditions configured by the trigger can be obtained through the GetTriggerDetails method, and then the obtained information is returned to the caller (the job control class). The job detail class JobDetail is used to configure job configuration information including the name, description, execution class, and its parameters of the task. The trigger is used to configure the execution time, execution frequency, etc. The executor starts and executes the task. For example, the job control class starts the scheduled task through Schedule. The event operation class ProcessingJobEventHandler starts the job module of the synchronization middleware, that is, starts the data collector, data converter, and data writer. The data log service is used to record the log data generated during the execution process. The results or data generated by each class are finally returned to the job control class. Developed based on Quartz, such as Figure 5 the various classes shown, and the interaction between the various classes is realized through various methods to complete the corresponding tasks (such as the target task).
[0096] In some other embodiments, the present disclosure also provides a data synchronization method applied to an electronic device, such as Figure 6 shown, the data synchronization method may include the following steps:
[0097] Step S610, in response to reaching the execution conditions configured for the target task, collect the source data corresponding to the target task from at least one source database associated with the target task;
[0098] Step S620, according to the data storage attributes of at least one destination database associated with the target task, perform data conversion processing on the source data to obtain the corresponding target data;
[0099] Step S630, send the corresponding target data to at least one destination database to complete data synchronization, where the at least one source database and the at least one destination database are respectively configured through the reserved plug-in interface of the synchronization middleware.
[0100] In this exemplary embodiment, the target task may be to synchronize data between a specified source database and a destination database, and the corresponding source data may be determined according to information such as the identifier or name of the data to be synchronized in the target task. Data storage attributes refer to the attribute information when the destination database stores data, and the data storage attributes may include data types (integer, floating point, fixed point, etc.) and data field attributes (whether it is empty, primary key, field description, etc.). According to the data storage attributes of the destination database, the source data may be converted to obtain the target data for storage in the destination database. The data conversion process may be a data matching process for heterogeneous data sources, for example, it may be a process of deleting, inserting, changing, etc. certain attribute fields of the source data. In this example, at least one source database and at least one destination database that need to be synchronized may be configured accordingly through the reserved plug-in interface of the synchronization middleware, such as configuring the data structure, database name, identifier, address, etc. of the source database and the destination database, so that the synchronization middleware is adapted to the data structure of the database, thereby realizing data synchronization between various heterogeneous data sources.
[0101] Above Figure 6 The technical solution of the illustrated embodiment can perform data conversion processing on the corresponding source data to be synchronized according to the data storage properties of the destination database, thereby realizing data synchronization from the source to the destination, and realizing the adaptation of various heterogeneous data sources through the reserved plug-in interface of the synchronization middleware, thereby realizing data synchronization of various heterogeneous data sources, unbinding the database and the corresponding synchronization tool, and improving the flexibility of synchronization; new data sources can be introduced quickly and conveniently, with strong scalability, greatly facilitating user use and improving user experience.
[0102] In some embodiments of the present disclosure, based on the above scheme, the data storage attribute includes a data field attribute, and the source data is converted, including:
[0103] According to the data field attributes of at least one target database, field attributes of the source data are matched to obtain target data for writing into the corresponding target database.
[0104] In this exemplary embodiment, the data field attributes of the destination database can be matched with the source data to convert the source data into the storage data of the destination database. For example, some field attributes of the source data can be deleted or new field attributes can be inserted to make it conform to the destination storage rules. This example can reduce data omissions and conversion errors by matching the field attribute dimensions, thereby improving synchronization accuracy.
[0105] In some embodiments of the present disclosure, based on the foregoing solution, performing data conversion processing on the source data further includes: performing at least one of desensitization processing, completion processing, and data filtering on the source data.
[0106] In this exemplary embodiment, the desensitization processing may be a process of encrypting and protecting or not disclosing sensitive information in the source data. The sensitive information may include user private information (such as user name, ID number, password, etc.), business sensitive information (such as internal price), and so on. The completion processing is to complete some missing values in a specified format, and the data filtering refers to filtering duplicate data or obviously incorrect data. The desensitization processing, completion processing, and data filtering can all be implemented by configuring in response to the data conversion processing. Exemplarily, for the desensitization processing, the desensitization effect can be achieved through string replacement. For example, the specified position characters of the fields of sensitive data can be replaced with **** through the dx_replace function; a custom Transformer can also be defined to encrypt the sensitive data through encoding to implement its function; the data can also be processed through groovy functions. In this example, various data processing processes are implemented through configuration. The desensitization processing can protect sensitive data and improve data security. The completion processing can reduce the synchronization error situation and improve the synchronization success rate; the data filtering can filter out unimportant or duplicate data and save the database performance overhead.
[0107] In some embodiments of the present disclosure, based on the foregoing solution, after the data synchronization is completed, the method further includes:
[0108] Performing consistency verification on the synchronized data of at least one destination database associated with the target task;
[0109] In response to the verification failure, triggering the re-execution of the target task.
[0110] In some embodiments of the present disclosure, based on the foregoing solution, after the synchronization is completed, the consistency of the synchronization result can be verified, that is, comparing whether the data content or quantity of the corresponding source database and destination database is the same. If they are the same, the synchronization is successful. If they are different, the failure reason can be determined according to the preset rules, and the data to be synchronized can be repaired based on the failure reason. Based on the repaired data, triggering the re-execution of the target task once. The previously synchronized data can be deleted based on the data writing mode.
[0111] The present disclosure can also perform statistical analysis on the data to be synchronized and the historical data after the synchronization is successful, obtain certain index trends or analyze production quality, sales products, etc., and provide a certain basis for the user's production decision-making.
[0112] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes: controlling the speed of data synchronization according to the configured traffic control mode.
[0113] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes:
[0114] Determine the resource occupancy corresponding to the current data flow rate;
[0115] In response to the resource occupancy exceeding a preset threshold, trigger the current working thread to enter the sleep state, and the sleep time of the current working thread is determined based on the current data flow rate corresponding to the traffic control mode and the preset flow rate.
[0116] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes:
[0117] According to the configured data writing mode, trigger the execution of the corresponding target execution statement, and the target execution statement is used to instruct the destination database to delete the historical data corresponding to the data to be written under the processing conditions corresponding to the data writing mode.
[0118] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes:
[0119] Real-time collect at least one of the processing performance data, synchronization failure data, and processing progress data during the execution of the target task, and record and display the real-time collected data in a log.
[0120] In some embodiments of the present disclosure, based on the foregoing solution, the method further includes: monitoring the dirty data generated during the data synchronization process; the dirty data is the error data generated during the data synchronization process;
[0121] In response to the data proportion of the dirty data being greater than a preset threshold, generate an alarm message.
[0122] The specific details of the above data synchronization methods have been described in detail in the corresponding data synchronization system, so they will not be repeated here.
[0123] In other embodiments, the present disclosure further provides a data backup method, which is applied to various production systems, such as the APS scheduling system, as Figure 7 shown, the data backup method may include the following steps:
[0124] Step S710, obtain the production data of each production link;
[0125] Step S720, in response to the current satisfaction of the preset data backup condition, send the data to be backed up in the production database to the target database through the synchronization middleware to complete the production data backup; the synchronization middleware includes: a scheduling module and a working module, and the scheduling module is used to, when the data backup condition is met currently, schedule the working module to back up the data to be backed up to the target database based on the configuration class; the production database and the target database are respectively configured through the reserved plug-in interfaces of the synchronization middleware.
[0126] In this exemplary embodiment, production data can be obtained from each production system or business system of a manufacturing enterprise. For example, the APS production scheduling system needs to schedule production according to sales demands and production factors of the factory (such as equipment, personnel, materials, etc.), and needs to collect data from each production link regularly (such as once a day or every few hours) for production scheduling. The production scheduling data also needs to be backed up regularly to ensure the data security of production data and production scheduling data. The synchronization middleware can be deployed in the production scheduling system. By configuring the source end (production database) and the destination end (backup database, i.e., the target database) in the reserved plug-in interface of the synchronization middleware, important data generated during production can be backed up regularly, which can improve data security.
[0127] In some embodiments, after the data backup is completed, the following steps can be performed for data verification:
[0128] Determine the number of backup data records backed up to the target database;
[0129] Perform consistency verification on the current data backup according to the number of backup data records and the number of data records to be backed up;
[0130] In response to the verification failure, re-perform data backup on the restored data, where the restored data is obtained by processing the data to be backed up according to the failure reason, and the failure reason is determined according to the preset rule and the consistency verification result.
[0131] In this exemplary embodiment, the data in the database is generally table data. After the backup is completed, the number of backed-up data can be recorded, and consistency verification is performed by comparing the number of backed-up data in the target database with the number of data to be backed up. If the two are consistent, the verification passes and the data backup is successful; if they are inconsistent, the data backup fails. The reason for failure can be determined according to the verification result and the preset rules. The preset rules can be mapping rules between the verification result and the reason for failure. For example, if the verification result is that a certain column of data is incorrect, the reason for failure can be that the data columns are inconsistent. After modifying the data columns of the data to be backed up, the data backup can be performed again. For the data that has been backed up in the target database, it can be deleted before or after re-backing up, and the specific deletion time can be determined according to the configured data writing mode.
[0132] The specific details of the above data backup methods have been described in detail in the corresponding data synchronization methods and systems, so they will not be elaborated here.
[0133] It should be noted that although several modules or units of the devices for execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more of the above-mentioned modules or units can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.
[0134] In this exemplary embodiment, an electronic device capable of implementing the above method is also provided.
[0135] Those skilled in the art of the relevant technology can understand that various aspects of the present invention can be implemented as a system, a method, or a program product. Therefore, various aspects of the present invention can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuitry", "module", or "system" here.
[0136] The following refers to Figure 8 to describe the electronic device 800 according to this embodiment of the present invention. Figure 8 The electronic device 800 shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.
[0137] As Figure 8 shown, the electronic device 800 is presented in the form of a general-purpose computing device. The components of the electronic device 800 may include, but are not limited to: at least one of the above-mentioned processing units 810, at least one of the above-mentioned storage units 820, a bus 830 connecting different system components (including the storage unit 820 and the processing unit 810), and a display unit 840.
[0138] Among them, the storage unit stores program code, which can be executed by the processing unit 810, so that the processing unit 810 executes the steps according to various exemplary embodiments of the present invention described in the data synchronization method above in this specification. For example, the processing unit 810 can execute the steps shown in Figure 3 or Figure 4 .
[0139] The storage unit 820 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 8201 and / or a cache storage unit 8202, and may further include a read-only storage unit (ROM) 8203.
[0140] The storage unit 820 may also include a program / utility 8204 having a set (at least one) of program modules 8205. Such program modules 8205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0141] The bus 830 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0142] The electronic device 800 can also communicate with one or more external devices (such as a keyboard, a pointing device, a Bluetooth device, etc.), can also communicate with one or more devices that enable the audience to interact with the electronic device 800, and / or can communicate with any device that enables the electronic device 800 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 850. And, the electronic device 800 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 860. As shown in the figure, the network adapter 860 communicates with other modules of the electronic device 800 through the bus 830. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 800, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0143] Based on the description of the above embodiments, those skilled in the art can easily understand that the exemplary embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a portable hard drive, etc.) or on a network, including several instructions to enable a computing device (such as a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the embodiments of the present disclosure.
[0144] In this exemplary embodiment, a computer-readable storage medium is also provided, on which a program product capable of implementing the above methods of this specification is stored. In some possible embodiments, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to enable the terminal device to execute the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification.
[0145] The program product for implementing the above method according to the embodiments of the present invention can use a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of the present invention is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0146] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0147] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.
[0148] The program code contained on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0149] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0150] In addition, the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, and are not for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the chronological order of these processes. Additionally, it is also easy to understand that these processes may be executed synchronously or asynchronously in, for example, multiple modules.
[0151] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed by the present disclosure. The specification and embodiments are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0152] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A data synchronization system, characterized in that, it includes: at least one source database, at least one destination database, and a synchronization middleware, and the synchronization middleware is respectively connected to at least one source database and at least one destination database; the synchronization middleware includes: a scheduling module and a working module, the scheduling module is used to schedule the working module to execute the target task based on the configured target task and the execution condition of the target task, and when the current execution condition is met, the scheduling module schedules the working module to execute the target task based on the configuration class of the target task, and the configuration class includes the working configuration information of the target task; the working module is used to start corresponding working threads according to the working configuration information to perform data synchronization of the data to be synchronized associated with the target task between at least one source database and the corresponding at least one destination database; the at least one source database and the at least one destination database are respectively configured through the reserved plug-in interfaces of the synchronization middleware.
2. The system according to claim 1, characterized in that, the scheduling module includes a trigger, and the trigger is used to configure the execution condition; the scheduling module is used to bind the trigger of the target task and the corresponding configuration class; scheduling the working module to execute the target task based on the configuration class of the target task includes: when the current execution condition is met, triggering the working module to execute the target task based on the configuration class through the trigger.
3. The system according to claim 1, characterized in that, the working configuration information includes the primary key sharding information of the table data, and the working module includes: a task splitting unit, which is used to split the target task into multiple subtasks according to the splitting strategy of at least one source database or the primary key sharding information of the table data, and each subtask is used to perform data synchronization of different parts of the source data.
4. The system according to claim 3, characterized in that, the working configuration information includes the channel concurrency number, and the working module includes: a subtask management unit, which is used to form a task group from multiple subtasks according to the channel concurrency number, and manage the execution process of each subtask based on the task group, and the channel concurrency number is the same as the concurrent thread number of the task group.
5. The system according to claim 3, characterized in that, the reserved plug-in interface includes a first reserved plug-in interface and a second reserved plug-in interface, and the working module corresponding to the subtask includes: a data collector, which is used to collect the corresponding source data from the source database and write the source data into a data converter; a data converter, which is used to perform data conversion processing on the source data to obtain the corresponding target data; a data writer, which is used to read the target data from the data converter and send the target data to the corresponding destination database; Among them, at least one source database corresponding to the target task is configured through the first reserved plug-in interface corresponding to the data collector, and at least one destination database corresponding to the target task is configured through the second reserved plug-in interface corresponding to the data writer.
6. The system according to claim 5, wherein, the working configuration information includes a flow control mode, and the data converter is configured to control the speed of the data synchronization according to the flow control mode.
7. The system according to claim 6, wherein, the data converter is configured to: determine the resource occupancy corresponding to the current data flow rate; in response to the resource occupancy exceeding a preset threshold, trigger the current working thread to enter a sleep state, and the sleep time of the current working thread is determined based on the current data flow rate corresponding to the flow control mode and a preset flow rate.
8. The system according to claim 5, wherein, the working configuration information includes a data writing mode, and the data converter is configured to trigger the execution of a corresponding target execution statement according to the data writing mode, and the target execution statement is used to instruct the destination database to delete the historical data corresponding to the data to be written under the processing conditions corresponding to the data writing mode.
9. The system according to claim 1, wherein, the working module further includes: a data monitoring unit configured to collect at least one of processing performance data, synchronization failure data, and processing progress data during the execution of the target task in real time, and record and display the data collected in real time.
10. The system according to claim 9, wherein, the data monitoring unit is further configured to monitor the dirty data generated during the data synchronization process; the dirty data is the error data generated during the data synchronization process; in response to the data proportion of the dirty data being greater than a preset threshold, generate an alarm message.
11. A data backup method, wherein, comprises: obtaining production data of each production link; in response to currently meeting a preset data backup condition, sending the data to be backed up in the production database to a target database through a synchronization middleware to complete the production data backup; the synchronization middleware includes: a scheduling module and a working module, the scheduling module is configured to, when the data backup condition is currently met, schedule the working module based on a configuration class to back up the data to be backed up to the target database; the production database and the target database are respectively configured through the reserved plug-in interfaces of the synchronization middleware.
12. The method according to claim 11, wherein, the method further comprises: determining the number of backup data records backed up to the target database; performing consistency verification on the current data backup according to the number of backup data records and the number of data records to be backed up; in response to the verification failing, re-performing data backup on the restored data, where the restored data is obtained by processing the data to be backed up according to a failure reason, and the failure reason is determined according to a preset rule and the consistency verification result.
13. A data synchronization method, wherein, Applied to the synchronization middleware, including: In response to reaching the execution condition configured for the target task, collect the source data corresponding to the target task from at least one source database associated with the target task; According to the data storage attributes of at least one destination database associated with the target task, perform data conversion processing on the source data to obtain the corresponding target data; Send the corresponding target data to at least one destination database to complete data synchronization; Wherein, the at least one source database and the at least one destination database are respectively configured through the reserved plug-in interface of the synchronization middleware.
14. The method according to claim 11, characterized in that, Performing data conversion processing on the source data further includes: Performing at least one of desensitization processing, completion processing, and data filtering on the source data.
15. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the data synchronization method according to any one of claims 11 to 14.