Data synchronization method and data processing method
By responding to multiple database log changes in the information construction of steel plants, the problem of data integration and real-time synchronization is solved, and an efficient, reliable and low-cost data synchronization and processing mechanism is achieved.
Patent Information
- Application Number
- CN202411862225.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-06
Smart Images

Figure CN119938782A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a data synchronization method and a data processing method. Background Art
[0002] In complex processes such as steel production, various production equipment, precision testing instruments and advanced control systems continuously generate massive amounts of real-time operational data and business information. These data play an irreplaceable role in achieving efficient production control, promoting significant improvements in energy efficiency and implementing accurate fault warnings. In view of this, building an integrated data center is particularly critical.
[0003] However, the informatization construction of steel plants usually takes a long time, with a wide variety of data sources and independent systems, which undoubtedly poses a severe challenge to the efficient integration of data. The data center needs to capture data changes as real-time as possible. The current methods for capturing data changes are mainly divided into query-based CDC (Change Data Capture) technology and log-based CDC technology. Query-based CDC is an offline scheduling query job that uses batch processing technology to obtain the latest data in the table each time through queries, and cannot guarantee data consistency and real-time performance. Log-based CDC technology uses stream processing to consume database logs in real time. Although it can ensure data consistency and real-time performance, the currently commonly used Debezium uses a single-machine architecture, with poor data conversion and cleaning capabilities, and insufficient support for downstream data sources.
[0004] Therefore, there is an urgent need for an efficient and reliable data synchronization mechanism to aggregate and deeply analyze multivariate data from different systems in real time. Summary of the invention
[0005] The purpose of the present invention is to provide a data synchronization method and a data processing method, which can aggregate and deeply analyze multivariate data from different systems in real time, thereby realizing an efficient, reliable and low-cost data synchronization and processing mechanism.
[0006] In order to achieve the above-mentioned purpose, the first aspect of the present invention provides a data synchronization method, which includes: in response to a change in a log in each of a plurality of databases, obtaining configuration information corresponding to the changed log in each of the databases, wherein the configuration information includes connection information and data reading information; based on the connection information, capturing the changed log from each of the databases; based on the data reading information, parsing the log captured from each of the databases to obtain the source data table of each of the databases; mapping the source data table of each of the databases to obtain the target data of each of the databases; and synchronizing the target data of each of the databases to a distributed publish-subscribe message system to obtain a target table corresponding to the source data table of each of the data; and synchronizing the target table corresponding to the source data table of each of the data to a data middle platform.
[0007] Preferably, the connection information includes: database address, user name and password.
[0008] Preferably, the data reading information includes an initial reading position, a reading mode and an end reading position.
[0009] Preferably, the data synchronization method further comprises: regularly storing data involved in the processes of capturing, parsing, mapping and synchronizing; and in case of failure of any process, continuing to execute any process starting from the stored data.
[0010] Preferably, the plurality of databases include at least two of the following databases involved in the steel plant: equipment management database, manufacturing execution system database, logistics database, metering database, energy database, cost database, enterprise resource planning database.
[0011] Preferably, the distributed publish-subscribe messaging system is Kafka.
[0012] Preferably, the data synchronization method is executed under the Flink framework.
[0013] Through the above technical scheme, the present invention creatively responds to changes in logs in each of a plurality of databases, obtains configuration information corresponding to the changed logs in each database; captures the changed logs from each database according to the connection information; parses the logs captured from each database according to the data reading information to obtain the source data table of each database; maps the source data table of each database to obtain the target data of each database; and synchronizes the target data of each database to a distributed publish-subscribe message system to obtain a target table corresponding to the source data table of each data, thereby enabling real-time aggregation and in-depth analysis of multivariate data from different systems, thereby realizing an efficient, reliable and low-cost data synchronization and processing mechanism.
[0014] A second aspect of the present invention provides a data processing method, which includes: receiving multiple data requests; according to each data request in the multiple data requests, obtaining data corresponding to the each data request from a target table obtained according to the data synchronization method; and processing the data corresponding to each data request to obtain a data result corresponding to the each data request.
[0015] Preferably, the acquiring of data corresponding to each data request comprises: acquiring data corresponding to each data request from the target table using load balancing and asynchronous message queue technology according to each data request in the multiple data requests.
[0016] A third aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data synchronization method and the data processing method are implemented.
[0017] Other features and advantages of the present invention will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following specific implementations, they are used to explain the embodiments of the present invention, but do not constitute a limitation on the embodiments of the present invention. In the accompanying drawings:
[0019] Figure 1 is a flow chart of a data synchronization method provided by an embodiment of the present invention;
[0020] Figure 2 It is a flow chart of a data synchronization method provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0021] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0022] Figure 1 FIG. 1 is a flow chart of a data synchronization method provided by an embodiment of the present invention. Figure 1 As shown, the data synchronization method includes: step S101, in response to a change in a log in each of a plurality of databases, obtaining configuration information corresponding to the changed log in each of the databases, wherein the configuration information includes connection information and data reading information; step S102, capturing the changed log from each of the databases according to the connection information; step S103, parsing the log captured from each of the databases according to the data reading information to obtain the source data table of each of the databases; step S104, mapping the source data table of each of the databases to obtain the target data of each of the databases; step S105, synchronizing the target data of each of the databases to a distributed publish-subscribe message system to obtain a target table corresponding to the source data table of each of the data; and step S106, synchronizing the target table corresponding to the source data table of each of the data to the data middle platform.
[0023] The data synchronization method can be executed under the Flink framework. This embodiment uses Flink CDC technology as the core to achieve efficient, low-latency, and high-fidelity synchronization of data. Through fine configuration and optimization, not only the development and operation costs are reduced, but also the real-time and reliability of data synchronization are significantly improved. The following describes and illustrates each of the above steps.
[0024] Before executing step S101, the archiving mode of the source database (for example, each of the multiple databases in step S101, such as MySQL, PostgreSQL, etc.) may be enabled to record the change history of the data. Figure 2 As shown. As needed, select table-level or library-level log enablement to control the granularity of data synchronization. For example, Oracle requires additional library-level and table-level supplementary logs. In the source data table configuration for monitoring, specify the source data table to be synchronized. You can configure multiple tables for synchronization. In addition, you can create a Flink CDC dedicated user (such as Figure 2 As shown in the figure), and grant it the permission to read archive logs and access required tables. That is, configure the source database, such as Figure 2 shown.
[0025] Specifically, create a Source object of Flink CDC to obtain the source data change log. For example, mainly for the Source object: 1. Configure database connection information (set database URL, port number, and database name); 2. Define the database and table to be monitored; 3. Set database access credentials (user name and password); 4. Configure data deserializer (data converted to JSON string format); 5. Set startup options (configure the data source function to start in the latest offset mode, this setting will skip the full data synchronization stage and directly start capturing incremental change data); 6. Use Debezium as the CDC engine to capture change data in the database.
[0026] Step S101 : in response to a change in a log in each of a plurality of databases, obtaining configuration information corresponding to the changed log in each of the databases.
[0027] Among them, the multiple databases include at least two of the following databases involved in the steel plant: equipment management database (such as SQL Server), manufacturing execution system (MES) database (such as Oracle), logistics database (such as MySQL), metering database (such as SQL Server), energy database, cost database (such as PostgreSQL), enterprise resource planning (ERP) database (such as Oracle).
[0028] The configuration information includes connection information and data reading information. Specifically, the connection information includes: database address (such as URL), user name and password; the data reading information includes initial reading position, reading mode and end reading position.
[0029] Specifically, data change events are monitored in real time by monitoring the log files of the database (such as MySQL's binlog). These change events include insert, update, and delete operations, which are read in real time and converted into streaming data. For any of the multiple data sources, when data is inserted into one of the logs, changes to the log file will be monitored, and the connection information of the source database will be created at this time, such as the database type, address, port, user name, and password. This is equivalent to specifying the source data tables that need to be monitored. These tables will serve as the source of data synchronization. Similar operations are performed for other databases, which will not be repeated here. At the same time, the startup options of the data source (i.e., data reading information) are also created, such as the initial reading position, reading mode (full or incremental), and end reading position. In addition, authentication information such as SSL / TLS certificates, Kerberos authentication, etc. can also be configured to ensure the security of data transmission.
[0030] The log can be a table-level or database-level log. Specify table-level or database-level log enablement: Based on the data scope to be synchronized, choose to enable the logging function for a specific table or the entire database. A technical solution based on database logs to ensure real-time data capture and processing.
[0031] In this embodiment, the data stream API (DataStream API) of Flink CDC can be used to create a Source object of source data according to the configuration information.
[0032] You can create a Flink job flow execution environment in advance to handle abnormal data interruptions and ensure data synchronization and stability. You can perform the following configurations based on actual needs. Then place the Source object in the following configured execution environment to execute steps S102-S104.
[0033] 1. Configure checkpoints (enable checkpoints: trigger every 300,000 milliseconds to ensure fault tolerance and state consistency of the job); set the checkpoint mode: specify the consistency semantics of the checkpoint as EXACTLY_ONCE, to ensure that during failure recovery, each piece of data will be processed exactly once to avoid duplication or omission).
[0034] 2. Configure externalized checkpoints (set the data of the last checkpoint to be retained when the job is canceled, so that the latest status information can still be accessed after the job stops).
[0035] 3. Configure the restart strategy (Configure a fixed delay restart strategy for the stream execution environment. When a job fails, it will try to restart up to 3 times, with an interval of 2,000 milliseconds between each restart. This helps to automatically recover jobs that are interrupted by short failures).
[0036] 4. Configure the state backend (select the file system state backend and store the state in the specified file path. It helps to persist the state information while the job is running so that the state can be restored during failure recovery).
[0037] Step S102: Capture the changed log from each database according to the connection information.
[0038] For example, a designated source database is obtained according to a database address (eg, URL), and then the source database is logged in according to a user name and password, so that a log of changes occurring in the source database can be captured.
[0039] This embodiment provides an efficient data capture mechanism that can obtain data changes from multiple data sources in real time without affecting the performance and stability of the source system.
[0040] Step S103: parsing the logs captured from each database according to the data reading information to obtain the source data table of each database.
[0041] Adding the Source object to the execution environment of the Flink job flow will get the serialized JSON data. A custom data deserializer is used to parse the change log captured from the source database. The parsing content mainly includes operation type - add, delete, modify, query, database, data table, field information, primary key and field value and other information. Specifically, according to the initial reading position, reading mode (full or incremental), and end reading position corresponding to the log of the change in each database, the log of the change in each database is parsed, so that the corresponding source data table can be obtained.
[0042] Step S104: Map the source data table of each database to obtain target data of each database.
[0043] Add the created Source object to the execution environment of the Flink job flow. Use the map function to process the Source data (including the parsed source data table), customize the conversion logic according to business needs, and implement the mapping from the source table to the target data (the specific mapping process is not the main improvement of this application and will not be described here). Figure 2 shown.
[0044] Specifically, data conversion is performed in the map function. The parsed data is processed to meet the needs of the data center or other downstream systems. According to actual needs, the following steps can be performed before data mapping: data cleaning (removing invalid values, format standardization); data conversion (data type conversion, data structure conversion); data aggregation (calculating the sum, average, etc. of data according to business needs); data filtering (filtering data according to specific conditions, retaining only data that meets the conditions). Then, data mapping is performed (mapping the fields in the source database to the fields of the data center or other downstream systems).
[0045] In this embodiment, the parallelism of the map function may also be set to improve the efficiency of data processing.
[0046] Step S105: synchronizing the target data of each database to the distributed publish-subscribe message system to obtain a target table corresponding to the source data table of each data.
[0047] Wherein, the distributed publish-subscribe messaging system is Kafka.
[0048] Specifically, Kafka can be introduced as a distributed publish-subscribe messaging system to reduce the writing pressure on the target table. Add Kafka middleware to prevent the data volume of some source tables from growing too fast and causing excessive pressure on the target table. Synchronize the source data to Kafka, then read the data from Kafka and distribute it to the downstream system (see step S106 for details). Create a corresponding topic (Topic) for each source data table in Kafka, such as Figure 2 As shown. Submit the encapsulated Flink CDC task to the Flink cluster for execution to achieve data synchronization. The Flink cluster is used to execute the above tasks because the program has high reliability and is easy to monitor.
[0049] In this embodiment, Kafka's high throughput, persistent storage and other characteristics are used to ensure efficient and reliable data transmission. In Kafka, a corresponding topic is created for each source data table so that the data center can subscribe to the corresponding data according to the topic.
[0050] Steps S104-S105 integrate the Source object and run the Flink CDC task.
[0051] Step S106, synchronizing the target table corresponding to the source data table of each data to the data middle platform.
[0052] Specifically, data is read from Kafka and distributed to downstream systems (such as the data center). Figure 2 The GaussDB in the data center is the target database in the data center. The source data changes can be synchronized to the GaussDB in the data center through the tasks on the Flink cluster.
[0053] In one embodiment, the data synchronization method further comprises: regularly storing data involved in the processes of capturing, parsing, mapping, and synchronizing; and in case of failure of any process, continuing to execute any process starting from the stored data.
[0054] For example, configure checkpoints and set the checkpoint interval to ensure that the job status can be saved regularly during data synchronization. Specify the consistency mode of the checkpoint (such as Exactly-Once Processing) to maintain data consistency. Enable the externalized checkpoint function to control the processing strategy of checkpoint data after job cancellation or failure. Configure the restart strategy so that the job can automatically restart and continue data synchronization when it fails. Select the file system as the system status backend, specify the storage path for the checkpoint data, and store the checkpoint data in the specified path. In other words, configure the Flink job flow execution environment. In this way, breakpoint resumption and abnormal recovery can be achieved.
[0055] Through the above steps, data verification, data error correction, and data filtering can ensure real-time, consistency, accuracy, and completeness: By optimizing data transmission and processing processes, ensure that data can be captured and processed within a short time after it is generated. During data capture and processing, ensure data consistency in all dimensions to avoid data ambiguity and errors. Strictly control data quality to ensure data correctness and reliability. Ensure data integrity, including the integrity of entity relationships and data domains, to avoid data missing or errors that lead to misleading analysis results.
[0056] In the above-mentioned embodiments, the database is the source database, and the two can be used interchangeably.
[0057] The present invention proposes a data synchronization and processing method for establishing a steel plant data center, which can significantly improve real-time performance and data synchronization efficiency, so that data accuracy, integrity and consistency are guaranteed. At the same time, it reduces development and maintenance costs, enhances system flexibility and scalability, and improves business decision-making efficiency. It brings new breakthroughs and progress in the field of data synchronization and real-time processing.
[0058] In summary, the present invention creatively responds to changes in logs in each of a plurality of databases, obtains configuration information corresponding to the changed logs in each database; captures the changed logs from each database according to the connection information; parses the captured logs from each database according to the data reading information to obtain the source data table of each database; maps the source data table of each database to obtain the target data of each database; and synchronizes the target data of each database to a distributed publish-subscribe message system to obtain the target table corresponding to the source data table of each data. The present invention can aggregate and deeply analyze multivariate data from different systems in real time, thereby realizing an efficient, reliable and low-cost data synchronization and processing mechanism.
[0059] An embodiment of the present invention provides a data processing method, which includes: receiving multiple data requests; according to each data request in the multiple data requests, obtaining data corresponding to the each data request from a target table obtained according to the data synchronization method; and processing the data corresponding to each data request to obtain a data result corresponding to the each data request.
[0060] These stream data are processed (any existing technology can be used for processing) and transmitted through Flink's stream processing engine, and finally reach the target system or data platform.
[0061] Specifically, the following data consumption and synchronization processing can be performed according to the data request.
[0062] Task initialization: Create a new task in the Flink environment, which is responsible for data consumption and synchronization.
[0063] Data source and Sink object creation: Configure the Kafka data source for the task, and create a Sink object pointing to the data center for data reception and storage.
[0064] Configuration reading and parsing: Read necessary configuration information from pre-generated configuration files to ensure the correct configuration of data sources and Sink objects.
[0065] Data reading and conversion: Read data from the specified Kafka topic and perform necessary parsing and conversion to meet the data format and business logic requirements of the data center.
[0066] Data synchronization operation: According to the type of change in the data (such as addition, modification, deletion), the corresponding data insertion, update or deletion operations are performed on the Sink object of the data center to ensure the real-time and accuracy of the data.
[0067] In one embodiment, the obtaining of data corresponding to each data request includes: according to each data request in the multiple data requests, using load balancing and asynchronous message queue technology to obtain data corresponding to each data request from the target table.
[0068] In addition, in one embodiment, the following exception handling and recovery strategies may also be performed.
[0069] During data synchronization, you need to pay attention to and properly handle the following abnormal scenarios to ensure the continuity and integrity of data processing:
[0070] Handling missing archive logs: If you cannot obtain archive logs from the source database, you should try to restore the missing logs from the backup system. If the recovery fails, you need to perform data initialization to restore data integrity.
[0071] Table space occupancy problem: When the table space of the source database is filled with archive logs, it is necessary to review the database's archive log retention policy and expand the table space if necessary to avoid data loss or processing interruption.
[0072] Synchronization of table structure changes: If the table structure of the source database changes, the corresponding synchronization change operation must be performed on the data center, and the data must be re-acquired from the time of the change to ensure data accuracy and consistency.
[0073] Source database exception handling: During the operation of the FlinkCDC process, if the source database is unexpectedly closed, causing the process to be abnormal, it is necessary to integrate an intelligent retry mechanism into the program and set a reasonable retry interval and number to restore the data processing flow.
[0074] Sink-side exception recovery: If an exception occurs on the Sink side, resulting in data loss or corruption, the offset value of the corresponding Topic in Kafka can be reset so that the Flink task can re-consume and write the data to the Sink side to ensure data integrity and accuracy.
[0075] In one embodiment, the following performance tuning and resource optimization may also be performed: To meet the requirements of data volume and writing speed, performance tuning and resource optimization of Flink tasks are required.
[0076] Parallelism adjustment: According to the actual needs of data processing and cluster resources, reasonably adjust the parallelism parameters of Flink tasks to improve data processing efficiency.
[0077] Batch processing optimization: By increasing the size of batch processing, the overhead of data transmission and processing is reduced, and the overall performance is improved.
[0078] Serialization and compression technology: Use efficient serialization methods and compression algorithms to further reduce the load and cost of data transmission.
[0079] The present invention provides flexible custom expansion functions, and users can add computing nodes or adjust system configurations according to actual needs. Through distributed architecture, load balancing, asynchronous message queues and other technical means, it is ensured that the system can efficiently handle high-concurrency requests. Efficient data processing algorithms and storage technologies are adopted to ensure that the system can cope with the data synchronization requirements of large amounts of data. Custom expansion is used to cope with high concurrency and large amounts of data.
[0080] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the data synchronization method and the data processing method are implemented.
[0081] The preferred embodiments of the present invention are described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the specific details in the above embodiments. Within the technical concept of the present invention, a variety of simple modifications can be made to the technical solution of the present invention, and these simple modifications all belong to the protection scope of the present invention.
[0082] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not further describe various possible combinations.
[0083] Those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a single-chip microcomputer, a chip or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0084] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0085] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0086] In addition, various embodiments of the present invention may be arbitrarily combined, and as long as they do not violate the concept of the present invention, they should also be regarded as the contents disclosed by the present invention.
Claims
1. A data synchronization method, characterized in that: The data synchronization method comprises: In response to a change in a log in each of the plurality of databases, obtaining configuration information corresponding to the changed log in each of the databases, wherein the configuration information includes connection information and data reading information; According to the connection information, capturing a log of changes from each database; Parsing the logs captured from each database according to the data reading information to obtain a source data table of each database; Mapping the source data table of each database to obtain the target data of each database; Synchronizing the target data of each database into a distributed publish-subscribe message system to obtain a target table corresponding to a source data table of each data; and The target table corresponding to the source data table of each data is synchronized to the data center.
2. The data synchronization method according to claim 1, characterized in that: The connection information includes: database address, user name and password.
3. The data synchronization method according to claim 1, characterized in that: The data reading information includes an initial reading position, a reading mode, and an end reading position.
4. The data synchronization method according to claim 1, characterized in that: The data synchronization method further includes: Regularly store the data involved in the capture, parsing, mapping, and synchronization processes; and In the event of a failure of any process, execution of the process is continued starting from the stored data.
5. The data synchronization method according to claim 1, characterized in that: The multiple databases include at least two of the following databases involved in the steel plant: an equipment management database, a manufacturing execution system database, a logistics database, a metering database, an energy database, a cost database, and an enterprise resource planning database.
6. The data synchronization method according to claim 1, characterized in that: The distributed publish-subscribe messaging system is Kafka.
7. The data synchronization method according to claim 1, characterized in that: The data synchronization method is executed under the Flink framework.
8. A data processing method, characterized in that: The data processing method comprises: receiving multiple data requests; According to each of the multiple data requests, acquiring data corresponding to each of the data requests from a target table acquired by the data synchronization method according to any one of claims 1 to 6; and The data corresponding to each data request is processed to obtain a data result corresponding to each data request.
9. The data processing method according to claim 8, characterized in that: The acquiring of data corresponding to each data request comprises: According to each data request among the multiple data requests, load balancing and asynchronous message queue technology are used to obtain data corresponding to each data request from the target table.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the data synchronization method described in any one of claims 1 to 7 and the data processing method described in any one of claims 8 to 9.