Data real-time synchronization method and related equipment
By real-time monitoring and cache source data changes, and processing and synchronizing data based on preset rules, the problem of inefficient and low accuracy of data synchronization on the big data platform is solved, and efficient, accurate and flexible data synchronization is achieved.
Patent Information
- Application Number
- CN202510108750.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-30
AI Technical Summary
The data synchronization processing of the prior art on the big data platform is inefficient and has low accuracy, resulting in insufficient data synchronization speed, accuracy and flexibility, and complex configuration and management work.
By monitoring the source data changes in real time, cache them to the message queue to form a change data queue, and data processing is performed on the change data queue based on preset processing rules to obtain the target data, and synchronize it to the target data set with the data structure of the target data set.
Real-time data capture, efficient processing and flexible synchronization are realized, improving the accuracy and reliability of data synchronization, and simplifying configuration and management work.
Smart Images

Figure CN120067209A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular, to a method for real-time data synchronization and related devices. Background Art
[0002] With the continuous improvement of the informatization level of enterprises, the requirements for the timeliness and consistency of data are also getting higher and higher. Traditional data replication and backup methods are no longer able to meet the needs of modern businesses. As an effective method, Change Data Capture (CDC) can monitor and record any changes that occur in a database and transmit this change information to other services for consumption in a timely manner. However, in the actual application process, the data synchronization processing efficiency is low and the accuracy is not high. Summary of the Invention
[0003] The present disclosure proposes a method for real-time data synchronization and related devices to solve the technical problems such as low data synchronization processing efficiency and low accuracy to a certain extent.
[0004] In a first aspect of the present disclosure, a method for real-time data synchronization is provided, including:
[0005] In response to detecting a change in source data, caching the corresponding changed data into a message queue to obtain a changed data queue;
[0006] Performing data processing on the changed data queue based on a preset processing rule to obtain target data;
[0007] Synchronizing the target data to the target data set in the data structure of the target data set.
[0008] In a second aspect of the present disclosure, a device for real-time data synchronization is provided, including:
[0009] A caching module, configured to cache the corresponding changed data into a message queue to obtain a changed data queue in response to detecting a change in source data;
[0010] A processing module, configured to perform data processing on the changed data queue based on a preset processing rule to obtain target data;
[0011] A synchronization module, configured to synchronize the target data to the target data set in the data structure of the target data set.
[0012] In a third aspect of the present disclosure, an electronic device is provided, including one or more processors, a memory; and one or more programs, where the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.
[0013] In a fourth aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium including a computer program, which, when executed by one or more processors, causes the processors to execute the method described in the first aspect.
[0014] In a fifth aspect of the present disclosure, there is provided a computer program product including computer program instructions, which, when executed on a computer, cause the computer to execute the method described in the first aspect.
[0015] As can be seen from the above, a data real-time synchronization method and related devices provided by the present disclosure realize real-time capture, efficient processing, and flexible synchronization of data by monitoring source data changes in real time, caching them into a message queue to form a change data queue, then processing the change data queue based on preset processing rules to obtain target data, and finally synchronizing the processed target data to a target data set in the data structure of the target data set, improving the accuracy and reliability of data synchronization. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a schematic diagram of the data real-time synchronization architecture according to an embodiment of the present disclosure.
[0018] Figure 2 It is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
[0019] Figure 3 It is a schematic diagram of the process of the data real-time synchronization method according to an embodiment of the present disclosure.
[0020] Figure 4 It is a schematic diagram of the data real-time synchronization method according to an embodiment of the present disclosure.
[0021] Figure 5 It is a schematic diagram of the index-level monitoring according to an embodiment of the present disclosure.
[0022] Figure 6 It is a schematic diagram of the setting interface according to an embodiment of the present disclosure.
[0023] Figure 7 It is a schematic diagram of the data real-time synchronization device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] To make the objectives, technical solutions, and advantages of the present disclosure more clearly understood, the present disclosure will be further described in detail below with reference to specific embodiments and the accompanying drawings.
[0025] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The terms "first", "second", and similar terms used in the embodiments of the present disclosure do not denote any order, quantity, or importance, but are only used to distinguish different components. The terms "including" or "comprising" and the like mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms "connected" or "coupled" and the like are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms "upper", "lower", "left", "right", etc. are only used to indicate relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0026] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0027] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, application program, server, or storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0028] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure. Other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0029] Figure 1 shows a schematic diagram of the data real-time synchronization architecture of the embodiments of the present disclosure. Refer to Figure 1, the real-time data synchronization architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 can be connected through the wired or wireless network 130. Among them, the server 110 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.
[0030] The terminal 120 can be implemented by hardware or software. For example, when the terminal 120 is implemented by hardware, it can be various electronic devices with a display screen and supporting page display, including but not limited to smartphones, tablets, e-book readers, laptop computers, and desktop computers, etc. When the terminal 120 device is implemented by software, it can be installed in the above-listed electronic devices; it can be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.
[0031] It should be noted that the real-time data synchronization method provided by the embodiments of the present application can be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 the numbers of terminals, networks, and servers in
[0032] Figure 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided by the embodiments of the present disclosure. As Figure 2 shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. Among them, the processor 202, the memory 204, the network module 206, and the peripheral interface 208 are communicatively connected to each other inside the electronic device 200 through the bus 210.
[0033] The processor 202 can be a central processing unit (CPU), a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 can be used to execute functions related to the technology described in the present disclosure. In some embodiments, the processor 202 may further include multiple processors integrated as a single logic component. For example, asFigure 2 As shown, the processor 202 may include a plurality of processors 202a, 202b, and 202c.
[0034] The memory 204 may be configured to store data (e.g., instructions, computer code, etc.). As Figure 2 shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the data real-time synchronization method of the embodiments of the present disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204 and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard disk, a solid state drive (SSD), a flash memory, a memory stick, etc.
[0035] The network module 206 may be configured to provide communication with other external devices to the electronic device 200 via a network. The network may be any wired or wireless network capable of transmitting and receiving data. For example, the network may be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples. In some embodiments, the network module 206 may include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.
[0036] The peripheral interface 208 may be configured to connect the electronic device 200 to one or more peripheral devices to implement information input and output. For example, the peripheral devices may include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc. and output devices such as a display, a speaker, a vibrator, an indicator light, etc.
[0037] The bus 210 may be configured to transmit information between various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (a USB port, a PCI-E bus), etc.
[0038] It should be noted that although the architecture of the above electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208 and the bus 210, in the specific implementation process, the architecture of the electronic device 200 may also include other components necessary for normal execution. In addition, those skilled in the art can understand that the architecture of the above electronic device 200 may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0039] With the continuous deepening of enterprise informatization, the timeliness and consistency of data have become key indicators for measuring enterprise competitiveness. Traditional data replication and backup methods have inherent defects such as slow processing speed and difficulty in maintaining data consistency, and it is difficult to meet the urgent needs of modern business for data immediacy and accuracy. The emergence of Change Data Capture (CDC) technology provides enterprises with an efficient and real-time data synchronization solution. CDC can capture any data changes in the database in real time, including operations such as addition, modification, and deletion, and then quickly transfer this change information to other services for subsequent processing or consumption. This technology not only significantly improves the timeliness of data, but also effectively guarantees the consistency of data between different systems, providing more accurate and timely data support for enterprise decision-making.
[0040] However, although the Change Data Capture (CDC) technology provides an efficient and real-time data synchronization solution for enterprises in theory, in practical applications, there are significant technical defects in the real-time synchronization of big data platforms in the existing technology, mainly including the deficiencies of middleware in data collection efficiency and flexibility, the bottlenecks of data processing and conversion tools in processing performance and conversion accuracy, the complexity of data lake update operations during data storage, and the challenges of data persistence and security. In addition, the configuration and management of the entire data synchronization process are also cumbersome and difficult. These defects limit the speed, accuracy, and flexibility of data synchronization, and increase the complexity of configuration and management work. Therefore, how to improve the efficiency, accuracy, and flexibility of data real-time synchronization in big data platforms and reduce configuration operations has become an urgent technical problem to be solved.
[0041] In view of this, the embodiments of the present disclosure provide a data real-time synchronization method and related devices. By monitoring source data changes in real time and caching them into a message queue to form a change data queue, then processing the change data queue based on preset processing rules to obtain target data, and finally synchronizing the processed target data to the target data set in the data structure of the target data set, it realizes real-time capture, efficient processing, and flexible synchronization of data, and improves the accuracy and reliability of data synchronization.
[0042] See Figure 3, Figure 3 A schematic flowchart of a data real-time synchronization method according to an embodiment of the present disclosure is shown. The data real-time synchronization method according to an embodiment of the present disclosure can be deployed on a terminal or a server side. Figure 3 In [it], the data real-time synchronization method 300 may further include the following steps.
[0043] In step S310, in response to detecting that the source data has changed, cache the corresponding changed data into a message queue to obtain a changed data queue.
[0044] Wherein, the source data may refer to the original data or information source. A change may refer to a change or update of data or status. For example, data addition, data deletion, data modification, etc. Changed data may refer to the part of the data that has changed, for example, the changed part in the source data. Caching may refer to storing data in a temporary storage area for quick access or processing. A message queue may refer to a queue data structure for storing and transmitting messages. The message queue allows the system to process messages asynchronously, that is, the sender and the receiver do not need to be online or process messages immediately at the same time. In a distributed system, the message queue can be used to transfer messages and data between different services. The changed data queue may refer to a queue data structure formed after caching the changed data into the message queue, which contains all the changed parts of the data and is arranged in the order of the changes. When the source data changes, the system can detect these changes and cache the changed data into the message queue to form a changed data queue for subsequent processing and synchronization.
[0045] In some embodiments, the source data comes from different types of databases; the method 300 further includes:
[0046] Obtain the log information of the database, and generate the changed data based on the change events in the log information.
[0047] Among them, different types of databases may include, but are not limited to, relational databases (such as MySQL, Oracle), non-relational databases (such as MongoDB, Cassandra), data warehouses (such as Hive, Redshift), etc. These different types of databases may store data with different structures. Generating change data through the log information of the database can ensure the accuracy and real-time nature of data synchronization. Specifically, for relational databases, the log information can be obtained through the database's trigger or change data capture (CDC) technology, and this log information records the historical records of all insert, update, and delete operations in the data table. For non-relational databases, the way to obtain log information may be different. For example, MongoDB provides oplog (operation log) to record all write operations of the database. These log information can be parsed to extract events related to data changes, such as data insertion, update, and deletion. According to these events, corresponding change data can be generated, and this change data may include information such as the data values before and after the change, the timestamp when the change occurred, and the type of change operation.
[0048] Specifically, refer to Figure 4 , Figure 4 which shows a schematic diagram of real-time data synchronization according to an embodiment of the present disclosure. Figure 4 Among them, at the data source side: Deploy a CDC proxy service inside each data center, which is responsible for monitoring the database transaction log file and extracting the incremental update records therein. Middleware layer: Build a distributed message queue as a buffer to receive the data streams from each CDC proxy and perform preliminary processing according to preset rules (such as filtering useless information, compressing the packet size, etc.). Target side: Develop an efficient writing engine that can quickly parse the received messages and apply the changes to the local copy; at the same time, it also needs to support the breakpoint resumption function to cope with unstable network conditions. Consistency guarantee mechanism: Introduce the concept of a global clock to ensure the unification of timestamps among all nodes; use the two-phase commit protocol to coordinate cross-site transactions and avoid phenomena such as dirty reads or lost updates. Fault tolerance and recovery strategy: Set up redundant paths so that when the main line fails, it can automatically switch to the backup channel; perform a full-scale comparison check regularly to detect and correct potential data inconsistency problems in a timely manner. The relevant tools and jar packages of kafka-connect can be integrated into the execution service, and users do not need to make any settings for the execution environment to complete the collection. Through the collection ability of debeziem in kafka-conenct, the collection of the entire database is realized, and the data connection is reduced to the database level, greatly reducing the pressure on the source-side database.
[0049] In step S320, data processing is performed on the change data queue based on preset processing rules to obtain target data.
[0050] Among them, the preset processing rules can refer to the pre-defined processing logics for performing specific conversions and other operations on the input data. They can be designed according to business requirements, data specifications, or system architectures. By using the preset processing rules, it can be ensured that the data meets the requirements of the target system during the synchronization process, improving the accuracy and consistency of the data, and optimizing the data storage and query efficiency. Specifically, the preset processing rules can be set based on a simple and unified operation interface. Through the guided configuration and batch configuration on the interface, the configuration difficulty for users is greatly reduced. Users only need to make a small number of settings to achieve the synchronization of the stock and incremental data of the database.
[0051] In some embodiments, the preset processing rules include preprocessing rules and data conversion rules;
[0052] Then, data processing is performed on the change data queue based on the preset processing rules to obtain target data, including:
[0053] Performing preprocessing on the change data queue based on the preprocessing rules to obtain an intermediate data queue, where the preprocessing rules include at least one of data filtering, removing duplicate data, handling missing values, and correcting incorrect data;
[0054] Performing data conversion processing on the intermediate data queue based on the data conversion rules to obtain the target data, where the data conversion rules include at least one of data type conversion, data aggregation, data mapping, and generating preset fields.
[0055] Among them, operations such as removing duplicate data, handling missing values, and correcting incorrect data help to ensure the quality and reliability of the data. Data filtering can refer to deleting unnecessary or irrelevant data from the data set, which helps to improve performance and reduce storage requirements. Data type conversion can convert data from one format or type to another format or type, such as data type conversion and encoding unification. Data aggregation can integrate data from multiple data sources to create a more comprehensive view. Data mapping can be a process of defining the relationship between source data and target data, which can identify the fields in the source data corresponding to the fields in the target data to ensure that the converted data is consistent with the required output format. Generating preset fields can refer to calculating new fields or metrics according to business rules and requirements.
[0056] In step S330, the target data is synchronized to the target data set in the data structure of the target data set.
[0057] Among them, the target data set can refer to the target location of data synchronization, which can be a storage structure such as a database, a data warehouse, or a data lake. The target data set can have a specific data structure, including data tables, fields, data types, etc., for storing and managing the synchronized data. For exampleFigure 4 As shown, Kafka Connect is used as the middleware to collect database change events, and Apache Flink is utilized for real-time data processing and transformation. Finally, the processed data is written into a data lake (such as Hudi or Iceberg) that supports update operations. This method not only improves the speed and accuracy of data synchronization but also simplifies the configuration and management work in the entire process.
[0058] Specifically, as Figure 4 shown, real-time data synchronization for common databases (such as MySQL, Oracle, PostgreSQL, etc.) is achieved. The Debezium connector captures the changed data and transmits them to the Kafka message queue. Then, the Cdc data is updated to Hudi through the Hudi connector, enabling users to query the Hudi data through different query engines.
[0059] In some embodiments, the target data is synchronized to the target dataset in the data structure of the target dataset, including:
[0060] Determining corresponding target fields based on the data structure of the target dataset;
[0061] Performing data extraction on the target data based on the target fields to obtain the target values of the target fields;
[0062] Synchronizing the target values to the corresponding target fields in the target dataset.
[0063] Among them, the data structure of the target dataset may include the data tables, fields (columns), data types, etc. it contains. According to the definitions of these data structures, the target fields can be determined. Once the target fields are determined, the target values corresponding to these fields can be extracted from the target data, and these target values are filled into the corresponding target fields in the target dataset.
[0064] In some embodiments, method 300 further includes:
[0065] Performing data extraction on the source data to obtain the source field information of the source data;
[0066] Generating the target field information of the target dataset based on the source field information.
[0067] Among them, source field information for extracting data from source data is provided. The source field information may include field names (column names), data types (such as integers, strings, dates, etc.), field lengths (if applicable), field descriptions (if provided), and possible field constraints (such as non-null constraints, uniqueness constraints, etc.). Target field information is created in the target dataset based on the source field information to ensure that data can be correctly synchronized from the source dataset to the target dataset. The target field information can be generated based on field name mapping, data type conversion, field length adjustment, and application of field constraints. For example, if the names of the source field and the target field are different, a field mapping relationship can be established through a configuration file, a mapping table, or programming logic. Sometimes the data types of the source field and the target field may be different, and data type conversion rules can be defined to ensure that data is not lost or damaged during synchronization. If the lengths of the source field information and the target field are different, the data can be truncated or padded to fit the length of the target field. The target field may need to apply the same constraints as the source field (such as non-null constraints, uniqueness constraints, etc.) to ensure data integrity and consistency.
[0068] The target field information may refer to Schema field information. Since the source database has a field structure, when implementing the background task, there is no need to configure the field information, and the field information will be automatically extracted from the source table, avoiding cumbersome user configuration.
[0069] In some embodiments, method 300 further includes:
[0070] Determining associated fields in the source data and / or the target data based on configuration parameters of the target dataset;
[0071] Generating configuration data corresponding to the configuration parameters based on associated data of the associated fields.
[0072] Among them, the configuration parameters define the structure, rules, and mapping relationship between the target dataset and the source data, and may include field names, data types, field lengths, field constraints, association relationships, etc. The associated fields may refer to fields in the source data and the target data used to establish data associations, and these fields may include primary keys, foreign keys, and association fields in business logic. By parsing and comparing the configuration parameters of the target dataset, it is possible to identify which fields in the source data and the target data are associated to determine the mapping relationship between the source field and the target field.
[0073] For example, configuration parameters may be stored in configuration files such as XML, JSON, YAML, etc. These configuration files can be read, and the field mapping relationships therein can be parsed. If the configuration parameters are stored in a database, database queries can be executed to obtain this information. In some cases, logical judgments can be made to determine the associated fields. For example, the association relationships between fields can be inferred based on business rules or data patterns. Once the associated fields are determined, the configuration data corresponding to the configuration parameters can be generated based on the associated data of the associated fields. Data values related to the associated fields can be extracted from the source data and / or target data. If the associated data needs to be converted in format or type to meet the requirements of the target data set, corresponding data conversion operations can be performed. The generated configuration data may need to be stored in a certain location for subsequent data synchronization or processing, such as writing the data to a configuration file, a database, or other storage media. For the configuration information of the data lake, such as primary key information, key generators, and index types, it is very troublesome for full-library collection, and these configuration information can be automatically inferred according to the table type of the collection table.
[0074] In some embodiments, method 300 further includes:
[0075] Obtaining the operation status data of the data synchronization, where the operation status data includes at least one of log data, execution data, and resource occupancy data;
[0076] Generating index data of a preset task metric based on the operation status data, where the preset task metric is associated with the operation status data;
[0077] Displaying the operation status data and the index data.
[0078] Among them, the operation status data may refer to various data generated during the execution of the data synchronization task, which is used to reflect the running status of the task. The operation status data may include log data, execution data, resource occupancy data, etc. The log data can record various events, errors, warnings, etc. during the data synchronization process, which helps to troubleshoot problems and understand the detailed process of task execution. The execution data may include the start time, end time, amount of synchronized data, synchronization speed, etc. of the synchronization task, which is used to evaluate the execution efficiency and effect of the task. The resource occupancy data can reflect the occupancy of system resources (such as CPU, memory, disk I / O, etc.) by the data synchronization task, which helps to evaluate the impact of the task on system performance. The log data can be collected through a logging framework (such as Log4j, logback, etc.) or a system logging service. The execution data can be collected by embedding timers and counters in the synchronization task. The resource occupancy data can be collected using system monitoring tools (such as top, vmstat, iostat, etc.) or dedicated resource monitoring services.
[0079] The preset task metrics can be metrics predefined for evaluating the task execution effect according to business requirements and the characteristics of the data synchronization task. These metrics may be directly related to the running status data, or may need to be obtained through certain calculations or conversions. The metric data can be the actual data values calculated based on the running status data and the preset task metrics. Specifically, before the task starts, appropriate task metrics can be defined according to business requirements and the characteristics of the data synchronization task. Based on the defined task metrics and the collected running status data, metric calculations (such as simple mathematical operations, statistical analysis, or complex algorithm processing) are performed. The calculated metric data can be stored in a certain location for subsequent analysis and reporting. The collected running status data and the calculated metric data are presented in an intuitive way so that users can easily understand the running status and execution effect of the data synchronization task. For example, a graphical interface (such as a dashboard, line chart, bar chart, etc.) can be used to display the running status data and the metric data. The running status data and the metric data can be output to a log file or the console for detailed viewing and analysis. In this way, the logs, execution status, resource occupancy, and task metrics of the cluster task can be monitored. Further, when the running status data or the metric data reaches a preset threshold, an alarm mechanism can also be triggered to alert the user.
[0080] It can be seen that by obtaining the running status data of data synchronization, generating the metric data of the preset task metrics based on these data, and then displaying them in an intuitive way, the running status and execution effect of the data synchronization task can be effectively monitored and evaluated. This helps to timely discover and solve problems, and improve the reliability and efficiency of data synchronization. At the same time, this process also provides important data support for subsequent data analysis and reporting.
[0081] In some embodiments, method 300 further includes:
[0082] In response to receiving a data request for the target data set, determining matching data from the target data set based on the query information in the data request;
[0083] Returning the matching data to the initiator of the data request.
[0084] Among them, the data request can be a data acquisition request initiated by the initiator, aiming to retrieve corresponding data from the target dataset. The data request can contain query information for specifying the specific conditions or scope of the data to be retrieved. The data request can be received and parsed to understand the data that the initiator wants to retrieve. For example, the data request can be received through an API interface, and the query information in the request can be parsed. The query information can refer to the information contained in the data request for specifying the retrieval conditions. The query information can include field names, operators (such as equal to, greater than, less than, etc.), values, and possible logical operators (such as AND, OR, etc.). Query operations are performed on the target dataset according to the query information to find matching data that meets the conditions. For example, SQL statements can be used to perform query operations. To improve query efficiency, indexing and caching technologies can be adopted to accelerate the data retrieval process. After retrieving the matching data, the retrieval result can be returned to the initiator of the data request. The returned matching data may include data in the target dataset, processed data in the target dataset, or a combination of both. The returned data can be presented to the initiator in an easy-to-understand and processable manner, such as JSON, XML, CSV, or other formats. In this way, by responding to the data request, parsing the query information, performing query operations on the target dataset, and returning the matching data, a flexible and powerful data retrieval function can be provided for users. This helps users quickly obtain the required data, thus supporting various business decision-making and analysis tasks.
[0085] Specifically, refer to Figure 4 , Figure 4Common databases can be connected: including common relational databases such as MySQL, Oracle, and PostgreSQL. The data in these databases will be synchronously replicated to the data lake in real time. Source Task: This is the first step of data collection. The Debezium connector is responsible for capturing all insert, update, and delete events from the above databases and converting them into messages in the Kafka message queue. Kafka message queue: This is a message queue used to temporarily store the CDC data obtained from the Source Task. This data will wait in the queue to be processed by the Sink Task. Sink Task: This is the second step of data collection. It extracts the CDC data from the Kafka message queue in real time and updates it to the Hudi data lake through the Hudi connector. Hudi data: This is a part of the data lake used to store the data synchronously replicated from the source database. Users can query and analyze the Hudi data through different query engines (such as Spark and Presto). Query engines: including two query engines, Spark and Presto, both of which can access the Hudi data and perform queries and analysis on it, thus providing more powerful analysis capabilities. In this way, users can easily synchronize data from the source database to the data lake in real time without worrying about complex configuration and maintenance work. Through features such as a simple and unified operation interface, deeply encapsulated tool integration capabilities, the ability to reduce the pressure on the source database, automated configuration inference, and metric-level monitoring, users can conveniently and quickly achieve real-time data synchronization and processing. Metric-level monitoring can monitor information such as the logs, execution status, resource occupancy, and task metrics of cluster tasks, helping users better understand the status and performance of data synchronization, as Figure 5 shown Figure 5 shows a schematic diagram of metric-level monitoring according to an embodiment of the present disclosure.
[0086] The configuration difficulty for users can be greatly reduced through interface-guided configuration and batch configuration. Users only need to make a small number of settings to achieve the synchronization of the stock and incremental data of the database, as Figure 6 shown Figure 6A schematic diagram of a configuration interface according to an embodiment of the present disclosure is shown. Specifically, source resources can be added to the configuration interface. For example, a Mysql resource can be added. After adding the resource, the source common attributes are automatically filled, and the default owner user is the database name of the resource. The source common attributes can also be set in the configuration interface. If no configuration is required, the default values can be used. Hovering the mouse over the attribute can view the attribute description. Target resources can also be added to the configuration interface. For example, the target resource can be set as a Kafka resource. The target common attributes can also be set in the configuration interface. For example, the common attributes are automatically filled. By default, it is the node name, and the topic name of the Kafka message can be selected from the dropdown box as the Kafka topic. Source objects can also be added to the configuration interface. For example, click to add a new one from the metadata object, filter by resource name and owner user, and view all table objects under this resource on the platform. Single addition of metadata objects is supported, or click on batch operation, check the specified objects, and batch addition of metadata objects is supported. Adding table objects from the database is supported. Enter to view the database object list. On the left are all the schemas with permissions under this resource, and on the right are all the tables under the current schema. Single addition and batch addition are also supported. The addition of target objects can also be set in the configuration interface. For example, after selecting the source metadata object, according to the same encoding and structure as the source table, the target resource entity object is automatically matched. If the matching is unsuccessful, the outer frame of the object is shown as a dotted line. Proxy resources can also be configured. For example, the optional execution node is the proxy node, and the proxy node should have the relevant environment required for configuring the CDC task. Engine parameters can also be configured. By default, the Kafka-connect engine is not shown. The CDC task provides a default list of engine parameters, and adding the broker.list, the broker address of Kafka (generally, the only engine parameter that needs to be modified in the engine parameter list is the broker.list), delaytime.check.service, and delaytime.check.task for the delay duration (in seconds) of successful detection of the connect service is supported. Figure 4 Among them, the source-side database can include MySQL / PostgreSQL, etc. CDC component: Debezium Connector is responsible for listening for changes in the source database. Middleware: Apache Kafka is used as a message queue to cache the changed data from different tables. Real-time computing engine: Apache Flink reads the data stream in Kafka and performs ETL operations. Target storage: Hudi or Iceberg, supporting upsert operations to maintain data consistency.
[0087] CDC data collection: Use the Debezium framework to uniformly access various types of databases. For each table that needs to be synchronized, create a corresponding Kafka Topic. Set the partition rules according to business requirements to ensure sequential consistency.
[0088] Data transmission and caching: The change logs generated by Debezium are directly sent to Kafka. The message queue feature of Kafka is used for data buffering.
[0089] Real-time data processing: Flink subscribes to topics from Kafka and receives change events. Corresponding data processing logics are defined in Flink, such as field mapping, type conversion, etc.
[0090] Data writing and management: The processed data is written into Hudi or Iceberg. The automatic modeling function is supported to automatically generate the target table structure according to the source table structure. A visual interface is provided for managing and monitoring the status of CDC tasks.
[0091] Introducing Kafka as a cache layer reduces the pressure on memory. Automatically inferring the table structure reduces the manual configuration workload of users. Batch-configuring the data synchronization tasks of multiple tables improves efficiency. Integrating the lineage tracking mechanism facilitates subsequent data analysis.
[0092] It can be seen that according to the method of the embodiments of the present disclosure, when the source data changes, it can be achieved through technologies such as database triggers and change data capture (CDC), and the changes in the data in the source system can be monitored in real time. After detecting the changes, the corresponding changed data can be cached in the message queue to form a changed data queue to ensure that the changed data is processed in the order of occurrence. Based on the preset processing rules, the changed data queue is processed to ensure that the target data conforms to the data structure and requirements of the target data set, and the valid target data can be extracted from the changed data queue. The processed target data is synchronized to the target data set in the data structure of the target data set. In this way, the changes in the source data can be captured in real time and synchronized to the target data set, ensuring that the data in the target data set is always consistent with the source data, and improving the timeliness and accuracy of data synchronization. By using the message queue to cache the changed data, batch processing of data can be achieved, reducing the direct access times to the data, thereby improving the efficiency of data synchronization. At the same time, the preset processing rules can automatically process the data, reducing the need for manual intervention and further improving the synchronization efficiency. Customizable processing rules can also be supported, enabling the system to flexibly process data according to different business requirements, enhancing the scalability and adaptability of data synchronization.
[0093] It should be noted that the method of the embodiments of the present disclosure can be executed by a single device, such as a computer or a server. The method of this embodiment can also be applied to a distributed scenario, and multiple devices cooperate with each other to complete it. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the embodiments of the present disclosure, and these multiple devices will interact with each other to complete the described method.
[0094] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired results. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0095] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides a data real-time synchronization device, see Figure 7 , the data real-time synchronization device, the device includes:
[0096] A cache module, configured to cache the corresponding changed data into a message queue to obtain a changed data queue in response to detecting a change in the source data;
[0097] A processing module, configured to perform data processing on the changed data queue based on a preset processing rule to obtain target data;
[0098] A synchronization module, configured to synchronize the target data to the target data set in the data structure of the target data set.
[0099] For the convenience of description, when describing the above device, it is divided into various modules according to functions and described separately. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.
[0100] The device of the above embodiment is used to implement the corresponding data real-time synchronization method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0101] Based on the same inventive concept, corresponding to the method of any of the above embodiments, the present disclosure further provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the computer to execute the data real-time synchronization method described in any of the above embodiments.
[0102] The computer-readable medium of this embodiment includes both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device.
[0103] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the data real-time synchronization method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.
[0104] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary, and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above. For the sake of brevity, they are not provided in detail.
[0105] In addition, for simplicity of explanation and discussion, and in order not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the devices may be shown in block diagram form in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In the case where specific details (such as circuits) are set forth to describe the exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure can be implemented without these specific details or with variations of these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0106] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0107] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the embodiments of the present disclosure should be included within the protection scope of the present disclosure.
Claims
1. A real-time data synchronization method, comprising: In response to detecting that the source data has changed, caching the corresponding changed data in a message queue to obtain a changed data queue; Performing data processing on the changed data queue based on preset processing rules to obtain target data; The target data is synchronized to the target data set in the data structure of the target data set.
2. The method according to claim 1, wherein: The preset processing rules include preprocessing rules and data conversion rules; Then, the changed data queue is processed based on the preset processing rules to obtain target data, including: Preprocessing the changed data queue based on the preprocessing rule to obtain an intermediate data queue, wherein the preprocessing rule includes at least one of data filtering, removing duplicate data, processing missing values, and correcting erroneous data; The intermediate data queue is subjected to data conversion processing based on the data conversion rule to obtain the target data, wherein the data conversion rule comprises at least one of data type conversion, data aggregation, data mapping, and generation of preset fields.
3. The method according to claim 2, wherein: The target data is synchronized to the target data set in the data structure of the target data set, including: Determine a corresponding target field based on the data structure of the target data set; Extracting data from the target data based on the target field to obtain a target value of the target field; The target value is synchronized to the corresponding target field in the target data set.
4. The method according to claim 1, further comprising: Extracting the source data to obtain source field information of the source data; Generate target field information of the target data set based on the source field information; and / or, Determining associated fields in the source data and / or the target data based on configuration parameters of the target data set; Configuration data corresponding to the configuration parameter is generated based on the associated data of the associated field.
5. The method according to claim 1, further comprising: Acquire the running status data of the data synchronization, wherein the running status data includes at least one of log data, execution data, and resource occupancy data; generating indicator data of a preset task indicator based on the operation status data, wherein the preset task indicator is associated with the operation status data; The operating status data and the indicator data are displayed.
6. The method according to claim 1, wherein: The source data comes from different types of databases; the method further includes: The log information of the database is obtained, and the change data is generated based on the change events in the log information.
7. The method according to claim 1, further comprising: In response to receiving a data request for the target data set, determining matching data from the target data set based on query information in the data request; The matching data is returned to the initiator of the data request.
8. A real-time data synchronization device, comprising: A cache module, configured to cache corresponding changed data to a message queue to obtain a changed data queue in response to detecting that a change has occurred in the source data; A processing module, used for performing data processing on the change data queue based on a preset processing rule to obtain target data; A synchronization module is used to synchronize the target data to the target data set in the data structure of the target data set.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 7 when executing the program. 10 . A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method according to claim 1 .