Data real-time synchronization method, system and device of distributed architecture and medium

By defining data source and subscription relationships in a distributed environment, and using incremental data listeners and message queues to achieve data synchronization, the real-time, accuracy and orderly data synchronization problems in distributed environments are solved, and the reliability of data synchronization is improved.

CN120162384APending Publication Date: 2025-06-17INSPUR GENERSOFT CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510254140.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

In a distributed environment, the prior art is difficult to achieve real-time, accurate and orderly data synchronization, and cannot effectively avoid data loss, duplication or incorrect synchronization.

Method used

By defining data sources and subscription relationship tables, configuring wildcards or specifying service units, using incremental data listeners to capture data changes in real time, generate data change sets, and sending them to the subscription end through message queues, ensuring that data is synchronized in real time, accurately and orderly in a distributed environment.

Benefits of technology

Real-time, accurate and orderly synchronization of data in a distributed environment is realized, and the problems of data loss, duplication or incorrect synchronization are avoided, and the reliability of data synchronization is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120162384A_ABST
    Figure CN120162384A_ABST
Patent Text Reader

Abstract

The invention provides a data real-time synchronization method, system and device of a distributed architecture and a medium. The method comprises the following steps: defining a subscription relationship between a data source and data, and judging a synchronization state; if all sub-libraries need to be synchronized, wildcard characters are used; if only partial synchronization is carried out, corresponding service units are configured; deploying an incremental data monitor according to the data source type; capturing changing data in the system in real time through an incremental data monitor; checking whether the data source has defined synchronous data or not; if yes, configuring a data change set for the synchronous data according to a preset structure; according to the subscription relationship of the data, sending the data change set to a subscription end through a message queue; and the subscription end analyzes the change set, organizes the change set into a standard SQL, and persists the standard SQL to a target end business data sub-library. According to the invention, real-time, accurate and orderly synchronization of data in a distributed environment is ensured, the problems of data loss, repetition or wrong synchronization and the like are effectively avoided, and the reliability of data synchronization is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data real-time synchronization, and particularly relates to a method, system, device and medium for data real-time synchronization in a distributed architecture. Background Art

[0002] Distributed systems can operate in different data centers and geographical locations, providing users with high availability, high performance and high scalability. By dispersing different components of the system in different physical or logical locations, distributed systems achieve decoupling and independent expansion of the system. In a distributed system, data synchronization can ensure data consistency and availability.

[0003] However, implementing data synchronization in a distributed environment is very complex because the architectures of related technologies are relatively fixed in data source definition and subscription relationship configuration, making it difficult to adapt to different business entity types and complex synchronization requirements. It is impossible to switch between full sub-library synchronization or partial service unit synchronization in terms of the data synchronization scope. Moreover, in terms of capturing changed data, it is impossible to ensure the timeliness, accuracy and processing sequence of data capture. Consequently, accurate data capture and processing sequence cannot be guaranteed. Summary of the Invention The present invention provides a method for data real-time synchronization in a distributed architecture. The method ensures real-time, accurate and orderly synchronization of data in a distributed environment, effectively avoiding problems such as data loss, duplication or incorrect synchronization, and improving the reliability of data synchronization.

[0004] The method includes: Define a data source, create a subscription relationship table to record the subscription relationship between the data source, the changed table name and the target service unit, and judge the synchronization status; If it is necessary to synchronize to all sub-libraries, use wildcards; if only partial synchronization is required, configure the corresponding service unit; According to the selected data source type, select an incremental data listener to capture the data change status in real time; Capture the changed data in the system in real time through the incremental data listener; Check whether there is defined synchronization data in the data source; If there is, configure a data change set for the synchronization data according to a preset structure; Send the data change set to the subscription end through a message queue according to the data subscription relationship, and record the synchronization log of the publishing end at the same time; After receiving the data change set, the subscription end parses the change set and organizes it into a standard SQL, and persists it to the target-side business data sub-library.

[0005] It should be further noted that the step of deploying the listener according to the data source type further includes: Configure JPA data listeners to listen for data changes; Configure business entity listeners to listen for data changes; Customize data listeners to support other types of data changes.

[0006] Furthermore, it should be noted that defining the data source and the subscription relationship of the data, and judging the synchronization status also includes: selecting the data source implementation method according to the business entity type; Listen to its persistence operations; Define the table structure and change rules based on the business entity framework; Implement the listening of non-standard persistence methods through the extended interface; Configure the service unit and preset wildcards to represent that the data needs to be synchronized to all sub-databases; In the method, judge whether the current data change needs to trigger synchronization according to the subscription relationship table of the data source; The logical condition for judgment is: if the synchronization data instruction ∈ subscription relationship table and the target service unit matches the subscription configuration, then mark it as needing to be synchronized.

[0007] Furthermore, it should be noted that the method also includes: Obtain the data change set; Generate the synchronization key data points of the synchronization data according to the synchronization data instruction and the data structure type; Generate a data change sequence according to the synchronization key data points and the data structure type; Obtain the synchronization data label from the preset message queue according to the synchronization key data points, and store the data change sequence in the message queue according to the synchronization data label; Extract the data change sequence from the message queue, and create a synchronization data channel between the data change sequence and the preset subscription end; Store the data change sequence in the subscription end according to the synchronization data channel.

[0008] Furthermore, it should be noted that generating the synchronization key data points of the synchronization data according to the synchronization data instruction and the data structure type also includes: During the data change process, define the synchronization key data points that uniquely identify the change operation. The synchronization key data points have a primary key field, an operation type, and a timestamp; Configure the synchronization data instruction, and the synchronization data instruction comes from the subscription relationship configuration in step S101; Define the precondition as that the data change has been captured and verified that it needs to be synchronized; Parse the synchronization instruction, extract the target service unit list of the current change table from the subscription relationship table, and determine the range of fields to be synchronized; Extract key data points and generate a set of key data points based on the primary key field, operation type, and changed fields; Synchronize the set of key data points as the input for generating the data change sequence, and generate the data change sequence based on the key data points; Among them, if the primary key is not defined for the table, it is marked as an exception, triggering an alarm and terminating the synchronization; If the subscription - side table structure lacks key fields, record an error log and enter the retry process.

[0009] Furthermore, it should be noted that the step of obtaining the synchronization data label from the preset message queue according to the synchronized key data points and storing the data change sequence in the message queue based on the synchronization data label further includes: Obtain the data structure type of the synchronized key data points, and determine the synchronization data format of the synchronized key data points according to the data structure type; Generate a synchronization task for the synchronized key data points according to the synchronization data format; Perform unified subscription - side configuration on the synchronization task to obtain the unified task configuration information; Obtain the initial synchronization data label of the message queue, and filter out the synchronization data label from the initial synchronization data label based on the unified task configuration information and the synchronization task.

[0010] Furthermore, it should be noted that the step of storing the data change sequence to the subscription - side according to the synchronization data channel further includes: Obtain the configured synchronization transmission instruction of the data change sequence according to the synchronization data channel; Perform parameter category configuration on the configured synchronization transmission instruction to obtain the configured synchronization data structure type; Execute the configured synchronization transmission instruction according to the configured synchronization data structure type and the synchronization data channel to determine that the data change sequence is stored to the subscription - side.

[0011] This application also provides a data real - time synchronization system with a distributed architecture. The system includes: A status judgment module, used to define the data source and the subscription relationship of the data, and judge the synchronization status; A synchronization execution module, used to use wildcards if synchronization to all sub - databases is required; configure the corresponding service units if only partial synchronization is needed; A listening deployment module, used to deploy an incremental data listener according to the data source type; A listening capture module, used to capture the changed data in the system in real - time through the incremental data listener; A data detection module, used to check whether there is defined synchronization data in the data source; if so, configure the data change set for the synchronization data according to the preset structure; A data sending module, configured to send a data change set to a subscription end through a message queue according to the subscription relationship of data, and record a publisher synchronization log at the same time; A data processing module, configured to, after the subscription end receives the data change set, parse the change set and organize it into a standard SQL, and persist it to a target-side business data sub-database.

[0012] According to another embodiment of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the data real-time synchronization method of the distributed architecture are implemented.

[0013] According to still another embodiment of the present application, there is further provided a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data real-time synchronization method of the distributed architecture are implemented.

[0014] It can be seen from the above technical solutions that the present invention has the following advantages: The data real-time synchronization method of the distributed architecture provided by the present application realizes flexible control of the data synchronization range through simple wildcards or specified service unit configurations. Whether it is to broadcast the data change set to all or only send it to the MQ Topic corresponding to a specific service unit, it can be completed.

[0015] This method specifies that each data source corresponds to a listener instance, and binds the data source ID and table name, and can accurately capture the changed data of a specific data source. The listener of this method can capture the incremental data in real time after the database transaction is committed, and store it in the memory queue in the order of events, ensuring the accurate capture and processing timing of the data. Moreover, this method captures the incremental data from after the database transaction is committed, judges whether to synchronize according to the subscription matching rules, and then organizes the synchronized data into a change set and sends it according to the preset structure, ensuring the real-time, accurate and orderly synchronization of the data in the distributed environment, effectively avoiding problems such as data loss, duplication or incorrect synchronization, and improving the reliability of data synchronization. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the data real-time synchronization method for the distributed architecture; Figure 2 It is a flowchart of an embodiment of the data real-time synchronization method for the distributed architecture; Figure 3 Schematic diagram of a data real-time synchronization system for a distributed architecture; Figure 4 Schematic diagram of an electronic device. Specific implementation manners

[0018] The following will describe in detail the specific steps of the data real-time synchronization method for a distributed architecture. For illustration rather than limitation, specific details such as specific system structures and technologies are proposed to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details.

[0019] Statements such as "in one embodiment" or "in some embodiments" described in the present application mean that specific features, structures, or characteristics described in the embodiment are included in one or more embodiments of the present application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" that appear in different places in the present application do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to Figure 1 The following is a flowchart of the data real-time synchronization method for a distributed architecture in a specific embodiment. The method includes: S101: Define a data source, create a subscription relationship table, record the subscription relationship between the data source, the changed table name, and the target service unit, and judge the synchronization status.

[0022] In an exemplary embodiment, the defined data source is the source that generates the data to be synchronized. Exemplarily, the data source can be a node where data such as users and organizations are stored. The target service unit refers to the destination of data synchronization, including microservice units and business data sub-databases. A microservice unit is a service unit that processes business logic in a microservice architecture; a business data sub-database, or sub-database, is a storage area where data is split and stored according to business modules.

[0023] For example, based on the JPA data source, the database table structure can be mapped through the Hibernate framework, and its persistence operations can be monitored. Define the table structure and change rules based on the business entity framework. The monitoring of non-standard persistence methods can be implemented through extended interfaces.

[0024] The configuration of the subscription relationship is based on the configured target microservice unit. A wildcard can be preset to indicate that the data needs to be synchronized to all sub-databases.

[0025] In this embodiment, according to the subscription relationship table of the data source, specifically storing the mapping of table names and service units, it is determined whether the current data change needs to trigger synchronization.

[0026] Optionally, the subscription relationship table created in this embodiment can include fields such as the data source name, the changed table name, and the list of target service units. Based on querying the subscription relationship table, it is determined whether the changed table name is within the subscription scope.

[0027] If there is a subscription relationship, extract the list of target service units; if the list of target service units is empty, discard the changed data. For example, when data changes occur in the user table in the data source, query the subscription relationship table to check if there is a service unit subscribing to the data of the user table. If so, synchronize the changed data to the corresponding service unit.

[0028] Optionally, the logical condition for judgment is: if the data synchronization instruction ∈ the subscription relationship table and the target service unit matches the subscription configuration, it is marked as needing to be synchronized.

[0029] S102: If it needs to be synchronized to all sub-databases, use wildcards; if only partial synchronization is required, configure the corresponding service unit.

[0030] In an exemplary embodiment, when the subscription end is all business sub-databases, the data change set will be broadcast to all MQTopics. Of course, in this embodiment, the service unit can also be specified, and the data change set will only be sent to the corresponding MQ Topic.

[0031] The processing method can be: target Topic = sync / ${subscribed service unit}. If the subscribed service unit is a preset wildcard character, then the Topic is sync / all.

[0032] In this embodiment, the Topic classifies the data. Different Topics represent different data structure types, which enables the subscription end to filter and change different inconsistent data. Through the Topic, loose coupling is achieved between the publishing end and the subscription end. The publishing end only needs to publish the message to the Topic, while the subscription end can subscribe to the Topic as needed without directly interacting with the publishing end.

[0033] S103: Deploy an incremental data listener according to the data source type.

[0034] In this embodiment, each data source corresponds to a listener instance, and the listener needs to be bound to the data source ID and the table name.

[0035] S104: Real-time capture of the changed data in the system through the incremental data listener.

[0036] For the change capture logic, the listener captures the incremental data after the database transaction is committed. Extract the changed table name, operation type, primary key value, and field value. Store the changed data in the memory queue in the order of events to ensure the timing.

[0037] S105: Check whether there is defined synchronization data in the data source.

[0038] In this embodiment, subscription matching rules can be defined. Based on querying the subscription relationship table, determine whether the changed table name is within the subscription scope. If there is a subscription relationship, extract the target service unit list. If the target service unit list is empty, discard the changed data; otherwise, proceed to S106.

[0039] S106: If so, configure the data change set for the synchronization data according to the preset structure.

[0040] In this embodiment, when it is determined that the captured data is the data that needs to be synchronized, the synchronization data is organized into a data change set according to the preset structure.

[0041] The preset structure in this embodiment includes but is not limited to: JSON data structure, XML structured data, etc. For the data change set with the preset structure, the business identifier or timestamp, etc. can be set according to the actual usage situation. S107: Send the data change set to the subscription end through the message queue according to the data subscription relationship, and record the synchronization log of the publishing end at the same time.

[0042] As an implementation manner of this embodiment, at the data transmission level, it is implemented by using the MQ message queue + Two Way high-reliability confirmation mode, and the transaction consistency between the data sender and the business persistence layer is fully considered.

[0043] Specifically, the logs of the data publishing end and the subscription end are consistent, and it is confirmed that the synchronization is successful.

[0044] Otherwise, retry according to the retry policy, for example, trigger the synchronization again at 5 seconds, 20 seconds, and 60 seconds.

[0045] In the case where multiple retries still fail, it will be marked as abnormal data, a warning will be sent, and after manual intervention to analyze and solve the problem, the synchronization logic will be triggered again.

[0046] As an implementation of step S107, the message sending logic can use the Two Way transaction mode: first commit the local transaction, record the publishing log, and then send the MQ message. If the MQ sending fails, roll back the local transaction to avoid data inconsistency.

[0047] The fields of the publishing end log table can be represented as: message ID, table name, operation type, sending time, target service unit, success status / failure.

[0048] S108: After the subscribing end receives the data change set, parse the change set and organize it into standard SQL, and persist it to the target-side business data sub-database.

[0049] In this embodiment, the corresponding SQL can be generated according to the operation type. And duplicate execution is avoided by the message ID or (tableName, keyField). After executing the SQL, record the subscribing end log table, and the fields include: message ID, execution time, status, error message. If the execution fails, trigger the retry mechanism, that is, return to step S105 for retry.

[0050] It can be seen that through the linkage of the local transaction and the MQ message, it is ensured that sending means persistence. Based on the MQ partitioning strategy, the change order of the same entity is guaranteed. Through the custom data source and listener, heterogeneous system access is supported. High-reliability and low-latency data real-time synchronization in a distributed environment are achieved.

[0051] In an embodiment of the present invention, based on step S103, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.

[0052] This embodiment is based on the framework interception technology, realizes non-intrusiveness to business functions, and supports custom extension for heterogeneous services; through the incremental data listener, the data that changes in the system is captured in real time.

[0053] The incremental data listener provides three types of listeners, including the JPA data listener, the business entity listener, and the custom data listener. Among them, the JPA data listener mainly listens to the data persisted based on Hibernate; the business entity listener mainly listens to the data persisted based on the business entity framework (BEF); the custom data listener provides a custom extension mechanism to listen to the data persisted by other means.

[0054] The specific execution steps are as follows: S1031: Determine whether the data captured by the listening needs to be synchronized, that is, there is corresponding definition configuration information in the synchronization data source.

[0055] S1032: Organize the streaming data captured by monitoring and needing to be synchronized into a data change set in the unified structure of {tableName: "", operation: "", keyfield: "", rowData: [{field: "", value: ""}]}.

[0056] S1033: Send the data change set to the subscription end through MQ according to the subscription relationship of the synchronized data, and record the synchronization log of the publishing end at the same time.

[0057] S1034: After receiving the change set, the subscription end records the processing log of the subscription end, then parses the change set, organizes it into a standard SQL, and persists it to the business data sub-library of the target end.

[0058] In this embodiment, step S1031 is used to judge whether the data captured by monitoring needs to be synchronized, ensuring that only the data with corresponding defined configuration information in the synchronization data source will be processed, avoiding the interference of irrelevant data, and ensuring that the data entering the subsequent processing process is accurate and needs to be synchronized. Step S1032 organizes the streaming data captured by monitoring and needing to be synchronized into a data change set in a unified structure, enabling data from different sources and of different types to be processed in a standardized format in the subsequent process, facilitating the transmission, storage, and parsing of data in the system. Step S1033 uses the message queue to send the data change set. The message queue has characteristics such as asynchronous decoupling and peak shaving and valley filling, enabling efficient data transmission, avoiding performance problems caused by the mismatch between the processing speeds of the sending end and the receiving end, and enabling data to be quickly transmitted from the publishing end to the subscription end. In step S1034, the subscription end parses the received change set, organizes it into a standard SQL, and persists it to the business data sub-library of the target end, ensuring that the data can be accurately stored in the target database.

[0059] Based on the above embodiment, in order to further improve the reliability of the data real-time synchronization method for the distributed architecture provided by the above embodiment, the following is an implementable way. In one embodiment, as Figure 2 shown, the data real-time synchronization method for the distributed architecture further includes the following steps: Step S211, obtain the data change set.

[0060] In this embodiment, the data change set includes a synchronization data instruction, the data structure type of the synchronization data instruction, and the synchronization data matching the data structure type of the synchronization data instruction.

[0061] The synchronization data instruction is an instruction used to direct the system to perform data synchronization operations and specify the specific requirements for synchronization. It can be set in the system configuration file to clearly specify which data needs to be synchronized. For example, according to the business module division, synchronization instructions for user module data, order module data, etc. can be set; it can also be set through a visual management interface, and in a graphical interaction manner, select the data source, target end, and synchronization trigger conditions (such as real-time synchronization, scheduled synchronization, etc.) that need to be synchronized.

[0062] When the system executes the data synchronization process, it will read the synchronization data instruction. When the incremental data listener captures data changes, it will determine whether these changed data belong to the category of data that needs to be synchronized according to the synchronization data instruction. If it meets the instruction requirements, it will be processed according to the subsequent process; if not, no synchronization operation will be performed.

[0063] In the method for real-time data synchronization in a distributed architecture, the message queue is used as follows: According to the data subscription relationship, the data publisher sends the sorted data change set to the message queue, and the subscriber then extracts these data change sets from the message queue to achieve asynchronous transmission of data from the publisher to the subscriber, decoupling the direct dependence between the publisher and the subscriber.

[0064] When data changes occur, the data change sequence can be stored in the message queue first. If the subscriber is temporarily unable to process or there are network fluctuations, etc., the message queue can cache this data and process it after the subscriber returns to normal, ensuring that the data will not be lost and can be processed in an orderly manner.

[0065] In the case of a large amount of data changes, the message queue can temporarily store a large number of data change requests, avoiding the system from crashing due to the subscriber receiving too much data instantly, playing a role in peak shaving and valley filling, and making the system run more stably.

[0066] Step S211 captures the changed data in the system in real time through the incremental data listener, and then checks whether these changed data belong to the data that needs to be synchronized defined previously. If it belongs to the category of synchronized data, it will be configured into a data change set according to the preset structure. And what is obtained in step S211 is this data change set with the configured structure.

[0067] The process of obtaining the data change set depends on a series of previous operations, including the definition of the data source and subscription relationship, the capture of data changes, and the screening and judgment of data, etc. Only through these pre-steps can the data change set that meets the synchronization requirements be accurately obtained.

[0068] Step S212 generates the synchronization key data points of the synchronized data according to the synchronization data instruction and the data structure type.

[0069] In some embodiments, a synchronization key data point refers to a core field or metadata that uniquely identifies a change operation, ensures data consistency, and drives synchronization logic during the data change process. Its functions include: setting a uniqueness identifier, that is, determining the target record of the change operation (such as a primary key field); operation semantics definition: clarifying the change type; timing control: ensuring the order of changes.

[0070] The setting parameters of the synchronization key data point include: primary key field, operation type, and timestamp.

[0071] The primary key field setting method is to specify the primary key of the table when defining the data source. In this way, the subscription end locates the target record through the primary key and performs update or delete operations.

[0072] The operation type setting method is automatically captured by the incremental listener. Its function is to determine the type of SQL generated by the subscription end (insert, delete, update).

[0073] The timestamp setting method is to add a `version` or `update_time` field to the business table, and the listener automatically captures its value. It solves concurrent conflicts (such as optimistic locking) or ensures the order of changes.

[0074] The use of the synchronization key data point can generate synchronization instructions. The primary key and operation type are included in the data change set. - The subscription end locates the target record according to `keyField` and `keyValue`, and generates SQL in combination with `operation`. If there is a version number, the subscription end needs to verify the version to avoid overwriting unexpected changes.

[0075] It should be noted that a data change sequence refers to a set of data change operations arranged in chronological order, and each operation contains complete change information (such as table name, operation type, primary key, field value, etc.). The data change sequence can ensure data consistency. By processing changes in sequence, it avoids data state errors caused by out-of-order (such as deleting before inserting). It supports transactional synchronization, combining multiple changes within the same transaction into one sequence to ensure atomic execution at the subscription end.

[0076] When implementing cross-sharding synchronization, the insert operation of the user table needs to be synchronized to multiple shards such as inventory. The sequence ensures that each shard executes in the same order. After the order status is updated, the logistics information needs to be cascaded and updated. The sequence ensures the order of the two operations. Combined with step S212, the synchronization data instruction comes from the subscription relationship configuration in step S211. The table structure of the data source can be in the form of a relational table or a document type collection.

[0077] Parse the synchronization instruction, extract the list of target service units of the current changed table from the subscription relationship table, and determine the range of fields to be synchronized.

[0078] Extract key data points. In the case of a relational table, the primary key field can be obtained from the table metadata. The operation type is captured by the listener. The changed fields include the fields that are actually modified.

[0079] This embodiment can synchronize the set of key data points as the input for generating the data change sequence. Step S213 generates the data change sequence based on the key data points.

[0080] For the abnormal state, if the primary key of the table is not defined, it is marked as abnormal, an alarm is triggered, and the synchronization is terminated. If the subscription - end table structure lacks key fields, an error log is recorded and the retry process is entered. It can be seen that the accurate positioning and execution of changes are ensured through fields such as the primary key and operation type; the data change sequence ensures consistency through sequential encapsulation; step S212 extracts key data points by parsing the instructions and structure type, providing a basis for subsequent sequence generation and synchronization.

[0081] Step S213 generates a data change sequence according to the synchronized key data points and the data structure type.

[0082] Step S214 obtains the synchronization data tag from the preset message queue according to the synchronized key data points, and stores the data change sequence in the message queue according to the synchronization data tag.

[0083] In some embodiments, step S214 includes the following steps: Step S2141 obtains the data structure type of the synchronized key data points, and determines the synchronization data format of the synchronized key data points according to the data structure type.

[0084] Step S2142 generates a synchronization task for the synchronized key data points according to the synchronization data format.

[0085] Step S2143 performs unified subscription - end configuration on the synchronization task to obtain the task unified configuration information.

[0086] Step S2144 obtains the initial synchronization data tag of the message queue, and filters out the synchronization data tag from the initial synchronization data tag based on the task unified configuration information and the synchronization task.

[0087] Step S2145 obtains the data synchronization address of the synchronization data tag.

[0088] Step S2146 obtains the synchronization mode of the data change sequence.

[0089] Step S2147 stores the data change sequence in the message queue according to the data synchronization address and the synchronization mode.

[0090] In some embodiments, the synchronization key data points generated in step S2142 are a set of key information identifying the characteristics of the synchronization data. For example, it may include key identifiers such as the business module to which the data belongs, the type of data change (add, delete, modify), and the table structure involved. These key data points provide the core basis for data processing in subsequent operations.

[0091] The preset message queue has a set of label management mechanisms. Based on the synchronization key data points, the system will search for the corresponding synchronization data labels in the message queue. It can be defined that the synchronization key data points indicate an add operation on the user information table, and the message queue will follow the predefined rules. This search process involves matching algorithms, such as the hash matching algorithm. The system first performs a hash calculation on the synchronization key data points to obtain a hash value, and then quickly locates the corresponding synchronization data label in the label index table of the message queue through this hash value.

[0092] After obtaining the synchronization data label, store the data change sequence generated in step S103 at the corresponding position in the message queue according to this label. In this way, the data in the message queue can be classified and stored according to certain rules, facilitating the subsequent subscriber to quickly and accurately extract the required data change sequence according to different labels, ensuring the orderliness and efficiency of data synchronization.

[0093] Step S215, extract the data change sequence from the message queue and create a synchronization data channel between the data change sequence and the preset subscriber.

[0094] The message queue in this embodiment maintains the storage structure and index of the data. Retrieve and extract the previously stored data change sequence from the message queue. For example, extract it according to the first-in, first-out rule, or according to the priority of the data (which can be defined by the synchronization data label or other metadata). After extracting the data change sequence, establish a communication channel, that is, a synchronization data channel, between the data change sequence and the preset subscriber.

[0095] To achieve this in this embodiment, a mechanism similar to the establishment of a network connection is adopted. First, the system will obtain connection information such as the network address and port number of the subscriber (the preset subscriber information is clearly defined in the system configuration). Then, establish a reliable connection through a process similar to the TCP three-way handshake. The system sends a connection request message to the subscriber, the subscriber returns a confirmation message after receiving the request, and the system then replies with a confirmation to complete the establishment of the channel. In this process, some security verification mechanisms may also be involved, such as authentication and data encryption, to ensure the security and accuracy of data transmission. Only when the synchronization data channel is successfully created can the data change sequence be transmitted to the subscriber smoothly and securely, preparing for subsequent data persistence operations.

[0096] Step S216, store the data change sequence into the subscriber end according to the synchronization data channel.

[0097] In step S215, the data change sequence has been extracted from the message queue and a synchronization data channel has been created. This synchronization data channel is similar to a data transmission channel, containing various information for establishing a connection with the subscriber end, such as network address, port number, communication protocol, etc.

[0098] Based on the established synchronization data channel, the data change sequence is transmitted to the subscriber end. When the data change sequence arrives at the subscriber end, the system will store the data into the corresponding business data sub-library according to the storage rules and requirements of the subscriber end. During the storage process, to ensure the integrity and consistency of the data, some data verification and error handling operations are performed, such as checking whether the data is complete and whether it conforms to the database table structure.

[0099] In some embodiments, step S216 further includes the following steps: Step S2161: Obtain the configuration synchronization transmission instruction of the data change sequence according to the synchronization data channel.

[0100] The synchronization data channel is not only a data transmission channel but also contains various information for interacting with the subscriber end. The system understands the relevant requirements and data processing methods of the subscriber end by analyzing the configuration information of the synchronization data channel.

[0101] Based on the understanding of the synchronization data channel, obtain the configuration synchronization transmission instruction for transmitting the data change sequence. If the configuration synchronization transmission instruction is a pre-compiled SQL statement, then a corresponding pre-compiled SQL statement will be generated according to the database type of the subscriber end and the content of the data change sequence. For example, if the subscriber end is a MySQL database and the data change sequence is an insert operation of user data, a pre-compiled SQL statement will be generated.

[0102] Step S2162, perform parameter category configuration on the configuration synchronization transmission instruction to obtain the configuration synchronization data structure type.

[0103] In this embodiment, the obtained configuration synchronization transmission instruction (such as a pre-compiled SQL statement) is parsed to identify the parameters. According to the corresponding data types in the data change sequence and the field type requirements of the subscriber end database, these parameters are configured by category. For example, user_id may be an integer type, user_name is a string type, and user_age is an integer type. Through this configuration, the data structure type of each parameter is clarified, and the configuration synchronization data structure type is obtained.

[0104] Step S2163: Execute the configuration synchronization transmission instruction according to the configuration synchronization data structure type and the synchronization data channel, and determine the data change sequence to be stored in the subscriber end.

[0105] After obtaining the configuration synchronization data structure type, fill the actual data in the data change sequence into the placeholder of the pre-compiled SQL statement according to the configured parameter category and data structure type. At the same time, create a target interface connected to the subscriber end using the connection object of the synchronization data channel, and this target interface is used to represent the pre-compiled SQL statement.

[0106] Execute the pre-compiled SQL statement filled with data at the subscriber end through the target interface. If the execution is successful, the data change sequence is successfully stored in the business data sub-library of the subscriber end, completing the entire data synchronization step; if the execution fails, the system will perform corresponding error handling according to the error type, such as rolling back the transaction, recording error logs, etc.

[0107] In this embodiment, by combining the synchronization data channel information, it can be ensured that the obtained instruction is accurately adapted to the subscriber end, improving the pertinence and accuracy of data transmission. For parameters of different data types such as numeric and character types, configure them according to the field type requirements of the subscriber end database to prevent errors caused by data type mismatches and ensure the consistency and integrity of data during transmission. The synchronization data structure type ensures that data is transmitted and stored in a format that meets the requirements of the subscriber end, while the synchronization data channel provides the specific information and channel for connecting to the subscriber end. The combination of these two enables the execution instruction to be adapted to the execution environment of the subscriber end, thereby storing the data accurately and without error in the target business data sub-library. The integrity and accuracy of the data are maintained.

[0108] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0109] The following is an embodiment of a data real-time synchronization system with a distributed architecture provided by the embodiments of the present disclosure. This system and the data real-time synchronization method with a distributed architecture in the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the data real-time synchronization system with a distributed architecture, reference can be made to the embodiments of the data real-time synchronization method with a distributed architecture above.

[0110] As Figure 3 shown, the system includes: A status judgment module, used to define the data source and the subscription relationship of the data, and judge the synchronization status; The synchronous execution module is used to use wildcards if synchronization to all sub-databases is required; if only partial synchronization is needed, the corresponding service unit is configured.

[0111] The monitoring and deployment module is used to deploy an incremental data listener according to the data source type.

[0112] The monitoring and capture module is used to capture the changed data in the system in real time through the incremental data listener.

[0113] The data detection module is used to check whether there is defined synchronization data in the data source; if so, the synchronization data is configured into a data change set according to a preset structure.

[0114] The data sending module is used to send the data change set to the subscription end through a message queue according to the data subscription relationship, and record the synchronization log of the publishing end at the same time.

[0115] The data processing module is used to parse the change set and organize it into a standard SQL after the subscription end receives the data change set, and persist it to the target-side business data sub-database.

[0116] As Figure 4 shown, the present application also provides an electronic device, including a display module 103, a memory 102, a processor 101, and a computer program stored on the memory and executable on the processor 101. When the processor 101 executes the program, the steps of the data real-time synchronization method for a distributed architecture are implemented.

[0117] In the embodiments of the present invention, the electronic device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.

[0118] In the embodiments of the present application, the processor 101 may be implemented by using at least one of an application specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), a processor, a controller, a microcontroller, a microprocessor, and an electronic unit designed to execute the functions described herein. In some cases, such an implementation may be implemented in a controller. For a software implementation, an implementation of a process or a function may be implemented with a separate software module that allows execution of at least one function or operation. The software code may be implemented by a software application (or program) written in any suitable programming language. The software code may be stored in a memory and executed by a controller.

[0119] The display module 103 is used to display information input by a user or information provided to the user. The display module 103 may include a display panel, and the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0120] The memory 102 may be used to store software programs and various data. The memory 102 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0121] The present application also provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data real-time synchronization method of the distributed architecture are implemented.

[0122] The storage medium may be any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0123] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A distributed architecture data real-time synchronization method, characterized in that: Methods include: Define the data source, create a subscription relationship table, record the data source, change the table name and the subscription relationship between the target service unit, and determine the synchronization status; If you need to synchronize to all sub-databases, use a wildcard; if you want to synchronize only some of them, configure the corresponding service unit; According to the selected data source type, select the incremental data listener to capture the data change status in real time; Capture the changing data in the system in real time through the incremental data listener; Check whether there is defined synchronization data in the data source; If yes, the synchronized data will be configured into a data change set according to the preset structure; According to the data subscription relationship, the data change set is sent to the subscriber through the message queue, and the publisher synchronization log is recorded at the same time; After receiving the data change set, the subscriber parses the change set and organizes it into standard SQL, which is persisted to the target business data sub-database.

2. The distributed architecture data real-time synchronization method according to claim 1, characterized in that: Steps to deploy the listener according to the data source type also include: Configure JPA data listener to monitor data changes; Configure business entity listeners to monitor data changes; Custom data listeners are used to support other methods of data changes.

3. The distributed architecture data real-time synchronization method according to claim 1, characterized in that: The steps of defining the data source and the subscription relationship of the data and determining the synchronization status also include: selecting a data source implementation method according to the business entity type; Monitor its persistence operations; Define table structure and change rules based on business entity framework; Implement non-standard persistence monitoring through extended interfaces; Configure the service unit and preset wildcards to indicate that data needs to be synchronized to all sub-databases; In the method, determine whether the current data change needs to trigger synchronization based on the subscription relationship table of the data source; The logical condition for judgment is: if the synchronization data instruction∈subscription relationship table, and the target service unit matches the subscription configuration, it is marked as requiring synchronization.

4. The distributed architecture data real-time synchronization method according to claim 1, characterized in that: The method also includes: Get data changeset; Generate synchronization key data points of synchronization data according to synchronization data instructions and data structure types; Generate data change sequence based on synchronization key data points and data structure types; Obtain synchronization data tags from a preset message queue according to synchronization key data points, and store data change sequences in the message queue according to the synchronization data tags; Extract the data change sequence from the message queue and create a synchronous data channel between the data change sequence and the preset subscriber; The data change sequence is stored in the subscriber according to the synchronous data channel.

5. The distributed architecture data real-time synchronization method according to claim 4, characterized in that: The step of generating synchronization key data points of synchronization data according to the synchronization data instruction and the data structure type also includes: During the data change process, define the synchronization key data point that uniquely identifies the change operation. The synchronization key data point has the primary key field, operation type, and timestamp. Configure a synchronization data instruction, where the synchronization data instruction comes from the subscription relationship configuration in step S101; The precondition is defined as having captured data changes and verifying that they need to be synchronized; Parse the synchronization instruction, extract the target service unit list of the current change table from the subscription relationship table, and determine the field range to be synchronized; Extract key data points and generate a set of key data points based on primary key fields, operation types, and change fields; Synchronize a set of key data points as input for generating a data change sequence, and generate a data change sequence based on the key data points; If the table does not have a primary key defined, it is marked as abnormal, an alarm is triggered, and synchronization is terminated; If the subscription-side table structure lacks key fields, an error log is recorded and a retry process is initiated.

6. The distributed architecture data real-time synchronization method according to claim 4, characterized in that: The step of obtaining a synchronization data tag from a preset message queue according to the synchronization key data point, and storing the data change sequence in the message queue according to the synchronization data tag also includes: Obtaining the data structure type of the synchronization key data point, and determining the synchronization data format of the synchronization key data point according to the data structure type; Generate synchronization tasks for synchronization key data points according to synchronization data formats; Perform unified subscription configuration on synchronization tasks to obtain unified task configuration information; Get the initial synchronization data tag of the message queue, and filter out the synchronization data tag from the initial synchronization data tag based on the unified configuration information of the task and the synchronization task.

7. The distributed architecture data real-time synchronization method according to claim 4, characterized in that: The step of storing the data change sequence to the subscriber according to the synchronous data channel also includes: Acquire configuration synchronization transmission instructions of a data change sequence according to a synchronization data channel; Perform parameter category configuration on the configuration synchronization transmission instruction to obtain the configuration synchronization data structure type; The configuration synchronization transmission instruction is executed according to the configuration synchronization data structure type and the synchronization data channel, and the data change sequence is determined and stored in the subscription end.

8. A distributed data real-time synchronization system, characterized in that: The system is used to execute the real-time data synchronization method of the distributed architecture as described in any one of claims 1 to 7; The system includes: The status judgment module is used to define the data source and data subscription relationship, and judge the synchronization status; Synchronous execution module, used to use wildcards if synchronization is required to all sub-databases; if only partial synchronization is required, the corresponding service unit is configured; The listener deployment module is used to deploy incremental data listeners according to the data source type; The monitoring and capturing module is used to capture the changing data in the system in real time through the incremental data monitor; The data detection module is used to check whether there is defined synchronization data in the data source; if so, the synchronization data is configured with a data change set according to a preset structure; The data sending module is used to send the data change set to the subscriber through the message queue according to the data subscription relationship, and record the synchronization log of the publisher at the same time; The data processing module is used to parse the data change set after the subscriber receives it, organize it into standard SQL, and persist it to the target business data sub-database.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the real-time data synchronization method of the distributed architecture as described in any one of claims 1 to 7 are implemented.

10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the real-time data synchronization method of the distributed architecture as claimed in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Real-time data synchronization method and system for distributed architecture, and device and medium

    WO2026184028A1