Data synchronization method and device, computer equipment and storage medium
By determining the baseline transaction sequence number in the distributed database and processing archived logs in groups and shards, the problems of low efficiency and inconsistency in existing technologies are solved, and efficient and accurate data synchronization is achieved.
Patent Information
- Application Number
- CN202511154660.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing data synchronization methods suffer from inefficiency, inconsistency, and reliability issues in distributed databases, especially in high-concurrency transactions and distributed environments, making it difficult to effectively improve the efficiency and accuracy of archived log synchronization.
By determining the baseline transaction sequence number of the target transaction in the source database, scanning the archived logs, grouping and sorting them by physical nodes, performing sharding based on the file density model, and migrating them sequentially to the target database, the node queue and sharding strategy are optimized to improve synchronization efficiency and accuracy.
It significantly improves the throughput and latency of archived log synchronization, ensuring data integrity and consistency, and is suitable for large-scale, high-concurrency distributed database environments.
Smart Images

Figure CN121029883A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data migration, and particularly relates to a data synchronization method and device, computer equipment and a computer readable storage medium. BACKGROUND
[0002] At present, in modern distributed database systems, data synchronization is a key task, especially in scenarios that require maintaining data consistency between multiple database instances. Traditional data synchronization methods mainly rely on full data synchronization or incremental synchronization based on timestamps. Full synchronization is simple but inefficient, especially when dealing with large-scale data. Incremental synchronization based on timestamps improves efficiency, but in high-concurrency transaction processing and distributed environments, there is a risk of data loss and inconsistency.
[0003] In recent years, with the development of distributed computing and storage technology, log-based incremental synchronization methods have emerged. These methods capture the archive logs of the database to achieve data synchronization, which can effectively reduce data transmission and improve synchronization efficiency. However, existing methods still have some problems such as poor data synchronization efficiency, consistency and reliability when dealing with archive logs in distributed databases.
[0004] Therefore, how to provide a data synchronization method, device, computer equipment and computer readable storage medium that can effectively improve the efficiency and accuracy of archive log synchronization is a problem that needs to be solved by the technical personnel in the field. SUMMARY
[0005] In view of the above deficiencies of the prior art, the purpose of the present application is to provide a data synchronization method, device, computer equipment and computer readable storage medium, which aims to solve the problem of how to effectively improve the efficiency and accuracy of archive log synchronization.
[0006] In order to achieve the above purpose, the present application adopts the following technical solutions:
[0007] In a first aspect, the present application provides a data synchronization method, comprising:
[0008] determining a reference transaction sequence number of a target transaction in a source database; wherein the source database comprises a plurality of physical nodes;
[0009] scanning all archive logs of the target transaction in the source database, and obtaining all archive files with transaction sequence numbers greater than the reference transaction sequence number in the archive logs;
[0010] grouping each of the archive files according to its corresponding physical node, and sorting each of the physical nodes according to a predetermined node arrangement strategy to obtain a node queue;
[0011] The computing module is configured to calculate the density of each of the archive files in the node queue based on a preset file density model, and perform fragmentation processing on each of the archive files according to a density calculation result.
[0012] The migration module is configured to sequentially migrate each of the archive files after the fragmentation processing to a target end database in order according to an arrangement order of each of the physical nodes in the node queue.
[0013] In a second aspect, the present application provides a data synchronization device, which comprises:
[0014] The determining module is configured to determine a reference transaction sequence number of a target transaction in a source end database; wherein the source end database comprises a plurality of physical nodes.
[0015] The scanning module is configured to scan all archive logs of the target transaction in the source end database, and acquire all archive files with a transaction sequence number greater than the reference transaction sequence number in the archive logs.
[0016] The grouping module is configured to group each of the archive files according to the corresponding physical node, and sort each of the physical nodes according to a preset node arrangement strategy to obtain a node queue.
[0017] The computing module is configured to calculate the density of each of the archive files in the node queue based on a preset file density model, and perform fragmentation processing on each of the archive files according to a density calculation result.
[0018] The migration module is configured to sequentially migrate each of the archive files after the fragmentation processing to a target end database in order according to an arrangement order of each of the physical nodes in the node queue.
[0019] In a third aspect, the present application provides a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the data synchronization method as described above when executing the computer program.
[0020] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, wherein the computer program is executable on a processor to implement the data synchronization method as described above.
[0021] Compared with the prior art, the present application provides a data synchronization method, device, computer equipment and computer readable storage medium, wherein the reference transaction sequence number of a target transaction in a source database is determined; wherein the source database comprises a plurality of physical nodes; all archive logs of the target transaction in the source database are scanned to obtain all archive files with a transaction sequence number greater than the reference transaction sequence number in the archive logs; each archive file is grouped according to its corresponding physical node, and each physical node is sorted according to a preset node arrangement strategy to obtain a node queue; the density of each archive file in the node queue is calculated based on a preset file density model, and each archive file is fragmented according to the density calculation result; each fragmented archive file is sequentially migrated to a target database according to the arrangement order of each physical node in the node queue; thereby the efficiency and accuracy of archive log synchronization can be effectively improved by the present application. BRIEF DESCRIPTION OF DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0023] Figure 1 An application environment schematic diagram of a data synchronization method provided by an embodiment of the present application.
[0024] Figure 2 A flowchart of a data synchronization method provided by an embodiment of the present application.
[0025] Figure 3 A program module schematic diagram of a data synchronization device provided by an embodiment of the present application.
[0026] Figure 4 A structure schematic diagram of a computer equipment provided by an embodiment of the present application.
[0027] Figure 5 Another structure schematic diagram of a computer equipment provided by an embodiment of the present application. DETAILED DESCRIPTION
[0028] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are a part rather than all of the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts should fall within the scope of the present application.
[0029] It should be understood that the term "comprising" as used in the specification and the appended claims indicates the presence of the described features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0030] It should also be understood that the term "and / or" as used in the specification and the appended claims indicates any combination of one or more of the associated listed items and all possible combinations of the items.
[0031] As used in the specification and the appended claims, the term "if" can be interpreted as meaning "when" or "once" or "in response to a determination" or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if detected [the described condition or event]" can be interpreted as meaning "once determined" or "in response to a determination" or "once detected [the described condition or event]" or "in response to detecting [the described condition or event]" depending on the context.
[0032] In addition, in the description of the specification and the appended claims, the terms "first", "second", "third", and the like are used only to distinguish descriptions, and cannot be understood as indicating or implying relative importance.
[0033] The reference "one embodiment" or "some embodiments" and the like described in the specification means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Therefore, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in yet some embodiments", and the like appearing in various places in the specification are not necessarily all referring to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically noted. The terms "comprise", "include", "have", and their variants mean "including but not limited to", unless otherwise specifically noted.
[0034] It should be understood that the size of the serial number of each step in the following embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0035] In order to illustrate the technical solutions of the present application, the following will be described by specific examples.
[0036] An embodiment of the present application provides a data synchronization method, which can be applied to an application environment as shown in the accompanying drawings. Figure 1 The client includes but is not limited to a palm computer, a desktop computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cloud computer device, a personal digital assistant (PDA) and the like. The server can be an independent server, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms and the like basic cloud computing services.
[0037] Please refer to Figure 2 An embodiment of the present application provides a data synchronization method, which includes the following steps.
[0038] S100, determining a reference transaction sequence number of a target transaction in a source database; wherein the source database includes a plurality of physical nodes;
[0039] S200, scanning all archive logs of the target transaction in the source database, and acquiring all archive files with a transaction sequence number greater than the reference transaction sequence number in the archive logs;
[0040] S300, grouping each archive file according to its corresponding physical node, and sorting each physical node according to a preset node arrangement strategy to obtain a node queue;
[0041] S400, performing density calculation on each archive file in the node queue based on a preset file density model, and performing sharding processing on each archive file according to the density calculation result;
[0042] S500, according to the arrangement order of each physical node in the node queue, sequentially migrating each archive file after sharding processing to a target database.
[0043] In specific implementation, the data synchronization method of the present embodiment significantly improves the efficiency and accuracy of archive log synchronization through a series of innovative steps. The specific analysis is as follows:
[0044] 1. Determine the reference transaction sequence number (S100)
[0045] By determining the reference transaction sequence number of the target transaction in the source database, this method provides a clear starting point for the data synchronization task. This step ensures that the synchronization process starts from the correct transaction position, avoiding the repeated processing of already synchronized data, thereby improving the accuracy and efficiency of synchronization;
[0046] Wherein, the transaction sequence number is a sequence number used to identify the order of transactions, which is usually a monotonically increasing number used to record the commit order of transactions, applicable to all database systems that require transaction management, including relational databases and NoSQL databases.
[0047] Specifically, it can be used to describe the transaction order recording mechanism in different database systems, clearly expressing its function of recording transaction order. Examples: Oracle's SCN (System Change Number), MySQL's transaction sequence number (implemented through an auto-increment field), PostgreSQL's XID (Transaction ID), etc.
[0048] 2. Scan the archive log (S200)
[0049] Scan all archive logs of the target transaction in the source database to obtain all archive files with transaction sequence number greater than the reference transaction sequence number. This step accurately filters the archive files that need to be synchronized, reducing unnecessary data processing and further improving synchronization efficiency. At the same time, through the filtering of transaction sequence numbers, the integrity and consistency of data are ensured.
[0050] 3. Archive file grouping and node sorting (S300)
[0051] Group the archive files according to their corresponding physical nodes, and sort the physical nodes according to the pre-set node arrangement strategy to obtain the node queue. This step optimizes the order of data processing through reasonable grouping and sorting, ensuring the efficiency of parallel processing. Especially in a distributed environment, through node sorting, nodes with lower load and smaller network delay can be processed first, thereby improving the overall synchronization efficiency.
[0052] 4. Sharding processing based on file density model (S400)
[0053] Based on the preset file density model, the density of the archived file is calculated, and the file is processed according to the density calculation result. This step dynamically adjusts the sharding strategy to generate appropriate size of the shard according to the transaction density of the file. High-density files are finely granularly sharded, and low-density files are coarsely granularly sharded, thereby achieving load balancing and further improving resource utilization efficiency and synchronization speed.
[0054] 5. Migrating to the target database in order (S500)
[0055] According to the arrangement order of the physical nodes in the node queue, the archived file after the sharding is sequentially migrated to the target database. This step ensures the order and consistency of the data through the ordered migration operation. At the same time, through the sharding, each shard can be processed independently in parallel, further improving the migration efficiency. In addition, through the dynamic optimization mechanism, the system can automatically adjust the migration strategy according to the real-time data distribution and system load, ensuring efficient operation in different scenarios.
[0056] Through the synergistic effect of the above steps, the data synchronization method of the present application can significantly improve the efficiency and accuracy of the archived log synchronization. Specifically:
[0057] High throughput: through dynamic sharding and parallel processing of shards, the throughput of data synchronization is significantly improved, which is 3-5 times higher than that of the traditional serial processing method;
[0058] Low latency: through optimization of node sorting and sharding strategy, the processing time of a single task is reduced, and the latency of data synchronization is significantly reduced;
[0059] High accuracy: through accurate filtering and ordered migration of transaction sequence numbers, the integrity and consistency of the data are ensured;
[0060] Dynamic optimization: through dynamic adjustment of real-time data distribution and system load, the maximum utilization of resources is ensured, and the overall performance and reliability of the system are improved;
[0061] Therefore, through a series of innovative technologies, the present application breaks through the bottleneck of the traditional data synchronization method, significantly improves the efficiency and accuracy of the archived log synchronization, and is suitable for large-scale, high-concurrency distributed database environment.
[0062] Further, in one embodiment, the data synchronization method, wherein the determining the reference transaction sequence number of the target transaction in the source database comprises the steps of:
[0063] receiving the input instruction of the user;
[0064] If the input instruction is a sequence number input value, the sequence number input value is taken as the reference transaction sequence number of the target transaction in the source database;
[0065] If the input instruction is automatic detection, the minimum uncommitted transaction sequence number of the target transaction in the source database is obtained and taken as the reference transaction sequence number.
[0066] In implementation, the specific implementation process of the steps of the embodiment is as follows:
[0067] Step 1: receiving user input instruction
[0068] The system (the software system corresponding to the method of the application) receives the instruction input by the user, which can be a transaction sequence number input manually by the user or an automatic detection command selected by the user.
[0069] Step 2: judging the type of input instruction
[0070] The system judges the type of the instruction input by the user. If the input instruction is a specific sequence number input value, the value is directly taken as the reference transaction sequence number. If the input instruction is automatic detection, the automatic detection process is entered.
[0071] Step 3: automatically detecting the minimum uncommitted transaction sequence number
[0072] If the user selects automatic detection, the system queries the source database, obtains the minimum uncommitted transaction sequence number in the target transaction, and takes the number as the reference transaction sequence number.
[0073] Step 4: setting the reference transaction sequence number
[0074] According to the sequence number input value input by the user or the result of automatic detection, the system sets the reference transaction sequence number of the target transaction and records the value for use in subsequent steps.
[0075] Step 5: verifying the reference transaction sequence number
[0076] The system verifies whether the set reference transaction sequence number is valid. If not, the user is prompted to re-input or re-perform automatic detection. If yes, the subsequent synchronization task is continued to be executed.
[0077] Through the above process, the system can flexibly determine the reference transaction sequence number of the target transaction in the source database according to the result of user input or automatic detection, and ensure the accuracy and reliability of the data synchronization task.
[0078] Further, in one embodiment, the data synchronization method, wherein the scanning all the archived logs of the target transaction in the source database, obtaining all the archived files in the archived logs whose transaction sequence numbers are greater than the reference transaction sequence number, specifically comprises the steps of:
[0079] traversing all the archived logs of the target transaction in the source database, and reading the metadata information of each of the archived logs;
[0080] According to the metadata information, filtering out all the archived files in the archived logs whose transaction sequence numbers are greater than the reference transaction sequence number.
[0081] In specific implementation, the specific implementation process of the steps of the embodiment is approximately as follows:
[0082] Step 1: initializing the scanning environment
[0083] Configuring scanning parameters: the system loads the configuration information of the source database, including the reference transaction sequence number of the target transaction, the storage path of the archived logs, etc.
[0084] Loading the index of the archived logs: if there is an index of the archived logs, the system loads the index to speed up the positioning process of the archived logs. The index contains the basic information of each archived log, such as the file path, size, transaction sequence number range, etc.
[0085] Step 2: traversing the archived logs and reading the metadata
[0086] Traversing the archived logs: the system traverses all the archived logs of the target transaction in the source database. For each archived log, the system reads its metadata information, including the file size, transaction sequence number range (start SCN and end SCN), etc.
[0087] Recording the metadata: the system records the metadata of each archived log to a temporary storage area for subsequent processing.
[0088] Step 3: filtering the archived files
[0089] Comparing the transaction sequence numbers: the system compares the start SCN and end SCN of each archived log with the reference transaction sequence number according to the transaction sequence number range in the metadata.
[0090] Filtering the files meeting the conditions: if the start SCN of the archived log is greater than the reference transaction sequence number, or the end SCN is greater than the reference transaction sequence number, the archived log is marked as an archived file that needs to be synchronized.
[0091] Recording the filtering results: the system records the paths and related information of the filtered archived files in a list for subsequent sharding processing and migration.
[0092] Step 4: Verify the screening results
[0093] Integrity check: The system performs an integrity check on the screened archive files to ensure that the files are not corrupted and are readable.
[0094] Record error information: If a file is found to be corrupted or unreadable, the system records the error information and decides whether to skip the file based on the configuration.
[0095] Through the above process, the system can efficiently scan the archive logs in the source database and screen all archive files with transaction sequence numbers greater than the reference transaction sequence number. This process not only ensures the accuracy and integrity of data synchronization, but also improves synchronization efficiency by optimizing the screening results.
[0096] Further, in one embodiment, the data synchronization method, wherein the grouping each of the archive files according to its corresponding physical node, and sorting each of the physical nodes according to a preset node arrangement strategy to obtain a node queue, specifically includes steps of:
[0097] Configuring a node arrangement strategy;
[0098] According to the metadata information, determining the physical node corresponding to each of the archive files, and grouping each of the archive files according to its corresponding physical node;
[0099] According to the node arrangement strategy, arranging each of the physical nodes in sequence to generate a node queue.
[0100] In specific implementation, the specific implementation process of the steps of the embodiment is approximately as follows:
[0101] Step 1: Configure the node arrangement strategy
[0102] Define strategy parameters: The system defines the parameters of the node arrangement strategy, which can include the load condition of the node, network delay, transaction processing capacity, etc.
[0103] Set priority rules: Set the priority rules of the nodes according to actual needs. For example, preferentially select nodes with lower load and smaller network delay, or preferentially select nodes with stronger transaction processing capacity.
[0104] Store the strategy configuration: Store the configured node arrangement strategy in the system configuration file or database for subsequent steps.
[0105] Step 2: Read the metadata of the archive files
[0106] Traverse archive files: The system traverses all filtered archive files and reads the metadata information of each archive file. Metadata usually includes file path, size, transaction serial number range, physical node belonging, etc.
[0107] Record metadata: Record the metadata of each archive file to the temporary storage area for subsequent processing.
[0108] Step 3: Group archive files by physical nodes
[0109] Determine physical nodes: Determine the physical node to which each archive file belongs according to the metadata of the archive file.
[0110] Group archive files: Group archive files according to their corresponding physical nodes. Each physical node corresponds to a group list, and the archive files belonging to the node are added to the corresponding list.
[0111] Step 4: Sort physical nodes by node arrangement strategy
[0112] Load strategy configuration: The system loads the pre-configured node arrangement strategy.
[0113] Evaluate node priority: Evaluate the priority of each physical node according to the node arrangement strategy. For example, calculate the load, network delay, transaction processing capacity and other indicators of each node.
[0114] Generate node queue: According to the evaluation results, arrange the physical nodes in order of priority to generate a node queue. The order in the node queue will determine the order of subsequent processing.
[0115] Step 5: Verify and optimize node queue
[0116] Verify node queue: The system verifies the generated node queue to ensure that the archive files of each node have been correctly grouped and the order of the nodes meets the pre-set node arrangement strategy.
[0117] Dynamic optimization: Dynamically adjust the order of the node queue according to real-time system load and resource usage. For example, if the load of a certain node suddenly increases, the system can automatically lower its priority and advance other nodes.
[0118] Record queue information: Record the final generated node queue information to the system log or configuration file for subsequent steps.
[0119] Through the above process, the system can group archive files according to their corresponding physical nodes and generate an ordered node queue according to the pre-set node arrangement strategy. This process not only ensures the efficiency and reliability of data synchronization tasks, but also improves the adaptability and flexibility of the system through dynamic optimization mechanism.
[0120] Further, in one embodiment, the data synchronization method, wherein the file density model is configured, and the file fragmentation strategy is configured, and the file density model is used to calculate the density of each archive file in the node queue, and the file fragmentation strategy is used to fragment each archive file according to the density calculation result, specifically comprising the steps of:
[0121] configuring a file density model and a file fragmentation strategy;
[0122] using the file density model to calculate the density of each archive file in the node queue to generate a density calculation result;
[0123] fragmenting each archive file according to the density calculation result and the file fragmentation strategy.
[0124] In specific implementation, the specific implementation process of the steps of the embodiment is approximately as follows:
[0125] Step 1: configuring a file density model and a file fragmentation strategy
[0126] defining a file density model: the system defines the parameters of the file density model, which include file size, SCN interval size, etc. The file density model is used to evaluate the distribution density of transactions in the archive file.
[0127] setting a file fragmentation strategy: setting the file fragmentation strategy according to actual needs, including fragmentation size, fragmentation number, etc. The file fragmentation strategy can be dynamically adjusted according to the file density to optimize processing efficiency.
[0128] storing configuration information: storing the configured file density model and file fragmentation strategy in the system configuration file or database for subsequent steps.
[0129] Step 2: loading the file density model and the file fragmentation strategy
[0130] loading the model and the strategy: the system loads the pre-configured file density model and file fragmentation strategy. Ensure that the parameters of the model and the strategy are correctly loaded into the memory for subsequent calculation and processing.
[0131] initializing the calculation environment: initializing the environment for density calculation and fragmentation processing, including necessary data structures and variables.
[0132] Step 3: calculating the density of the archive file
[0133] traversing the archive file: the system traverses each archive file in the node queue, reads the metadata information of the file, including file size, transaction sequence number range (start SCN and end SCN), etc.
[0134] Calculate file density: Calculate the density of each archive file using the file density model. The file density can be calculated by the following formula:
[0135] AFD = Normalized density function (file size / SCN interval size), where AFD (Archive File Density) represents the density of the archive file.
[0136] Record density calculation results: Record the density calculation results of each archive file to the temporary storage area for subsequent fragmentation processing.
[0137] Step 4: Fragmentation processing according to density calculation results
[0138] Evaluate file fragmentation strategy: Evaluate the fragmentation method of each archive file according to the density calculation results and the preset file fragmentation strategy. For example:
[0139] If AFD > threshold 0.7, trigger fine-grained fragmentation, divide the file into more small fragments.
[0140] If AFD ≤ threshold 0.3, trigger coarse-grained fragmentation, divide the file into fewer large fragments.
[0141] If 0.3 < AFD ≤ 0.7, keep the default fragmentation parameters unchanged.
[0142] Perform fragmentation operation: Perform fragmentation processing on each archive file according to the evaluation results. The fragmentation operation can include splitting the file into multiple subfiles and recording the metadata information of each subfile.
[0143] Store fragmentation results: Store the fragmented archive files and their metadata information to the temporary storage area for subsequent migration operations.
[0144] Step 5: Verify and optimize fragmentation results
[0145] Verify fragmentation results: The system verifies the fragmented archive files to ensure that the size and transaction distribution of each fragment meet the preset file fragmentation strategy.
[0146] Dynamic optimization: Dynamically adjust the file fragmentation strategy according to the real-time system load and resource usage. For example, if the system detects that the current memory resource is tight, it will generate smaller fragments to reduce memory occupancy; if the CPU resource is sufficient, it will generate larger fragments to improve processing speed.
[0147] Record fragmentation information: Record the final generated fragmentation information to the system log or configuration file for subsequent steps.
[0148] Through the above process, the system can calculate the density of the archived files based on the pre-set file density model and perform sharding processing according to the density calculation result. This process not only ensures the efficiency and reliability of the data synchronization task, but also improves the adaptability and flexibility of the system through dynamic optimization mechanism.
[0149] Further, in one embodiment, the data synchronization method, wherein the migrating the sharded archived files to the target database in sequence according to the arrangement order of the physical nodes in the node queue, specifically comprises the steps of:
[0150] sending the sharded archived files to the message middleware in sequence according to the arrangement order of the physical nodes in the node queue;
[0151] reading the archived files in the message middleware and converting them into the format required by the target database;
[0152] writing the converted archived files back to the target database.
[0153] In specific implementation, the specific implementation process of the steps of the present embodiment is generally as follows:
[0154] Step 1: Initialize the message middleware
[0155] Configure the message middleware: the system loads the configuration information of the message middleware (such as the kafka middleware), including connection parameters, queue name, topic, etc.
[0156] Initialize the connection: the system establishes a connection with the message middleware to ensure stable and available connection.
[0157] Prepare the queue or topic: create or prepare the queue or topic for storing archived files in the message middleware.
[0158] Step 2: Send the archived files to the message middleware in sequence
[0159] Traverse the node queue: the system processes each physical node in sequence according to the arrangement order of the physical nodes in the node queue.
[0160] Send the archived files: for each physical node, the system sends the sharded archived files to the message middleware in sequence.
[0161] Record the sending status: the system records the sending status of each archived file, including sending time, success or failure flag, etc., for subsequent monitoring and retry mechanism.
[0162] Step 3: Read and convert the format of the archived files
[0163] Reading archived files from the message middleware: The system reads archived files from the message middleware, ensuring the orderliness and completeness of the reading operation.
[0164] Format conversion: The system converts the read archived files into the format required by the target database. The conversion process may include data cleaning, field mapping, encoding conversion, etc.
[0165] Recording the conversion status: The system records the format conversion status of each archived file, including conversion time, success or failure flag, etc., for subsequent monitoring and retry mechanism.
[0166] Step 4: Playback and write to the target database
[0167] Connecting to the target database: The system establishes a connection to the target database, ensuring stable and available connection.
[0168] Playback and write: The system plays back and writes the format-converted archived files to the target database in sequence. During playback, ensure the orderliness and consistency of data.
[0169] Recording the write status: The system records the write status of each archived file, including write time, success or failure flag, etc., for subsequent monitoring and retry mechanism.
[0170] Step 5: Verify and confirm synchronization results
[0171] Verify data integrity: The system verifies the data in the target database, ensuring that all archived files have been correctly written, and the integrity and consistency of the data are guaranteed.
[0172] Confirm synchronization completion: The system confirms the completion of the synchronization task and records the final state of the synchronization task. If any problems are found, the system can trigger a retry mechanism or alarm mechanism.
[0173] Clean up temporary data: The system cleans up temporary data in the message middleware, releasing resources, ensuring efficient operation of the system.
[0174] Through the above process, the system can migrate the sharded archived files to the target database in sequence according to the arrangement order of the physical nodes in the node queue. This process not only ensures the efficiency and reliability of data synchronization, but also improves the fault tolerance and scalability of the system through the buffering mechanism of the message middleware.
[0175] Further, in one embodiment, the data synchronization method, wherein the source database and the target database are heterogeneous databases of different database types.
[0176] In specific implementation, after the archived files processed by the fragmentation are sent to the message middleware, the method continues to process the case where the source end database and the target end database are heterogeneous databases through the steps of determining the data mapping rule, reading and converting the archived files, processing the data type difference, playing back and writing into the target end database, and verifying and confirming the synchronization result, thereby ensuring the efficiency and reliability of data synchronization. These steps collectively ensure the correct conversion and synchronization of data between different database types.
[0177] As can be seen from the above method embodiment, the data synchronization method provided by the application comprises the following steps: determining a reference transaction sequence number of a target transaction in a source end database; wherein the source end database comprises a plurality of physical nodes; scanning all archived logs of the target transaction in the source end database to obtain all archived files with a transaction sequence number greater than the reference transaction sequence number in the archived logs; grouping each of the archived files according to the corresponding physical node, and sorting each of the physical nodes according to a preset node arrangement strategy to obtain a node queue; performing density calculation on each of the archived files in the node queue based on a preset file density model, and performing fragmentation processing on each of the archived files according to the density calculation result; and sequentially migrating each of the archived files processed by the fragmentation to a target end database according to the arrangement order of each of the physical nodes in the node queue. In this way, the method of the application can effectively improve the efficiency and accuracy of archived log synchronization.
[0178] It should be understood that, although the application provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps can be included based on conventional or non-inventive labor, and the operation steps are not necessarily executed in the order of the embodiments or flowcharts. The order of the steps listed in the embodiments or flowcharts is only one of the many execution orders, and does not represent the only execution order. It should be noted that there is no certain sequence between the above steps, and a person skilled in the art can understand from the description of the embodiments of the application that the above steps can have different execution orders in different embodiments, that is, they can be executed in parallel, or they can be exchanged and executed, and the like. Moreover, at least part of the steps in the embodiments or flowcharts can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation, alternation or synchronization with other steps or sub-steps or stages of other steps.
[0179] Based on the above method embodiment, please refer to Figure 3 Another embodiment of the application also provides a data synchronization device, wherein the device comprises:
[0180] determining a reference transaction sequence number of a target transaction in a source database, wherein the source database comprises a plurality of physical nodes;
[0181] scanning all archive logs of the target transaction in the source database to obtain all archive files with a transaction sequence number greater than the reference transaction sequence number;
[0182] grouping each of the archive files according to a corresponding physical node, and sorting each of the physical nodes according to a preset node arrangement strategy to obtain a node queue;
[0183] calculating a density of each of the archive files in the node queue based on a preset file density model, and performing a sharding process on each of the archive files according to a density calculation result;
[0184] migrating each of the archive files after the sharding process to a target database in sequence according to an arrangement order of each of the physical nodes in the node queue.
[0185] Further, in an embodiment, the data synchronization device, wherein the determining a reference transaction sequence number of a target transaction in a source database specifically comprises:
[0186] receiving an input instruction of a user;
[0187] if the input instruction is a sequence number input value, taking the sequence number input value as the reference transaction sequence number of the target transaction in the source database;
[0188] if the input instruction is automatic detection, obtaining a minimum uncommitted transaction sequence number of the target transaction in the source database, and taking the minimum uncommitted transaction sequence number as the reference transaction sequence number.
[0189] Further, in an embodiment, the data synchronization device, wherein the scanning all archive logs of the target transaction in the source database to obtain all archive files with a transaction sequence number greater than the reference transaction sequence number specifically comprises:
[0190] traversing all archive logs of the target transaction in the source database, and reading metadata information of each of the archive logs;
[0191] according to the metadata information, screening all archive files with a transaction sequence number greater than the reference transaction sequence number in the archive logs.
[0192] Further, in one embodiment, the data synchronization device, wherein the grouping of each of the archive files according to the corresponding physical node, and arranging each of the physical nodes according to a preset node arrangement strategy to obtain a node queue, specifically includes:
[0193] configuring a node arrangement strategy;
[0194] determining the corresponding physical node of each of the archive files according to the metadata information, and grouping each of the archive files according to the corresponding physical node;
[0195] arranging each of the physical nodes in sequence according to the node arrangement strategy to generate a node queue.
[0196] Further, in one embodiment, the data synchronization device, wherein the density calculation of each of the archive files in the node queue based on a preset file density model, and the file fragmentation processing of each of the archive files according to the density calculation result, specifically includes:
[0197] configuring a file density model and a file fragmentation strategy;
[0198] using the file density model to perform density calculation on each of the archive files in the node queue to generate a density calculation result;
[0199] performing file fragmentation processing on each of the archive files according to the density calculation result and the file fragmentation strategy.
[0200] Further, in one embodiment, the data synchronization device, wherein the according to the arrangement order of each of the physical nodes in the node queue, sequentially migrating each of the archive files after the fragmentation processing to the target end database, specifically includes:
[0201] according to the arrangement order of each of the physical nodes in the node queue, sequentially sending each of the archive files after the fragmentation processing to a message middleware;
[0202] reading the archive files in the message middleware and converting them into a format required by the target end database;
[0203] writing the archive files after the format conversion into the target end database.
[0204] Further, in one embodiment, the data synchronization device, wherein the source end database and the target end database are heterogeneous databases of different database types.
[0205] It should be noted that the information interaction, execution process and the like between the modules in the device embodiments of the present application are based on the same concept as the method embodiments of the present application, and the specific functions and the technical effects brought by the same can be referred to the aforementioned method embodiments, which will not be described here.
[0206] Based on the above method embodiments, another embodiment of the present application further provides a computer device, which can be a server, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the functions or steps of the data synchronization method server side in any one of the above method embodiments.
[0207] Based on the above method embodiments, another embodiment of the present application further provides a computer device, which can be a client, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements the functions or steps of the data synchronization method client side in any one of the above method embodiments.
[0208] Those skilled in the art can understand that Figure 4 and Figure 5 The structural schematic diagram shown in the above
[0209] The processor can be a CPU, and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0210] The memory includes a readable storage medium, an internal memory, etc. The internal memory can be the memory of the computer device, and the internal memory provides an environment for running the operating system and the computer readable instructions in the readable storage medium. The readable storage medium can be the hard disk of the computer device, and in other embodiments, can also be the external storage device of the computer device, for example, the plug-in hard disk, the smart media card (SMC), the secure digital (SD) card, the flash card, etc. Further, the memory can include both the internal storage unit of the computer device and the external storage device. The memory is used to store the operating system, the application program, the boot loader, the data, and other programs, such as the program code of the computer program, etc. The memory can also be used to temporarily store the data that has been output or will be output.
[0211] Based on the above method embodiments, another embodiment of the present application further provides a computer readable storage medium storing a computer program, wherein the computer program is executed by a processor to implement the data synchronization method in any one of the above method embodiments. The computer readable storage medium can be non-volatile or volatile.
[0212] It should be noted that the functions or steps that the computer readable storage medium or the computer device can achieve and the technical effects brought by the functions / steps can be referred to the related description in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0213] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, database or other medium used in each embodiment provided in the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The disclosed memory components or memories of the operating environment described herein are intended to include one or more of these and / or any other suitable type of memory.
[0214] It can be clearly understood by those skilled in the art that, for the convenience and brevity of description, in the device embodiment of the present application, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the above-described functions. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the unit and module in the above-mentioned device can refer to the corresponding process in the above-mentioned method embodiment, which will not be repeated here. If the integrated unit is realized in the form of software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium.
[0215] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0216] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0217] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0218] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A data synchronization method, characterized in that, include: Determine the base transaction sequence number of the target transaction in the source database; wherein, the source database includes multiple physical nodes; Scan all archived logs of the target transaction in the source database to obtain all archived files in the archived logs whose transaction sequence number is greater than the baseline transaction sequence number; Each archived file is grouped according to its corresponding physical node, and the physical nodes are sorted according to a preset node arrangement strategy to obtain a node queue. Based on a preset file density model, the density of each archived file in the node queue is calculated, and the archived files are fragmented according to the density calculation results. According to the order of the physical nodes in the node queue, the archived files after sharding are migrated to the target database in sequence.
2. The data synchronization method according to claim 1, characterized in that, Determining the base transaction sequence number of the target transaction in the source database includes: Receive user input commands; If the input instruction is a sequence number input value, then the sequence number input value is used as the base transaction sequence number of the target transaction in the source database; If the input instruction is automatic detection, then the minimum uncommitted transaction sequence number of the target transaction in the source database is obtained and used as the baseline transaction sequence number.
3. The data synchronization method according to claim 1, characterized in that, The step of scanning all archived logs of the target transaction in the source database to obtain all archived files in the archived logs whose transaction sequence number is greater than the baseline transaction sequence number includes: Traverse all archived logs of the target transaction in the source database and read the metadata information of each archived log; Based on the metadata information, all archived files in the archived logs whose transaction sequence number is greater than the baseline transaction sequence number are selected.
4. The data synchronization method according to claim 3, characterized in that, The step of grouping the archived files according to their corresponding physical nodes and sorting the physical nodes according to a preset node arrangement strategy to obtain a node queue includes: Configure node arrangement strategy; Based on the metadata information, the physical node corresponding to each archive file is determined, and the archive files are grouped according to their corresponding physical nodes; According to the node arrangement strategy, the physical nodes are arranged in order to generate a node queue.
5. The data synchronization method according to claim 1, characterized in that, The process of calculating the density of each archived file in the node queue based on a preset file density model, and then fragmenting each archived file according to the density calculation results, includes: Configuration file density model and file fragmentation strategy; Using the file density model, the density of each archived file in the node queue is calculated, and the density calculation result is generated. Based on the density calculation results and the file fragmentation strategy, each archived file is fragmented.
6. The data synchronization method according to claim 1, characterized in that, The step of sequentially migrating the fragmented archive files to the target database according to the order of the physical nodes in the node queue includes: According to the order of the physical nodes in the node queue, the fragmented archive files are sent to the message middleware in sequence. Read the archived file in the message middleware and convert it into the format required by the target database; The archived file, after format conversion, is replayed and written to the target database.
7. The data synchronization method according to any one of claims 1-6, characterized in that, The source database and the target database are heterogeneous databases of different database types.
8. A data synchronization device, characterized in that, include: The determination module is used to determine the base transaction sequence number of the target transaction in the source database; wherein, the source database includes multiple physical nodes; The scanning module is used to scan all archived logs of the target transaction in the source database to obtain all archived files in the archived logs whose transaction sequence number is greater than the baseline transaction sequence number; The grouping module is used to group the archived files according to their corresponding physical nodes, and sort the physical nodes according to a preset node arrangement strategy to obtain a node queue. The calculation module is used to perform density calculation on each archive file in the node queue based on a preset file density model, and to perform fragmentation processing on each archive file according to the density calculation results. The migration module is used to migrate the sharded archive files to the target database in sequence according to the arrangement order of the physical nodes in the node queue.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the data synchronization method as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the data synchronization method as described in any one of claims 1-7.
Citation Information
Patent Citations
A data synchronization method and a data synchronization device
CN109241185A
Data synchronization method, device, electronic device and storage medium
CN119760021A
Providing access to data within a migrating data partition
US10657154B1
Document store export / import
US20190340278A1
Method for index data processing, method for index generation, medium and device
US20250110923A1