Data synchronization method and apparatus, computing device, and storage medium
Patent Information
- Application Number
- CN202510397356.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-09-29
AI Technical Summary
但是,为保障数据一致性,在无锁变更期间,其他数据库节点上不允许针对该无锁变更的表中的数据进行更新、插入以及删除等业务操作,若执行无锁变更的耗时较长,会严重影响其他数据库节点上的业务进展
[0009]若仅通过预设名单进行判断,则可仅应用第一名单进行判断,也可以仅应用第二名单进行判断,还可以同时应用第一名单和第二名单进行判断。相应地,对于任一操作记录数据,在该操作记录数据关联的表名与第一名单中的任一表名匹配的情况下,表示该操作记录数据指示的操作针对临时表,该条数据无需进行传输,过滤掉即可。这可以实现精准过滤,将针对临时表的这类数据从后续处理流程中排除,减少了数据处理量。在表名与第二名单中的各个表名均不匹配的情况下,表示该操作记录数据指示的操作并不是针对指定的数据表,该条数据不允许进行传输,过滤掉该操作记录数据。这可以确保后续处理的数据都来自于指定的表,避免了对无关或不可信数据的处理,提高了数据同步的准确性、可靠性以及效率。在表名与第一名单中的任一表名匹配且与第二名单中的各个表名均不匹配的情况下,过滤掉操作记录数据。这进一步提高了数据过滤的准确性和严格性,更加灵活,有利于适应复杂多变的数据处理场景需求。上述这几种通过预设名单进行过滤的方式,有利于实现对数据的精准过滤,只需判断表名是否命中名单,即可快速判断是否过滤相应的操作记录数据,提高了处理效率。
Smart Images

Figure CN122838503A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a data synchronization method, apparatus, computing device, and storage medium. Background Technology
[0002] In data management scenarios, to ensure data security and high availability, multiple database nodes are typically deployed in different locations to synchronize data. Each database node contains multiple databases, and each database stores multiple tables. Currently, lock-free modification operations are commonly used to change the table structure. When performing a lock-free modification on a single database node, a temporary table with the same structure as the original table is first created. The table structure is then modified on the temporary table. Data from the original table is gradually copied to the temporary table, and finally, the temporary table replaces the original table. After the above process is completed, the operation records during the lock-free modification are synchronized to other database nodes to achieve data synchronization across all database nodes. However, to ensure data consistency, during the lock-free modification period, other database nodes are not allowed to perform update, insert, or delete operations on the data in the table that has undergone the lock-free modification. If the lock-free modification takes a long time, it will seriously affect the business progress on other database nodes. Summary of the Invention
[0003] This application provides a data synchronization method, apparatus, computing device, and storage medium, which can reduce latency and ensure data consistency across database nodes. The technical solution is as follows.
[0004] Firstly, a data synchronization method is provided. Upon receiving at least one operation record from a first database node, operation record data targeting temporary tables is filtered out from the at least one operation record, and the remaining data is sent to a second database node. The aforementioned at least one operation record indicates an operation on a data table in the first database node, which can be a table or a temporary table. In this way, there is no need to transmit operation record data targeting temporary tables between database nodes, significantly reducing latency, improving data transmission efficiency, and ensuring data consistency.
[0005] In some embodiments, only operation records of temporary tables within a preset filtering range are filtered out, while operation records of temporary tables outside the preset filtering range are retained. The preset filtering range can be set at multiple levels, such as instance level, database level, or table level. Accordingly, the preset filtering range can be all data tables under the first database node, all data tables in the target database, or the target data table. By supporting multi-level filtering range configuration, flexibility is improved, enabling flexible responses to various complex scenarios and requirements.
[0006] This application does not limit how the preset filtering range can be configured. For example, the preset filtering range can be a pre-set filtering range corresponding to a lockless change operation; or, the preset filtering range can be a filtering range determined based on a configuration operation. By supporting multiple configuration methods for preset filtering ranges, flexibility and user experience are greatly improved.
[0007] In some embodiments, when filtering at least one operation record, the associated table name is first parsed from the operation record data, that is, the data table to which the operation in the operation record data is targeted is determined. This allows subsequent processing to determine whether to filter out operation record data based on the table name. If the table name indicates that the operation in the operation record data targets a temporary table, it means that the operation record data is intermediate data generated during a lock-free change process and does not need to be transmitted between database nodes; therefore, this operation record data is filtered out. If the table name indicates that the operation in the operation record data targets a data table other than a temporary table, it means that the operation record data needs to be synchronized between database nodes; therefore, this operation record data is retained. By filtering out operation record data associated with temporary tables in this way, the amount of data transmitted between database nodes is greatly reduced, data transmission efficiency is improved, and data consistency is ensured.
[0008] When determining which data tables are associated with the operation instructions in the operation log data, multiple processing methods are supported. For example, the determination can be made using only a preset list, only a preset regular expression, or a combination of both. The preset list includes at least one of a first list and a second list. The first list records the table names associated with the operation log data for the temporary table, and the second list records the table names associated with the remaining data.
[0009] If only a preset list is used for judgment, the first list can be used alone, the second list alone, or both lists can be used simultaneously. Accordingly, for any operation record data, if the table name associated with the operation record data matches any table name in the first list, it indicates that the operation indicated by the operation record data is for a temporary table, and this data does not need to be transmitted and can be filtered out. This achieves precise filtering, excluding such data for temporary tables from subsequent processing, reducing the amount of data processing. If the table name does not match any table name in the second list, it indicates that the operation indicated by the operation record data is not for the specified data table, and this data is not allowed to be transmitted and should be filtered out. This ensures that subsequent processed data comes from the specified table, avoiding the processing of irrelevant or unreliable data, and improving the accuracy, reliability, and efficiency of data synchronization. If the table name matches any table name in the first list but does not match any table name in the second list, the operation record data is filtered out. This further improves the accuracy and strictness of data filtering, making it more flexible and adaptable to complex and ever-changing data processing scenarios. The above-mentioned filtering methods using preset lists facilitate accurate data filtering. By simply checking whether the table name matches the list, it is possible to quickly determine whether to filter the corresponding operation record data, thus improving processing efficiency.
[0010] If only a preset regular expression is used for judgment, then for any operation record data, if its associated table name matches any preset regular expression, the operation record data will be filtered out. The preset regular expression indicates the element in the temporary table name. This method avoids complex string comparisons, and can efficiently and accurately complete data filtering by identifying elements in the table name. Furthermore, the regular expression can be formulated according to different temporary table naming rules, and can achieve accurate matching even in complex scenarios, making it more flexible and versatile.
[0011] If a pre-defined list is used for judgment, and the table name of a certain operation record does not match the pre-defined list, it may be difficult to determine whether to filter it. To address this, a combination of pre-defined lists and pre-defined regular expressions is used for judgment. For any operation record, if its associated table name does not match any table name in the first list or the second list, the table name is matched based on at least one pre-defined regular expression. The first list records the table names associated with operation record data for temporary tables, and the second list records the table names associated with the remaining data. The pre-defined regular expressions indicate the elements in the temporary table name. If the table name matches any pre-defined regular expression, the operation record data is filtered out. This multi-level judgment method further improves the accuracy of identifying temporary table operation records, enabling more comprehensive and detailed data filtering to ensure that only the data that truly needs processing is retained. The pre-defined list can quickly identify common temporary tables and specified tables, while regular expressions can handle some specially named or difficult-to-cover cases. This approach utilizes both the explicitness of lists and the flexibility of regular expressions, adapting to more complex data environments.
[0012] Building upon the aforementioned scheme, updating the preset list based on already matched table names further improves filtering efficiency. Specifically, if a table name matches any preset regular expression, the operation records associated with that table name should be filtered out, and the table name is added to the first list. If a table name does not match any of the preset regular expressions, the operation records associated with that table name should be retained, and the table name is added to the second list. This real-time updating of the preset list improves data filtering efficiency. Table names already matched by preset regular expressions do not need to be matched again; applying the preset list for identification is faster and more efficient. Furthermore, this method dynamically adapts to data changes. For example, with the creation of new tables or modifications to table names, relevant table names can be added to the appropriate list promptly without manual updates. This ensures the accuracy and effectiveness of data filtering and improves the automation and efficiency of data processing.
[0013] Similarly, this application does not limit how the preset regular expressions and preset lists are configured. For example, the preset regular expression can be a pre-set regular expression corresponding to a lock-free change operation or a regular expression determined based on the configuration operation. By providing pre-set regular expressions and supporting various custom configuration methods, flexibility and user experience are greatly improved.
[0014] Secondly, a data synchronization apparatus is provided for executing the aforementioned data synchronization method. Specifically, the data synchronization apparatus includes a functional module for executing the data synchronization method provided in the first aspect or any optional embodiment of the first aspect.
[0015] Thirdly, a computing device or cluster of computing devices is provided, the computing device including a processor for executing program code, causing the computing device or cluster of computing devices to perform operations as described above in the data synchronization method.
[0016] Fourthly, a computer-readable storage medium is provided, which stores at least one piece of program code that is read by a processor to cause a computing device to perform operations as described in the data synchronization method above.
[0017] Fifthly, a computer program product or computer program is provided, the computer program product or computer program including program code stored in a computer-readable storage medium, a processor of a computing device reading the program code from the computer-readable storage medium, the processor executing the program code, causing the computing device to perform the methods provided in the first aspect or various alternative implementations of the first aspect.
[0018] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a system architecture provided in an embodiment of this application;
[0020] Figure 2 This is a schematic diagram of a system architecture in a dual-master disaster recovery scenario provided in an embodiment of this application;
[0021] Figure 3 This is a flowchart illustrating the configuration of filtering parameters provided in an embodiment of this application;
[0022] Figure 4 This is a flowchart of a data synchronization method provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of a lock-free change at both ends provided in an embodiment of this application;
[0024] Figure 6 This is a schematic diagram of a lock-free change filtering method provided in an embodiment of this application;
[0025] Figure 7 This is a schematic diagram of the structure of a data synchronization device provided in an embodiment of this application;
[0026] Figure 8 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below with reference to the accompanying drawings. It should be noted that all information and data involved in this application are authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the operation log data, preset regular expressions, preset filtering ranges, first lists, and second lists involved in this application were all obtained under fully authorized conditions.
[0028] To facilitate understanding, the key terms and concepts involved in this application will be explained below.
[0029] Active-active database: This refers to a distributed database system where multiple database nodes are simultaneously active, sharing the workload and providing services. In this scenario, data synchronization is performed between the multiple database nodes to ensure data consistency.
[0030] Database node: Used for storing, processing, and managing data. A single database node may include an independent database server and instances running on that database server; a single database node may also include a server cluster consisting of multiple database servers and instances running on that server cluster. Instances include various processes and data files, etc. Each database node contains multiple databases, and each database stores multiple tables. Different databases within each database node may be distributed across the same or different database servers; this application embodiment does not impose such limitations.
[0031] Active-active disaster recovery: This is a special type of active-active database scenario where two active database nodes run simultaneously in two different geographical locations or data centers. Both database nodes can independently perform database operations, and data is synchronized between them in real time. Dual-primary disaster recovery is a special type of active-active disaster recovery where both database nodes are primary nodes, and either primary node can act as a disaster recovery node for the other. When one primary node fails or experiences a disaster, the other primary node can automatically take over its tasks.
[0032] Relational databases are databases that store data in a tabular format, consisting of rows and columns. Commonly used relational databases include MySQL and PostgreSQL. In MySQL, the binary log (Bin log) is used to record all changes made to the database.
[0033] Data definition language (DDL) refers to statements used to define the structure of tables in a database, such as creating, modifying, or deleting rows or columns. Data manipulation language (DML) refers to statements used to manipulate data in a database, such as inserting, updating, or deleting data.
[0034] Lock-free changes refer to modifying a table structure without locking the table, allowing DML operations to execute normally during the change process and avoiding impact on business operations. A common approach is using online data definition languages (Online DDL). This method allows DDL operations to be executed without blocking DML operations. Lock-free changes are typically implemented using lock-free change applications, such as GitHub Online Schema Change (GH-OST) and Percona Toolkit Online Schema Change (PT-OSC).
[0035] Currently, in data management scenarios, to ensure data security and high availability, multiple database nodes are typically deployed in different locations to synchronize data. Each database node contains multiple databases, and each database stores multiple tables. These tables can store application-related data or enterprise business-related data. For example, table A in the database stores transaction order information, table B stores item inventory information, and table C stores item logistics information. Business operations such as updating, inserting, and deleting can be performed on the tables in the database. For example, deleting the inventory information of an item from table B, writing new transaction order information to table A, or updating the logistics information of an item in table C.
[0036] When a database table needs structural modifications, a typical change process would involve locking the table to prevent read and write operations during the change. For example, adding a global lock would prevent any insert, update, or delete operations from proceeding until the change is complete. This can lead to business interruptions and impact database availability, especially when dealing with large tables, where this approach is time-consuming and has a significant impact.
[0037] To ensure business continuity, lock-free modification operations are commonly used to alter table structures. When performing lock-free modifications on a single database node, a temporary table with the same structure as the original table is first created. The structure is then modified on the temporary table. For example, if a column is added to the original table, it will be added to the temporary table at the corresponding location. A progressive replication method is used to copy data row by row from the original table to the temporary table. During replication, read and write operations are allowed on the original table, meaning update, insert, and delete operations are permitted. A binary log is used to record changes occurring on the original table, ensuring consistency between the two tables during data replication. Once replication is complete, both the original and temporary tables are renamed. The temporary table is renamed to the original table's name, and the original table is renamed to another name. The renamed original table can then be deleted, effectively replacing the original table with the temporary table. During the lock-free modification operation on that database node, all operations performed on each table within that node are recorded in the operation log data. After the lock-free change operation is completed, the operation record data is synchronized to other database nodes to ensure data synchronization across all database nodes. To guarantee data consistency, during the lock-free change, other database nodes are not allowed to perform business operations on the data in the table that underwent the lock-free change.
[0038] However, the above method has significant drawbacks. During lock-free change operations, relying on technical personnel to control database nodes to prevent business operations from occurring not only incurs substantial manpower costs but also poses potential risks due to human error. Furthermore, during lock-free change operations, other database nodes are prohibited from performing business operations. If the lock-free change takes too long, it can severely impact business progress on other database nodes, potentially leading to serious business interruptions, especially in scenarios where business writes are essential. Additionally, during lock-free change operations, data from the original table needs to be copied to a temporary table. This process generates a large amount of operation record data, resulting in significant incremental latency on the synchronization link between database nodes, affecting data synchronization efficiency and data consistency.
[0039] Based on this, this application provides a data synchronization method that reduces latency on the synchronization link, enabling database nodes in a multi-active database scenario to perform lock-free change operations simultaneously. During this process, business operations can be performed on these database nodes, thereby ensuring business continuity and data consistency.
[0040] The following is an exemplary system architecture, such as Figure 1 As shown. Figure 1This is a schematic diagram of a system architecture provided in an embodiment of this application, used to implement the data synchronization method provided in this embodiment. The system includes a database node 101, a business server 102, and a data synchronization server 103. Different database nodes 101 correspond to their respective business servers 102. Data can be transmitted between the database nodes 101 and their corresponding business servers 102. Different database nodes 101 transmit data through the data synchronization server 103.
[0041] Database node 101 is used to store, process, and manage data. Optionally, a single database node 101 may include an independent database server and instances running on that database server; a single database node 101 may also include a server cluster consisting of multiple database servers and instances running on that server cluster. Each database node 101 has multiple databases, and each database stores multiple tables. Different databases in each database node 101 may be distributed on the same or different database servers, and this embodiment does not impose any restrictions on this. Business server 102 is used to process operations on tables in the database node. For example, business server 102 provides background services for business operations on data or lock-free change operations on table structures. Data synchronization server 103 is used to filter operation record data for temporary tables. For example, data synchronization server 103 obtains the filtering parameters for lock-free change filtering, filters the operation record data from a certain database node 101 based on the filtering parameters, and transmits the remaining operation record data to other database nodes 101. The filtering parameters include at least one of a preset filtering range, a preset regular expression, a preset list, and a lock-free change type.
[0042] Optionally, each synchronization link corresponds to its own data synchronization server 103. The data synchronization server 103 for any synchronization link receives operation record data from the starting database node of that synchronization link and, after filtering, sends the remaining data to the target database node of that synchronization link. For example, the database synchronization server 103 and its corresponding target database node are located in the same location. Typically, the database synchronization server 103 and the target database node perform multiple data interactions. Placing the database synchronization server 103 and the target database node within a short distance can improve data transmission and synchronization efficiency, which is beneficial for ensuring data consistency. Alternatively, there may be only one independent data synchronization server 103 in the system. Each database node 101 sends operation record data to the data synchronization server 103, which filters this data and sends the remaining data to other database nodes 101. The above description of data transmission and synchronization implemented by the data synchronization server 103 is merely an illustrative example. In some embodiments, a new functional module is added to each database node 101. This functional module is used to receive operation record data from other database nodes 101 and process the current database node 101 based on the remaining data after filtering.
[0043] The above briefly introduces the system architecture provided in the embodiments of this application. The following describes various exemplary functional modules in this system architecture in a dual-master disaster recovery scenario, such as... Figure 2 As shown. Figure 2 This is a schematic diagram of a system architecture in a dual-master disaster recovery scenario provided in this application embodiment. The system includes a first database node and a second database node, which are the two master nodes in the dual-master disaster recovery scenario. There are two synchronization links between the two master nodes, and the data transmission directions on these two synchronization links are opposite. The system is illustrated by the example where each synchronization link corresponds to its own data synchronization server. The system includes a data synchronization server M and a data synchronization server N. Data synchronization server M is used to filter the operation record data sent by the first database node and send the remaining data to the second database node; data synchronization server N is used to filter the operation record data sent by the second database node and send the remaining data to the first database node. Exemplarily, data synchronization server M and the second database node are located in the same location, and data synchronization server N and the first database node are located in the same location.
[0044] Each database server deploys a database function module, which stores and processes data within the database node. Each business server deploys a business module, which handles read / write operations and structural change operations on tables within the database node. Each data synchronization server deploys a lock-free change configuration module, a lock-free change filtering module, and a data replay module. The lock-free change configuration module retrieves filtering parameters, the lock-free change filtering module filters out operation records from temporary tables generated by lock-free changes, and the data replay module converts the remaining data into a suitable format for the database node to write.
[0045] In multi-active scenarios, to achieve data synchronization between different database nodes, multi-active synchronization tasks are created for multiple database nodes. Taking a dual-master disaster recovery scenario as an example, a dual-active synchronization task is created to achieve data synchronization between the two master nodes. To facilitate filtering of operation record data during data synchronization, it is supported to configure lock-free change filtering parameters. The following two steps are described in detail.
[0046] In some embodiments, a dual-active synchronization task is created on a data synchronization application. For example, a first database node and a second database node are selected in the data synchronization application, and a dual-active synchronization task is created for these two database nodes. After the task starts, two synchronization links are enabled, allowing the two database nodes to synchronize data with each other. The two synchronization links are used to respectively realize data transmission from the first database node to the second database node and data transmission from the second database node to the first database node. In this case, both database nodes are in a disaster recovery state, that is, incremental synchronization is performed between the database nodes. The data synchronization application can be a functional module of a data management application, or it can be a standalone application; this embodiment does not limit this. The data management application is used to manage the database nodes.
[0047] In some embodiments, at least one of the following filtering parameters is configured for lock-free change filtering: lock-free change type, preset filtering range, preset regular expression, and preset list. Optionally, this parameter configuration process is implemented in the aforementioned data synchronization application or in other control applications targeting lock-free changes; this application embodiment does not impose any limitations on this. The filtering parameters are explained below through (1)-(3).
[0048] (1) The lock-free change type refers to the type of lock-free change application that implements lock-free change operations. Lock-free change applications are used to enable change operations in the database without a lock mechanism. Different lock-free change applications indicate different lock-free change types, such as GH-OST and PT-OSC; or, each lock-free change application may indicate a different lock-free change type when using different temporary table naming rules, which will not be elaborated here. Optionally, the data synchronization application provides a default lock-free change application and supports modification of the lock-free change application; or, the data synchronization application provides multiple candidate lock-free change applications, supporting selection from among them.
[0049] (2) The preset filtering scope is used to limit the scope of lock-free change filtering. The data table is the basic unit of the preset filtering scope. During the subsequent execution of lock-free change filtering, operation record data associated with all tables within the preset filtering scope are filtered. In some embodiments, three levels of configuration for the filtering scope are supported: table-level, database-level, and instance-level. Table-level configuration indicates that the lock-free change filtering scope is at the table level, filtering only the configured table (i.e., the target data table), and has no filtering effect on other tables. That is, only operation record data associated with temporary tables corresponding to the target data table is filtered out. Database-level configuration indicates that the lock-free change filtering scope is at the database level, filtering only the data tables in the configured database (i.e., the target database), and has no filtering effect on data tables in other databases. That is, only operation record data associated with all temporary tables in the target database is filtered out. Instance-level configuration indicates that the lock-free change filtering scope is at the instance level, filtering all temporary tables corresponding to all tables under the database node. That is, all operation record data associated with temporary tables in the database node are filtered out. By allowing configuration of multiple levels of filtering range, it can be applied to a variety of scenarios, improving the flexibility and applicability of the solution.
[0050] Optionally, the preset filtering range is the filtering range corresponding to a pre-defined lock-free change operation, which can be the filtering range corresponding to the lock-free change tool that implements the lock-free change operation; or, the preset filtering range is the filtering range determined based on the configuration operation. The above configuration operation can be an adjustment operation for the pre-defined filtering range, or a custom operation for the filtering range. This provides diverse configuration methods and improves the user experience. The customization process is briefly described below.
[0051] When configuring a preset filtering range, the data synchronization application responds to the selection of a filtering level by providing the names of candidate filter objects at that filtering level, and determines the preset filtering range in response to the selection of a specific name; alternatively, the data synchronization application determines the preset filtering range in response to both the selection of a filtering level and the specification of filter objects at that filtering level. Optionally, the specification of filter objects can be an input operation of the name of the filter object or a selection operation of the storage path of the filter object. When the filtering level is at the database level, the filter object is the database; when the filtering level is at the table level, the filter object is the data table.
[0052] (3) Both the preset regular expressions and the preset lists are used to filter out specific operation record data.
[0053] Optionally, a preset regular expression can be used to indicate elements in the name of a temporary table. In this case, operation records matching the preset regular expression will be filtered out during the lock-free change filtering process. Alternatively, a preset regular expression can be used to indicate elements in the name of a data table other than a temporary table. In this case, operation records matching the preset regular expression will be retained and transmitted to other database nodes during the lock-free change filtering process. There can be one or more preset regular expressions.
[0054] Optionally, the preset list can be a first list (i.e., a blacklist), a second list (i.e., a whitelist), or both. The first list records the table names associated with operation records of temporary tables, and the second list records the table names associated with operation records of data tables other than temporary tables. When the preset list is the first list, operation records whose associated table names match the first list are filtered out during the lock-free change filtering process. When the preset list is the second list, operation records whose associated table names do not match the second list are filtered out during the lock-free change filtering process. When the preset list includes both the first and second lists, operation records whose associated table names match the first list but do not match the second list are filtered out during the lock-free change filtering process.
[0055] The aforementioned preset regular expression is either a pre-defined regular expression corresponding to the lock-free change operation or a regular expression determined based on the configuration operation. The aforementioned preset list is either a pre-defined list corresponding to the lock-free change operation or a list determined based on the configuration operation. The regular expression corresponding to the lock-free change operation can be the regular expression corresponding to the lock-free change tool implementing the lock-free change operation, and the list corresponding to the lock-free change operation can be the list corresponding to the lock-free change tool implementing the lock-free change operation. The configuration operation here can be an adjustment operation based on the pre-defined regular expression or the pre-defined list, or it can be a custom operation, such as an input operation, which will not be elaborated further here. This provides diverse configuration methods and improves the user experience.
[0056] In some embodiments, multiple preset filtering ranges can be configured, and different filtering parameters can be configured for different preset filtering ranges, such as configuring different lock-free change types, preset regular expressions, or at least one of preset lists. For example, two preset filtering ranges have been configured, one of which indicates filtering operation record data for a certain table, and its corresponding lock-free change type is the type indicated by GH-OST; the other preset filtering range indicates filtering operation record data for a certain database, and its corresponding lock-free change type is the type indicated by PT-OSC.
[0057] In a multi-active scenario, the filtering parameters for the two synchronization links between any two database nodes remain consistent. Taking the first and second database nodes as examples, the data synchronization servers corresponding to these two database nodes use the same filtering parameters for filtering. That is, they use the same filtering parameters to process the operation record data sent from the first database node to the second database node and the operation record data sent from the second database node to the first database node.
[0058] The configuration order of the above-mentioned filtering parameters is not limited in this embodiment. For example, the preset filtering range can be configured first, followed by the lock-free change type, preset regular expression, or preset list; or, the configuration of the lock-free change type, preset regular expression, or preset list can be completed first, followed by the configuration of the preset filtering range. The names of the temporary tables generated after different lock-free change applications perform lock-free change operations are different. In some embodiments, the lock-free change application is first selected to determine the pre-set regular expression or list corresponding to that application, and a preset regular expression or preset list is configured based on this. In subsequent processes, operation record data associated with specific table names generated by that lock-free change application is filtered out.
[0059] The following is an exemplary configuration process for lock-free change filtering parameters. See [link to documentation]. Figure 3 As shown, Figure 3 This is a flowchart illustrating the configuration of filtering parameters provided in an embodiment of this application. First, the lock-free change type is configured, such as selecting Tool 1 (GH-OST) or Tool 2 (PT-OSC) for lock-free change applications. Then, the regular expression for this lock-free change type is configured, such as adjusting a pre-set regular expression corresponding to this lock-free change type or creating a custom regular expression. Finally, the preset filtering range is configured, supporting configuration at three levels: database level, table level, and instance level. When selecting the database level, the name of the database to be filtered can be configured; when selecting the table level, the name of the data table can be configured.
[0060] This application does not restrict the order of the two steps: creating the active-active synchronization task and configuring the filtering parameters for lock-free changes. In some embodiments, the active-active synchronization task is created first, and the parameters for lock-free change filtering are configured only after all database nodes are in disaster recovery mode. This allows relevant personnel to flexibly and accurately configure the filtering parameters according to the actual operating conditions of the database nodes, improving the adaptability of the parameters and the practicality of the solution. Alternatively, the parameters for lock-free change filtering can be configured first for specific database nodes, and the active-active synchronization task can be created later if there is a data synchronization requirement for lock-free change scenarios among the aforementioned database nodes. In this approach, the parameters are pre-configured, and there is no need to temporarily adjust the parameters when the corresponding requirement occurs. This allows for a rapid response and initiation of the data synchronization and filtering process, effectively shortening the processing cycle and improving subsequent processing efficiency.
[0061] Based on the preparatory work described above, such as creating a multi-active synchronization task and configuring filtering parameters, the following describes the detailed steps for filtering and transmitting operation record data using the data synchronization method provided in this application embodiment after completing these preparatory work. For example... Figure 4 As shown, Figure 4 This is a flowchart illustrating a data synchronization method provided in an embodiment of this application. The method is explained using a data synchronization server executing the data synchronization method as an example. This data synchronization server can provide subsequent... Figure 8 The method includes the following steps: (The text abruptly ends here, likely due to an incomplete sentence or a formatting error.)
[0062] 401. The data synchronization server receives at least one operation record data from the first database node. The operation record data indicates the operation on a data table in the first database node. The data table includes a table and a temporary table. The temporary table is used to perform lock-free change operations based on the data in the table.
[0063] In this scenario, the first database node is the database node in a multi-active environment. All database nodes in the multi-active environment are performing lock-free change operations, i.e., modifying the table structure within the database node. Optionally, lock-free change operations can be implemented on all database nodes in the multi-active environment through a lock-free change application. For example, the lock-free change operation to be performed can be set in the lock-free change application, and then the relevant data can be distributed to the business server associated with each database node, thereby implementing the lock-free change operation on each database node. An example of a multi-active environment including a first database node and a second database node that are in a data synchronization state is provided below. Figure 5 As shown, Figure 5 This is a schematic diagram of a lock-free change operation provided in an embodiment of this application. Lock-free change operations are implemented on both the first and second database nodes using a lock-free change application. Data synchronization is achieved between the two database nodes via two synchronization links, using data synchronization applications deployed on the two database synchronization servers.
[0064] The aforementioned at least one operation record includes operation record data for DML operations and DDL operations on the first database node. Specifically, the DML operation record data includes operation record data for the original table and operation record data for temporary tables during lock-free changes. Similarly, the DDL operation record data also includes operation record data for temporary tables during lock-free changes and operation record data for the original table during lock-free changes.
[0065] In some embodiments, the data synchronization server obtains at least one operation record from the incremental log. Accordingly, the data synchronization server obtains the incremental log from the first database node, which records operations occurring on the first database node; it then parses the incremental log to obtain the operation record data. For example, the incremental log is a Bin log pulled from the first database node, and the data synchronization server parses the log using both a DDL parser and a DML parser to obtain operation record data for DML operations and DDL operations, respectively.
[0066] 402. The data synchronization server filters out operation record data for temporary tables from at least one operation record data.
[0067] In traditional multi-active scenarios, the incremental logs transmitted between database nodes include operation records for temporary tables, as well as operation records generated by business operations during lock-free changes. This results in a large volume of incremental log data, leading to significant latency. Other database nodes, except those performing lock-free changes, cannot receive the incremental logs in a timely manner, thus failing to achieve timely data synchronization or adjust to a state that allows business operations. This significantly impacts the business operations of the entire database system.
[0068] In this embodiment, lock-free change operations can be performed simultaneously on database nodes in a multi-active scenario. Under this architecture, there is no need to transmit operation record data related to lock-free change operations between database nodes to achieve synchronization of this operation. Therefore, a large amount of operation record data for temporary tables is filtered out in the data to be transmitted to other database nodes. This not only greatly improves data transmission efficiency and reduces latency, ensuring data consistency between database nodes, but also avoids resource waste and performance loss caused by invalid transmission, thus improving overall performance.
[0069] In some embodiments, operation record data for temporary tables is filtered out from operation record data within a preset filtering range. The operation record data within the preset filtering range includes at least one of the following: operation record data for all data tables in the first database node; operation record data for all data tables in the target database within the first database node; and operation record data for the target data table in the first database node.
[0070] In some embodiments, the process of filtering out operation record data for temporary tables from at least one operation record data is achieved through the following steps (1)-(2).
[0071] (1) For any operation record, the data synchronization server determines the table name associated with the operation record. The table name associated with any operation record indicates the data table to which the operation in the operation record targets. For example, the table name is parsed from the operation record data of the DML operation and the operation record data of the DDL operation mentioned above by the table name resolver.
[0072] (2) If the table name indicates that the operation in the operation record data is for a temporary table, the data synchronization server will filter out the operation record data.
[0073] The following provides three exemplary implementation methods for determining the table name, indicating the operation record data, targeting a temporary table, and filtering out the operation record. These are shown in Method A, Method B, and Method C.
[0074] Method A: The data synchronization server determines whether an operation in the operation record data targets a temporary table based on a preset list. Accordingly, if the table name matches any table name in the first list (which records the table names associated with operation record data targeting temporary tables), the data synchronization server filters out the operation record data. If the table name does not match any table name in the second list, the data synchronization server filters out the operation record data. The second list records the table names associated with the remaining data. If the table name matches any table name in the first list but does not match any table name in the second list, the data synchronization server filters out the operation record data.
[0075] The first list is also called the blacklist, and the second list is also called the whitelist. Optionally, the cache center stores the first and second lists, and the parsed table names are matched in the cache center to determine whether to filter the operation record data.
[0076] Using only the first list for judgment can accurately determine whether the operation record data pertains to a temporary table, thus achieving precise filtering and excluding such data from subsequent processing, reducing the amount of data processed. Using only the second list ensures that subsequent processed data comes from the specified table, avoiding the processing of irrelevant or unreliable data and improving the accuracy, reliability, and efficiency of data synchronization. A dual judgment mechanism combining the first and second lists further enhances the accuracy and rigor of data filtering, offering greater flexibility and adaptability to complex and ever-changing data processing scenarios.
[0077] The above-described filtering method using a preset list facilitates precise data filtering. By simply checking whether the table name matches the list, it is possible to quickly determine whether to filter the corresponding operation record data, thus improving processing efficiency.
[0078] Method B: The data synchronization server determines whether the operation in the operation record data targets a temporary table based on a preset regular expression. Accordingly, if the table name matches any of the preset regular expressions, the data synchronization server filters out the operation record data; in this case, the preset regular expression indicates an element in the temporary table's name. Alternatively, if the data synchronization server does not match any of the preset regular expressions, the data synchronization server filters out the operation record data; in this case, the preset regular expressions indicate elements in the names of data tables other than temporary tables.
[0079] The aforementioned elements are specific prefixes, suffixes, or contain specific characters in the middle, which will not be elaborated further in this embodiment. Optionally, a temporary table regular expression matcher is used to match the table name, and the temporary table regular expression matcher is configured with preset regular expressions.
[0080] If the preset regular expression specifies elements in the name of a temporary table, operations on the temporary table can be quickly and accurately identified in the operation record data, achieving efficient data filtering. If the preset regular expression specifies elements in the name of a data table other than a temporary table, it can further ensure that the data processed subsequently comes from the specified table, achieving precise retention of important data.
[0081] The above-described filtering method using pre-defined regular expressions avoids complex string comparisons. By identifying elements in the table name, data filtering can be completed efficiently and accurately. The regular expressions can be customized according to different temporary table naming rules, and can achieve accurate matching even in complex scenarios, making it more flexible and versatile.
[0082] Method C: The data synchronization server determines whether an operation in the operation record data targets a temporary table based on a preset list and preset regular expressions. Accordingly, the preset list includes a first list and a second list as an example. If the table name does not match any of the table names in the first list or the second list, the data synchronization server matches the table name based on at least one preset regular expression, which indicates the elements in the temporary table's name. If the table name matches any of the preset regular expressions, the operation record data is filtered out. Specifically, if the table name matches any preset regular expression, it means that the operation in the operation record data associated with that table name targets a temporary table, and this operation record data is to be filtered out. Alternatively, if the table name does not match any of the table names in the first list or the second list, the data synchronization server matches the table name based on at least one preset regular expression, which indicates the elements in the names of data tables other than temporary tables. If the table name does not match any of the preset regular expressions, the operation record data is filtered out. If the table name does not match any of the preset regular expressions, it means that the operation in the operation record data associated with the table name is not for the specified data table, and the operation record data is redundant.
[0083] This multi-level judgment method, combining a pre-defined list and pre-defined regular expressions, further improves the accuracy of identifying temporary table operation records by applying regular expressions only when the table name does not match the pre-defined list. It allows for more comprehensive and detailed data filtering, ensuring that only the data truly needed for processing is retained. The pre-defined list can quickly identify common temporary tables and specified tables, while regular expressions can handle cases with special naming conventions or those difficult to cover with a list. This approach leverages both the explicitness of lists and the flexibility of regular expressions, adapting to more complex data environments.
[0084] Optionally, the data synchronization server updates a preset list during the lock-free change filtering process. Accordingly, if a preset regular expression indicates an element in a temporary table, and the table name matches any preset regular expression, it means that the operation record data associated with that table name should be filtered out, and the data synchronization server adds the table name to the first list; if the table name does not match any of the preset regular expressions, it means that the operation record data associated with that table name should be retained, and the data synchronization server adds the table name to the second list. Alternatively, if a preset regular expression indicates an element in a data table other than a temporary table, and the table name matches any preset regular expression, it means that the operation record data associated with that table name is data to be retained, and the data synchronization server adds the table name to the second list; if the table name does not match any of the preset regular expressions, it means that the operation record data associated with that table name is not data to be retained and should be filtered out, and the data synchronization server adds the table name to the first list.
[0085] This method of real-time updating of the preset list improves data filtering efficiency. For table names already matched by the preset regular expression, there's no need to match them again using the preset regular expression; applying the preset list for identification is faster and more efficient. Furthermore, this method dynamically adapts to data changes. For example, with the creation of new tables or modifications to table names, relevant table names can be added to the appropriate list promptly without manual updates. This ensures the accuracy and effectiveness of data filtering and improves the automation and efficiency of data processing.
[0086] 403. The data synchronization server sends the remaining data from at least one operation record to the second database node. The remaining data is the data from at least one operation record excluding the filtered operation record data. The second database node and the first database node are in a data synchronization state.
[0087] If data filtering is completed for at least one operation record, the remaining data that was not filtered out is sent to the second database node, which is the target database node of the synchronization link where the current data synchronization server is located.
[0088] To facilitate the description of the above process, we will take the example of a data synchronization server using method C with a preset regular expression indicating the elements in the temporary table name. See [link to documentation]. Figure 6 , Figure 6This is a schematic diagram of a lock-free change filtering method provided in an embodiment of this application. First, the data synchronization server pulls incremental logs (i.e., Bin logs) from the first database node. It then parses these logs using a first parser (i.e., a DDL parser) and a second parser (i.e., a DML parser) to obtain operation record data for DML and DDL operations. The data synchronization server then uses a third parser (i.e., a table name parser) to parse the table names from the operation record data. The data synchronization server matches the parsed table names in the cache center. Operation record data whose table names are in the second list are not filtered; that is, operation record data matching the second list is retained and subsequently written to the second database node. Operation record data whose table names are in the first list are filtered out; that is, operation record data matching the first list is filtered out and not subsequently written to the second database node. Operation record data that does not match either the first or second list is processed in a temporary table regular expression matcher. Operation record data whose table names can match a preset regular expression are filtered out. Optionally, the above table names are added to the first list in the cache center to update the first list. Operation record data whose table names do not match the preset regular expression are retained. Optionally, the above table names are added to the second list in the cache center to update the second list.
[0089] Optionally, the data synchronization server compresses the remaining data and writes the compressed data to a data compression storage pool, awaiting data playback. During data playback, the data synchronization server reads data from the data compression storage pool and converts the format of the read data into a format recognizable by the second database node, obtaining the converted data. The data synchronization server then sends the converted data to the second database node, thus completing the data playback. The data compression storage pool uses a queue structure for data storage and retrieval, ensuring the orderly processing of data. Alternatively, the data synchronization server performs real-time format conversion and transmission of the remaining data; this embodiment will not be elaborated further.
[0090] The above explanation focuses on the first and second database nodes in a data synchronization state within a multi-active scenario. The second database node is a target database node corresponding to the first database node, meaning it receives data sent by the first database node. In a multi-active scenario with other database nodes, these nodes can also be used as target database nodes corresponding to the first database node to receive remaining data from it; alternatively, they can be used as target database nodes corresponding to the second database node to facilitate receiving remaining data from it after the data synchronization server has performed steps similar to steps 401 to 403.
[0091] In the above solution, by filtering out data related to temporary tables from the operation log data, invalid transmission of data related to lock-free change operations is avoided, reducing latency on the synchronization link and enabling rapid data synchronization between database nodes, thus ensuring data consistency. Furthermore, in multi-active scenarios, this solution allows lock-free change operations to be performed simultaneously on database nodes while still permitting business operations, ensuring normal business progress and preventing business blockage. The solution also supports configuration of lock-free change types, regular expressions, lists, and filtering ranges, greatly increasing its flexibility and applicability, and enabling it to handle various complex requirements and scenarios.
[0092] The methods of the embodiments of this application have been described above; the apparatus of the embodiments of this application will be described below. It should be understood that the apparatus described below has any of the functions of the computing device in the above methods. (The above is in conjunction with...) Figures 1 to 6 The data synchronization method provided according to the embodiments of this application is described in detail. Based on the same inventive concept, the following will be combined with Figure 7 This application describes a data synchronization apparatus provided according to embodiments of the present application. It should be understood that the technical features described in the method embodiments are also applicable to the following apparatus embodiments.
[0093] See Figure 7 This application provides a data synchronization device, which includes:
[0094] The receiving module 701 is used to receive at least one operation record data from the first database node. The operation record data indicates an operation on a data table in the first database node. The data table includes a table and a temporary table. The temporary table is used to perform lock-free change operations based on the data in the table.
[0095] Filtering module 702 is used to filter out operation record data for temporary tables from at least one operation record data;
[0096] The sending module 703 is used to send the remaining data in at least one operation record data to the second database node. The remaining data is the data in at least one operation record data excluding the filtered operation record data. The second database node and the first database node are in a data synchronization state.
[0097] In some embodiments, the filtering module 702 is used to filter out operation record data for temporary tables from operation record data within a preset filtering range.
[0098] The operation record data within the preset filtering range includes at least one of the following: operation record data for all data tables in the first database node; operation record data for all data tables in the target database in the first database node; and operation record data for the target data table in the first database node.
[0099] In some embodiments, the preset filtering range is the filtering range corresponding to a pre-set lockless change operation or the filtering range determined based on the configuration operation.
[0100] In some embodiments, the filtering module 702 includes:
[0101] The determining unit is used to determine the table name associated with any operation record data, wherein the table name associated with any operation record data indicates the data table to which the operation in the operation record data is targeted.
[0102] The filtering unit is used to filter out operation record data when the table name indicates that the operation in the operation record data is for a temporary table.
[0103] In some embodiments, the filtering unit is configured to perform any of the following: filtering out operation record data when the table name matches any table name in the first list, the first list being used to record the table names associated with the operation record data for the temporary table; filtering out operation record data when the table name does not match any of the table names in the second list, the second list being used to record the table names associated with the remaining data; and filtering out operation record data when the table name matches any of the table names in the first list but does not match any of the table names in the second list.
[0104] In some embodiments, the filtering unit described above is used to filter out operation record data if the table name matches any preset regular expression, where the preset regular expression indicates the element in the table name of the temporary table.
[0105] In some embodiments, the filtering unit is configured to match table names based on at least one preset regular expression when the table name does not match any of the table names in the first list and the second list. The first list is used to record the table names associated with the operation record data of the temporary table, and the second list is used to record the table names associated with the remaining data. The preset regular expression indicates the elements in the table name of the temporary table. When the table name matches any preset regular expression, the operation record data is filtered out.
[0106] In some embodiments, the device further includes:
[0107] The add module is used to add the table name to the first list if the table name matches any of the preset regular expressions, and to add the table name to the second list if the table name does not match any of the preset regular expressions.
[0108] In some embodiments, the preset regular expression is a regular expression corresponding to a pre-set lock-free change operation or a regular expression determined based on a configuration operation.
[0109] It should be understood that the data synchronization device corresponds to the computing device in the above method embodiments. Each module and the other operations and / or functions in the device are respectively for implementing various steps and methods performed by the computing device in the method embodiments. Specific details can be found in the above method embodiments, and for simplicity, they will not be repeated here. It should be understood that the division of modules and units in the above data synchronization device is merely illustrative; the data synchronization device can also be configured in other ways. Figure 2 The functional modules shown in the diagram are divided, and will not be elaborated further here.
[0110] Based on the aforementioned data synchronization device, a structural diagram of a computing device is given below as an example. (See also...) Figure 8 As shown, Figure 8 This is a schematic diagram of a computing device provided in an embodiment of this application. It should be understood that the computing device described below can implement any of the functions described in the above methods. Typically, the computing device 800 includes a processor 801 and a memory 802.
[0111] Processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 801 may be implemented using at least one hardware form selected from digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). Processor 801 may also include a main processor and a coprocessor. The main processor, also known as the central processing unit (CPU), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may integrate a graphics processing unit (GPU), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 801 may also include an artificial intelligence (AI) processor, which is used to handle computational operations related to machine learning.
[0112] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 are used to store at least one program segment for execution by the processor 801 to implement the data synchronization method provided in the method embodiments of this application.
[0113] In some embodiments, the computing device 800 may also optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 can be connected via a bus or signal lines. Each peripheral device can be connected to the peripheral device interface 803 via a bus, signal lines, or a circuit board.
[0114] In some embodiments, the aforementioned computing device can be an independent physical server, or implemented as a cluster of computing devices, that is, a server cluster composed of multiple physical servers or a distributed file system, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. Taking a computing device as a cloud server as an example, a computing device can also be called a cloud platform (short for cloud computing platform), which refers to a service based on hardware and software resources that provides computing, network, and storage capabilities. Through the network "cloud," massive amounts of data are processed and analyzed remotely before being returned to the user, featuring large scale, distributed nature, virtualization, high availability, scalability, on-demand service, and security. Cloud platforms can achieve rapid deployment and release of configurable computing resources with relatively low management costs or low interaction complexity between users and service providers.
[0115] In some embodiments, the computing device 800 may be a portable mobile terminal, such as a smartphone, tablet, laptop, or desktop computer. The computing device 800 may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0116] In this embodiment, the computing device 800 can be configured as a data synchronization server, which acts as the execution entity to implement the technical solution provided in this embodiment. Optionally, the computing device 800 deploys a lock-free change configuration module, a lock-free change filtering module, and a data playback module. The lock-free change module is used to obtain filtering parameters, such as at least one of a preset filtering range, a preset regular expression, a preset list, and a lock-free change type, as shown in step 401 above; the lock-free change filtering module is used to filter operation record data based on the filtering parameters, as shown in step 402 above; the data playback module is used to perform data playback based on the remaining data, as shown in step 403 above. Alternatively, the computing device 800 can also be configured as a database server to store and manage data in database nodes. Optionally, a database management module is deployed in the database server. Alternatively, the computing device can also be configured as a business server to handle operations on data tables in database nodes. Optionally, a business module is deployed in the business server.
[0117] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a data synchronization device to perform the data synchronization method described above. For example, the computer-readable storage medium is a non-transitory computer-readable storage medium, such as read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.
[0118] This application also provides a computer program product or computer program, which includes program code. The computer instructions are stored in a computer-readable storage medium. A processor in a computing device reads the program code from the computer-readable storage medium and executes the program code, causing the computing device to perform the above-described data synchronization method.
[0119] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the data synchronization methods in the above-described method embodiments.
[0120] In this embodiment, the apparatus, device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0121] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the data synchronization method embodiments provided above belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0122] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially or in other words, the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0126] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0127] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0128] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the sensitive words involved in this application were obtained with full authorization.
[0129] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0130] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data synchronization method, characterized in that, The method includes: Receive at least one operation record data from a first database node, the operation record data indicating an operation on a data table in the first database node, the data table including a table and a temporary table, the temporary table being used to perform lock-free change operations based on data in the table; Filter out operation record data for the temporary table from the at least one operation record data; The remaining data in the at least one operation record data is sent to the second database node. The remaining data is the data in the at least one operation record data excluding the filtered operation record data. The second database node and the first database node are in a data synchronization state.
2. The method according to claim 1, characterized in that, The step of filtering out operation record data for the temporary table from the at least one operation record data includes: Filter out operation record data targeting the temporary table from the operation record data within the preset filtering range; The operation record data within the preset filtering range includes at least one of the following: Operation record data for all data tables in the first database node; Operation record data for all data tables within the target database in the first database node; Operation record data for the target data table in the first database node.
3. The method according to claim 2, characterized in that, The preset filtering range is either a pre-set filtering range corresponding to the lockless change operation or a filtering range determined based on the configuration operation.
4. The method according to claim 1, characterized in that, The step of filtering out operation record data for the temporary table from the at least one operation record data includes: For any operation record data, determine the table name associated with the operation record data. The table name associated with any operation record data indicates the data table to which the operation in the operation record data is targeted. If the table name indicates that the operation in the operation record data is for the temporary table, then the operation record data is filtered out.
5. The method according to claim 4, characterized in that, When the table name indicates that the operation in the operation record data is for the temporary table, filtering out the operation record data includes any of the following: If the table name matches any table name in the first list, the operation record data is filtered out. The first list is used to record the table names associated with the operation record data of the temporary table. If the table name does not match any of the table names in the second list, the operation record data is filtered out. The second list is used to record the table names associated with the remaining data. If the table name matches any table name in the first list but does not match any table name in the second list, the operation record data is filtered out.
6. The method according to claim 4, characterized in that, When the table name indicates that the operation in the operation record data is for the temporary table, filtering out the operation record data includes: If the table name matches any preset regular expression, the operation record data is filtered out, where the preset regular expression indicates the element in the table name of the temporary table.
7. The method according to claim 4, characterized in that, When the table name indicates that the operation in the operation record data is for the temporary table, filtering out the operation record data includes: If the table name does not match any of the table names in the first list or the second list, the table name is matched based on at least one preset regular expression. The first list is used to record the table names associated with the operation record data of the temporary table, and the second list is used to record the table names associated with the remaining data. The preset regular expression indicates the elements in the table name of the temporary table. If the table name matches any preset regular expression, the operation record data is filtered out.
8. The method according to claim 7, characterized in that, The method further includes: If the table name matches any preset regular expression, the table name is added to the first list; If the table name does not match any of the at least one preset regular expressions, the table name is added to the second list.
9. The method according to any one of claims 6-8, characterized in that, The preset regular expression is either a pre-set regular expression corresponding to the lock-free change operation or a regular expression determined based on the configuration operation.
10. A data synchronization device, characterized in that, The device includes: A receiving module is configured to receive at least one operation record data from a first database node, the operation record data indicating an operation on a data table in the first database node, the data table including a table and a temporary table, the temporary table being used to perform lock-free change operations based on the data in the table; The filtering module is used to filter out operation record data for the temporary table from the at least one operation record data; The sending module is used to send the remaining data in the at least one operation record data to the second database node. The remaining data is the data in the at least one operation record data excluding the filtered operation record data. The second database node and the first database node are in a data synchronization state.
11. A computing device, characterized in that, The computing device includes a processor and a memory, the processor being configured to execute at least one piece of program code stored in the memory to enable the computing device to perform the method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store at least one piece of program code, which, when executed by a computing device, causes the computing device to perform the method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, When the computer program product is run on a computing device, the computing device performs the method as described in any one of claims 1 to 9.