Data comparison method and device, equipment, medium and program product
By hash modulus operation of the data alignment keys in heterogeneous databases and parallel comparisons in buckets, the problems of low data comparison efficiency and poor accuracy in the prior art are solved, and efficient and highly accurate data comparison effects are achieved.
Patent Information
- Application Number
- CN202510192531.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
In heterogeneous database scenarios, the existing data consistency verification methods are inefficient and have poor accuracy, and cannot take into account both efficient and highly accurate data comparisons.
By comparing the keys, the bucket values of each row of data are determined, the data is allocated to multiple preset buckets, and the data in each bucket is compared in parallel through multiple comparison threads.
In heterogeneous database scenarios, efficient and highly accurate data comparison is achieved, avoiding errors and omissions in comparison results caused by sorting differences, and improving data processing performance.
Smart Images

Figure CN120123339A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a data comparison method, device, equipment, medium, and program product. Background Art
[0002] In the scenario of data synchronization between a source database and a target database, it is necessary to perform consistency verification based on the data in the source and the target to determine which data in the target is consistent with the data in the source and which is inconsistent. For the inconsistent data, the target is updated to achieve data synchronization between the source and the target.
[0003] In the existing data consistency verification methods, the rows of data in the source data table and the target data table are sorted, and the rows of data are compared in the sorting order to determine the consistency of the data in the two data tables. This method sorts based on the comparison keys in each row of data. The comparison key can be understood as one or more columns that are non-null and can uniquely identify each row of data.
[0004] When comparing in the sorting order in sequence, the comparison efficiency is low. And in the scenario where the source and the target are heterogeneous databases, due to the sorting differences of the comparison keys between the source and the target, if the above method is still used for data table comparison, the accuracy of the comparison result is low. Therefore, the existing methods cannot balance high comparison accuracy and high comparison efficiency. Summary of the Invention
[0005] Embodiments of this application provide a data comparison method, device, equipment, medium, and program product, which are used to achieve the technical effect of balancing high comparison accuracy and high comparison efficiency when comparing the source and the target in the case where the source and the target are heterogeneous databases.
[0006] In a first aspect, an embodiment of the present application provides a data comparison method, which is applied to a source data table and a target data table in the case where the source database and the target database are heterogeneous databases. The method includes: according to the comparison keys of each row of data in the source data table and the target data table, by performing a hash modulo operation on the comparison keys, determining the bucket values of each row of data, where the comparison keys are used to distinguish each row of data, and the bucket values are used to indicate that any row of data is allocated to one of a preset plurality of buckets; allocating each row of data in the source data table to the corresponding bucket based on the bucket values of each row of data in the source data table, and allocating each row of data in the target data table to the corresponding bucket based on the bucket values of each row of data in the target data table; configuring a plurality of comparison threads for the preset plurality of buckets, and performing parallel data comparison on each row of data in the preset plurality of buckets through the plurality of comparison threads to obtain a data comparison result.
[0007] In a possible implementation manner, the determining the bucket value of each row of data by performing a hash modulo operation on the comparison key includes: for any row of data in each row of data, determining the bucket value in the following manner: obtaining the field value corresponding to the comparison key of the any row of data; calculating the hash value corresponding to the field value through a hash function, and performing a hash modulo operation on the hash value to obtain the bucket value of the any row of data.
[0008] In a possible implementation manner, the performing a hash modulo operation on the hash value to obtain the bucket value of any row of data includes: obtaining that the number of comparison threads to be used in the comparison thread pool is a first number, presetting a plurality of buckets with the number of buckets being the second number according to a second number greater than or equal to the first number, where the comparison thread pool is a thread pool preset with the plurality of comparison threads; performing a hash modulo operation on the hash value according to the second number to obtain the bucket value of any row of data.
[0009] In a possible implementation manner, the performing parallel data comparison on each row of data in the preset plurality of buckets through the plurality of comparison threads includes: for any one of the preset plurality of buckets, performing data comparison in the following manner: performing a hash join on the row data of the source data table allocated to the any one of the buckets and the row data of the target data table allocated to the any one of the buckets through the comparison thread corresponding to the any one of the buckets to obtain an association set representing the hash join result, and performing data comparison according to the association set.
[0010] In a possible implementation, the data comparison according to the association set includes: in the association set, for any pair of row data for which a hash join is established, respectively compare the field values of each column in the pair of row data. In the case where at least one pair of different field values is compared, determine the pair of row data as a difference row, and the difference row is used to indicate the data comparison result of the row data having a field value difference; and / or, in the association set, for any row data for which a hash join is not established, determine the row data as an extra row, and the extra row is used to indicate the data comparison result of the newly added row data.
[0011] In a possible implementation, the method further includes: respectively partitioning the row data in the source-side data table and the target-side data table to obtain a partition list, where the partition list includes a preset plurality of partitions, and the partition list is used to indicate the correspondence between the preset plurality of partitions and the row data included in each partition; the determining the respective bucket values of the row data by performing a hash modulo operation on the comparison key includes: configuring a plurality of reading threads according to the partition list, parallelly reading the comparison keys of the row data in the preset plurality of partitions through the plurality of reading threads, and determining the respective bucket values of the row data by performing a hash modulo operation on the comparison key.
[0012] In a second aspect, an embodiment of the present application provides a data comparison device, which is applied to the source-side data table and the target-side data table in the case where the source-side database and the target-side database are heterogeneous databases. The device includes: a determining module, configured to determine the respective bucket values of the row data by performing a hash modulo operation on the comparison key according to the comparison keys of the row data in the source-side data table and the target-side data table, where the comparison key is used to distinguish the row data, and the bucket value is used to indicate that any row data is assigned to one of a preset plurality of buckets; an assignment module, configured to assign the row data in the source-side data table to the corresponding buckets based on the bucket values of the row data in the source-side data table, and assign the row data in the target-side data table to the corresponding buckets based on the bucket values of the row data in the target-side data table; a comparison module, configured to configure a plurality of comparison threads for the preset plurality of buckets, and perform parallel data comparison on the row data in the preset plurality of buckets through the plurality of comparison threads to obtain a data comparison result.
[0013] In a possible implementation, the determining module is specifically configured to: for any row data among the row data, determine the bucket value in the following manner: obtain the field value corresponding to the comparison key of the row data; calculate the hash value corresponding to the field value through a hash function, and perform a hash modulo operation on the hash value to obtain the bucket value of the row data.
[0014] In a possible implementation, the determination module is specifically used to: obtain the number of comparison threads to be used in the comparison thread pool as a first number, and based on a second number greater than or equal to the first number, preset a number of buckets of the second number, wherein the comparison thread pool is a thread pool preset with the multiple comparison threads; perform a hash modulo operation on the hash value according to the second number to obtain the bucket value of any row of data.
[0015] In a possible implementation, the comparison module is specifically used to perform data comparison for any of the preset multiple buckets in the following manner: through the comparison thread corresponding to any of the buckets, a hash connection is performed on the row data of the source data table allocated to the any of the buckets and the row data of the target data table allocated to the any of the buckets, an association set representing the hash connection result is obtained, and data comparison is performed based on the association set.
[0016] In one possible implementation, the comparison module is specifically used to: in the associated set, for any pair of row data that establish a hash connection, respectively compare the field values of each column in the any pair of row data, and when at least one pair of different field values are compared, determine the any pair of row data as a difference row, and the difference row is used to indicate a data comparison result in which there is a difference in field values in the row data; and / or, in the associated set, for any row data that does not establish a hash connection, determine the any row data as multiple rows, and the multiple rows are used to indicate a data comparison result in which newly added row data exists.
[0017] In a possible implementation, the device further includes a partitioning module, which is used to: partition each row of data in the source data table and the target data table respectively to obtain a partition list, wherein the partition list includes a plurality of preset zones, and the partition list is used to indicate the correspondence between the plurality of preset zones and the row data in each zone; the determination module is specifically used to: configure a plurality of reading threads according to the partition list, read the comparison keys of the row data in the plurality of preset zones in parallel through the plurality of reading threads, and determine the bucket value of each row of data by performing a hash modulo operation on the comparison key.
[0018] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a memory and a processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect above and / or various possible implementations of the first aspect.
[0019] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the first aspect above and / or various possible implementations of the first aspect.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the above first aspect and / or various possible implementation methods of the first aspect.
[0021] The data comparison method, device, equipment, medium and program product provided by the embodiment of the present application are applied to the source data table and the target data table when the source database and the target database are heterogeneous databases. The method buckets each data according to the comparison key, and the bucket value based on the bucketing is obtained by hashing the comparison key. Due to the bucket value after hashing, each row of data can be distributed to each bucket more evenly, so that the number of row data in each bucket is less different. Based on this, on the one hand, since each row of data is divided into multiple buckets, multiple buckets are compared in parallel through multiple comparison threads, parallel data processing can be realized, the time required for data comparison can be shortened, and the comparison efficiency can be improved; on the other hand, since the number of row data in each bucket is relatively average, the workload of multiple comparison threads for parallel comparison can be relatively balanced, thereby improving the comparison speed as a whole, shortening the time required to compare all data, and improving the comparison efficiency. In addition, the method of the embodiment of the present application is to first bucket each row of data according to the comparison key, and then compare the row data allocated in each bucket in the bucket, so as to circumvent the method of sorting the comparison key and then comparing in sequence, which can avoid the defect of the difference in the sorting order of the comparison key caused by the differences in heterogeneous databases, and then avoid the problem of errors and omissions in the comparison results of the sequential comparison caused by the sorting difference, and reduce the probability of errors in the comparison. In addition, since each row of data is allocated to the corresponding bucket according to the bucket value, a pair of row data that should be compared can be allocated in the same bucket, so that the comparison in each bucket can be more accurate, not only will there be no errors and omissions, but also the comparison can be completed quickly. Therefore, the method provided by the embodiment of the present application can achieve the technical effect of taking into account both high comparison accuracy and high comparison efficiency in the scenario where the source end and the target end are heterogeneous databases. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0023] Figure 1 A schematic diagram of a data synchronization scenario provided in an embodiment of the present application;
[0024] Figure 2 It is a schematic diagram of the data comparison process in the existing method;
[0025] Figure 3 It is a schematic flowchart of the data comparison method provided by the embodiment of the present application;
[0026] Figure 4 It is a schematic diagram of bucketing and data comparison provided by the embodiment of the present application;
[0027] Figure 5 It is a flowchart of a data comparison method provided by the embodiment of the present application;
[0028] Figure 6 It is a schematic structural diagram of the data comparison device provided by the embodiment of the present application;
[0029] Figure 7 It is a schematic structural diagram of an electronic device provided by the embodiment of the present application.
[0030] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0031] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0032] In the technical solution of the embodiment of the present application, the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0033] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0034] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0035] In the embodiments of the present application, if words such as "first" and "second" are used, they are for distinguishing identical or similar items with basically the same functions and roles. For example, the first electronic device and the second electronic device are only for distinguishing different electronic devices, and do not limit their sequence. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit being different.
[0036] In the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0037] In the scenario of data synchronization between databases, data synchronization tools such as synchronization software can be used for real-time data synchronization. Data synchronization mainly includes the following three stages: The first stage is to perform the initialization loading of the stock data to obtain the basic point of data synchronization. The second stage is to perform incremental data synchronization based on the synchronization basic point established by the initialization data loading. The third stage is to periodically compare and verify the source data and target data of the data synchronization to determine whether there is any lost data during the data synchronization process. Among them, the second stage and the third stage can be in a state of long-term parallel execution.
[0038] Figure 1 For the schematic diagram of the data synchronization scenario provided by the embodiments of the present application, during the incremental data synchronization in the second stage, incremental data can be obtained by analyzing the database log, so as to achieve real-time data synchronization. As Figure 1As shown, by parsing data such as the online log or archived log of the source database, the add, delete, and modify changes of the data can be obtained. Then, these changes are converted into a specific message format within the synchronization software in units of transactions and sent to the synchronization software of the target database through the private transmission protocol of the synchronization software. After that, the synchronization software at the target end restores the obtained transaction log into a Structured Query Language (SQL) statement supported by the target database and executes it on the target database, thereby enabling real-time data synchronization and maintaining data consistency between the source end and the target end.
[0039] As Figure 1 shown, during the operation of the second stage, in real-time data synchronization, to ensure the consistency of the full amount of data after synchronization, it is necessary to perform data verification on the full amount of data at the source end and the target end to verify whether the full amount of data at the source end and the target end is consistent. When performing data verification, the source end data and the target end data can be obtained first, and then the obtained source end data and target end data are subjected to data verification.
[0040] In the data verification stage, the existing data comparison methods usually sort each row of data first, and then compare the row data row by row in order to obtain the comparison result. Specifically, the existing data comparison methods will first perform ascending sorting on the source end data table and the target end data table through the comparison key respectively to obtain the ResultSet data sets. There is a cursor in the ResultSet data sets of the source end data table and the target end data table respectively, and each row of data is compared row by row through the cursor.
[0041] Figure 2 is a schematic diagram of the data comparison process in the existing method. As Figure 2 shown, after sorting with the primary key of each row of data as the comparison key, the ResultSet data set of the source end data table is data table 1 and the ResultSet data set of the target end data table is data table 2. The cursors of data table 1 and data table 2 both start from the first row and move down row by row as the data comparison progresses.
[0042] After the comparison starts, if the primary keys at the source end and the target end are equal and the other values in the row data are also equal, it is marked as I, and the cursors at the source end and the target end both move down one row. If the primary key at the source end is greater than the primary key at the target end, it is marked as N, and the cursor at the target end continues to move down. If the primary key at the source end is less than the primary key at the target end, it is marked as D, and the cursor at the source end continues to move down. If the primary keys at the source end and the target end are equal but the other values in the row data are not equal, it is marked as C, and the cursors at the source end and the target end continue to move down. And so on, until all the row data in data table 1 and data table 2 are compared, the comparison ends, and the comparison result is obtained.
[0043] As can be seen from the above, in the existing methods, it is necessary to sort according to the comparison keys of the source - end data table and the target - end data table, and perform data comparison row by row according to the sorted order. However, when there are differences in the sorting implementation or sorting rules between the source end and the target end, even if the comparison keys of the source - end data table and the target - end data table are the same, different sorting results may be obtained, and thus comparison errors may occur due to differences in the sorting order.
[0044] For example, when the source end and the target end are heterogeneous databases, due to differences in the implementation of different data character sets between heterogeneous databases, the order of the same data after sorting may be different. It can be understood that even for the same set of comparison keys, due to differences in the implementation of the character set at the source end and the character set at the target end, the sorting obtained at the source end is completely different from the sorting obtained at the target end, which will directly render the existing data comparison method unavailable, that is, it is impossible to compare the real differential data.
[0045] After analyzing the existing data comparison methods, it can also be known that the existing methods perform comparison sequentially based on the sorting order, which is a serial comparison method and has the defect of low comparison efficiency. In addition, both the source - end data table and the target - end data table rely on comparison keys for sorting, which will increase the additional resource overhead of the database. If there is no index on the comparison keys of the data table, it will also cause relatively serious query performance problems and reduce the speed and efficiency of data comparison.
[0046] In view of this, the embodiments of the present application provide a data comparison method. For the source - end data table and the target - end data table in the case where the source - end database and the target - end database are heterogeneous databases, by evenly bucketing the rows of data in the source - end data table and the target - end data table, a large number of row data are allocated to their respective corresponding buckets, and for each bucket, multiple comparison threads are used to perform data comparison in parallel within each bucket. In this way, data comparison can be performed in parallel to improve the comparison efficiency; it can also bypass the method of sorting according to comparison keys and comparing row by row, reduce the probability of comparison errors, and improve the comparison accuracy.
[0047] The method provided by the embodiments of the present application can be based on a single - machine comparison service system, bypass the mechanism of database sorting for comparison, make full use of the computing resources of the server, and use the divide - and - conquer algorithm to quickly complete the comparison of all data in the scenario of massive data comparison, improving data processing performance.
[0048] The technical solution of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0049] Figure 3It is a schematic flowchart of the data comparison method provided by an embodiment of the present application. The execution subject of this method can be an electronic device with corresponding data storage capabilities and computing capabilities, such as a computer, a server, or a server cluster, etc. This method is applied to the source-side data table and the target-side data table in the case where the source-side database and the target-side database are heterogeneous databases, such as Figure 3 shown. This method includes:
[0050] S301, according to the comparison keys of each row of data in the source-side data table and the target-side data table, perform a hash modulo operation on the comparison keys to determine the respective bucket values of each row of data. The comparison key is used to distinguish each row of data, and the bucket value is used to indicate that any row of data is allocated to one of a preset number of buckets.
[0051] Exemplarily, the source-side database and the target-side database being heterogeneous databases can be understood as that the source-side database and the target-side database have differences in at least one aspect such as data model, query language, data type and structure, storage and access mechanism, etc. If the source-side database and the target-side database adopt different database management systems, it may result in the two being heterogeneous databases.
[0052] The source-side data table can be a data table obtained from the source-side database, and the target-side data table can be a data table obtained from the target-side database. In data verification and data comparison, the source-side data table and the target-side data table are usually two data tables that are corresponding or have a correlation with each other.
[0053] Row data can be understood as a complete row of data in a data table. For example, if the source-side data table is a data table with p rows and q columns, where both p and q are positive integers, then there are a total of p row data in this data table, and each row of data includes the field values of q fields. Any field is a certain column in the q columns, and the target-side data table is similar.
[0054] The comparison key can be data in the data table that can distinguish each row of data. For example, it can be one column or multiple columns that are non-empty and can uniquely identify each row of data in each row of data. The comparison key can be determined according to the data characteristics of the data table and may not be the only one. The comparison key can be the primary key or the field value of other fields, specifically depending on the requirements of data comparison.
[0055] For example, the source data table and the target data table are two data tables of user information. Each row represents the information of a user, such as User A, User B, User C, and User D, etc.; each column represents a type of user information, such as name, gender, native place, ID number, and phone number, etc. Since the ID number is the unique identifier of each user, it can uniquely identify each row of data, and the ID numbers are all non-empty. Therefore, the ID number can be used as the comparison key. Or, different natural number codes are assigned to each row of users in the data table. For example, the codes of User A, User B, User C, and User D, etc. are 1, 2, 3, and 4, etc. Then this code can also be used as the comparison key.
[0056] The bucketing value can be indication information used to indicate that any row of data is assigned to one of a preset multiple buckets. Herein, a bucket can be understood as a data space, and multiple data within a numerical range can be stored in this data space. A bucket can also be referred to as a slot. By performing a hash modulo operation on the comparison key of any row of data, the bucketing value of that row of data can be obtained.
[0057] The hash modulo operation can be used to map data to a fixed-size range. The hash modulo operation can be to convert the comparison key into an integer value through a hash function, and then perform a modulo operation on this integer value to ensure that the result falls within a predefined range.
[0058] For example, n buckets can be preset, where n is a positive integer. The n buckets can be numbered as Bucket 1, Bucket 2, Bucket 3 up to Bucket n. By performing a hash modulo operation based on the comparison key of a row of data in the data table, a value between 1 and n can be obtained, and this value is the bucketing value corresponding to that row of data. In this way, by performing the same hash modulo operation on each row of data in the source data table and on each row of data in the target data table respectively, the bucketing values of each row of data in the source data table and the target data table can be obtained.
[0059] S302, allocate each row of data in the source data table to the corresponding bucket based on the bucketing values of each row of data in the source data table, and allocate each row of data in the target data table to the corresponding bucket based on the bucketing values of each row of data in the target data table.
[0060] Exemplarily, after determining the bucketing values of each row of data, each row of data in the source data table can be respectively allocated to the corresponding bucket, or each row of data in the target data table can be respectively allocated to the corresponding bucket.
[0061] Figure 4 This is a schematic diagram of bucketing and data comparison provided by an embodiment of the present application, such as Figure 4As shown, read the data of each row in the source - end data table. For example, the comparison key of row data 1 is 1, and the bucket value obtained by performing a hash modulo operation on it is 1. The comparison key of row data 3 is 3, and the bucket value obtained by performing a hash modulo operation on it is 3. The comparison key of row data 5 is 5, and the bucket value obtained by performing a hash modulo operation on it is 5, and so on. After obtaining the bucket values of all row data, according to the bucket values and the bucket values of each bucket, allocate the row data to the buckets that match the bucket values. Here, the bucket value can be understood as information such as bucket number, bucket sequence number, or bucket ID used to indicate the bucket, and it can be a numerical value that matches the bucket value. The same operation is performed on the target end, and the row data in the target - end data table can be allocated to the corresponding buckets.
[0062] S303, configure multiple comparison threads for a preset number of buckets, and perform parallel data comparison on the row data in the preset number of buckets through the multiple comparison threads to obtain a data comparison result.
[0063] Exemplarily, a comparison thread can be understood as a thread used to perform data comparison on the row data in a bucket. The number of multiple comparison threads can be set according to the computing resources of the electronic device.
[0064] As Figure 4 shown, m comparison threads can be set, where m is a positive integer. One comparison thread can be responsible for the data comparison work of the row data in one bucket, and m comparison threads can perform parallel data comparison on the row data in m buckets. If a comparison thread finishes the data comparison of the row data in a bucket, then this comparison thread can be allocated to process the bucket where the comparison has not started yet, which can speed up the completion of the data comparison of all row data in the buckets. After the data comparison of all row data in the buckets is completed, the data comparison result can be obtained.
[0065] Exemplarily, when the computing resources of the electronic device have high computing power, the number of multiple comparison threads can be set to be greater than or equal to the number of multiple buckets, so that parallel data comparison of multiple buckets can be completed in a shorter time. Or, when the computing resources of the electronic device do not have high computing power, the number of multiple comparison threads can be set to be less than the number of multiple buckets, so that the computing resources can be reasonably utilized to carry out parallel data comparison without bringing computing pressure to the electronic device.
[0066] The data comparison method, device, equipment, medium and program product provided by the embodiment of the present application are applied to the source data table and the target data table when the source database and the target database are heterogeneous databases. The method buckets each data according to the comparison key, and the bucket value based on the bucketing is obtained by hashing the comparison key. Due to the bucket value after hashing, each row of data can be distributed to each bucket more evenly, so that the number of row data in each bucket is less different. Based on this, on the one hand, since each row of data is divided into multiple buckets, multiple buckets are compared in parallel through multiple comparison threads, parallel data processing can be realized, the time required for data comparison can be shortened, and the comparison efficiency can be improved; on the other hand, since the number of row data in each bucket is relatively average, the workload of multiple comparison threads for parallel comparison can be relatively balanced, thereby improving the comparison speed as a whole, shortening the time required to compare all data, and improving the comparison efficiency. In addition, the method of the embodiment of the present application is to first bucket each row of data according to the comparison key, and then compare the row data allocated in each bucket in the bucket, so as to circumvent the method of sorting the comparison key and then comparing in sequence, which can avoid the defect of the difference in the sorting order of the comparison key caused by the differences in heterogeneous databases, and then avoid the problem of errors and omissions in the comparison results of the sequential comparison caused by the sorting difference, and reduce the probability of errors in the comparison. In addition, since each row of data is allocated to the corresponding bucket according to the bucket value, a pair of row data that should be compared can be allocated in the same bucket, so that the comparison in each bucket can be more accurate, not only will there be no errors and omissions, but also the comparison can be completed quickly. Therefore, the method provided by the embodiment of the present application can achieve the technical effect of taking into account both high comparison accuracy and high comparison efficiency in the scenario where the source end and the target end are heterogeneous databases.
[0067] Exemplarily, the hash modulo operation can be implemented based on a hash function, and can be used to distribute data more evenly to a preset number of buckets (or slots). The hash modulo operation has the effect of uniform distribution. Through the hash modulo operation, the input data (such as the comparison key) can be mapped to integers in a limited range. These integers can be used to represent the index positions of multiple buckets, so that each row of data can be distributed as evenly as possible at these index positions, avoiding the data concentration in a local position and causing the workload of multiple comparison threads to be unbalanced. When the row data is evenly distributed, it helps to improve the overall performance and reliability of the comparison process. The hash modulo operation also has the effect of fast search. In the hash table, the hash modulo operation is used to calculate the storage location of the key, so that the search operation can achieve a constant time complexity in the average case, which is more important for application scenarios that require fast data comparison. In addition, since the hash modulo operation is simple and efficient, determining the bucket value through the hash modulo operation can also simplify the implementation process.
[0068] In a possible implementation, when determining the respective bucketing values of each row of data by performing a hash modulo operation on the comparison keys, the bucketing value of any row of data among the rows of data can be determined in the following manner: obtain the field value corresponding to the comparison key of any row of data; calculate the hash value corresponding to the field value through a hash function, and perform a hash modulo operation on the hash value to obtain the bucketing value of any row of data.
[0069] Exemplarily, the field value corresponding to the comparison key can be understood as the specific data of the comparison key. For example, when the identity card number is used as the comparison key, the identity card number "610xxxxxxx" of user A is the field value corresponding to the comparison key of the row data of user A. Another example is that when the code of each row of data is used as the comparison key, the code "1" of user A is the field value corresponding to the comparison key of the row data of user A.
[0070] When calculating the hash value corresponding to the field value through a hash function, any hash function can be used. For example, it can be the Secure Hash Algorithm 1 (SHA-1), the Secure Hash Algorithm 2 (SHA-2) and its variants, or the Message-Digest Algorithm 5 (MD5), etc.
[0071] After calculating the hash value corresponding to the field value through a hash function, a hash modulo operation can be performed on the hash value to obtain the bucketing value of the row data. Specifically, the hash modulo operation can be performed through the formula: bucketId = hash(key) % numBuckets, where bucketId represents the calculated bucketing value, hash(key) represents the hash value calculated for the field value of the comparison key, % represents taking the remainder after the division operation, and numBuckets represents the modulo value, and this modulo value can be the number of preset multiple buckets.
[0072] Based on this, for any row of data, calculating the hash value for the field value of the comparison key of the row data and performing a hash modulo operation based on the hash value can conveniently obtain the bucketing value, simplify the implementation process, and make the obtained bucketing value idempotent, thereby improving the reliability of allocating the row data into the corresponding bucket.
[0073] In a possible implementation, when obtaining the bucket value of any row of data through a hashing modulo operation on the hash value, specifically, it can be: obtaining the number of comparison threads to be used in the comparison thread pool as the first number, and presetting a plurality of buckets with the second number greater than or equal to the first number, where the comparison thread pool is a thread pool preset with a plurality of comparison threads; performing a hashing modulo operation on the hash value according to the second number to obtain the bucket value of any row of data.
[0074] Exemplarily, the comparison thread pool can be a thread pool preset with a plurality of comparison threads. The number of comparison threads to be used in the thread pool can be the preset number of comparison threads, that is, the first number. According to the first number, the second number greater than or equal to the first number can be determined as the number of a plurality of buckets.
[0075] Exemplarily, the number of the preset plurality of buckets is greater than the number of comparison threads, and the number of the plurality of buckets can be a multiple of the number of comparison threads. For example, if the obtained first number is k, the N times of k can be determined as the second number, and then N·k buckets can be preset. Since the bucket value is obtained based on the hashing modulo operation, the difference in the number of rows of data allocated to each bucket is small, and the amount of data in each bucket is relatively uniform. Therefore, setting the second number as N times the first number can more evenly use k threads to process the rows of data in N·k buckets in parallel, and more reasonably utilize the computing resources.
[0076] Exemplarily, if the data volume of the data table is large, the number of a plurality of buckets can be appropriately increased, that is, each row of data is allocated to more buckets. In this way, the amount of data in each bucket will be more balanced, and the overall efficiency during parallel processing can be improved. In addition, increasing the number of buckets can also prevent the risk of memory overflow when the comparison threads are running. If the computing unit and memory and other resources of the electronic device are sufficient, the number of comparison threads can be increased to improve the efficiency of data comparison.
[0077] In this embodiment, the number of buckets is the second number, which is greater than or equal to the number of comparison threads in the thread pool. In this way, the comparison threads already configured in the comparison thread pool can be utilized more efficiently, and the utilization rate of computing resources can be improved. In addition, when the number of buckets is large, each row of data can be bucketed more evenly, reducing the situation of data skew, which helps to improve the overall efficiency during parallel data comparison.
[0078] In a possible implementation, when parallel data comparison is performed on the row data in each of a preset plurality of buckets through a plurality of comparison threads, data comparison for any one of the preset plurality of buckets can be performed in the following manner: Through the comparison thread corresponding to any one bucket, perform a hash join on the row data of the source-side data table allocated to any one bucket and the row data of the target-side data table allocated to any one bucket, obtain an association set representing the hash join result, and perform data comparison based on the association set.
[0079] Exemplarily, hash join is a join method used to implement join operations in relational databases and can be used for equijoin of large-scale data sets, that is, a join based on equal conditions. Hash join can be used in a database management system to optimize join operations and is more efficient than nested loop join and sort-merge join.
[0080] When any comparison thread processes the row data from the source side and the row data from the target side in a bucket, a hash join can be performed through the hash value of the comparison key. Specifically, the hash values of the comparison keys in the bucket can be compared for equality, and a source-side row data and a target-side row data with the same hash value can be joined, then a hash join result can be obtained. After comparing the hash values of all the row data in the bucket and establishing hash joins, all the obtained hash join results form a set, and thus an association set representing the hash join result is obtained. Data comparison can be further performed based on the association set.
[0081] In this embodiment, since hash join is a relatively fast join method in a data table, in this embodiment, an association set is obtained for the row data in any one bucket through hash join, so that the association relationship of each row data can be established relatively quickly, providing conditions for the specific comparison of row data. Based on this, the speed of data comparison can be increased and the comparison efficiency can be improved.
[0082] In a possible implementation, when performing data comparison based on the association set, specifically, it can be: In the association set, for any pair of row data for which a hash join is established, respectively compare the field values of each column in any pair of row data. In the case where at least one pair of different field values is compared, determine any pair of row data as a difference row, and the difference row is used to indicate the data comparison result of the existence of field value differences in the row data; and / or, in the association set, for any one row data for which a hash join is not established, determine any one row data as an extra row, and the extra row is used to indicate the data comparison result of the existence of newly added row data.
[0083] Exemplarily, the field values of each column in the row data can be understood as the specific data in each column of this row. Any pair of row data for which a hash join is established can be understood as the row data from the source side and the row data from the target side for which a hash join is established.
[0084] For any pair of row data for establishing a hash join, when comparing the field values of each column in any pair of row data, it can be to separately compare the field values of the corresponding columns.
[0085] For example, for the row data of user A from the source data table and the row data of user A from the target data table, since their ID numbers are the same, the hash values calculated based on the field values of the ID number as the comparison key are also the same. When performing a hash join, these two row data will be joined to obtain a pair of row data for establishing a hash join. For this pair of row data, the field values in the name column, gender column, native place column, ID number column, and phone number column, etc. can be separately compared. This process can also be understood as a process of performing a full-scale comparison of the field values of all columns.
[0086] In the case where at least one pair of different field values is compared, this any pair of row data can be determined as a difference row. A difference row can be understood as a data comparison result used to indicate that there are differences in the field values of the row data. For example, in the comparison, the field value of the phone number column in the row data of user A from the source data table being compared is "1234", while the field value of the phone number column in the row data of user A from the target data table is "0023", and the two field values are different. Then the row data of user A is a difference row, and this pair of row data can be recorded in the data comparison result. Subsequently, the field value of the phone number column in this row data in the target data table can be modified to achieve data synchronization between the target end and the source end.
[0087] In the association set, any row data that has not established a hash join can also be recorded. In the association set, for any row data that has not established a hash join, this any row data can be determined as a multiple row, and the multiple row is a data comparison result used to indicate that there are newly added row data.
[0088] For example, the row data of user G from the source data table is a newly added row data after the last data synchronization, and this row data does not exist in the target data table. Therefore, in this bucket, there is no row data that can establish a hash join with it, and this row data is a single row data and cannot be joined into a pair of row data. All row data that has not established a hash join can also be recorded through the association set. Recording each multiple row in the data comparison result, subsequently, the row data of each multiple row newly added in the target data table can be processed to achieve data synchronization between the target end and the source end.
[0089] In this embodiment, according to whether a hash join is established in the association set, the difference rows and multiple rows can be determined simply and quickly. The difference rows and multiple rows can indicate two types of comparison results, and thus a comparison result with relatively high accuracy can be obtained quickly.
[0090] In a possible implementation, the method further includes: partitioning each row of data in the source data table and the target data table respectively to obtain a partitioning list, where the partitioning list includes a plurality of preset partitions, and the partitioning list is used to indicate the correspondence between the plurality of preset partitions and the rows of data included in each partition; determining the respective bucketing values of each row of data by performing a hash modulo operation on the comparison keys, including: configuring a plurality of reading threads according to the partitioning list, parallelly reading the comparison keys of the rows of data in the plurality of preset partitions through the plurality of reading threads, and determining the respective bucketing values of each row of data by performing a hash modulo operation on the comparison keys.
[0091] Exemplarily, for the source data table and the target data table, partitioning each row of data in the data table can be understood as grouping the data, and the partitioning list can be understood as a total list including all the groups and the information within the groups. The partitioning list can preset a plurality of partitions in advance, and each partition can respectively correspond to a numerical value or a numerical range. According to the field values of one or several columns in the row data, the row data can be assigned to the partition corresponding to the data range to which its field value belongs.
[0092] The same partitioning method can be adopted when partitioning the source data table and the target data table. Taking the source data table as an example, assuming that the codes of each row of data in the source data table are 1 - 3000, several partitions can be preset in the partitioning list to divide each row of data in the source data table into these preset partitions. For example, 3 partitions can be set. The numerical range corresponding to partition 1 is [1, 1000), the numerical range corresponding to partition 2 is [1000, 2000), and the numerical range corresponding to partition 3 is [2000, 3000]. Then, each row of data can be assigned to the partition corresponding to the numerical range of its code according to the code. After each row of data is assigned to each partition in the partitioning list, the partitioning list can indicate the correspondence between the plurality of preset partitions and the rows of data included in each partition.
[0093] The purpose of partitioning the data table is to disassemble the data table into multiple sub - tables, so as to use the thread pool technology to parallelly read the comparison keys, determine the bucketing values, and perform bucketing on the multiple sub - tables, and thus can quickly complete the bucketing of the data table containing a large number of rows of data.
[0094] Exemplarily, when partitioning each row of data in the data table, the partitioning key used for partitioning can be determined according to the data characteristics of the data table, and then partitioning is performed according to the partitioning key. The partitioning key can be understood as the direct object based on which partitioning is carried out during partitioning. The partitioning key can be any one or more columns in the row data that can group the row data, and when partitioning the data table, partitioning can be performed according to the partitioning strategy applicable to the data table.
[0095] Optionally, the partitioning strategy can be a range-based partitioning strategy. For example, partitioning by the range of data (such as time, numbers), given the range of the partitioning key and the number of partitions, the step size can be calculated. Finally, a partition list like [1, 1000), [1000, 2000),... can be obtained. To prevent data skew, the number of partitions can be set to a relatively large value, that is, the more partitions there are, the smaller the difference in the number of in-row data in each partition.
[0096] Optionally, the partitioning strategy can also be based on the original partitioning implementation of the data table. For example, if the data table of some databases is already a partitioned table, the partition list can be directly obtained according to the original partitioning of the data table.
[0097] Optionally, the partitioning strategy can also be a list-based strategy. For example, for some data tables with business characteristics, the partitioning key can be a field within a specific set of fields. For instance, if the specific set of fields is the set of all region numbers, then one region number corresponds to one partition. By obtaining the value of the field in the row data that is the partitioning key as the region number, any row data can be assigned to the corresponding partition, and finally the partition list for each partition can be obtained.
[0098] Another example is that the partitioning key can be a column representing the encoding or sorting of the row data, or a column representing a certain attribute of the row data, etc. Taking the data table of user information as an example, it can be partitioned according to the gender of the users. Then the gender column can be determined as the partitioning key, and two partitions are preset. Partition 1 represents the partition corresponding to males, and Partition 2 represents the partition corresponding to females. By reading the specific value of the gender column in each row data of the user information data table, it can be assigned to Partition 1 or Partition 2.
[0099] Alternatively, it can also be based on a part of the specific value in a certain column of the row data as the field value of the partitioning key. For example, the first 6 digits of the user's ID card number represent the region. Then the values of the ID card number column of all row data can be obtained. Based on these values, multiple partitions can be preset, and these partitions cover all the regions to which the ID card numbers in the data table belong. Further, according to the first 6 digits of the ID card number of each row data, each row data is assigned to the corresponding partition, and then the partition list can be obtained.
[0100] After obtaining the partition list, when determining the respective bucket values of each row data through the hash modulo operation on the comparison key, specifically, it can be: configure multiple reading threads according to the partition list, parallelly read the comparison keys of the row data in the preset multiple partitions through the multiple reading threads, and determine the respective bucket values of each row data through the hash modulo operation on the comparison key.
[0101] Exemplarily, a read thread pool including multiple read threads can be preset. Each read thread can be responsible for reading the comparison keys of each row of data in a partition in the partition list, and performing a hash modulo operation on the comparison keys to determine the respective bucket values of each row of data. If a read thread has completed the data processing of each row of data in a partition, the read thread can be assigned to process other partitions that have not started reading, which can speed up the processing of all rows of data in all partitions. After all rows of data in all partitions have been read, bucketing can be performed according to the bucket values of the row data.
[0102] In this embodiment, before determining the bucket value, the rows of data can be partitioned first, and then multiple read threads can be used to read the rows of data in multiple partitions in parallel, which can improve the speed of determining the bucket values of each row of data, enable a large amount of row data to quickly enter the bucketing link, and can speed up the obtaining of the data comparison result and improve the efficiency of data comparison.
[0103] To overcome the defects and deficiencies of the prior art, the embodiments of the present application provide a data comparison method, which realizes the comparison of a large amount of data on a single machine based on the divide-and-conquer strategy, and significantly improves the performance of detailed data verification of a large amount of data in heterogeneous database data tables under a single machine service. Figure 5 It is a flowchart of a data comparison method provided by an embodiment of the present application. Next, in combination with Figure 5 A specific implementation manner is used to further describe the data comparison method provided by the embodiments of the present application.
[0104] As Figure 5 shown, the source end and the target end can each include a read module for reading data. Through the read module, each row of data in the source data table and the target data table can be read. The read row data can be stored in the buckets at the source end and the target end respectively. The comparison module can be a module for comparing the row data in the buckets. The comparison module can read each row of data from any bucket at the source end and the corresponding bucket at the target end for data comparison.
[0105] The read module can include a sharder, a shard queue, and a read thread pool. The comparison module can include a list of buckets to be processed and a comparison thread pool. Specifically, when the read module starts to work, the sharder can split the data table into multiple partitions according to the partitioning strategy, that is, the rows of data are partitioned. When partitioning, the values of the partition list can be stored in the shard queue.
[0106] After starting the read thread pool, the working thread (worker) of each read thread in the read thread pool can sequentially obtain the row data of a partition from the shard queue for reading to determine the bucket value of each row of data.
[0107] For each line of data read, its corresponding bucket value can be calculated according to the comparison key. The bucket value can be obtained by calculating the hash modulo of the comparison key, which can ensure the idempotency of calculating the bucket value for the same comparison key and reduce data skew. The number of buckets can be set to a larger value when the amount of data in the data table is larger to balance the line data in each bucket.
[0108] When all the line data at the source end and the target end are read and stored in the corresponding buckets, the comparison thread pool of the comparison module can be started. Each worker of the comparison thread can load a bucket at the source end and a bucket at the target end with the same serial number as the bucket at the source end, and load the line data in the two buckets into the memory. By directly performing a hash join according to the comparison key, an associated set can be obtained. The list of buckets to be processed can be used to indicate the serial numbers of the buckets to be processed. After obtaining the associated set, the field values of each pair of line data that have performed the hash join can be compared line by line, and then the different lines can be obtained. The line data of the different lines can be output to the comparison result library. For multiple extra lines, their line data can also be output to the comparison result library.
[0109] For example, there are three lines of data in bucket x at the source end, which are: {[1, a1], [8, a8], [9, a9]}; there are three lines of data in bucket x at the target end, which are: {[1, a1], [8, a7], [10, a10]}. The first field value in each line of data is the comparison key. For example, the first field value 1 in [1, a1] represents the comparison key, and a1 represents the field values of other columns in this line of data. After performing a hash join on the line data in the two buckets x according to the comparison key, the associated set obtained is: [ {[1, a1], [1, a1]}; {[8, a8], [8, a7]}; {[9, a9], null}; {null, [10, a10]} ], where null represents null.
[0110] From the above associated set, it can be determined that the different lines in bucket x are {[8, a8], [8, a7]}, the extra lines at the source end are {[9, a9], null}, and the extra lines at the target end are {null, [10, a10]}, thus obtaining the data comparison result of bucket x. After a worker of a certain comparison thread finishes comparing a certain bucket, it can start comparing the line data in the next uncompared bucket number in the list of buckets to be processed, and so on. After all the line data in all buckets have been compared, the data comparison results of the source data table and the target data table can be obtained.
[0111] Exemplarily, the number of partitions can be greater than the number of reading threads in the reading thread pool. The number of partitions can be N times the number of reading threads, so that the hardware computing resources of the electronic device can be reasonably utilized and the computing overhead can be saved. If there is data skew in the data rows of the data table, that is, the data rows may be overly concentrated in a few partitions, the partitioning strategy can be changed or the number of partitions can be appropriately increased to make the value ranges corresponding to each partition more fine-grained, which can overcome the problem of data skew to a certain extent and make the data volume in each partition more uniform. When the number of data rows in the data table is large, the number of partitions can also be appropriately increased, so as to prevent the risk of memory overflow for the reading threads. When the resources such as the processing unit and memory of the electronic device are sufficient, the number of reading threads in the reading thread pool can be increased, which can increase the degree of parallelism and improve the data processing efficiency.
[0112] In the embodiments of the present application, the reading threads do not need to pay attention to the sorting of the data rows in the data table. Through various partitioning strategies such as the data range, partition table, value list, etc., the speed of the data rows entering the comparison link can be accelerated, and the problem of decreased comparison accuracy caused by the sorting difference of character sets between heterogeneous databases can be avoided. By reading the data and bucketing the data rows, the data table can be split for comparison. The comparison threads can directly connect to the datasets to be compared in memory, effectively utilizing the computing resources to improve the data comparison efficiency. The reading threads and comparison threads in the embodiments of the present application both use the thread pool technology, which can improve the parallel processing ability of reading and comparing data, and thus can quickly complete the data comparison.
[0113] Figure 6 It is a schematic structural diagram of the data comparison device provided by the embodiments of the present application, as Figure 6 shown. The embodiments of the present application provide a data comparison device, which is applied to the source data table and the target data table in the case where the source database and the target database are heterogeneous databases. The device includes:
[0114] A determination module 601, configured to determine the bucketing value of each row of data according to the comparison keys of the data rows in the source data table and the target data table by performing a hash modulo operation on the comparison keys. The comparison keys are used to distinguish each row of data, and the bucketing value is used to indicate that any row of data is assigned to one of a preset plurality of buckets;
[0115] An allocation module 602, configured to allocate the data rows in the source data table to the corresponding buckets based on the bucketing values of the data rows in the source data table, and allocate the data rows in the target data table to the corresponding buckets based on the bucketing values of the data rows in the target data table;
[0116] The comparison module 603 is used to configure multiple comparison threads for the preset multiple buckets, and perform parallel data comparison on each row of data in the preset multiple buckets through the multiple comparison threads to obtain data comparison results.
[0117] In one possible implementation, the determination module 601 is specifically used to determine the bucket value for any row of data in each row of data in the following manner: obtain the field value corresponding to the comparison key of any row of data; calculate the hash value corresponding to the field value through a hash function, and perform a hash modulus operation on the hash value to obtain the bucket value of any row of data.
[0118] In a possible implementation, the determination module 601 is specifically used to: obtain the number of comparison threads to be used in the comparison thread pool as a first number, and based on a second number greater than or equal to the first number, preset the number of buckets to be multiple buckets of the second number, wherein the comparison thread pool is a thread pool preset with multiple comparison threads; perform a hash modulo operation on the hash value according to the second number to obtain the bucket value of any row of data.
[0119] In a possible implementation, the comparison module 603 is specifically used to perform data comparison for any of the preset multiple buckets in the following manner: through the comparison thread corresponding to any bucket, a hash connection is performed on the row data of the source data table allocated to any bucket and the row data of the target data table allocated to any bucket, an association set representing the hash connection result is obtained, and data comparison is performed based on the association set.
[0120] In one possible implementation, the comparison module 603 is specifically used for: in an associated set, for any pair of row data that establish a hash connection, respectively comparing the field values of each column in any pair of row data, and when at least one pair of different field values are compared, determining any pair of row data as a difference row, and the difference row is used to indicate a data comparison result in which there is a difference in field values in the row data; and / or, in an associated set, for any row data that does not establish a hash connection, determining any row data as an additional row, and the additional row is used to indicate a data comparison result in which newly added row data exists.
[0121] In a possible implementation, the device further includes a partitioning module, which is used to: partition each row of data in the source data table and the target data table respectively to obtain a partition list, wherein the partition list includes a plurality of preset zones, and the partition list is used to indicate the correspondence between the plurality of preset zones and the row data in each zone; the determination module 601 is specifically used to: configure a plurality of reading threads according to the partition list, read the comparison keys of the row data in the plurality of preset zones in parallel through the plurality of reading threads, and determine the bucket value of each row of data by performing a hash modulo operation on the comparison keys.
[0122] The data comparison device provided in the embodiments of the present application can be used to implement the technical solutions of the data comparison methods in any of the above embodiments of the present application. Its implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0123] Figure 7 FIG. is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 7 shown, the electronic device in this embodiment may include: at least one processor 701; and a memory 702 communicatively connected to the at least one processor; wherein, the memory 702 stores instructions executable by the at least one processor 701, and the instructions are executed by the at least one processor 701 to enable the electronic device to execute the method in any of the above embodiments.
[0124] Optionally, the memory 702 can be either independent or integrated with the processor 701.
[0125] The implementation principle and technical effects of the electronic device provided in this embodiment can be referred to the foregoing embodiments, and will not be elaborated here.
[0126] The embodiments of the present application further provide a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the method described in any of the foregoing embodiments is implemented.
[0127] The embodiments of the present application further provide a computer program product, including a computer program, which when executed by a processor implements the method described in any of the foregoing embodiments.
[0128] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0129] The integrated modules implemented in the form of software function modules as described above can be stored in a computer-readable storage medium. The above software function modules are stored in a storage medium, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the embodiments of the present application.
[0130] It should be understood that the above-mentioned processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The memory may include random access memory (RAM), and may also include non-volatile memory (NVM), such as at least one disk memory, and can also be a USB flash drive, a portable hard drive, a read-only memory, a magnetic disk or an optical disc, etc.
[0131] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read only memory (PROM), read only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0132] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application specific integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0133] It should be noted that in this document, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising that element.
[0134] The serial numbers of the embodiments of the present application above are for description only and do not represent the superiority or inferiority of the embodiments.
[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.
[0136] The above are only the preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present application.
[0137] After considering the specification and practicing the application disclosed herein, those skilled in the art will readily conceive of other implementations of the embodiments of the present application. The present application is intended to cover any variations, uses or adaptations of the embodiments of the present application, which follow the general principles of the embodiments of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the embodiments of the present application.
[0138] It should be understood that the embodiments of the present application are not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the embodiments of the present application is only limited by the appended claims.
Claims
1. A data comparison method, characterized in that: Applied to source-end data tables and target-end data tables when the source-end database and the target-end database are heterogeneous databases, the method includes: According to the comparison key of each row of data in the source data table and the target data table, a bucket value of each row of data is determined by performing a hash modulo operation on the comparison key, wherein the comparison key is used to distinguish each row of data, and the bucket value is used to indicate that any row of data is allocated to one of the preset multiple buckets; Allocate each row of data in the source data table to a corresponding bucket based on the bucket value of each row of data in the source data table, and allocate each row of data in the target data table to a corresponding bucket based on the bucket value of each row of data in the target data table; A plurality of comparison threads are configured for the plurality of preset buckets, and the plurality of comparison threads are used to perform parallel data comparison on each row of data in the plurality of preset buckets to obtain a data comparison result.
2. The method according to claim 1, characterized in that The determining the bucket values of each row of data by performing a hash modulo operation on the comparison key includes: For any row of data in the above-mentioned rows, the bucket value is determined in the following manner: Get the field value corresponding to the comparison key of any row of data; The hash value corresponding to the field value is calculated by a hash function, and a hash modulus operation is performed on the hash value to obtain the bucket value of any row of data.
3. The method according to claim 2, characterized in that The performing a hash modulo operation on the hash value to obtain the bucket value of any row of data includes: Acquire the number of comparison threads to be used in the comparison thread pool as a first number, and according to a second number greater than or equal to the first number, preset the number of buckets to be a plurality of buckets of the second number, wherein the comparison thread pool is a thread pool preset with the plurality of comparison threads; A hash modulo operation is performed on the hash value according to the second number to obtain a bucket value for any row of data.
4. The method according to claim 1, characterized in that: The performing parallel data comparison on each row of data in the preset multiple buckets by using the multiple comparison threads includes: For any of the preset multiple buckets, data comparison is performed in the following manner: Through the comparison thread corresponding to any one of the buckets, a hash connection is performed on the row data of the source data table allocated to any one of the buckets and the row data of the target data table allocated to any one of the buckets to obtain an association set representing the hash connection result, and data comparison is performed based on the association set.
5. The method according to claim 4, characterized in that The performing data comparison according to the association set includes: In the association set, for any pair of row data for which a hash connection is established, the field values of each column in the pair of row data are compared respectively, and when at least one pair of different field values are found out through comparison, the pair of row data is determined as a difference row, and the difference row is used to indicate a data comparison result in which the row data has a field value difference; and / or, In the association set, for any row data for which a hash connection is not established, the any row data is determined as multiple rows, and the multiple rows are used to indicate the data comparison result that there is newly added row data.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Partitioning each row of data in the source data table and the target data table respectively to obtain a partition list, wherein the partition list includes a plurality of preset zones, and the partition list is used to indicate a corresponding relationship between the plurality of preset zones and the row of data in each zone; The determining the bucket values of each row of data by performing a hash modulo operation on the comparison key includes: Multiple reading threads are configured according to the partition list, and the comparison keys of the row data in the preset multiple zones are read in parallel by the multiple reading threads, and the bucket values of the respective row data are determined by performing a hash modulo operation on the comparison keys.
7. A data comparison device, characterized in that: Applicable to source data tables and target data tables when the source database and the target database are heterogeneous databases, the device includes: A determination module, configured to determine the bucket value of each row of data according to the comparison key of each row of data in the source data table and the target data table by performing a hash modulo operation on the comparison key, wherein the comparison key is used to distinguish the rows of data, and the bucket value is used to indicate that any row of data is allocated to one of the preset multiple buckets; an allocation module, configured to allocate each row of data in the source data table to a corresponding bucket based on the bucket value of each row of data in the source data table, and allocate each row of data in the target data table to a corresponding bucket based on the bucket value of each row of data in the target data table; The comparison module is used to configure multiple comparison threads for the preset multiple buckets, and perform parallel data comparison on each row of data in the preset multiple buckets through the multiple comparison threads to obtain data comparison results.
8. An electronic device, characterized in that: include: Memory and processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when being executed by a processor.
Citation Information
Cited By
Distributed project BOM data transmission system and device
CN120725821A
Data comparison method and device, equipment and storage medium
CN120849409A