A HASH-based data comparison method and device
By using HASH algorithm and region comparison technology in the data comparison method, the problem of slow data comparison speed under large data volume is solved, and a more efficient data comparison process is achieved.
Patent Information
- Application Number
- CN202111506096.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2041-12-10
AI Technical Summary
In the case of large amount of data, the existing data comparison method seriously affects the speed of data comparison because the MD5 values take a long time to sort.
Using the HASH-based data comparison method, the row MD5 value data in the first replica is stored in the HASH table through the HASH algorithm, and the row MD5 value data of the second replica is collided with the HASH table, and the regions are divided and compared to determine the consistency of the data of the two replicas.
By dividing the region comparison method, the time spent in sorting the row MD5 value data in traditional methods is significantly shortened, and the efficiency of data comparison is improved.
Smart Images

Figure CN114297193B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of databases, and particularly to a data comparison method and device based on HASH. Background Art
[0002] With the rapid development of informatization construction, the business is becoming more and more complex, and the coupling between various businesses is also becoming stronger. It often happens that different businesses need to access the same data. In this case, the data accessed by each business becomes hot data, and the competition for data access will cause some businesses to be in a waiting state, resulting in a decline in business performance. To solve this problem, the usual processing method is to make multiple copies of this data, and each business accesses different data copies. In this case, the data consistency of each copy is particularly important.
[0003] Data comparison is an important means for data consistency check. The usual data comparison method is divided into four steps: 1. Obtain the data in each data copy in units of tables; 2. Calculate the MD5 (Message-Digest Algorithm 5) value of each row of data in units of rows; 3. Sort the MD5 codes of the data in the table in units of tables; 4. Compare the MD5s in the two copies one by one in units of tables. If there are inconsistent MD5 values, it is determined that there is data inconsistency.
[0004] In the above data comparison method, when sorting the MD5 generated from the table data, quicksort / heap sort is usually used. This comparison method is relatively fast when the data volume is small, but as the data scale increases, the time spent on sorting the MD5 values becomes more and more, seriously affecting the data comparison speed.
[0005] In view of this, shortening the time spent on sorting during the data comparison process is an urgent problem to be solved in this technical field. Summary of the Invention
[0006] The embodiments of the present invention solve the problems existing in the above-mentioned prior art, and provide a data comparison method and device based on HASH. This method selects the first copy as a reference, arranges the row MD5 value data in the first copy in the HASH table through the HASH algorithm, and then collides the row MD5 value of the second copy with the row MD5 value data stored in the HASH table of the first copy. When the data volume to be compared is very large, the data to be compared is partitioned and compared zone by zone, and then the consistency of the two copy data is determined; the present invention shortens the time spent on sorting the row MD5 value data by the traditional method through the partitioned comparison method, and greatly improves the efficiency of data comparison.
[0007] The present invention adopts the following technical solutions:
[0008] In a first aspect, the present invention provides a HASH-based data comparison method, including:
[0009] Obtain all copies of the hot data;
[0010] Extract the data of each respective copy database and calculate the MD5 value of each row;
[0011] Set the MD5 buffer of each respective copy, and store the row MD5 value data of each respective copy in the corresponding buffer to obtain the data comparison table corresponding to each respective copy;
[0012] Select any one of all the copies as the first copy. Taking the buffer as a unit, use the HASH algorithm to store the row MD5 value data in the comparison table into the corresponding position area of the HASH table, and arrange the data in the same position area in the order of data size, so as to complete the insertion of the data comparison table of the first copy into the HASH table;
[0013] Select any one of the copies other than the first copy as the second copy. By determining the size of the data volume of the comparison table of the second copy, select the corresponding comparison method for data comparison to determine the consistency of the data between the first copy and the second copy.
[0014] Preferably, the specific steps for obtaining all copies of the hot data include:
[0015] Set the data accessed by multiple services as the hot data;
[0016] Create data copies corresponding to the number of services according to the number of services accessing the hot data.
[0017] Preferably, the setting of the MD5 buffer for each respective copy specifically includes: dividing the buffer according to the value range of the first byte number of the row MD5 value data and numbering the buffer.
[0018] Preferably, the storing of the row MD5 value data of each respective copy in the corresponding buffer specifically includes:
[0019] Obtain the memory capacity of each buffer;
[0020] Store the row MD5 value data in the corresponding buffer. When the memory occupied by the row MD5 value data is less than the memory capacity of the buffer, directly store the row MD5 value data in the buffer; when the memory occupied by the row MD5 data value is greater than the memory capacity of the buffer and the buffer is full, the data tool saves the row MD5 data file in the buffer to the local disk, and numbers the data file saved on the local disk with the same number as the buffer. At this time, release the buffer for storing subsequent data.
[0021] Preferably, the data comparison table includes: a line MD5 value data comparison table in the buffer and a line MD5 value data comparison table stored on the local disk when the memory occupied by the line MD5 value data is greater than the memory capacity of the buffer.
[0022] Preferably, the size order of the data is actually the arrangement order of the line MD5 value data strings.
[0023] Preferably, after obtaining the data comparison tables corresponding to their respective copies, indirectly judge the data volume size of the second copy comparison table by identifying whether the line MD5 value data in the first copy is stored on the local disk. The specific steps are: identify whether the local disk data stores the line MD5 value data of the first copy. If so, it is determined that the data volume is large; if not, it is determined that the data volume is small.
[0024] Preferably, the selection of the corresponding comparison method is based on the data volume size of the second copy comparison table, specifically:
[0025] When the data volume of the second copy comparison table is small, insert all the line MD5 value data in the buffers of the first copy into the HASH together to obtain the HASH table; then, one by one, collide the line MD5 value data in the second copy with the line MD5 value data in the HASH table of the first copy. After the collision, delete the same line MD5 value data in the first copy and the second copy until all the line MD5 value data in the second copy has been collided, and delete the line MD5 value data that has successfully collided in the first copy and the second copy.
[0026] When the data volume of the second copy comparison table is large, select any buffer of the first copy as the first buffer, and insert the first buffer and the line MD5 data file on the local disk with the same number as the first buffer into the HASH table; then, one by one, collide the line MD5 value data in the buffer of the second copy and the file on the local disk with the same number as the buffer with the line MD5 value data in the HASH table of the first copy until the line MD5 value data in the first buffer and the file on the local disk with the same number as the first buffer in the second copy has been collided. Then, select other buffers except the first buffer for collision, and delete the line MD5 value data that has successfully collided in the HASH table until all the line MD5 value data in the first copy and the second copy has been collided.
[0027] Preferably, the determination of the consistency of the data in the first copy and the second copy is based on the completion of the collision. The specific steps of the determination process are:
[0028] Determine whether there is any remaining MD5 value data in the rows of the first copy HASH table or whether the second copy has an unsuccessful collision; if there is no remaining MD5 value data in the rows of the first copy HASH table and all the collisions of the second copy are successful, it is determined that the data of the first copy and the second copy are consistent; otherwise, it is determined that the data of the first copy and the second copy are inconsistent.
[0029] In a second aspect, the present invention further provides a HASH-based data comparison device for implementing the HASH-based data comparison method described in the first aspect. The device includes:
[0030] At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to perform the HASH-based data comparison method described in the first aspect.
[0031] In a third aspect, the present invention further provides a non-volatile computer storage medium storing computer-executable instructions, which are executed by one or more processors to complete the HASH-based data comparison method described in the first aspect.
[0032] The present invention can not only compare the data between two copies, but also compare the data between multiple copies. Arbitrarily select a reference copy, and then use the method of the present invention to compare other copies with the reference copy one by one, and then obtain the inconsistent copies; if all are inconsistent, it may be that some data is lost / added during the process of generating the copy of the hot data by the reference copy. At this time, a new reference copy can be replaced, and then the method of the present invention can be used to compare the new reference copy and other copies one by one, so as to achieve the purpose of multi-copy comparison. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0034] Figure 1 is a flowchart of a HASH-based data comparison method provided by an embodiment of the present invention;
[0035] Figure 2 is a flowchart of a method for obtaining a copy comparison table provided by an embodiment of the present invention;
[0036] Figure 3It is a schematic diagram of a method for partitioning a data buffer based on HASH provided by an embodiment of the present invention;
[0037] Figure 4 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0038] Figure 5 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0039] Figure 6 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0040] Figure 7 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0041] Figure 8 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0042] Figure 9 It is a flowchart of a method for data comparison based on HASH provided by an embodiment of the present invention;
[0043] Figure 10 It is a schematic diagram of the process of storing row MD5 values in a HASH table provided by an embodiment of the present invention;
[0044] Figure 11 It is a schematic diagram of the structure of a data comparison device based on HASH provided by an embodiment of the present invention. Detailed implementation manners
[0045] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0046] In the description of the present invention, the orientation or positional relationships indicated by the terms "inner", "outer", "longitudinal", "transverse", "upper", "lower", "top", "bottom", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention rather than requiring the present invention to be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.
[0047] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0048] Embodiment 1:
[0049] Embodiment 1 of the present invention provides a data comparison method based on HASH, as follows: Figure 1 As shown below, the specific steps are as follows:
[0050] Step 101: Obtain all copies of the hotspot data.
[0051] In the embodiments of the present invention, hotspot data refers to data that is simultaneously accessed by multiple services. The process of obtaining hotspot data is actually a process of monitoring hotspot data. By monitoring the processes of hotspot data and related services and the interaction between the two, it is determined whether the monitored data is hotspot data. Obtaining all copies of the hotspot data is a preparatory step for subsequent data comparison between copies.
[0052] Step 102: Extract the data of each respective copy database and calculate the MD5 value of each row.
[0053] Before comparing the data between copies, a traceable variable needs to be selected, and the output values displayed by different data should be different, while the same data in different copies should display the same output value. The MD5 algorithm can generate a unique "digital fingerprint" for any file (regardless of its size, format, quantity). By checking whether the "digital fingerprints" in different copies are consistent, the consistency of the two copies is determined. In the actual comparison process, rows are selected as the minimum unit for data comparison between each copy. The row MD5 value data of the copies is obtained, and then the row MD5 value data in different copies is compared one by one to achieve the comparison purpose between different copies.
[0054] Step 103: Set the MD5 buffer for each respective copy and store the row MD5 value data of each respective copy in the corresponding buffer to obtain the data comparison table corresponding to each respective copy.
[0055] As Figure 2 shown, it represents the flowchart for obtaining the copy comparison table. Generally, if the copy data volume is relatively small (the determination of the data volume size will be described later), only the row MD5 data obtained from each copy needs to be sorted, and then the row MD5 value data between different copies is compared one by one. At this time, the comparison can be carried out between two copies. Through the data comparison method between two copies, until all copy comparisons are completed; or all copies can be compared together, and the consistency between different copies is determined by whether the row MD5 value data of different copies is exactly the same. When the copy data volume is very large, in order to reduce the time spent on sorting the row MD5 value data, the embodiments of the present invention set an MD5 buffer for the copies and place the row MD5 value data into the corresponding buffer according to the corresponding rules. By setting the buffer and numbering it, the blockification of the copy data is achieved. By comparing the row MD5 value data in the same numbered area of different copies, the time spent on sorting the row MD5 value data is shortened.
[0056] Step 104: Select any one of all the copies as the first copy. Taking the buffer as a unit, use the HASH algorithm to store the row MD5 value data in the comparison table into the corresponding location area of the HASH table, and arrange the data in the same location area in the order of data size, so as to complete the insertion of the data comparison table of the first copy into the HASH table.
[0057] During the process of sorting the row MD5 value data of the copies, the row MD5 value data that meets the same area rule will be assigned to the same area, but the MD5 value data in the same area is not sorted, which increases the difficulty of the comparison process, increases the comparison time, and reduces the comparison efficiency. In the embodiment of the present invention, the row MD5 value data of each buffer in the copy is sorted and stored in the HASH table in units of areas through the HASH algorithm. What is actually completed is that the row MD5 value data in the comparison table is stored in the corresponding area of the HASH table.
[0058] Step 105: Select any copy other than the first copy as the second copy. By determining the size of the data volume of the comparison table of the second copy, select the corresponding comparison method to perform data comparison, and determine the consistency of the data between the first copy and the second copy.
[0059] The size of the copy data volume determines the selection of the comparison method. When the data volume is very small, the row MD5 value data of all buffers of the first copy is sorted in a single HASH table using the HASH algorithm, and then the row MD5 value data in the second copy is collided with the data in the HASH table of the first copy one by one until all the row MD5 value data is collided; when the data volume is very large, select the first copy as the reference, use the HASH algorithm to obtain the HASH table of the first copy in units of areas and sort it. Then select the row MD5 value data in the second copy to collide with the row MD5 value data in the HASH table of the first copy with the same area number. After all the row MD5 value data in the same area is collided one by one, proceed area by area until all areas are compared. Determine the consistency of the data of the two copies by whether all the row MD5 value data in the second copy exists in the HASH table of the first copy.
[0060] When multiple services access a data at the same time, in this preferred embodiment, there is a further display of the implementation method for obtaining hot data. For example, step 101 can be further refined as Figure 3 shown, and specifically includes the following steps:
[0061] Step 1011: Set the data accessed by multiple services as hot data.
[0062] Step 1012: Create data copies corresponding to the number of services based on the number of services accessing the hotspot data.
[0063] Among them, hotspot data refers to data accessed by multiple services simultaneously. To enable each service to access the hotspot data simultaneously without crosstalk or lag between services, copies equal to the number of services are usually created in advance.
[0064] The specific process of setting the MD5 buffer for each copy includes: dividing the buffer according to the value range of the first byte number of the row MD5 value data and numbering the buffer.
[0065] As Figure 4 shown, setting the copy MD5 buffer requires storing all the row MD5 value data in the copy into the buffer. Based on the representation range of the first byte number of the row MD5 value (the representation range of one byte is 0 - 255), set the corresponding number of buffers. In the embodiment of the present invention, the buffer is set to 256 areas, numbered 0 - 255, and stored in the corresponding buffer according to the first byte number of the row MD5 value.
[0066] Further, as Figure 5 shown, the specific process of storing the row MD5 value data of each copy in the corresponding buffer includes:
[0067] Step 201: Obtain the memory capacity of each buffer.
[0068] Step 202: Store the row MD5 value data in the corresponding buffer. When the memory occupied by the row MD5 value data is less than the memory capacity of the buffer, directly store the row MD5 value data in the buffer; when the memory occupied by the row MD5 value data is greater than the memory capacity of the buffer and the buffer is full, the data tool saves the row MD5 value data file in the buffer to the local disk, numbers the data file saved on the local disk with the same number as the buffer, and at this time, releases the buffer for storing subsequent data.
[0069] In the embodiments of the present invention, the capacity of each area of the MD5 buffer can be configured according to the system (usually the default is 256M), and the MD5 buffer functions to cache the MD5 value data of the rows. Step 202 can be understood as the rule for storing / writing data into the buffer or the disk. For the convenience of the subsequent description of this solution, a comparison is introduced between the memory occupied by the MD5 value data of the rows and the memory capacity of the buffer. Here, the "memory occupied by the MD5 value data of the rows" refers to the data corresponding to the buffer number before the copy is stored in the buffer. Assume that before being stored in the buffer, the data of the copy has been divided into sub-data according to the buffer number. By comparing the memory occupied by the sub-data with the buffer capacity, the size relationship between the memory occupied by the MD5 value data of the rows and the memory capacity of the buffer can be determined. In the actual operation process, it is difficult to measure the data corresponding to the buffer number before the copy is stored in the buffer. Generally, the size relationship between the amount of MD5 value data of the copy that needs to be stored in the buffer with the corresponding number and the buffer capacity is determined through dynamic storage. When the MD5 value data of the rows stored / written in a certain numbered buffer reaches the memory capacity of the buffer (the length of each MD5 value of the row is 16 bytes, and the ratio of the memory of the buffer to the memory occupied by each MD5 value is the maximum number of MD5 values of the row that each buffer can accommodate), it means that the memory occupied by the MD5 value data of the rows is greater than the memory capacity of the buffer. At this time, the MD5 value data file in the buffer is saved in the local disk. After the buffer is released, the buffer stores / writes the subsequent data.
[0070] Furthermore, the data comparison table includes: the MD5 value data comparison table in the buffer and the MD5 value data comparison table stored in the local disk when the memory occupied by the MD5 value data of the rows is greater than the memory capacity of the buffer.
[0071] Among them, the comparison table composed of the MD5 value data of the rows includes a part in the buffer and a part in the local disk. However, it should be noted that when the amount of copy data is small (the data in the copy exists in the buffer and no data is stored on the local disk), at this time, the comparison table only has the part in the buffer.
[0072] Furthermore, the size order of the data is actually the arrangement order of the MD5 value data strings of the rows.
[0073] Taking the first copy as the reference, the comparison table of the first copy is stored in the HASH table through the HASH algorithm (the storage rule will be further described in the embodiments later), and the MD5 values of the rows at the same position in the HASH table are arranged according to the arrangement order of the MD5 value strings of the rows. For example: in the HASH table, in the area numbered 3, there are MD51, MD52, and MD53 of the rows, and the corresponding strings are: 0xbaf685bd7a2b7c2d93ecaa1a59e95d,
[0074] 0xe9587e52f6e83e873f33fa6cd2d36d and 0xf98117c5f3c733e09b6c6ad5b72896. From the numbers in the string, it can be seen that MD51 < MD52 < MD53. Therefore, the arrangement order of the three line MD5 values in the three regions is: MD51, MD52, MD53.
[0075] Furthermore, after obtaining the data comparison tables corresponding to their respective copies, indirectly determine the size of the data volume of the second copy comparison table by identifying whether the line MD5 value data in the first copy is stored on the local disk. As Figure 6 shown, the specific steps are as follows:
[0076] Step 301: Identify whether the local disk data stores the line MD5 value data of the first copy.
[0077] Step 302: If so, determine that the data volume is large; if not, determine that the data volume is small.
[0078] Both the first copy and the second copy are made of the same hot data, and they have the same data magnitude (basically the same). In the embodiment of the present invention, the first copy is used as a reference, a buffer is set for the first copy, and the line MD5 value data with the corresponding buffer number is sorted in the HASH table by using the HASH algorithm. Before comparing the two copies, the size of the data volume of the first copy can be known, and the size of the data volume of the second copy can be further determined through the size of the data volume of the first copy.
[0079] Furthermore, the selected corresponding comparison method is based on the size of the data volume of the second copy comparison table, specifically:
[0080] As Figure 7 shown, it is a flowchart of the HASH-based data comparison method when the data volume is small. When the data volume of the second copy comparison table is small, insert the line MD5 value data of all buffers of the first copy into the HASH together to obtain the HASH table; then, one by one, collide the line MD5 value data in the second copy with the line MD5 value data in the HASH table of the first copy. After the collision, delete the same line MD5 value data of the first copy and the second copy until all the line MD5 value data in the second copy is collided, and delete the line MD5 value data that has successfully collided between the first copy and the second copy.
[0081] As Figure 8As shown in the figure, it is a flowchart of the HASH-based data comparison method when the data volume is large. When the data volume of the second replica comparison table is large, any buffer of the first replica is selected as the first buffer, and the first buffer and the row MD5 data files in the local disk with the same number as the first buffer are inserted into the HASH table; then, the buffers in the second replica and the row MD5 value data in the local disk containing files with the same number as the buffer are collided with the row MD5 value data in the HASH table of the first replica one by one. After colliding the first buffer of the second replica and the row MD5 value data in the local row MD5 value data with the same number as the first buffer with the row MD5 value data in the disk file, then select other buffers except the first buffer for collision, and delete the row MD5 value data that has successfully collided in the HASH table until all the row MD5 value data in the first replica and the second replica are collided.
[0082] The present invention indirectly determines the data volume of the second replica by using the first replica. During the process of setting the buffer, for the replica in which the row MD5 value data in the replica is not stored / written on the local disk (determined that the data volume is small), it is selected to insert all the data of the first replica into the HASH table, and then the row MD5 value data in the second replica is collided with the row MD5 value data in the HASH table of the first replica one by one; if the row MD5 value data in the replica is stored / written on the local disk (determined that the data volume is large), it is necessary to partition and collide and compare. At this time, the HASH table consists of a series of sub-HASH tables with numbers. The row MD5 value data in one area of the second replica is collided and compared with the sub-HASH table in the first replica with the same number, and then collided and compared area by area until all the row MD5 value data are collided and compared.
[0083] Further, the determination of the consistency of the data between the first replica and the second replica is based on the completion of the collision. As Figure 9 shown, the specific steps of the determination process are as follows:
[0084] Step 401: Determine whether there is any remaining row MD5 value data in the HASH table of the first replica or whether there is any unsuccessful collision in the second replica.
[0085] Step 402: If there is no remaining row MD5 value data in the HASH table of the first replica and all the collisions in the second replica are successful, it is determined that the data between the first replica and the second replica is consistent; otherwise, it is determined that the data between the first replica and the second replica is inconsistent.
[0086] After all the line MD5 value data in the second copy has completed the collision, the line MD5 value data with successful collision (the same line MD5 value data in the first copy and the second copy) will be deleted. When the collision in the second copy is completed, there will be two result judgments. The first one is: after the HASH table of the first copy has completed the collision and deleted the same line MD5 value data in the two copies, whether there is still remaining line MD5 value data. The second one is: after the collision in the second copy, whether there is any line MD5 value data with unsuccessful collision in the second copy. When there is no remaining line MD5 value data in the HASH table of the first copy, and all the line MD5 value data in the second copy has successful collision with the line MD5 value data in the HASH table of the first copy, it indicates that the data in the first copy and the second copy is consistent; otherwise, in other cases, it indicates that the data in the first copy and the second copy is inconsistent.
[0087] Since all copies of the present invention are derived from the same popular data, the consistency between the copies can be determined by comparing between the copies. In the embodiment of the present invention, the type of comparison is determined by the size of the data volume. For a copy with a small data volume, when comparing, the inserted line MD5 value data is collided with the data in the HASH table of the reference copy one by one. Although it is different from the traditional comparison method, the time spent on comparison is basically the same (the comparison speed is very fast). When the copy data volume is extremely large, by dividing the area and comparing each area one by one, the time spent on sorting the line MD5 value data by the traditional method is greatly shortened, and the efficiency of data comparison is improved.
[0088] Embodiment 2
[0089] To better clarify the present invention, in this embodiment, further elaboration is made on how to store the line MD5 value data into the corresponding position of the HASH table by using the HASH algorithm in step 104 of Embodiment 1.
[0090] Such as Figure 10As shown in the figure, it is a schematic diagram showing the storage of row MD5 value data in the corresponding positions of the HASH table. First, set the number of HASH table buckets according to the configuration of the device (different systems can set different total numbers of buckets). During use, it is necessary to obtain the number HZ of system HASH table buckets in advance (a HASH table bucket represents a block set in the HASH table, and a skip list is stored in the bucket for storing row MD5 value data obtained through the HASH algorithm), and number the buckets from 0 to HZ - 1 (0 occupies one number, so the maximum number is HZ - 1). This numbering method corresponds to the composition of the key value numbers (ensuring that each key value can be stored in the skip list of the HASH table after being calculated through the corresponding rules); then obtain the byte value of the row MD5 value, and use the last four - byte number as the key value (K) to substitute into the position relationship function: f(K)=K%HZ. In the position relationship function, K represents the key value digital value of the last four bytes of the row MD5 value, HZ represents the maximum number of buckets, and % is the modulo operation. Through the position relationship calculation, the number after the decimal point obtained by the position relationship function corresponds to the number of the HASH table bucket. When there is no number or zero after the decimal point in the position relationship function, store this row MD5 value in the HASH table bucket numbered 0. Obtain the HASH table bucket position corresponding to each row MD5 value data by taking the remainder; then sort the row MD5 value data in the same bucket according to the size of the string. For example: the number of buckets HZ configured in the system is 10000 (the maximum bucket number is: 9999), and the last four bytes of the row MD5 value are 305419896(0x12345678), f(0x12345678)= 305419896%10000 = 9896. Through the operation of %, (take the number after the decimal point), store this row MD5 value in the skip list of the HASH bucket numbered 9896. In Embodiment 1, the sorting of different row MD5 values in the same bucket has been described, and the row MD5 values in the skip list corresponding to the corresponding numbers are sorted and stored according to the order of the strings.
[0091] Embodiment 3
[0092] As Figure 11 shown, it is a schematic diagram of the device architecture for HASH - based data comparison according to an embodiment of the present invention. The HASH - based data comparison device of this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 11 Take one processor 21 as an example.
[0093] The processor 21 and the memory 22 can be connected through a bus or other means, Figure 11 Take the connection through the bus as an example.
[0094] The memory 22 serves as a non-volatile computer-readable storage medium and can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the HASH-based data comparison method in Embodiment 1. The processor 21 executes the HASH-based data comparison method by running the non-volatile software programs and instructions stored in the memory 22.
[0095] The memory 22 may include high-speed random access memory and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 optionally includes a memory remotely located relative to the processor 21, and these remote memories can be connected to the processor 21 through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0096] The program instructions / modules are stored in the memory 22 and, when executed by the one or more processors 21, execute the HASH-based data comparison method in the above Embodiment 1. For example, execute each of the Figures 1 - 10 steps shown above.
[0097] It should be noted that the content such as information interaction and execution process between the modules and units in the above device and system, since it is based on the same concept as the method embodiment of the present invention, the specific content can be referred to the description in the method embodiment of the present invention and will not be elaborated here.
[0098] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium may include: a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, an optical disc, etc.
[0099] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A HASH-based data comparison method, characterized in that, Including: Obtain all copies of the hotspot data; Extract the data of each respective copy database and calculate the MD5 value of each row; Set the MD5 buffer of each respective copy and store the row MD5 value data of each respective copy in the corresponding buffer to obtain the data comparison table corresponding to each respective copy; Select any one of all the copies as the first copy. Taking the buffer as the unit, use the HASH algorithm to store the row MD5 value data in the comparison table into the corresponding position area of the HASH table, and arrange the data in the same position area in the order of data size, so as to complete the insertion of the data comparison table of the first copy into the HASH table; Select any one of the copies other than the first copy as the second copy. By judging the size of the data volume of the comparison table of the second copy, select the corresponding comparison method for data comparison to judge the consistency of the data between the first copy and the second copy; When the data volume of the comparison table of the second copy is small, insert the row MD5 value data of all buffers of the first copy into the HASH together to obtain the HASH table; one by one, collide the row MD5 value data in the second copy with the row MD5 value data in the HASH table of the first copy. After the collision, delete the same row MD5 value data between the first copy and the second copy until all the row MD5 value data in the second copy are collided, and delete the row MD5 value data with successful collision between the first copy and the second copy; when the data volume of the comparison table of the second copy is large, select any buffer of the first copy as the first buffer, and insert the first buffer and the row MD5 data file in the local disk with the same number as the first buffer into the HASH table; one by one, collide the row MD5 value data in the buffer of the second copy and the row MD5 value data in the file with the same number as the buffer in the local disk with the row MD5 value data in the HASH table of the first copy. After colliding the row MD5 value data in the first buffer of the second copy and the row MD5 value data in the local row MD5 value data file with the same number as the first buffer in the disk file, then select other buffers except the first buffer for collision, and delete the row MD5 value data with successful collision in the HASH table until all the row MD5 value data in the first copy and the second copy are collided; 2. The HASH-based data comparison method according to claim 1, characterized in that, The specific steps of the obtaining all copies of the hotspot data include: Set the data accessed by multiple services as the hotspot data; Create data copies corresponding to the number of services according to the number of services accessing the hotspot data.
3. The HASH-based data comparison method according to claim 1, characterized in that, The setting of the MD5 buffer of each respective copy specifically includes: divide the buffer according to the value range of the first byte number of the row MD5 value data and number the buffer.
4. The HASH-based data comparison method according to claim 1, characterized in that, The storing the row MD5 value data of each respective copy in the corresponding buffer specifically includes: Obtain the memory capacity of each buffer; Store the line MD5 value data in the corresponding buffer. When the memory occupied by the line MD5 value data is less than the memory capacity of the buffer, directly store the line MD5 value data in the buffer; when the memory occupied by the line MD5 value data is greater than the memory capacity of the buffer and the buffer is full, the data tool saves the line MD5 value data file in the buffer to the local disk, numbers the data files saved on the local disk with the same number as the buffer, and at this time, releases the buffer for storing subsequent data.
5. The HASH-based data comparison method according to claim 1, characterized in that, The data comparison table includes: the line MD5 value data comparison table in the buffer and the line MD5 value data comparison table stored on the local disk when the memory occupied by the line MD5 value data is greater than the memory capacity of the buffer.
6. The HASH-based data comparison method according to claim 1, characterized in that, The size order of the data is actually the arrangement order of the line MD5 value data strings.
7. The HASH-based data comparison method according to claim 4, characterized in that, After obtaining the data comparison tables corresponding to their respective copies, indirectly judge the size of the data volume of the second copy comparison table by identifying whether the line MD5 value data in the first copy is stored on the local disk. The specific steps are as follows: identify whether the line MD5 value data of the first copy is stored in the local disk data. If so, it is determined that the data volume is large; if not, it is determined that the data volume is small.
8. The HASH-based data comparison method according to claim 1, characterized in that, The determination of the consistency of the data in the first copy and the second copy is based on the completion of the collision. The specific steps of the determination process are as follows: Judge whether there is any remaining line MD5 value data in the first copy HASH table or whether there is an unsuccessful collision in the second copy; if there is no remaining line MD5 value data in the first copy HASH table and all the collisions in the second copy are successful, it is determined that the data in the first copy and the second copy is consistent; otherwise, it is determined that the data in the first copy and the second copy is inconsistent.
9. A HASH-based data comparison device, characterized in that, Includes: At least one processor; At least one memory; Wherein, the at least one processor and the at least one memory are communicatively connected to each other. The at least one memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the HASH-based data comparison method provided by any one of claims 1-8.
Citation Information
Patent Citations
Data comparison method and device, storage medium and electronic equipment
CN110362574A
File data comparison method and device, equipment and storage medium
CN113342750A