A data storage method, a method and a device for verifying data consistency

By adopting a multi-level HASH table method in data storage, the problem of inefficient data storage and comparison in the prior art is solved, and more efficient data search and comparison are achieved.

CN115617808BActive Publication Date: 2025-06-17WUHAN DAMENG DATABASE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211387907.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2025-06-17
Estimated Expiration
2042-11-08

AI Technical Summary

Technical Problem

The prior art has problems with inefficiency in data storage and comparison, especially when memory consumption is high and CPU multi-core advantages cannot be exploited.

Method used

The data storage method of a multi-level HASH table is adopted, and the data is stored in the corresponding data page by calculating the multi-level HASH value of the data, and an index record chain pointing to the target data page is generated to improve the efficiency of data search.

Benefits of technology

It reduces the resources and time required in the data search process, improves the efficiency of data search and comparison, and can effectively utilize the advantages of CPU multi-core.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115617808B_ABST
    Figure CN115617808B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of big data technology, and provides a data storage method, a method and a device for verifying data consistency. The data storage method includes: successively calculating the 1st-level HASH value, 2nd-level HASH value, ..., nth-level HASH value corresponding to the data according to the Key value of the data until the number of the first data with the same nth-level HASH value does not exceed the number of data that can be stored in a single data page; storing the first data in the target data page so that all the data stored in the target data page have the same nth-level HASH value; generating an index record chain pointing to the target data page in the 1st-level HASH table, 2nd-level HASH table, ..., nth-level HASH table. The present invention enables, when searching for the stored data, to find the target data page through the index record chain, thereby narrowing the data search range to a single data page, reducing the resources and time consumed in the data search process, improving the data search efficiency, and thus improving the data comparison efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data technology, and in particular, to a data storage method, a method and device for verifying data consistency. Background Art

[0002] With the development of Internet technology, more and more large information systems are applied in various industries. These large applications will use multiple databases simultaneously. These databases may be homogeneous or heterogeneous. In some application scenarios, it is necessary to compare the table data located in different databases, and as the scale of the comparison data is getting larger and larger, an efficient and fast data comparison solution is required.

[0003] Currently, traditional comparison tools first need to query the data to be compared from the databases respectively, then sort them according to rules, cache the sorted results in memory, and finally traverse the result set data for item-by-item comparison. This method poses a great challenge to memory consumption, and the performance gradually decreases as the memory consumption increases. When the memory cannot cache all the result sets, part of the result sets need to be persisted, and at this time, the performance will decrease significantly; at the same time, due to the huge memory consumption of this method, it is also impossible to take advantage of the multi-core of the CPU to perform concurrent comparison of multiple table data, and the comparison efficiency is increasingly unable to meet the challenges of today's large amounts of data.

[0004] At the same time, traditional comparison tools require a large number of data search operations during the comparison process. In the prior art, data is usually stored through conventional data tables, and each data search operation needs to traverse all the data in the table, making the data search efficiency extremely low, resulting in a large amount of resources and time consumed in the data comparison process, and thus unable to meet the requirements for comparison efficiency under today's large amounts of data.

[0005] In view of this, overcoming the defects of the prior art is an urgent problem to be solved in this technical field. Summary of the Invention

[0006] The technical problem to be solved by the present invention is that the prior art stores data through conventional data tables, resulting in low data search efficiency when storing a large amount of data.

[0007] The further technical problem to be solved by the present invention is that in the prior art, data comparison needs to store the sorted data results in memory and compare the data item by item in memory, resulting in high memory occupancy and poor comparison efficiency.

[0008] In a first aspect, the present invention provides a data storage method. The data storage structure includes a multi-level HASH table. The data storage method includes:

[0009] According to the Key value of the data, calculate the first-level HASH value, second-level HASH value, ..., n-level HASH value corresponding to the data in sequence until the number of the first data does not exceed the number of data that can be stored in a single data page; wherein, the first data is data with the same n-level HASH value;

[0010] Store the first data into the corresponding target data page so that all the data stored in the target data page have the same n-level HASH value;

[0011] Generate an index record chain pointing to the target data page in the first-level HASH table, second-level HASH table, ..., n-level HASH table; where n is a positive integer.

[0012] Preferably, the step of calculating the first-level HASH value, second-level HASH value, ..., n-level HASH value corresponding to the data according to the Key value of the data specifically includes:

[0013] When calculating the first-level HASH value, use a preset value as the target quantity, select the first target quantity of bytes from the Key value of the data, and calculate the corresponding first-level HASH value according to the first target quantity of bytes;

[0014] When calculating any k-level HASH value among the second to n-level HASH values, use the target quantity used when calculating the (k - 1)-level HASH value plus a preset increment as the target quantity used when calculating the k-level HASH value;

[0015] Select the first target quantity of bytes from the Key value of the data, and calculate the corresponding k-level HASH value according to the first target quantity of bytes; where k is an integer greater than or equal to 2 and less than or equal to n.

[0016] Preferably, the step of generating an index record chain pointing to the target data page in the first-level HASH table, second-level HASH table, ..., n-level HASH table specifically includes:

[0017] Insert the corresponding index records into the n-level HASH table, (n - 1)-level HASH table, ..., first-level HASH table in sequence. Specifically,

[0018] When inserting the corresponding index record into the n-level HASH table, establish an n-level target index record in the corresponding index page of the n-level HASH table according to the n-level HASH value of the first data, and the n-level target index record points to the target data page;

[0019] When inserting the corresponding index record into any k-level HASH table in the 1 to n-1 level HASH tables, a k-level target index record is established in the corresponding index page of the k-level HASH table according to the k-level HASH value of the first data, and the k-level target index record points to the k+1-level target index page;

[0020] Wherein, the k+1-level target index page is the index page where the k+1-level target index record is located in the k+1-level HASH table; wherein, the value of k is an integer greater than or equal to 1 and less than or equal to n-1.

[0021] Preferably, the method further includes a data search operation, and the data search operation specifically includes:

[0022] Search for the corresponding index record chain according to the Key value of the data;

[0023] If the corresponding index record chain cannot be found, the final search result is that the data cannot be found;

[0024] If the corresponding index record chain is found, obtain the corresponding target data page according to the index record chain;

[0025] Search for the data in the target data page, and use the result of searching for the data in the target data page as the final search result.

[0026] Preferably, the searching for the corresponding index record chain according to the Key value of the data specifically includes:

[0027] According to the Key value of the data, calculate the 1-level HASH value, 2-level HASH value,..., n-level HASH value corresponding to the data in sequence, and search for the corresponding index record. Specifically,

[0028] When calculating any k-level HASH value in the 1 to n-level HASH values and searching for the corresponding index record, search for the corresponding k-level target index record in the k-level target index page according to the k-level HASH value. If the corresponding k-level target index record is found and the k-level target index record points to the corresponding k+1-level index page, use the k+1-level index page as the k+1-level target index page to search for the k+1-level target index record; until the corresponding index record cannot be found, or the found index record points to the corresponding data page; wherein, the value of k is an integer greater than or equal to 1 and less than or equal to n;

[0029] If the corresponding index record cannot be found, the search result is that the corresponding index record chain cannot be found;

[0030] If a corresponding index record is found and the index record points to the corresponding data page, the data page is used as the target data page corresponding to the index record chain.

[0031] Preferably, the method further includes an insertion operation for data, and the insertion operation for data specifically includes:

[0032] According to the Key value of the data, find the corresponding index record chain;

[0033] If the corresponding index record chain cannot be found, store the data in the corresponding first target data page;

[0034] Calculate the 1-level HASH value corresponding to the data, and based on the 1-level HASH value, establish a corresponding index record in the 1-level HASH table, and the index record points to the first target data page;

[0035] If the corresponding index record chain is found, obtain the corresponding second target data page according to the index record chain;

[0036] If the second target data page is not full, store the data in the second target data page;

[0037] If the second target data page is full, select an empty data page as the third target data page, store some of the data in the second target data page in the third target data page, and select to store the data in the second target data page or the third target data page;

[0038] Update the index record chain pointing to the second target data page and the index record chain pointing to the third target data page.

[0039] Preferably, the update of the index record chain pointing to the second target data page and the index record chain pointing to the third target data page specifically includes:

[0040] If the index record pointing to the second target data page is in any k-level HASH table among the 1st to nth level HASH tables, add a new empty index page in the (k + 1)-level HASH table, and use the empty index page as the target index page; where the value of k is an integer greater than or equal to 1 and less than or equal to n; when k = n, create an (n + 1)-level HASH table;

[0041] According to the (k + 1)-level HASH value corresponding to the data in the second target data page, establish a corresponding second index record in the target index page, and the second index record points to the second target data page;

[0042] Based on the (k + 1)-level HASH value corresponding to the data in the third target data page, establish a corresponding third index record in the target index page, where the third index record points to the third target data page;

[0043] In the k-level HASH table, update the index record pointing to the second target data page to point to the target index page.

[0044] In a second aspect, the present invention further provides a method for verifying data consistency, including:

[0045] Generate corresponding tuples according to each row record in multiple data tables; wherein, the tuple includes the check code MD5 of the corresponding row record, the row ROWID of the corresponding row record, and the flag number FLAG of the data table where the corresponding row record is located;

[0046] Perform collision detection on all tuples in multiple data tables in a difference table;

[0047] After collision detection on all tuples in multiple data tables is completed, use the formed difference table as the final difference table, and obtain a data consistency result according to the final difference table; wherein, at least one of the multiple data tables is stored using the data storage method described in the first aspect, and / or the difference table in the method for verifying data consistency is stored using the data storage method described in the first aspect.

[0048] The performing collision detection in the difference table specifically includes:

[0049] Respectively use each tuple as the first tuple, and search in the difference table to find whether there is a second tuple whose MD5 is the same as that of the first tuple and whose FLAG is different; if the second tuple exists in the difference table, delete the second tuple from the difference table, otherwise, insert the first tuple into the difference table.

[0050] In a third aspect, the present invention further provides a device for verifying data consistency, which is used to implement the method for verifying data consistency described in the second aspect. The device includes:

[0051] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the processor to execute the method for verifying data consistency described in the first aspect.

[0052] In a fourth aspect, the present invention further provides a non-volatile computer storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by one or more processors to complete the method for verifying data consistency described in the first aspect.

[0053] The present invention sets up a multi-level HASH table and generates an index record chain pointing to the target data page in the multi-level HASH table, so that when searching for stored data, the corresponding target data page can be found through the index record chain, thereby narrowing the search range of the final data to a data page, reducing the resources and time consumed in the data search process, improving the data search efficiency, and thus improving the data comparison efficiency.

[0054] The present invention also calculates the checksum MD5 by taking the row record as a whole, so that when performing data comparison, it can be determined whether the entire row record is consistent only by comparing the checksum, thereby avoiding comparing each piece of data in the row record, and improving the comparison efficiency. And by performing collision detection on the tuples in different data tables in the difference table, it is possible to perform comparison using the same storage space in the difference table through data insertion and deletion in the difference table, thereby reducing memory occupancy and improving the data comparison efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments of the present invention. Obviously, the following described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0056] Figure 1 is a flowchart of a data storage method provided by an embodiment of the present invention;

[0057] Figure 2 is a storage structure diagram of a data storage method provided by an embodiment of the present invention;

[0058] Figure 3 is a storage structure diagram of a data storage method provided by an embodiment of the present invention;

[0059] Figure 4 is a flowchart of a data storage method provided by an embodiment of the present invention;

[0060] Figure 5 is a flowchart of a data storage method provided by an embodiment of the present invention;

[0061] Figure 6 is a flowchart of a data storage method provided by an embodiment of the present invention;

[0062] Figure 7 is a flowchart of a data storage method provided by an embodiment of the present invention;

[0063] Figure 8 It is a schematic flowchart of a method for verifying data consistency provided by an embodiment of the present invention;

[0064] Figure 9 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0065] Figure 10 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0066] Figure 11 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0067] Figure 12 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0068] Figure 13 It is a schematic flowchart of a method for verifying data consistency provided by an embodiment of the present invention;

[0069] Figure 14 It is a schematic flowchart of a method for verifying data consistency provided by an embodiment of the present invention;

[0070] Figure 15 It is a schematic flowchart of a method for verifying data consistency provided by an embodiment of the present invention;

[0071] Figure 16 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0072] Figure 17 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0073] Figure 18 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0074] Figure 19 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0075] Figure 20 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0076] Figure 21 It is a data schematic diagram in the application of a method for verifying data consistency provided by an embodiment of the present invention;

[0077] Figure 22It is a schematic structural diagram of a device for verifying data consistency provided by an embodiment of the present invention. Detailed implementation manners

[0078] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0079] In the description of the present invention, the orientation or positional relationship indicated by terms such as "inner", "outer", "longitudinal", "transverse", "upper", "lower", "top", "bottom", etc. is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention rather than requiring the present invention to be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.

[0080] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0081] Embodiment 1:

[0082] During the data comparison process, a large number of data search operations are usually required. However, in the prior art, data is usually stored through a conventional data table, and each data search operation needs to traverse all the data in the table, making the efficiency of data search extremely low, resulting in low efficiency in the data comparison process. To solve this problem, Embodiment 1 of the present invention provides a data storage method. The data storage structure used in the data storage method includes multiple levels of HASH tables. Among them, each level of HASH table consists of at least one index page, and the index page is used to store index records. Each index record corresponds to a HASH bucket, and the HASH bucket points to the corresponding data page or the index page in the corresponding lower-level HASH table. The data page is used to store the actual data to be stored. Among them, the sizes of each index page and data page are fixed. The number of HASH tables in the data storage structure is not constant, but dynamically changes according to the data to be stored. As Figure 1 shown, the data storage method includes:

[0083] In step 201, according to the Key value of the data, calculate the 1st-level HASH value, 2nd-level HASH value,..., nth-level HASH value corresponding to the data in sequence until the number of the first data does not exceed the number of data that can be stored in a single data page; wherein, the first data is the data with the same nth-level HASH value.

[0084] Among them, the data page is a container with a fixed size, and those skilled in the art allocate corresponding storage space for it according to requirements analysis. When the size of each piece of data occupying the storage space is the same, the number of data that each data page can store is fixed.

[0085] The first data does not refer to a specific piece of data, but represents a general term for data with the same HASH value when calculating the HASH value of the corresponding level. For example, if there are 4 pieces of data, and when calculating the 3 - level HASH value, these 4 pieces of data have the same 3 - level HASH value, then these 4 pieces of data are all called the first data. When calculating HASH values of different levels, the corresponding first data may be different.

[0086] The above - mentioned step 201 can be regarded as the process of finally determining the value of n. That is, let n be 1, 2, … in turn until the number of data with the same n - level HASH value calculated does not exceed the number of data that a single data page can store, then the value of n is determined. In the subsequent steps 202 and 203, the value of n is a fixed value.

[0087] After the value of n is fixed, n represents the level number of the HASH value corresponding to when the number of the calculated first data does not exceed the number of data that a single data page can store. For example, assume that the number of data that a single data page can store is 5. When calculating the 4 - level HASH value, if the number of the first data with the same 4 - level HASH value calculated is 3, which does not exceed the number of data that a single data page can store, then the value of n is 4.

[0088] Each level of HASH value has different calculation rules, so that as the 1 - level HASH value to the n - level HASH value, the number of the corresponding first data calculated can decrease in turn. For example, select the last n1 bytes in the Key value of the data, and use the HASH value of the last n1 bytes as the HASH value of each level, where n1 increases as the level number of the HASH value increases. For example, use the HASH value of the last 1 byte in the Key value as the 1 - level HASH value, and use the HASH value of the last 3 bytes in the Key value as the 3 - level HASH value; in some special scenarios, such as when storing phone numbers, the combined bytes of the first n1 bytes and the last n2 bytes in the Key value of the data can also be selected to calculate the HASH values of each level.

[0089] In step 202, store the first data into the corresponding target data page so that all the data stored in the target data page have the same n - level HASH value.

[0090] One optional implementation manner is to store the first data into an empty data page and use this empty data page as the target data page.

[0091] In step 203, an index record chain pointing to the target data page is generated in the 1st-level HASH table, 2nd-level HASH table, ..., nth-level HASH table; where n is a positive integer.

[0092] Among them, the calculation rule of each level of HASH value also needs to match the storage space of the index page, so as to ensure that among the first data with the same k-level HASH value, the number of data with the same k+1-level HASH value does not exceed the number of index records that can be stored in a single index page, where k here represents any level number from 1 to n, and the HASH value of each level number needs to meet the above conditions. For example, in Figure 3 Among the data with the same 1st-level HASH value HASH(13), the number of its corresponding 2nd-level HASH values cannot exceed the number of index records that can be stored in PAGE

[201] , so that the index record of HASH(13) points to the only index page in the 2nd-level HASH table.

[0093] When the number of the first data with the same 1st-level HASH value calculated from the data with different Key values does not exceed the number of data that can be stored in a single data page, that is, when n is equal to 1, the data storage structure described in this embodiment includes a 1st-level HASH table and multiple pages.

[0094] For example Figure 2 As shown, at most 4 pieces of data can be stored in each data page, and the HASH value of the first 2 characters in the Key value of the data is used as the 1st-level HASH value; if it is calculated that the number of the first data with the same 1st-level HASH value does not exceed the number of data that can be stored in each data page, then the first data with the same 1st-level HASH value is stored in the same data page, and an index record pointing to the corresponding data page is stored in the 1st-level HASH table. For example, the data (1211, Value) is located in PAGE

[102] , and its 1st-level HASH value is HASH(12), then the index record of HASH(12) in the 1st-level HASH table points to PAGE

[102] . It should be noted that in order to show the hierarchical relationship between the index page and the data page, Figure 2 The index page and the data page are presented in two columns. In fact, PAGE

[200] is the index page of the 1st-level HASH table, and PAGE

[101] , PAGE

[102] and PAGE

[103] are the data pages of the 1st-level HASH table.

[0095] When continuing to store data, if the number of the first data with the same 1st-level HASH value exceeds the number of data that can be stored in a single data page, then on the basis of the 1st-level HASH table, a lower-level HASH table, that is, a 2nd-level HASH table, will be added. For exampleFigure 3 as shown, thus forming a dynamic multi-level HASH table.

[0096] Figure 3 and Figure 2 the same as, it adopts the form of separating the index page and the data page for presentation. In actual implementation, PAGE

[200] is the index page of the 1st-level HASH table, PAGE

[101] and PAGE

[102] are the data pages of the 1st-level HASH table, PAGE

[201] is the index page of the 2nd-level HASH table, and PAGE

[103] and PAGE

[104] are the data pages of the 2nd-level HASH table.

[0097] It should be noted that when n equals 1, the data storage structure described in this embodiment is equivalent to a general HASH table, and its storage method and operations such as data searching, insertion, and deletion are the same as those of a general HASH table. Only when n is greater than or equal to 2 and this embodiment has at least two levels of HASH tables, the particularity of the storage method described in this embodiment is reflected. Therefore, in subsequent embodiments, if not specially stated, it is described on the premise that n is greater than or equal to 2.

[0098] For example, when storing data (1311, value), where the Key value of this data is 1311. At this time, the number of the first data with the 1st-level HASH value of HASH(13) is 5, exceeding the number of data that can be stored in a single data page. Then, the 2nd-level HASH value is further calculated for the first data. The HASH value of the first 3 characters in the Key value of the data is used as the 2nd-level HASH value. Among them, the number of the first data with the 2nd-level HASH value of HASH(130) is 2, and the number of the first data with the 2nd-level HASH value of HASH(131) is 3, both of which do not exceed the number of data that can be stored in a single data page. Then, the data with the 2nd-level HASH value of HASH(130) and the data with the 2nd-level HASH value of HASH(130) are stored in different pages respectively, and a 2nd-level HASH table is added. Index records pointing to the corresponding data pages are stored in the 2nd-level HASH table, and in the 1st-level HASH table, index records pointing to the corresponding index page of the 2nd-level HASH table are generated, as Figure 3 shown, the index record of HASH(13) points to the index page PAGE

[201] , the HASH(130) index record in the index page PAGE

[201] of the 2nd-level HASH table points to PAGE

[103] , and the HASH(131) index record points to PAGE

[104] .

[0099] In this embodiment, by setting up a multi-level HASH table and generating an index record chain pointing to the target data page in the multi-level HASH table, when searching for the stored data, the corresponding target data page can be found through the index record chain, thereby narrowing the search range of the final data to one data page, reducing the resources and time consumed in the data search process, and improving the data search efficiency.

[0100] In actual use, for the convenience of searching, the data is usually stored in descending order according to the Key value. In this case, due to different Key value calculation rules, data with similar Key values may have different n-level HASH values, resulting in the occupation of more data pages. To address this problem, this embodiment provides the following preferred implementation manner, that is, according to the Key value of the data, the 1-level HASH value, 2-level HASH value,..., n-level HASH value corresponding to the data are calculated in sequence, specifically including:

[0101] When calculating the 1-level HASH value, a preset value is used as the target quantity, and the first target quantity of bytes is selected from the Key value of the data, and the corresponding 1-level HASH value is calculated according to the target quantity of bytes.

[0102] When calculating any k-level HASH value among the 2-n level HASH values, the target quantity used when calculating the k-1 level HASH value plus the preset increment is used as the target quantity when calculating the k-level HASH value.

[0103] The first target quantity of bytes is selected from the Key value of the data, and the corresponding k-level HASH value is calculated according to the target quantity of bytes.

[0104] Wherein, k is any positive integer value from 2 to n, that is, k is greater than or equal to 2 and less than or equal to n.

[0105] The target quantity does not refer to a specific value, but changes with the change of the level number of the calculated HASH value. For example, when calculating the 5-level HASH value, the target quantity corresponding to the 4-level HASH value plus the preset increment is used as the target quantity when calculating the 5-level HASH value.

[0106] The preset value and the preset increment are obtained by those skilled in the art through the analysis of the Key value of the data. Taking a 16-byte MD5 code 80C1854645C026F6 as an example, assuming the preset value is 8 and the preset increment is 1, the 1-level HASH value is HASH(80C18546), and the k-level HASH value is to select the first 8 + k bytes from the MD5 code, and use the HASH value of these 8 + k bytes as the k-level HASH value.

[0107] This embodiment selects a corresponding number of bytes at the front and calculates HASH values ​​at each level, so that the HASH values ​​at each level are strongly correlated with the Key value of the data. Data with similar Key values ​​usually have the same first few bytes, and the calculated HASH values ​​of the corresponding levels are also the same, so that similar data can be allocated to one data page as much as possible to avoid unnecessary waste of data pages.

[0108] As an optional implementation, the index record chain pointing to the target data page is generated in the 1st level HASH table, the 2nd level HASH table, ..., the nth level HASH table, such as Figure 4 As shown, specifically including:

[0109] Insert corresponding index records into the n-level HASH table, n-1-level HASH table, ..., 1-level HASH table in sequence. Specifically:

[0110] In step 301, when inserting a corresponding index record into the n-level HASH table, an n-level target index record is established in the corresponding index page of the n-level HASH table according to the n-level HASH value of the first data, and the n-level target index record points to the target data page.

[0111] The corresponding index page is the index page pointed to by the index record of the n-1 level HASH value of the first data in the n-1 level HASH table. When n is 1, there is no n-1 level HASH table, or the n-1 level HASH table does not store the index record of the n-1 level HASH value of the first data, then the index record of the n-1 level HASH value of the first data is stored in the same index page (this index page is usually a newly allocated empty index page), and this index page is the corresponding index page.

[0112] In step 302, when inserting a corresponding index record into any one of the k-level HASH tables of the 1 to n-1 level HASH tables, a k-level target index record is established in the corresponding index page of the k-level HASH table according to the k-level HASH value of the first data, and the k-level target index record points to the k+1-level target index page;

[0113] The corresponding index page here is the index page pointed to by the index record of the k-2 level HASH value of the first data in the k-2 level HASH table. When k is 2, there is no k-2 level HASH table, or the k-2 level HASH table has not yet stored the index record of the k-2 level HASH value of the first data, then the index record of the k-1 level HASH value of the first data is stored in the same index page (this index page is usually a newly allocated empty index page), and this index page is the corresponding index page.

[0114] Wherein, the (k + 1)-level target index page is the index page where the (k + 1)-level target index record is located in the (k + 1)-level HASH table.

[0115] Wherein, k is any positive integer value from 1 to n - 1, that is, k is greater than or equal to 1 and k is less than or equal to n.

[0116] In actual use, after data is stored, data usage is usually also involved, including data search operations, data insertion operations, data deletion operations, etc. For data search operations, the present embodiment provides the following optional implementation manners, that is, the method further includes a data search operation, and the data search operation specifically includes:

[0117] According to the Key value of the data, search for the corresponding index record chain.

[0118] If the corresponding index record chain cannot be found, the final search result is that the data cannot be found.

[0119] If the corresponding index record chain is found, then according to the index record chain, obtain the corresponding target data page.

[0120] Search for the data in the target data page, and use the result obtained by searching for the data in the target data page as the final search result.

[0121] Wherein, the step of searching for the corresponding index record chain according to the Key value of the data specifically includes:

[0122] According to the Key value of the data, calculate the 1-level HASH value, 2-level HASH value,..., n-level HASH value corresponding to the data in sequence, and search for the corresponding index record. As Figure 5 shown, specifically includes:

[0123] In step 401, when calculating any k-level HASH value among the 1 to n-level HASH values and searching for the corresponding index record, according to the k-level HASH value, search for the corresponding k-level target index record in the k-level target index page. If the corresponding k-level target index record is found and the k-level target index record points to the corresponding (k + 1)-level index page, then use the (k + 1)-level index page as the (k + 1)-level target index page to search for the (k + 1)-level target index record.

[0124] In step 402, until no corresponding index record can be found, or the corresponding index record found points to the corresponding data page.

[0125] In step 403, if the corresponding index record cannot be found, the search result is that the corresponding index record chain cannot be found; if the corresponding index record is found and the index record points to the corresponding data page, the data page is used as the target data page corresponding to the index record chain.

[0126] Where n represents the highest level number of the HASH tables currently available in the data storage structure. For example, if there are 3 HASH tables in the current data storage structure, the value of n is 3.

[0127] After providing the specific implementation of the data insertion operation, this implementation also provides an alternative implementation for the data insertion operation, that is, the method further includes the data insertion operation, and the data insertion operation is as Figure 6 shown, and specifically includes:

[0128] In step 501, according to the Key value of the data, the corresponding index record chain is searched.

[0129] In step 502, if the corresponding index record chain cannot be found, the data is stored in the corresponding first target data page; calculate the 1-level HASH value corresponding to the data, and establish a corresponding index record in the 1-level HASH table according to the 1-level HASH value, and the index record points to the first target data page; the first target data page is usually a newly allocated empty data page, so that the data stored in the first target data page has the same 1-level HASH value.

[0130] In step 503, if the corresponding index record chain is found, the corresponding second target data page is obtained according to the index record chain; if the second target data page is not full, the data is stored in the second target data page.

[0131] In step 504, if the second target data page is full, an empty data page is selected as the third target data page, part of the data in the second target data page is stored in the third target data page, and it is selected to store the data in the second target data page or the third target data page; update the index record chain pointing to the second target data page and the index record chain pointing to the third target data page.

[0132] The above step 504 is only an optional implementation. In the actual implementation process, there is also an optional implementation, that is, two new third target data pages can be created. If the index record pointing to the second target data page is located in the k-level HASH table, then according to the (k + 1)-level HASH value of the data in the second target data page and the (k + 1)-level HASH value of the data to be inserted, each data is respectively stored in the above two third target data pages, and the corresponding index records are stored in the (k + 1)-level HASH table. The index record pointing to the second target data page is modified to point to the corresponding index page in the (k + 1)-level HASH table, so as to form a corresponding index record chain. In this case, the second target data page is equivalent to being replaced by two third target data pages, and the original second target data page is no longer used.

[0133] Among them, the first target data page, the second target data page, and the third target data page do not refer to specific data pages, but are used to make the corresponding defined objects stand out from the same category, and are added for the convenience of describing two or more different objects in the same category, and should not be interpreted as having a further restrictive meaning. Here, it is to distinguish different ways of finding the three and different subsequent processes for the three.

[0134] The updating of the index record chain pointing to the second target data page and the index record chain pointing to the third target data page, as Figure 7 shown, specifically includes:

[0135] In step 601, if the index record pointing to the second target data page is located in any k-level HASH table among the 1st to nth level HASH tables, then a new empty index page is added to the k-level HASH table, and the empty index page is used as the target index page; where k is any positive integer value from 1 to n, that is, k is greater than or equal to 1 and less than or equal to n. When k = n, it means that the index record is located in the HASH table with the highest level number, and there is no lower-level HASH table, so a new lower-level HASH table is created, that is, a new (n + 1)-level HASH table is created, and the empty index page in the (n + 1)-level HASH table is used as the target index page.

[0136] Steps 601 and 602 can also be described as: find the k-level HASH table where the index record pointing to the second target data page is located. If k is equal to n, create a new (n + 1)-level HASH table, and use the index page in the (n + 1)-level HASH table as the target index page; if k is less than n, add a new empty index page to the (k + 1)-level HASH table, and use the empty index page as the target index page; here, k represents the level number of the found HASH table.

[0137] In step 603, according to the level-k HASH value corresponding to the data in the second target data page, a corresponding second index record is established in the target index page, and the second index record points to the second target data page.

[0138] In step 604, according to the level-(k + 1) HASH value corresponding to the data in the third target data page, a corresponding third index record is established in the target index page, and the third index record points to the third target data page.

[0139] In step 605, in the level-k HASH table, the index record pointing to the second target data page is updated to point to the target index page.

[0140] The storing part of the data in the second target data page to the third target data page and selecting to store the data in the second target data page or the third target data page specifically includes:

[0141] According to the Key value of the data, calculate the level-k HASH value corresponding to the data.

[0142] Store the data with the same level-k HASH value in one target data page.

[0143] As Figure 2 and Figure 3 shown, when a new data (1311, value) is inserted in the state shown in Figure 2 shown, the second target data page found for this data is PAGE

[103] , and PAGE

[103] is already occupied by existing data and cannot store data anymore.

[0144] And the index record pointing to PAGE

[103] is located in the level-1 HASH table, then a new level-2 HASH table is created, calculate the level-2 HASH values for the data in the second target data page. There are 2 data with the calculated level-2 HASH value of HASH(130) and 2 data with the calculated level-2 HASH value of HASH(131). At the same time, the level-2 HASH value corresponding to the newly inserted data (1311, value) is HASH(131). Therefore, the data with these two different level-2 HASH values are stored in PAGE

[103] and PAGE

[104] respectively, and index records pointing to these two data pages are established in the level-2 HASH table respectively, and the index record in the level-1 HASH table that originally pointed to PAGE

[103] is updated to point to the index page where HASH(130) and HASH(131) are located in the level-2 HASH table, that is, as Figure 3 shown as PAGE

[201] .

[0145] After providing the specific implementation of the data insertion operation, this implementation also provides an optional implementation for the data deletion operation, that is, the method further includes the data deletion operation, and the data deletion operation specifically includes:

[0146] According to the Key value of the data, find the corresponding index record chain.

[0147] According to the found index record chain, obtain the corresponding target data page, and delete the data from the target data page.

[0148] Determine whether there is any other data stored in the target data page.

[0149] If there is no other data stored in the target data page, then in the k-level HASH table storing the target index record pointing to the target data page, delete the target index record pointing to the target data page; here, k refers to a specific value, that is, the level number of the HASH table storing the target index record pointing to the target data page.

[0150] When k is greater than 1, further determine whether the target index page still stores other index records, whether it only stores 1 index record, and whether there are other index pages in the k-level HASH table except the target index page.

[0151] If the target index page does not store other index records, then delete the index record pointing to the target index page in the (k - 1)-level HASH table.

[0152] If the target index page only stores 1 first index record, then update the index record pointing to the target index page in the (k - 1)-level HASH table to point to the data page pointed to by the first index record.

[0153] Furthermore, if there are other index pages in the k-level HASH table except the target index page, then delete the target index page from the k-level HASH table.

[0154] If there are no index pages in the k-level HASH table except the target index page, then delete the k-level HASH table. k is a positive integer, and k is greater than 1 and less than n.

[0155] Thus, through the change of the data volume, dynamically change the number of HASH tables to achieve a multi-level dynamic HASH table.

[0156] It should be noted that for the convenience of description, the parameter k is used multiple times in this embodiment. k is used to represent a single round in an iterative process with multiple rounds, or a single object among multiple objects. Depending on the different execution processes to which it is applied, the value range or meaning represented by k may also be different. The different value ranges of k in different execution processes should not be regarded as contradictory to each other. For example, in some execution processes of the above embodiment, the value range of k is 2 to n, while in other execution processes, the value range of k is 1 to n - 1. They are all set accordingly considering the clarity of the corresponding execution process description, and the k values in different execution processes do not affect each other.

[0157] Embodiment 2:

[0158] In the prior art, data comparison requires storing the sorted data results in memory and comparing the data item by item in memory, resulting in high memory occupancy and poor comparison efficiency. To solve this problem, this embodiment provides a method for verifying data consistency, as Figure 8 shown, including:

[0159] In step 701, corresponding tuples are generated according to each row record in multiple data tables; wherein, the tuple includes the check code MD5 of the corresponding row record, the ROWID of the corresponding row record, and the flag number FLAG of the data table where the corresponding row record is located.

[0160] In step 702, for all tuples in multiple data tables, collision detection is performed in the difference table. Specifically, each tuple is used as the first tuple, and it is searched in the difference table whether there is a second tuple whose MD5 is the same as that of the first tuple and the FLAG is different; if the second tuple exists in the difference table, the second tuple is deleted from the difference table, otherwise, the first tuple is inserted into the difference table.

[0161] In step 703, after the collision detection of all tuples in multiple data tables is completed, the formed difference table is used as the final difference table, and according to the final difference table, a data consistency result is obtained. Wherein, at least one of the multiple data tables is stored using the data storage method described in Embodiment 1, and / or the difference table in the method for verifying data consistency is stored using the data storage method described in Embodiment 1.

[0162] Among them, the collision detection in the difference table specifically includes:

[0163] Using each tuple as the first tuple, search in the difference table to check if there is a second tuple whose MD5 is the same as that of the first tuple but the FLAG is different. If the second tuple exists in the difference table, delete the second tuple from the difference table; otherwise, insert the first tuple into the difference table.

[0164] For example, there are two data tables, namely the left table and the right table. The data in these two tables is as Figure 9 shown. Calculate the MD5 value as a whole based on the contents of all columns in the row record to obtain the checksum MD5 of the corresponding row record. Set corresponding flag numbers FLAG for the left table and the right table. For example, set the FLAG of the left table to L and the FLAG of the right table to R. Then, based on the row where the row record is located, that is, Figure 9 the ROWID in, generate the corresponding tuple. Taking the calculation of the 16-byte MD5 value as the checksum of the row record as an example, according to Figure 9 the row records in the left table and the right table shown, calculate the corresponding tuples as Figure 10 shown. It can be seen that when the row record with ROWID = 1 in the left table is the same as the row record with ROWID = 1 in the right table, the checksum MD5 in the respective generated tuples is also the same. When the row record with ROWID = 2 in the left table is different from the row record with ROWID = 2 in the right table, the checksum MD5 in the respective generated tuples is also different.

[0165] Before performing data consistency verification, create an empty difference table to store tuples. When performing collision detection, such as performing collision detection on multiple tuples with ROWID = 1 in the difference table. When searching for the tuple (99896AF09545985A, 1, L) in the left table as the first tuple in the difference table, at this time, the difference table is empty and no second tuple can be found, so insert (99896AF09545985A, 1, L) into the difference table, that is, the tuple (99896AF09545985A, 1, L) exists in the difference table. Then, continue to search for the tuple (99896AF09545985A, 1, R) in the right table as the first tuple in the difference table. At this time, since the tuple (99896AF09545985A, 1, L) already exists in the difference table, search in the difference table to obtain a second tuple whose MD5 is the same as that of the first tuple but the FLAG is different, and then delete the second tuple in the difference table, that is, delete the tuple (99896AF09545985A, 1, L) from the difference table, and the difference table returns to an empty table, thus detecting that the two data with ROWID = 1 in the left table and the right table are consistent.

[0166] After continuing to perform collision detection on multiple tuples with ROWID 2 in the difference table and on multiple tuples with ROWID 3 in the difference table, the final difference table obtained is as Figure 11 shown. By looking up in the corresponding data table according to the ROWID and FLAG, the data consistency result is as Figure 12 shown, that is, (2, 11, USA, 2) in the left table is inconsistent with the data (2, 11, Russia, 1) with the same ROWID in the right table.

[0167] In this embodiment, first, by calculating the checksum MD5 with the row record as a whole, when performing data comparison, it is possible to obtain whether the whole row record is consistent only through the checksum comparison, thus avoiding comparing each piece of data in the row record and improving the comparison efficiency. Second, in this embodiment, by performing collision detection on the tuples in different data tables in the difference table, it is possible to use the same storage space in the difference table for comparison through data insertion and deletion in the difference table, thereby reducing memory occupancy and improving the efficiency of data comparison.

[0168] In actual situations, with the development of computer technology, multi-core CPUs are gradually applied. On this basis, in order to further improve the efficiency of data comparison, combined with the above embodiments, there is also a preferred implementation method, that is, the tuples corresponding to the row records in multiple data tables are collaterally compared in different threads.

[0169] For example, in a dual-core CPU environment, the collateral comparison of multiple tuples corresponding to the left table is placed in thread A, and the collateral comparison of multiple tuples corresponding to the right table is placed in thread B. Thread A and thread B are respectively controlled by different CPU cores, so that thread A and thread B can execute concurrently, thereby reducing the time required for data comparison.

[0170] Next, taking the difference table stored by using the data storage method described in Embodiment 1 as a specific example, the key steps in the method for verifying data consistency described in this embodiment will be specifically elaborated.

[0171] Among them, the process of converting the row record into a tuple has been described in the foregoing content and will not be elaborated here.

[0172] In the difference table, the checksum MD5 value of each tuple is used as the Key value of the data for storing the tuple. When performing collision detection on the first tuple in the difference table, as Figure 13 shown, it specifically includes:

[0173] In step 801, according to the checksum MD5 of the first tuple, the corresponding index record chain is searched.

[0174] In step 802, if the corresponding index record chain cannot be found, there is no second tuple in the difference table that has the same MD5 as the first tuple and a different FLAG.

[0175] In step 803, if the corresponding index record chain is found, the corresponding target data page is obtained according to the index record chain.

[0176] In step 804, in the target data page, according to the MD5 checksum of the first tuple, the corresponding data is searched. If the data with the same MD5 is found, it is further determined whether the FLAG of the data is different from that of the first tuple. If the FLAG is different from that of the first tuple, the data is the second tuple; otherwise, there is no second tuple in the difference table that has the same MD5 as the first tuple and a different FLAG.

[0177] Among them, the process of searching for the corresponding index record chain according to the MD5 checksum of the first tuple is implemented based on the same concept as in Embodiment 1 and will not be elaborated here.

[0178] When it is detected by collision that there is no corresponding second tuple in the difference table, the process of inserting the first tuple into the difference table is implemented based on the same concept as in Embodiment 1 and will not be elaborated here.

[0179] When it is detected by collision that there is a corresponding second tuple in the difference table, the process of deleting the second tuple from the difference table is implemented based on the same concept as in Embodiment 1 and will not be elaborated here.

[0180] After obtaining the final difference table, it is also necessary to reverse-lookup the data table through the final difference table to obtain the data with differences. When at least one of multiple data tables is stored using the data storage method described in Embodiment 1, the reverse-lookup of the data table specifically includes:

[0181] Determine the data table corresponding to the tuple according to the FLAG in the tuple in the final difference table.

[0182] According to the MD5 of the tuple, search for the corresponding index record chain, obtain the corresponding target data page according to the index record chain, and obtain the data corresponding to the tuple from the target data page, thereby obtaining the differential data.

[0183] Embodiment 3:

[0184] Based on the method described in Embodiment 2, the present invention combines specific application scenarios and elaborates the implementation process in the characteristic scenarios of the present invention through technical expressions in related scenarios.

[0185] This embodiment uses the method for verifying data consistency described in Embodiment 2 to verify data consistency, and takes the case where the difference table in the method for verifying data consistency is stored using the data storage method described in Embodiment 1 as an example for specific elaboration.

[0186] Taking Figure 9 the data comparison scenario between the left table and the right table shown as an example, the method for verifying data consistency, as Figure 14 shown, specifically includes:

[0187] In step 901, query the left table and the right table simultaneously, perform data conversion on the retrieved records one by one, calculate the MD5 value for the content of all columns of each record as a whole, and convert it together with the ROWID of this record and the FLAG of the data table to which this record belongs into a new tuple (MD5, ROWID, FLAG), and enter step 902.

[0188] In step 902, traverse the set of converted tuples. For each accessed tuple, obtain a tuple (MD5, ROWID, FLAG), and enter step 903 to perform collision detection in the difference table.

[0189] In step 903, initialize m = p, k = 1, perform collision detection on the tuple t1 (MD5, ROWID, FLAG) in the difference table, and enter step 904.

[0190] In step 904, calculate the HASH value based on the value of the first m bytes of the MD5 value to obtain the corresponding bucket (equivalent to the index record in Embodiment 1). If it points to the index page in the HASH table at the k + 1 level, then set m = p + 1, k = k + 1, and return to the beginning of step 904 to execute this step; if the bucket points to a data page or no corresponding bucket is obtained, then jump to step 905.

[0191] In step 905, if no corresponding bucket is obtained, then insert the tuple t1 into a new data page, and generate a bucket pointing to the new data page in the corresponding index page in the k-level HASH table. The corresponding index page is a new index page added to the k-level HASH table according to the k-level HASH value corresponding to the tuple t1; return to step 902 to continue traversing until all tuples in the set have been traversed, and then enter step 908.

[0192] If the bucket points to a data page, query in the data page whether there is a tuple with the same MD5 value using the MD5 value as the KEY. If there is a tuple with the same MD5 value, then enter step 906; if not, then jump to 907.

[0193] In step 906, after querying the tuple t2 with the same MD5 in the data page, continue to check whether the FLAG of tuple t2 is the same as that of tuple t1. If they are the same, go to step 107; otherwise, delete the found tuple t2, return to step 902, and continue the traversal. When all tuples in the set have been traversed, go to step 908.

[0194] In step 907, store tuple t1 into the data page, return to step 902, and continue the traversal. When all tuples in the set have been traversed, go to step 908.

[0195] In step 908, when all tuple collision detections are completed, the data comparison ends. At this time, the data in the difference table is the tuple corresponding to the difference data between the left table and the right table. Reverse query the corresponding table records through the ROWID and FLAG in the tuple to obtain the complete difference data records.

[0196] Among them, in step 907, storing tuple t1 into the data page is as Figure 15 shown, and specifically includes:

[0197] In step 1001, check whether the data page is full. If it is full, go to step 1002; otherwise, directly insert tuple t1 into the data page.

[0198] In step 1002, if the data page is full, find the lower-level HASH table of the k-level HASH table where the index record pointing to the data page is located, that is, find the (k + 1)-level HASH table.

[0199] If there is a (k + 1)-level HASH table, add an index page to the (k + 1)-level HASH table.

[0200] If there is no such (k + 1)-level HASH table, create a new (k + 1)-level HASH table. Go to step 1003.

[0201] In step 1003, after recalculating the HASH value of the first (m + 1) bytes of the MD5 values of all tuples in the data page (equivalent to the (k + 1)-level HASH value), store the tuples with the same (k + 1)-level HASH value in the same data page, and insert tuple t1 into the corresponding data page, then go to step 1004.

[0202] At this time, part of the data in the second data page is allocated to the third data page, and tuple t1 is inserted into one of the second data page and the third data page according to the corresponding (k + 1)-level HASH value.

[0203] In step 1004, in the corresponding index pages of the k + 1 level HASH table, buckets pointing to the second data page and buckets pointing to the third data page are respectively established. The buckets in the k level HASH table pointing to the data page are modified to point to the corresponding index pages of the k + 1 level HASH table.

[0204] The following will be specifically described with the left and right tables Figure 9 shown below. The left and right tables each have 3 pieces of data, as Figure 9 shown.

[0205] In the above step 901, the left and right tables are queried in parallel. The values of columns C1, C2, and C3 in the tables are calculated to obtain the MD5 value, and the corresponding ROWID and FLAG are added. The new tuples after conversion are as Figure 10 shown.

[0206] Collision detection is performed on the tuple t1(99896AF09545985A, 1, L) in the difference table. At this time, the HASH table is at the first level. Then, the HASH value is calculated using the first 8 bits 99896AF0 of the MD5 value of t1 to obtain the data page pointed to by the corresponding bucket. In the data page, query whether there is a tuple with the same MD5 value as the MD5 value 99896AF09545985A of t1. At this time, there is none, so t1 is inserted into the PAGE. At this time, the records in the difference table are as Figure 16 shown.

[0207] Collision detection is performed on the tuple t2(99896AF09545985A, 1, R) in the difference table. At this time, the HASH table is at the first level. Then, the HASH value is calculated using the first 8 bits 99896AF0 of the MD5 value of t2 to obtain the PAGE pointed to by the corresponding bucket. In the PAGE, query whether there is a tuple with the same MD5 value as the MD5 value 99896AF09545985A of t2. At this time, the tuple t1 is queried.

[0208] Check whether the FLAGs of tuple t2 and t1 are the same. The result is different, so tuple t1 is deleted from the PAGE. At this time, the difference table is empty.

[0209] After performing collision detection on the tuples: t3(80C1854645C026F6, 2, L), t4(DC37A6CC6856DDDE, 2, R), t5(26CF54564D5F2237, 3, L), t6(26CF54564D5F2237, 3, L) in the same way, the final difference table is as Figure 11 shown.

[0210] The original records are retrieved based on the ROWID and FLAG of the differential data. For the tuple (80C1854645C026F6, 2, L), since the values of FLAG and ROWID are L and 2 respectively, the original record (2, 11, USA) is retrieved from the left table.

[0211] After all the differential data has been retrieved in reverse, all the original differential data records are obtained as Figure 12 shown.

[0212] This embodiment also provides an application scenario for the left and right tables. As Figure 17 shown, each of the left and right tables has millions of data records, and the data in the tables is listed in tabular form.

[0213] The left and right tables are queried in parallel, and the MD5 values are calculated for the values in columns C1, C2, and C3 in the tables. After adding the corresponding ROWID and FLAG, the new tuples after conversion are as Figure 18 shown.

[0214] Tuple t1 (99896AF09545985A, 1031, L) is subjected to collision detection in the HASH table. Starting from the first-level HASH table, the HASH value is calculated using the first 8 bits 99896AF0 of the MD5 value of t1 to obtain the data page pointed to by the corresponding bucket. If the data page is full, a new second-level HASH table is generated, and the address pointed to by the corresponding bucket of the first-level HASH table to which the PAGE belongs is modified to the address of the newly created HASH table; after recalculating the HASH values of the first 9 bits of the MD5 values of all the tuples inside the original data page, they are inserted into the newly created HASH table.

[0215] The HASH value is calculated using the first 9 bits 99896AF09 of the MD5 value of tuple t1 to obtain the first-level PAGE pointed to by the corresponding bucket of this first-level HASH table.

[0216] In this first-level PAGE, query whether there is a tuple with the same MD5 value using the MD5 value 99896AF09545985A of t1. At this time, there is none, so t1 is inserted into the data page. At this time, the records in the differential table are as Figure 19 shown.

[0217] Tuple t2 (99896AF09545985A, 216, R) is subjected to collision detection in the differential table. Starting from the first-level HASH table, the HASH value is calculated using the first 8 bits 99896AF0 of the MD5 value of t2 to obtain the corresponding bucket, and the corresponding bucket points to a second-level HASH table.

[0218] Then, the HASH value is calculated using the first 9 bits 99896AF0 of the MD5 value of t2 to obtain the corresponding bucket, and the corresponding bucket points to a data page.

[0219] In the data page, query whether there are tuples with the same MD5 value as the MD5 value 99896AF09545985A of t2. At this time, tuple t1 is queried.

[0220] Check whether the FLAGs of tuple t2 and t1 are the same. If the results are different, delete tuple t1 from the data page.

[0221] After all tuples are detected, the tuples in the difference table are the tuples corresponding to the difference data, as Figure 20 shown.

[0222] Based on the ROWID and FLAG of the difference data, all the original records are retrieved by reverse lookup, as Figure 21 shown.

[0223] Example 4:

[0224] As Figure 22 shown, it is a schematic structural diagram of the device for verifying data consistency according to an embodiment of the present invention. The device for verifying data consistency in this embodiment includes one or more processors 21 and a memory 22. Among them, Figure 22 Take one processor 21 as an example.

[0225] The processor 21 and the memory 22 can be connected through a bus or other means. Figure 11 Take the connection through the bus as an example.

[0226] The memory 22, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs and non-volatile computer-executable programs, such as the method for verifying data consistency in Embodiment 2 or Embodiment 3. The processor 21 executes the method for verifying data consistency by running the non-volatile software programs and instructions stored in the memory 22.

[0227] The memory 22 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 22 may optionally include a memory remotely disposed relative to the processor 21, and these remote memories may be connected to the processor 21 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0228] The program instructions / modules are stored in the memory 22, and when executed by the one or more processors 21, execute the method for verifying data consistency in the above Embodiment 2 or Embodiment 3.

[0229] It should be noted that, regarding the information interaction, execution process, etc. among the modules and units in the above-mentioned device and system, since they are based on the same concept as the method embodiment of the present invention, for specific content, reference can be made to the description in the method embodiment of the present invention, and details will not be elaborated here.

[0230] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium can include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.

[0231] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A data storage method, characterized in that, The data storage structure includes multi-level HASH tables, and the data storage method includes: According to the Key value of the data, calculate the 1st-level HASH value, 2nd-level HASH value,..., nth-level HASH value corresponding to the data in sequence until the number of the first data does not exceed the number of data that can be stored in a single data page; wherein, the first data are the data with the same nth-level HASH value; Store the first data into the corresponding target data page so that all the data stored in the target data page have the same nth-level HASH value; Generate an index record chain pointing to the target data page in the 1st-level HASH table, 2nd-level HASH table,..., nth-level HASH table; wherein, n is a positive integer; The step of calculating the 1st-level HASH value, 2nd-level HASH value,..., nth-level HASH value corresponding to the data according to the Key value of the data specifically includes: When calculating the 1st-level HASH value, use a preset value as the target quantity, select the target quantity of bytes from the Key value of the data, and calculate the corresponding 1st-level HASH value according to the target quantity of bytes; when calculating any kth-level HASH value among the 2nd to nth-level HASH values, use the target quantity used when calculating the (k - 1)th-level HASH value plus a preset increment as the target quantity used when calculating the kth-level HASH value; select the target quantity of bytes from the Key value of the data, and calculate the corresponding kth-level HASH value according to the target quantity of bytes; wherein, the value of k is an integer greater than or equal to 2 and less than or equal to n; The step of generating an index record chain pointing to the target data page in the 1st-level HASH table, 2nd-level HASH table,..., nth-level HASH table specifically includes: insert the corresponding index records into the nth-level HASH table, (n - 1)th-level HASH table,..., 1st-level HASH table in sequence. Specifically, when inserting the corresponding index record into the nth-level HASH table, establish an nth-level target index record in the corresponding index page of the nth-level HASH table according to the nth-level HASH value of the first data, and the nth-level target index record points to the target data page; when inserting the corresponding index record into any kth-level HASH table among the 1st to (n - 1)th-level HASH tables, establish a kth-level target index record in the corresponding index page of the kth-level HASH table according to the kth-level HASH value of the first data, and the kth-level target index record points to the (k + 1)th-level target index page; wherein, the (k + 1)th-level target index page is the index page where the (k + 1)th-level target index record is located in the (k + 1)th-level HASH table; wherein, the value of k is an integer greater than or equal to 1 and less than or equal to (n - 1).

2. The data storage method according to claim 1, characterized in that, The method further includes a data search operation, and the data search operation specifically includes: Search for the corresponding index record chain according to the Key value of the data; If the corresponding index record chain cannot be found, obtain the final search result that the data cannot be found; If the corresponding index record chain is found, obtain the corresponding target data page according to the index record chain; Search for the data in the target data page, and use the result of finding the data in the target data page as the final search result.

3. The data storage method according to claim 2, characterized in that, The searching for the corresponding index record chain according to the Key value of the data specifically includes: According to the Key value of the data, calculate the 1st-level HASH value, 2nd-level HASH value,..., nth-level HASH value corresponding to the data in sequence, and search for the corresponding index record. Specifically, When calculating any kth-level HASH value among the 1st to nth-level HASH values and searching for the corresponding index record, search for the corresponding kth-level target index record in the kth-level target index page according to the kth-level HASH value. If the corresponding kth-level target index record is found and the kth-level target index record points to the corresponding (k + 1)th-level index page, then use the (k + 1)th-level index page as the (k + 1)th-level target index page to search for the (k + 1)th-level target index record; until no corresponding index record is found, or the found index record points to the corresponding data page; where the value of k is an integer greater than or equal to 1 and less than or equal to n; If no corresponding index record is found, the search result is that no corresponding index record chain is found. If a corresponding index record is found and the index record points to the corresponding data page, then use the data page as the target data page corresponding to the index record chain.

4. The data storage method according to claim 1, characterized in that, The method further includes the insertion operation of the data. The insertion operation of the data specifically includes: Search for the corresponding index record chain according to the Key value of the data; If no corresponding index record chain is found, store the data in the corresponding first target data page; Calculate the 1st-level HASH value corresponding to the data, and establish a corresponding index record in the 1st-level HASH table according to the 1st-level HASH value. The index record points to the first target data page; If a corresponding index record chain is found, obtain the corresponding second target data page according to the index record chain; If the second target data page is not full, store the data in the second target data page; If the second target data page is full, select an empty data page as the third target data page, store some data in the second target data page into the third target data page, and select to store the data in the second target data page or the third target data page; Update the index record chain pointing to the second target data page and the index record chain pointing to the third target data page.

5. The data storage method according to claim 4, characterized in that, The updating of the index record chain pointing to the second target data page and the index record chain pointing to the third target data page specifically includes: If the index record pointing to the second target data page is in any kth-level HASH table among the 1st to nth-level HASH tables, add a new empty index page in the (k + 1)th-level HASH table, and use the empty index page as the target index page; where the value of k is an integer greater than or equal to 1 and less than or equal to n; when k = n, create an (n + 1)th-level HASH table. Establish corresponding second index records in the target index page according to the (k + 1)-level HASH values corresponding to the data in the second target data page, where the second index records point to the second target data page; Establish corresponding third index records in the target index page according to the (k + 1)-level HASH values corresponding to the data in the third target data page, where the third index records point to the third target data page; Update the index records in the k-level HASH table that point to the second target data page to point to the target index page.

6. A method for verifying data consistency, characterized in that, Comprising: Generate corresponding tuples according to each row record in multiple data tables; wherein, the tuple contains the checksum MD5 of the corresponding row record, the row ROWID of the corresponding row record, and the flag number FLAG of the data table where the corresponding row record is located; Perform collision detection on all tuples in multiple data tables in the difference table; Take the difference table formed after the collision detection of all tuples in multiple data tables is completed as the final difference table, and obtain the data consistency result according to the final difference table; wherein, at least one of the multiple data tables is stored using the data storage method described in any one of claims 1-5, and / or the difference table in the method for verifying data consistency is stored using the data storage method described in any one of claims 1-5.

7. The method for verifying data consistency according to claim 6, wherein The performing collision detection in the difference table specifically includes: Respectively use each tuple as the first tuple to search in the difference table for whether there is a second tuple whose MD5 is the same as that of the first tuple and whose FLAG is different; if the second tuple exists in the difference table, delete the second tuple from the difference table, otherwise, insert the first tuple into the difference table.

8. An apparatus for verifying data consistency, wherein The apparatus includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions, when executed by the processor, are used to execute the method for verifying data consistency described in claim 6 or 7.

Citation Information

Patent Citations

  • Method for storing data by adopting HASH chain and data writing and reading method

    CN112667858A

  • System and method for searching strings of records

    US20090012957A1