An improved method based on memory data comparison

Through the multi-threaded memory data comparison method, the hash sharding and rocksdb database are used to solve the problem of slow speed and system crash in massive data comparison, and efficient and accurate isomorphic and heterogeneous data comparison is achieved.

CN115408425BActive Publication Date: 2025-07-25BEIJING JET-TECH ZHICHENG TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210952302.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-09
Publication Date
2025-07-25
Estimated Expiration
2042-08-09

AI Technical Summary

Technical Problem

The existing technology has problems such as slow speed, system crash, high hardware resource requirements and heterogeneous data in the comparison of massive data.

Method used

The multi-threaded memory data comparison method is adopted, hash sharding and paging technology is used, and rocksdb file database and dynamic proxy technology is combined to achieve efficient comparison of isomorphic and heterogeneous data.

Benefits of technology

It improves the efficiency and accuracy of data comparison, avoids service downtime, supports massive data comparison under limited resources, and supports heterogeneous data comparison.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115408425B_ABST
    Figure CN115408425B_ABST
Patent Text Reader

Abstract

The present invention relates to an improved method based on in-memory data comparison. The method comprises the following steps: Step 1: Preset a data reading script, specify a comparison index column, and preset a data reading threshold; Step 2: Configure hash sharding parameters and configure a data storage threshold; Step 3: Create a temporary storage unit for target data in memory and read the target data; Step 4: Calculate the hash of the target data and store it in the corresponding storage unit in memory in a sharded manner; Step 5: Consume the target data and store it in the local file database rocksdb; Step 6: Create a temporary storage unit for source data in memory; Step 7: Monitor the source data in memory; Step 8: Dynamically create an instance of the result table using dynamic proxy technology and write the data comparison details. This technical solution uses multi-threading and an efficient in-memory queue to load data into memory in batches for comparison multiple times, and at the same time utilizes technologies such as data paging and sharding to achieve efficient data comparison and effectively avoid service downtime.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an improved method, specifically to a method for completing massive data comparison based on HASH calculation in cooperation with a file database, belonging to the field of computer technology. Background Art

[0002] In the context of the rapid development of information technology, all walks of life are facing an avalanche of massive data. Such a large amount of big data has brought both pressure and opportunities to various industries. How to use this important asset of data to carry out effective data analysis, and on the basis of deeply understanding and grasping one's own and even the market situation, more scientifically evaluate the business performance of enterprises and assess business risks is one of the important challenges faced by most enterprises at present, and massive data comparison is an important part of data analysis.

[0003] For the comparison of massive data, the general techniques are as follows:

[0004] 1. Conducting table association queries through database dblink, with slow speed;

[0005] 2. Processing through data file sharding, with high requirements for physical hardware CPU and memory, often resulting in system crashes and unable to meet the high-performance requirements of large data volumes;

[0006] 3. Conducting data comparison processing by calculating the md5 value of each row of data in the table, with a long time;

[0007] 4. Unable to achieve data comparison for heterogeneous data (different types of databases). Therefore, there is an urgent need for a new solution to solve the above technical problems. Summary of the Invention

[0008] The present invention precisely aims at the problems existing in the prior art, and provides an improved method based on in-memory data comparison. This technical solution can complete the method and system for massive data comparison of homogeneous and heterogeneous data under limited resources (CPU, memory), and to a certain extent solve the above technical problems.

[0009] To achieve the above purpose, the technical solution of the present invention is as follows. An improved method based on in-memory data comparison, characterized in that the method includes the following steps:

[0010] Step 1: Presetting a data reading script, specifying a comparison index column, and presetting a data reading threshold;

[0011] Step 2: Configuring hash sharding parameters and configuring a data storage threshold;

[0012] Step 3: Creating a target data temporary storage unit in memory and reading the target data;

[0013] Step 4: Calculate the hash of the target data and store it in the corresponding storage unit in memory in a sharded manner;

[0014] Step 5: Consume the target data and store it in the local file database rocksdb;

[0015] Step 6: Create a temporary storage unit for the source data in memory, read the source data, calculate its hash, and store it in the corresponding storage unit;

[0016] Step 7: Monitor the source data in memory. When the amount of data in the storage unit equals the threshold, compare the target data with the source data;

[0017] Step 8: Dynamically create an instance of the result table using dynamic proxy technology and write the details of the data comparison.

[0018] Among them, Step 1 is specifically as follows: Preset the SQL query script and file reading script for the target data source data, and specify the comparison index column (which can be a single column or a combination of multiple columns), and preset the data reading threshold. By using the data reading script preset in Step 1, it is possible to support the comparison of heterogeneous data.

[0019] Among them, Step 2 is specifically as follows: Configure the hash sharding parameter with a default value of 1000, configure the data storage threshold with a default value of 10000, and configure the data reading threshold with a default value of 1000000. Then, when the program starts, 1000 storage units will be created in memory by default, and the data storage threshold for each storage unit will be set to 10000. Each time, 1000000 data will be read. At the same time, the above three parameters can be dynamically adjusted according to the server resource situation. By dynamically configuring the hash sharding parameter, the storage unit threshold, and the data reading threshold in Step 2, the parameters can be flexibly adjusted according to different server resources, and it is possible to complete the comparison of massive data under limited resources.

[0020] Among them, Step 3 is specifically as follows: Start the target data reading thread; create the corresponding number of data temporary storage units in memory as the sharding parameter in Step 2. Read the target data through the preset script in Step 1, which can be database data or file data. For database data, the stream reading mode is adopted, and the data is read in batches according to the server resource situation, and the amount of data read each time does not exceed the data reading threshold in Step 2, so as to reduce the pressure of reading database data; if it is file data, the IO buffer is used for reading, and the configuration of the reading buffer size is adjusted according to the resource situation. This solution can dynamically configure different parameters according to different server memory resources to reduce the risk of program crash caused by insufficient server memory.

[0021] Among them, step 4 is specifically as follows: For the read target data, calculate the hash value according to the comparison index column specified in step 1, and perform sharding through this hash value, and store it in the storage unit corresponding to step 3 in a key-value manner. When the data volume in a certain storage unit is equal to the storage threshold in step 2, wait for the thread in step 5 to start for data consumption. After data consumption, continue with step 3 until all the target data is read. The production-consumption mode is adopted to improve the data comparison efficiency.

[0022] Among them, step 5 is specifically as follows: Start the target data storage thread, create the number of RocksDB file database instances corresponding to the hash sharding parameters configured in step 2, monitor the target data volume in each storage unit in the memory. If it is equal to the threshold, take it out from the memory and store it in the corresponding file database in a key-value manner. Storing the data in the local file database can effectively prevent the problem of memory crash caused by excessive data volume. At the same time, this file database supports batch query by key value, which can improve the data query and comparison efficiency.

[0023] Among them, step 6 is specifically as follows: Start the source data reading thread, and the process of reading the source data is the same as that of step 3.

[0024] Among them, step 7 is specifically as follows: Start the data comparison thread, monitor the source data volume in each storage unit of the source data in the memory. If it is equal to the threshold, take it out from the memory, calculate the hash of the specified comparison index column and perform a loop comparison with the data in the corresponding RocksDB file database. The comparison logic preferentially performs data comparison row by row. If the comparison is successful, record the data. If the comparison fails, perform column-by-column comparison on the data until all column data is compared and then record the data. Through step 7 of this solution, the data is first compared horizontally, and the abnormal data is compared vertically again, which can effectively improve the data comparison efficiency and at the same time improve the accuracy of data comparison.

[0025] Among them, step 8 is specifically as follows: To solve the problem of storing the details of the comparison results of massive data. The present invention will use the dynamic proxy technology cglib to dynamically create MongoDB instances according to the case IDs of the preset scripts, and write the comparison results into the corresponding MongoDB database. Improve the query efficiency of the result details.

[0026] Compared with the prior art, the present invention has the following advantages: 1) The technical solution uses multi-threading and an efficient memory queue to load data into memory in batches for comparison multiple times, and at the same time utilizes technologies such as data paging and sharding to achieve efficient data comparison and effectively avoid service downtime; 2) By presetting the data reading script in step 1, it can support the comparison of heterogeneous data, preset the target data and source data reading scripts, and specify the comparison index column; 3) By dynamically configuring hash sharding parameters, storage unit thresholds, and data reading thresholds in step 2, the parameters can be flexibly adjusted according to different server resources, and it can achieve the comparison of massive data under limited resources; 4) Through step 7 of this solution, the data is first compared horizontally, and the abnormal data is compared vertically again, which can effectively improve the efficiency of data comparison and at the same time improve the accuracy of data comparison; 5) This solution supports the reading and comparison of homogeneous (same type of database) and heterogeneous (different types of databases, databases and files) data; 6) The hash sharding technology is used to store the source data content into an efficient memory queue, and at the same time the hash sharding technology is used to store the target data into a local file database. The multi-threading technology is used to compare the source data in memory with the target data in the file database, reducing the consumption of memory resources; 7) The dynamic proxy technology is used to dynamically generate mongodb instances and store the corresponding case comparison detailed data. Thus, the problem of low query efficiency caused by the unified storage of massive comparison result details is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the overall process of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] To deepen the understanding of the present invention, the following detailed description of this embodiment will be made with reference to the accompanying drawings.

[0029] Embodiment 1: Refer to Figure 1 An improved method for memory data comparison, the method includes the following steps:

[0030] Step 1: Preset the data reading script, specify the comparison index column, and preset the data reading threshold;

[0031] Step 2: Configure hash sharding parameters and configure the data storage threshold;

[0032] Step 3: Create a temporary storage unit for target data in memory and read the target data;

[0033] Step 4: Calculate the hash of the target data and store it in the corresponding storage unit in memory by sharding;

[0034] Step 5: Consume the target data and store it in the local file database rocksdb;

[0035] Step 6: Create a temporary storage unit for source data in memory, read the source data, perform hash calculation, and store it in the corresponding storage unit.

[0036] Step 7: Monitor the source data in memory. When the data volume in the storage unit equals the threshold, compare the target data with the source data.

[0037] Step 8: Dynamically create an instance of the result table using dynamic proxy technology and write the data comparison details.

[0038] Among them, Step 1 is as follows: Preset the SQL query script and file reading script for the target data source data, and specify the comparison index column (which can be a single column or a combination of multiple columns), and preset the data reading threshold. By means of the data reading script preset in Step 1, the comparison of heterogeneous data can be supported.

[0039] Among them, Step 2 is as follows: Configure the hash sharding parameter to be 1000 by default, configure the data storage threshold to be 10000 by default, and configure the data reading threshold to be 1000000 by default. Then, when the program starts, 1000 storage units will be created in memory by default, and the data storage threshold for each storage unit will be set to 10000, and 1000000 data will be read each time. At the same time, the above three parameters can be dynamically adjusted according to the server resource situation. By dynamically configuring the hash sharding parameter, the storage unit threshold, and the data reading threshold in Step 2, the parameters can be flexibly adjusted according to different server resources, and the comparison of massive data can be completed under limited resources.

[0040] Among them, Step 3 is as follows: Start the target data reading thread; create the corresponding number of data temporary storage units in memory as the sharding parameter in Step 2, and read the target data through the preset script in Step 1, which can be database data or file data. For database data, the stream reading mode is adopted, and the data is read in batches according to the server resource situation, and the data volume read each time does not exceed the data reading threshold in Step 2, so as to reduce the database data reading pressure; if it is file data, the IO buffer is used for reading, and the configuration of the read buffer size is adjusted according to the resource situation.

[0041] Among them, Step 4 is as follows: Calculate the hash value for the read target data according to the comparison index column specified in Step 1, and perform sharding through this hash value, and store it in the corresponding storage unit in Step 3 in a key-value manner. When the data volume in a certain storage unit equals the storage threshold in Step 2, wait for the thread in Step 5 to start for data consumption. After data consumption, continue with Step 3 until all the target data is read.

[0042] Among them, step 5 is specifically as follows: Start the target data storage thread, create the number of RocksDB file database instances corresponding to the hash sharding parameters configured in step 2, monitor the target data volume in each storage unit in the memory, if it is equal to the threshold, then take it out from the memory and store it in the corresponding file database.

[0043] Among them, step 6 is specifically as follows: Start the source data reading thread, and the process of reading the source data is the same as that in step 3.

[0044] Among them, step 7 is specifically as follows: Start the data comparison thread, monitor the source data volume in each storage unit of the source data in the memory, if it is equal to the threshold, then take it out from the memory, perform a hash calculation on the specified comparison index column and perform a loop comparison with the data in the corresponding RocksDB file database. The comparison logic preferentially performs data comparison row by row. If the comparison is successful, record the data. If the comparison fails, perform column-by-column comparison on the data until all column data is compared and then record the data. Through step 7 of this solution, the data is first compared horizontally, and the abnormal data is compared vertically again, which can effectively improve the efficiency of data comparison and at the same time improve the accuracy of data comparison;

[0045] Among them, step 8 is specifically as follows: To solve the problem of storing the details of the comparison results of massive data. The present invention will use the dynamic proxy technology cglib to dynamically create a MongoDB instance according to the case ID of the preset script, and write the comparison results into the corresponding MongoDB database.

[0046] It should be noted that the above embodiments are not used to limit the protection scope of the present invention. Equivalent transformations or substitutions made on the basis of the above technical solutions all fall within the protection scope of the claims of the present invention.

Claims

1. An improved method based on memory data comparison, characterized in that, The method includes the following steps: Step 1: Preset a data reading script, specify a comparison index column, and preset a data reading threshold; Step 2: Configure hash sharding parameters and configure a data storage threshold; Step 3: Create a temporary storage unit for target data in memory and read the target data; Step 4: Calculate the hash of the target data and store it in the corresponding storage unit in memory in a sharded manner; Step 5: Consume the target data and store it in the local file database rocksdb; Step 6: Create a temporary storage unit for source data in memory, calculate the hash of the read source data and store it in the corresponding storage unit; Step 7: Monitor the source data in memory. When the data volume in the storage unit is equal to the threshold, compare the target data with the source data; Step 8: Dynamically create an instance of the result table using dynamic proxy technology and write the data comparison details; Among them, Step 4 is specifically as follows: For the read target data, calculate the hash value according to the comparison index column specified in Step 1, and perform sharding through this hash value, and store it in the corresponding storage unit in Step 3 in a key-value manner. When the data volume in a storage unit is equal to the storage threshold in Step 2 or the end flag is read, wait for the thread in Step 5 to start for data consumption. After data consumption, continue with Step 3 until all the target data is read and consumed; Step 5 is specifically as follows: Start the target data storage thread, create an instance of the rocksdb file database corresponding to the number of hash sharding parameters configured in Step 2, monitor the target data volume in each storage unit in memory. If it is equal to the threshold or the end flag is read, take it out from memory and store it in the corresponding file database until all the target data is stored in the rocksdb file database; Step 7 is specifically as follows: Start the data comparison thread, monitor the source data volume in each storage unit in memory. If it is equal to the threshold or the data end flag is read, take it out from memory, calculate the hash of the specified comparison index column and perform a loop comparison with the data in the corresponding rocksdb file database. The comparison logic preferentially performs data comparison row by row. If the comparison is successful, record the data. If the comparison fails, perform column-by-column comparison until all column data is compared and then record the data.

2. The improved method based on in-memory data comparison according to claim 1, wherein, Step 1 is specifically as follows: Preset the sql query script for the target data source and the file reading script, specify the comparison index column, and preset the data reading threshold.

3. The improved method based on memory data comparison according to claim 2, characterized in that, Step 2 is specifically as follows: Configure the hash sharding parameter to default 1000, configure the data storage threshold to default 10000, and configure the data reading threshold to default 1000000. Then when the program starts, 1000 storage units will be created in memory by default, and the data storage threshold for each storage unit is set to 10000, and 1000000 data is read each time.

4. The improved method based on memory data comparison according to claim 3, wherein Step 3 is specifically as follows: Start the target data reading thread; Create the corresponding number of data temporary storage units in the memory according to the sharding parameters in Step 2, and read the target data through the preset script in Step 1. The target data can be database data or file data. For database data, the stream reading mode is adopted, and the data is read in batches according to the server resource situation, and the amount of data read each time does not exceed the data reading threshold in Step 2, so as to reduce the database data reading pressure; If it is file data, then the IO buffer is used for reading, the configuration of the reading buffer size is adjusted according to the resource situation, and at the same time, the target data reading end flag is set.

5. The improved method based on in-memory data comparison according to claim 4, wherein Step 6 is specifically as follows: Start the source data reading thread, and the source data reading process is the same as that in Step 3.

6. The improved method based on memory data comparison according to claim 5, wherein, Step 8 is specifically as follows: Use the dynamic proxy technology cglib to dynamically create a mongodb instance according to the case ID of the preset script, and write the comparison result into the corresponding mongodb database.

Citation Information

Patent Citations

  • Data parallel processing method, device and system suitable for big data set

    CN111198847A

  • Data comparison method and related equipment

    CN112395276A