Data pre-query method, device and equipment and computer readable storage medium
Patent Information
- Application Number
- CN202311386156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-24
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-10-24
AI Technical Summary
[0002]在对大规模数据的查询场景中(例如在大型数据库中查找某项数据、网络爬虫对待爬取网页通过查询已爬取网址数据库来避免重复爬取、邮箱系统对接收到的邮件通过查询垃圾邮件库来判断是否为垃圾邮件等),直接在相应数据库中精确查找待查询数据,耗时极长,效率较低
[0060] The data pre-query method, apparatus, device, and computer-readable storage medium provided in this invention, by adaptively adjusting the length of the query array based on the amount of data stored in the current database, can make the hash mapping positions of the stored data in the query array at the current level more sparse compared to a Bloom filter that does not adjust the query array length. This reduces the probability of hash collisions between the data to be queried and the stored data during the query process, thus lowering the pre-query false positive rate for the data to be queried. Furthermore, by adaptively increasing the total number of levels of the query array, the different hash mapping positions of the same stored data in the query arrays at more levels change the distribution of hash mapping positions of each stored data from a macroscopic perspective. This makes it difficult for data to be queried that does not exist in the database to simultaneously cause hash collisions with the stored data in multiple levels of the query array during the pre-query process, thus filtering out data that does not exist in the database and concluding that it definitely does not exist in the database, further reducing the pre-query false positive rate for the data to be queried. Therefore, this effectively improves the pre-query efficiency for the data to be queried.
Smart Images

Figure CN117390022B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more particularly to a data pre-query method, apparatus, device, and computer-readable storage medium. Background Technology
[0002] In scenarios involving large-scale data queries (such as searching for specific data in a large database, web crawlers avoiding duplicate crawling of web pages by querying databases of already crawled URLs, and email systems determining whether received emails are spam by querying spam databases), directly and precisely searching for the data to be queried in the corresponding database is extremely time-consuming and inefficient. Summary of the Invention
[0003] This invention provides a data pre-query method, apparatus, device, and computer-readable storage medium to improve the query efficiency of data to be queried.
[0004] This invention provides a data pre-query method, including:
[0005] For any L-th level preset hash function, determine the L-th level hash value corresponding to the data to be queried; the initial value of L is 1; there are multiple L-th level preset hash functions;
[0006] Based on the query position of each Lth-level hash value in the Lth-level query array, determine whether the data to be queried necessarily does not exist in the database;
[0007] If it cannot be determined whether the data to be queried must not exist in the database, and it is determined that the query termination condition is not met, then increment L by 1, and return to the step of determining the L-th level hash value corresponding to the data to be queried for any L-th level preset hash function;
[0008] The query arrays at each level are determined as follows:
[0009] Based on the amount of data in the database, determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array, and set the newly added query position in the query array as an invalid query value;
[0010] For any stored data in the database, starting from the first level and continuing to the last level L0 corresponding to the stored data, the following steps are executed level by level to set multiple query positions corresponding to the stored data in each level of the query array:
[0011] For any of the Lth-level preset hash functions, determine the Lth-level hash value corresponding to the stored data based on the Lth-level preset hash function;
[0012] For any of the Lth level hash values, set the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value;
[0013] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z}, L0≤L max L max To query the total number of levels in the array.
[0014] Optionally, determining whether the data to be queried necessarily does not exist in the database based on the query position corresponding to each of the Lth-level hash values in the Lth-level query array includes:
[0015] If each of the Lth level hash values has an invalid query value in the corresponding query position of the Lth level query array, it is determined that the data to be queried must not exist in the database.
[0016] Otherwise, it cannot be determined whether the data to be queried necessarily does not exist in the database.
[0017] As an optional implementation, the preset hash functions corresponding to any two levels are different.
[0018] As another optional implementation, for the same target data, the function input values corresponding to the target data are different at different levels; wherein, the L-th level hash value corresponding to the target data is determined according to the L-th level function input value corresponding to the target data;
[0019] The target data is either the data to be queried or the stored data.
[0020] Furthermore, as an optional implementation, the preset hash functions corresponding to any two levels are the same;
[0021] For the same preset hash function, the input value of the first-level function corresponding to the target data is the target data; the input value of the non-first-level function corresponding to the target data is the hash value of the previous level corresponding to the target data.
[0022] Furthermore, as another optional implementation, the preset hash function is the same for any two levels;
[0023] For the same preset hash function, the Lth level input value corresponding to the target data is obtained by performing an algebraic transformation on the target data, and the algebraic transformation rules are different for any two levels.
[0024] Optionally, determining the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database specifically includes:
[0025] Determine the first amount of data currently stored in the database, and the second amount of data previously stored in the database;
[0026] If the difference between the first data volume and the second data volume is less than a preset first increment threshold, then the total number of levels is increased;
[0027] If the difference between the first data volume and the second data volume is greater than or equal to the preset first incremental threshold, then the length is increased;
[0028] The method further includes:
[0029] If the length is increased, the number of preset hash functions is increased.
[0030] Optionally, if the total number of levels determined this time is greater than the total number of levels determined last time, the amount of increase in the total number of levels is determined based on the total number of levels before the increase in the query array levels;
[0031] Alternatively, the increase in the total number of stages may satisfy the following relationship:
[0032]
[0033] Where △L is the increase in the total number of levels, and L0 is the total number of levels in the query array before the increase;
[0034] Optionally, if the length determined this time is greater than the length determined last time, the amount of increase in length is determined based on the length of the query array before the increase, the first data volume, and the number of preset hash functions after the query array length is increased.
[0035] Alternatively, the increase in length may satisfy the following relationship:
[0036]
[0037] Wherein, △m is the length increase, n′ is the first data volume, k′ is the number of preset hash functions after the length of the query array is increased, and m0 is the length of the query array before the length is increased;
[0038] The preset hash function hash(x) added each time is:
[0039] hash(x) = x mod m0
[0040] Where x is the data to be queried or the stored data.
[0041] Optionally, the preset first incremental threshold is determined based on at least one of the device hardware performance and the first data volume.
[0042] Optionally, determining that the query termination condition is not met includes:
[0043] If each of the Lth-level hash values is determined to be a flag query position in the corresponding query position of the Lth-level query array, then the query termination condition is not met; wherein, the flag query position is the query position in the Lth-level flag array where the corresponding flag position is a valid flag value;
[0044] The flag arrays at each level are determined as follows:
[0045] Determine the flag arrays at each level, and set the newly added flag positions in the flag arrays to invalid flag values, wherein each query position in the L-th level query array uniquely corresponds to a flag position in the L-th level flag array;
[0046] For any stored data in the database, after setting the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value for any Lth level hash value, if it is determined that the Lth level is not the last level L0 corresponding to the stored data, then the flag position corresponding to the Lth level hash value is set as a valid flag value.
[0047] The last level L0 corresponding to any stored data is determined in the following way:
[0048] At least a portion of the stored data from all L-level mapping stored data that participated in determining the corresponding L-level hash value is selected as the L-level ending mapping stored data; where the starting value of L is 1; and the L-level mapping stored data is all stored data in the database;
[0049] For the Lth level ending the mapping storage data, determine that the Lth level is the last level L0 corresponding to the ending mapping storage data; if the Lth level is not the last level L0... max The step involves selecting at least a portion of the stored data from all L-level mapping data that are not the L-level end mapping data as the next level mapping data, incrementing L by 1, and returning the step of selecting at least a portion of the stored data from all L-level mapping data that participated in determining the corresponding L-level hash value as the L-level end mapping data.
[0050] Optionally, the query array is a bit array.
[0051] Optionally, the flag array is a bit array.
[0052] Based on the same inventive concept, embodiments of the present invention also provide a data pre-query device, comprising:
[0053] The setting unit is used to determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database, and to set newly added query positions in the query array as invalid query values; for any stored data in the database, starting from the first level to the last level L0 corresponding to the stored data, the following steps are executed level by level to set multiple query positions corresponding to the stored data in each level of the query array: for any L-th level preset hash function, the L-th level hash value corresponding to the stored data is determined based on the L-th level preset hash function; there are multiple L-th level preset hash functions; for any L-th level hash value, the query position corresponding to the L-th level hash value in the L-th level query array is set as a valid query value;
[0054] The pre-query unit is used to determine the L-th level hash value corresponding to the data to be queried for any of the L-th level preset hash functions; the initial value of L is 1; based on the query position corresponding to each L-th level hash value in the L-th level query array, it is determined whether the data to be queried must not exist in the database; if it cannot be determined that the data to be queried must not exist in the database, and it is determined that the query termination condition is not met, then L is incremented by 1, and the step of determining the L-th level hash value corresponding to the data to be queried for any of the L-th level preset hash functions is returned;
[0055] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z}, L0≤L max L max To query the total number of levels in the array.
[0056] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, including: a processor and a memory for storing processor-executable instructions;
[0057] The processor is configured to execute the instructions to implement the data pre-query method.
[0058] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing computer program code, which, when executed on a computer, causes the computer to perform the data pre-query method.
[0059] The beneficial effects of this invention are as follows:
[0060] The data pre-query method, apparatus, device, and computer-readable storage medium provided in this invention, by adaptively adjusting the length of the query array based on the amount of data stored in the current database, can make the hash mapping positions of the stored data in the query array at the current level more sparse compared to a Bloom filter that does not adjust the query array length. This reduces the probability of hash collisions between the data to be queried and the stored data during the query process, thus lowering the pre-query false positive rate for the data to be queried. Furthermore, by adaptively increasing the total number of levels of the query array, the different hash mapping positions of the same stored data in the query arrays at more levels change the distribution of hash mapping positions of each stored data from a macroscopic perspective. This makes it difficult for data to be queried that does not exist in the database to simultaneously cause hash collisions with the stored data in multiple levels of the query array during the pre-query process, thus filtering out data that does not exist in the database and concluding that it definitely does not exist in the database, further reducing the pre-query false positive rate for the data to be queried. Therefore, this effectively improves the pre-query efficiency for the data to be queried. Attached Figure Description
[0061] Figure 1 A schematic diagram illustrating the process of setting the bit array for a Bloom filter;
[0062] Figure 2 This is one of the schematic diagrams illustrating the pre-query process of a Bloom filter on the data to be queried.
[0063] Figure 3 The second diagram illustrates the pre-query process for data to be queried using a Bloom filter.
[0064] Figure 4 The third diagram illustrates the pre-query process for data to be queried using a Bloom filter.
[0065] Figure 5 This is one of the flowcharts illustrating the preparation process of the data pre-query method provided in this embodiment of the invention;
[0066] Figure 6 This is one of the flowcharts illustrating the pre-query process of the data pre-query method provided in this embodiment of the invention;
[0067] Figure 7 This is one of the partial flowcharts illustrating the pre-query process of the data pre-query method provided in this embodiment of the invention;
[0068] Figure 8 A second schematic flowchart illustrating the preparation process of the data pre-query method provided in this embodiment of the invention;
[0069] Figure 9 The second flowchart illustrates the pre-query process of the data pre-query method provided in this embodiment of the invention.
[0070] Figure 10 The third flowchart illustrating the preparation process of the data pre-query method provided in this embodiment of the invention;
[0071] Figure 11 This is a schematic diagram of the query array and flag array in the data pre-query method provided in the embodiments of the present invention;
[0072] Figure 12 This is a second partial flowchart illustrating the pre-query process of the data pre-query method provided in this embodiment of the invention.
[0073] Figure 13 The third flowchart illustrates the pre-query process of the data pre-query method provided in this embodiment of the invention.
[0074] Figure 14 This is one of the structural schematic diagrams of the data pre-query device provided in an embodiment of the present invention;
[0075] Figure 15 This is a second schematic diagram of the structure of the data pre-query device provided in an embodiment of the present invention;
[0076] Figure 16 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0077] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, the present invention will be further described below in conjunction with the accompanying drawings and embodiments. However, the exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided to make the present invention more comprehensive and complete, and to fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the figures denote the same or similar structures, and therefore repeated descriptions of them will be omitted. Terms describing position and direction in the present invention are illustrative based on the accompanying drawings, but changes can be made as needed, and all such changes are included within the scope of protection of the present invention. The accompanying drawings of the present invention are for illustrative purposes only and do not represent actual proportions.
[0078] It should be noted that specific details are set forth in the following description to provide a full understanding of the invention. However, the invention can be practiced in many ways other than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below. The following description is a preferred embodiment for carrying out the present application; however, the description is for the purpose of illustrating the general principles of the application and is not intended to limit the scope of the application. The scope of protection of this application shall be determined by the appended claims.
[0079] In scenarios involving large-scale data queries, the data to be queried includes both stored data existing in the corresponding database and data not existing in the database. For example, web crawlers avoid duplicate crawling by querying databases of already crawled URLs, and email systems check spam databases to determine if received emails are spam. In some cases, the probability that the data to be queried is not in the database is far higher than the probability that it is stored data. Therefore, performing precise queries on all the data to be queried in the database is time-consuming and inefficient.
[0080] For the above query scenarios, a Bloom filter can be used to pre-query the data to be queried. This can quickly identify most of the data to be queried that does not exist in the database and return non-existent results. It can also directly cancel the exact query operation for the data that does not exist in the database, thereby effectively reducing the workload of performing exact queries in the database and improving query efficiency.
[0081] A Bloom filter is a data pre-query algorithm used to quickly determine whether a query exists in a database. A Bloom filter consists of a bit array and multiple hash functions. All bits of the bit array are initially set to the invalid value 0. For example... Figure 1 As shown, when a piece of data D1 is added to the database, multiple hash functions (hash1, hash2, hash3) are used to map the stored data to multiple positions in the bit array, and the bits at these positions are set to the valid value 1. When it is necessary to query whether a piece of data D4 exists in the database, the same multiple hash functions are used to map the data to be queried to the corresponding positions in the bit array, as shown. Figure 2 As shown, if any corresponding bit is 0, then it is determined that the data to be queried definitely does not exist in the database; for example... Figure 3 As shown, if all bits at these positions are 1, it is determined that the data to be queried may exist in the database. Since the bit array used in the Bloom filter only occupies a small amount of memory, and the query time complexity of the Bloom filter is O(k), where k is the number of hash functions and is independent of the amount of data stored in the database, the Bloom filter has the advantages of high space efficiency and fast query time.
[0082] However, due to the potential hash collisions between the data to be queried and the stored data, Bloom filters have a certain false positive rate regarding the existence of data in the database. That is, they cannot definitively conclude that some data that does not exist in the database is non-existent. For example, ... Figure 4As shown, bits C1, C2, and C3 in the bit array of the Bloom filter are set to 1 by hash functions mapped to data D1, D2, and D3 stored in the database. During a query, a piece of data D4 that does not actually exist in the database happens to be mapped to bits C1, C2, and C3 in the bit array using multiple hash functions. The Bloom filter cannot determine whether the data D4 must exist in the database. Furthermore, as the amount of data stored in the database increases, the number of bits set to 1 in the bit array will continue to increase, and the probability of hash collisions between the data to be queried and the stored data will increase. In other words, the false positive rate of the Bloom filter will also increase with the amount of stored data.
[0083] This invention provides a data pre-query method, apparatus, device, and computer-readable storage medium to further reduce the false judgment rate during the pre-query process, reduce the amount of data to be queried in the database for subsequent precise queries, and improve query efficiency. The data pre-query method, apparatus, device, and computer-readable storage medium provided by this invention will be described in detail below with reference to the accompanying drawings.
[0084] This invention provides a data pre-query method, which consists of a preparation process and a pre-query process. For example... Figure 5 As shown, the preparation process includes:
[0085] S110. Determine the preset hash function, the total number of levels of the query array, and the length of each level of the query array according to the amount of data in the database, and set the newly added query position in the query array as an invalid query value.
[0086] In specific implementation, step S110 can be executed when the amount of data stored in the database changes; or certain execution conditions can be set (for example, setting the execution cycle of step S110, setting a threshold for changes in the amount of database data, or setting the execution command triggered by the user as an execution condition), and step S110 can be executed when the execution conditions are met.
[0087] In practice, to reduce the storage space occupied by the query array, it can be optionally set as a bit array, where each query position is represented by a binary bit. Accordingly, for each query bit, 0 can be used as an invalid query value, and 1 as a valid query value. The value of the query bit can be set in subsequent steps using a union operation.
[0088] In practice, if the initial amount of data stored in the database is small, a level 1 query array can be set up based on the initial amount of data stored in the database. Subsequently, as the amount of data stored in the database increases, the length and / or the number of levels of the query array can be adaptively increased based on the current amount of data stored in the database.
[0089] In practical implementation, for MySQL databases, the following can be done: Obtain the MySQL database log file (binlog) (including index files with the .index filename and log files with the .* filename, where * represents the sequence number). Extract Insert and Delete events from all Data Manipulation Language (DML) and Data Definition Language (DDL) statement events (excluding SELECT statements) within the log files. Store events containing only Insert events and none containing Delete events as data in the MySQL database and determine the corresponding data volume. Initially, the current data volume in the database can be determined by iterating through all binlog.* files, recording the position of the last binlog.* file, extracting the total number of Insert events (count1) and the total number of Delete events (count2), and using the difference (count1 - count2) as the initial data volume in the database. When adding new data to the database, the incremental binlog.* file can be determined based on the current position' of the last binlog.* file and the position' of the last binlog.* file determined last time. The total number of Insert events (count1') and the total number of Delete events (count2') in the incremental binlog.* file are counted. The difference between these two (count1' - count2') is added to the previously determined second data volume in the database to obtain the current determined first data volume in the database. Parsing the MySQL database log files (including full and incremental parsing) to count the data volume in the database can be achieved by setting up an adaptation proxy module. This module pre-configures the database connection configuration file, database driver, database Uniform Resource Locator (URL), username and password for encryption, path of the database log files, and other configuration information. It also sets parameters such as the time interval event (interval) that triggers the execution of step S110, and the variable Total for the first data volume. Furthermore, it writes a Java Database Connectivity Application Programming Interface (JDBC API) and uses a DriverManager to implement different drivers for different databases.The adaptation proxy module then loads the database driver into the Java Virtual Machine (JVM), manages the implemented registration class in the driver manager, and creates a Connection object to establish a connection with the database. Adjustments to the total number of levels in the query array and the length of each level can be made through the adjustment module. After determining the initial amount of data currently stored in the database, the adaptation proxy module periodically provides this initial amount of data to the adjustment module using a set interval event. This reduces code intrusion into the database project and improves versatility.
[0090] S120. Select one piece of stored data from the database without repeating any data.
[0091] If step S120 is successfully selected, then step S130 is executed; if all stored data has been selected, then the data pre-query preparation process ends. Afterwards, when data to be queried needs to be pre-queried, step S210 of the query process is executed.
[0092] S130, let L = 1.
[0093] S140. For any Lth-level preset hash function, determine the Lth-level hash value corresponding to the stored data based on the Lth-level preset hash function.
[0094] Among them, there are multiple preset hash functions at the Lth level.
[0095] In practice, the preset hash functions for different levels can be the same or not. For example, if the query array has 5 levels, the preset hash functions for all levels are hash1, hash2, and hash3; or, the preset hash functions for level 1 are hash1, hash2, and hash3, for level 2 they are hash4, hash5, and hash6, for level 3 they are hash7, hash8, and hash9, and so on.
[0096] S150. For any of the Lth level hash values, set the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value.
[0097] In the specific implementation process, the query position corresponding to the coordinates of the Lth level hash value can be directly used as the query position corresponding to the Lth level hash value. Of course, other methods can also be used to determine the query position corresponding to the Lth level hash value (for example, the query position corresponding to the coordinates obtained by multiplying the coordinates of the Lth level hash value by a certain fixed value can be used as the query position corresponding to the Lth level hash value). This embodiment of the invention does not impose too many limitations here.
[0098] In the specific implementation process, if the amount of stored data in the database increases, the mapping process of the Lth-level hash value corresponding to the Lth-level preset hash function can be skipped for the existing stored data. For example, for the first level, there are three preset hash functions hash1, hash2, and hash3. The existing stored data D1 has already been determined to correspond to three hash values hash1(D1), hash2(D1), and hash3(D1) in the first level during the previous preparation process, and the corresponding query positions have been set as valid query values. Therefore, the steps of determining hash values hash1(D1), hash2(D1), and hash3(D1) and setting the corresponding query positions as valid values can be skipped in this implementation. For example, for Level 1, in the previous preparation process, Level 1 corresponded to 3 preset hash functions hash1, hash2, and hash3. In the current preparation process, it corresponds to 4 preset hash functions hash1, hash2, hash3, and hash4. The existing stored data D1 has already been determined to correspond to 3 hash values hash1(D1), hash2(D1), and hash3(D1) in Level 1 during the previous preparation process, and the corresponding query positions have been set as valid query values. Therefore, in this case, we can directly skip the steps of determining hash values hash1(D1), hash2(D1), and hash3(D1) and setting the corresponding query positions as valid values. We only need to perform the steps of determining hash value hash4(D1) and setting the query position corresponding to hash value hash4(D1) as a valid query value.
[0099] S160. Determine whether the current level L is the last level L0 corresponding to the stored data.
[0100] If the result of step S160 is negative, proceed to step S170; if the result of step S160 is positive, proceed to step S120.
[0101] S170, Increment L by 1. Return to step S140.
[0102] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z}, L0≤L max L max To query the total number of levels in the array.
[0103] When performing a pre-query on the data to be queried, such as Figure 6 As shown, the pre-query process specifically includes the following steps:
[0104] S210, Let L = 1.
[0105] S220. For any of the Lth-level preset hash functions, determine the Lth-level hash value corresponding to the data to be queried.
[0106] S230. Based on the query position of each Lth-level hash value in the Lth-level query array, determine whether the data to be queried necessarily does not exist in the database.
[0107] Optionally, if each of the Lth-level hash values has an invalid query value at its corresponding query position in the Lth-level query array, it is determined that the data to be queried must not exist in the database; otherwise, it is determined that it is impossible to determine whether the data to be queried must not exist in the database. That is, if the query array is a multi-level query array, then in any level of the query array, not all query positions in the hash mapping of the data to be queried are valid query values. As long as there is an invalid query value at any query position, it can be directly determined that the data to be queried must not exist in the database.
[0108] If step S230 cannot determine whether the data to be queried must not exist in the database, proceed to step S240; if step S230 determines that the data to be queried must not exist in the database, proceed to step S260.
[0109] S240. Determine whether the query termination condition is met.
[0110] If the result of step S240 is yes, the pre-query process ends. Then, a precise query is performed on the data to be queried in the database. The process of performing a precise query on the data to be queried can be implemented using existing technology; since this part is not the focus of this embodiment, it will not be elaborated further.
[0111] If the result of step S240 is negative, proceed to step S250.
[0112] S250, Increment L by 1. Then return to step S220.
[0113] S260. Return the query result if the data to be queried does not exist.
[0114] The data pre-query method provided in this invention, by adaptively adjusting the length of the query array based on the current amount of data stored in the database, compared to a Bloom filter that does not adjust the query array length, can make the hash mapping positions of the stored data in the query array at the current level more sparse. This reduces the probability of hash collisions between the data to be queried and the stored data during the query process, thus lowering the pre-query false positive rate for the data to be queried. Furthermore, by adaptively increasing the total number of levels of the query array, the different hash mapping positions of the same stored data in the query arrays at more levels change the distribution of hash mapping positions of each stored data from a macroscopic perspective. This makes it difficult for data to be queried that does not exist in the database to simultaneously cause hash collisions with the stored data in multiple levels of the query array during the pre-query process, thus filtering out data that does not exist in the database and concluding that it definitely does not exist in the database. This also reduces the pre-query false positive rate for the data to be queried. Therefore, it effectively improves the pre-query efficiency for the data to be queried.
[0115] In this embodiment of the invention, the process of increasing the length of the query array is called horizontal expansion of the query array, and the process of increasing the total number of levels of the query array is called vertical expansion of the query array. During horizontal expansion, in addition to increasing the length of the query array, the number of preset hash functions is also increased accordingly. In step S110, the total number of levels and the length of the query array are determined based on the amount of data in the database, and can be set according to actual needs. For example, only horizontal expansion of the query array may be allowed, and vertical expansion may not be allowed (i.e., the query array always has only one level, and the length of the query array is continuously increased as the amount of data stored in the database increases, i.e., the length of the Bloom filter is expanded); or only vertical expansion of the query array may be allowed, and horizontal expansion may not be allowed; or both horizontal and vertical expansion of the query array may be allowed.
[0116] For implementations that allow horizontal and vertical expansion of the query array, different methods can be used to determine whether the current update of the query array should involve horizontal expansion, vertical expansion, or simultaneous horizontal and vertical expansion.
[0117] As an optional implementation, each time step S110 is executed, the query array is alternately expanded horizontally or vertically. For example, the query array is expanded horizontally the second time step S110 is executed, expanded vertically the third time step S110 is executed, expanded horizontally the fourth time step S110 is executed, expanded vertically the fifth time step S110 is executed, and so on.
[0118] As another optional implementation, depending on the increase in the amount of data stored in the database, the query array can be expanded horizontally or vertically.
[0119] Optionally, such as Figure 7 As shown, in step S110, determining the preset hash function for each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database specifically includes:
[0120] S111. Determine the first amount of data currently stored in the database, and the second amount of data previously stored in the database.
[0121] S113. Determine whether the difference between the first data volume and the second data volume is less than a preset first incremental threshold.
[0122] If the result of step S113 is yes, proceed to step S114; if the result of step S113 is no, proceed to step S115.
[0123] S114. Increase the total number of levels in the query array. That is, expand the query array vertically.
[0124] S115. Increase the length of each level of the query array and increase the number of preset hash functions. That is, horizontally expand the query array.
[0125] In this way, a better expansion method can be selected to expand the query array based on the increase in the amount of data stored in the database.
[0126] Furthermore, as an optional implementation, the preset first incremental threshold can be set to an empirical value based on expert experience. For example, the preset first incremental threshold can be set to 100,000. Alternatively, the preset first incremental threshold can be determined based on at least one of the hardware performance of the device implementing the data pre-query method and the first data volume. The device hardware performance includes, but is not limited to, the size of the storage space used to store the query array and the device's processing speed.
[0127] Furthermore, it's also possible to consider not expanding the query array horizontally or vertically when the amount of data stored in the database doesn't increase significantly, thereby reducing the operational pressure on the device. Accordingly, before step S113, such as... Figure 7 As shown, the method further includes:
[0128] S112. Determine whether the difference between the first data volume and the second data volume is greater than a preset second increment threshold. The preset second increment threshold is less than the preset first increment threshold.
[0129] If the result of step S112 is yes, proceed to step S113; if the result of step S112 is no, proceed to step S116.
[0130] S116. Maintain the current preset hash function, the total number of levels of the query array, and the length of the query array at each level.
[0131] Furthermore, as an optional implementation, the preset first incremental threshold can be set to an empirical value based on expert experience. For example, the preset first incremental threshold can be set to 100,000. Alternatively, the preset first incremental threshold can be determined based on at least one of the hardware performance of the device implementing the data pre-query method and the first data volume.
[0132] The following sections will provide detailed explanations of the horizontal and vertical expansion of the query array during the preparation process.
[0133] (a) Horizontal expansion of the query array
[0134] Regarding the implementation of adding levels when vertically expanding the query array:
[0135] Optionally, the increase in length is determined based on the length of the query array before the length increase, the first data volume, and the number of preset hash functions after the query array length increase.
[0136] For a single-level query array, the false positive rate P is related to the length m of the query array. A query array that is too short will quickly set all query positions to valid query values, resulting in a higher false positive rate and failing to filter query data that does not exist in the database. Conversely, a longer query array takes longer to set all query positions to valid query values, resulting in a lower false positive rate. Secondly, the number of preset hash functions k is also related to the false positive rate P. More preset hash functions result in faster setting of all query positions to valid query values, a higher false positive rate, and lower query efficiency. However, if the number of preset hash functions is too small, or if the output range of the preset hash functions is unevenly distributed, the false positive rate of multi-level bit arrays will also increase.
[0137] For the L-th level query array, when the L-th level corresponds to k hash functions and the length of the L-th level query array is m, when setting the n·k corresponding query positions as valid query values based on the n·k hash values corresponding to the n stored data, the probability that a query position in the L-th level query array is not a valid query value is:
[0138]
[0139] Then the false positive rate P satisfies:
[0140]
[0141] According to the limit have:
[0142]
[0143] From the above formula, it can be seen through mathematical derivation that the false positive rate P is... The minimum time.
[0144] Optionally, the length m′ of the horizontally expanded query array satisfies:
[0145]
[0146] That is, the increase in length satisfies:
[0147]
[0148] Where △m is the length increase, m0 is the length of the query array before horizontal expansion (i.e., the length of the query array before the length increase), n′ is the first data volume, and k′ is the number of preset hash functions after horizontal expansion (i.e., the number of preset hash functions after the length of the query array increases). The calculation result is rounded up.
[0149] Accordingly, each time the query array is horizontally expanded, a preset hash function is added. The preset hash function hash(x) added each time is:
[0150] hash(x) = x mod m0
[0151] Where x is the data to be queried or the stored data.
[0152] (ii) Vertical expansion of the query array
[0153] (1) Regarding the implementation that for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i Implementation of ∈Z}:
[0154] As an optional implementation, the preset hash functions corresponding to any two levels can be set to be different. For example, the query array has 5 levels. The preset hash functions corresponding to level 1 are hash1, hash2, hash3; the preset hash functions corresponding to level 2 are hash4, hash5, hash6; the preset hash functions corresponding to level 3 are hash7, hash8, hash9, and so on. Then, for a certain stored data D1, its hash value corresponding to level 1 is hash1(D1), hash2(D1), hash3(D1); its hash value corresponding to level 2 is hash4(D1), hash5(D1), hash6(D1); its hash value corresponding to level 3 is hash7(D1), hash8(D1), hash9(D1), and so on. In this way, by using different preset hash functions on the same stored data in the query array at different levels to determine the corresponding hash value, and setting the query position corresponding to the hash value as the valid query value, the query position of the hash mapping of the same stored data is different at different levels. From a macro perspective, the distribution pattern of query positions set as valid query values differs at each level of the query array. Because it's less likely that query data not existing in the database will cause hash collisions with stored data across multiple levels of the query array, more query data not present in the database can be filtered out.
[0155] As another optional implementation, for the same target data, the function input values corresponding to different levels are different. Specifically, the L-th level hash value corresponding to the target data is determined based on the L-th level function input value. The target data is either the data to be queried or the stored data. For example, the query array has 5 levels, and the preset hash functions corresponding to all levels are hash1, hash2, and hash3. Therefore, for a certain stored data D1, its hash value corresponding to level 1 is hash1(D...). 1-1 ), hash2(D 1-1 ), hash3(D 1-1 Its hash value at level 2 is hash1(D). 1-2 ), hash2(D 1-2 ), hash3(D 1-2 ), ..., its corresponding hash value at level 5 is hash1(D 1-5 ), hash2(D 1-5 ), hash3(D 1-5 ), where D 1-1 D 1-2 ...D 1-5Both values are related to stored data D1, but no two values are equal. Thus, when determining the corresponding hash value for the same stored data across different levels of query arrays, and setting the query position corresponding to the hash value as the valid query value, by controlling the input values to the preset hash function to be different, the hash values mapped to the same stored data at different levels are different, resulting in different query positions for the same stored data at different levels. From a macroscopic perspective, this makes the distribution pattern of query positions set as valid query values in each level of query array different. Data not existing in the database is less likely to cause hash collisions with stored data across multiple levels of query arrays, thus enabling the filtering out of more data not existing in the database.
[0156] Furthermore, under the condition that the preset hash functions corresponding to any two levels are the same, as an optional implementation, for the same preset hash function, the input value of the first-level function corresponding to the target data is the target data; the input value of the non-first-level function corresponding to the target data is the hash value of the previous level corresponding to the target data. For example, the query array has 3 levels, and the preset hash functions corresponding to all levels are hash1, hash2, and hash3. Then, for a certain stored data D1, its hash value corresponding to the first level is hash1(D1), hash2(D1), hash3(D1), its hash value corresponding to the second level is hash1(hash1(D1)), hash2(hash2(D1)), hash3(hash3(D1)), and its hash value corresponding to the third level is hash1(hash1(hash1(D1))), hash2(hash2(hash2(D1))), hash3(hash3(hash3(D1))).
[0157] Furthermore, as another optional implementation, under the condition that the preset hash functions corresponding to any two levels are the same, for the same preset hash function, the Lth level input value corresponding to the target data is obtained by performing an algebraic transformation on the target data, and the algebraic transformation rules corresponding to any two levels are different. For example, the query array has 3 levels, and the preset hash functions corresponding to all levels are hash1, hash2, and hash3. Then, for a certain stored data D1, its hash value corresponding to the 1st level is hash1(D1), hash2(D1), hash3(D1), and its hash value corresponding to the 2nd level is hash1(D1). 2 ), hash2(D1) 2 ), hash3(D1) 2), whose corresponding hash values at level 3 are hash1(lg(D1)), hash2(lg(D1)), and hash3(lg(D1)).
[0158] It is understood that the above-described implementation methods can be implemented individually or in combination. For example, different preset hash functions can be set for any two levels, and for the same target data, the function input values corresponding to the target data at different levels can be set to be different. For example, the query array has three levels. The preset hash functions corresponding to the first level are hash1, hash2, and hash3; the preset hash functions corresponding to the second level are hash4, hash5, and hash6; and the preset hash functions corresponding to the third level are hash7, hash8, and hash9. Then, for a certain stored data D1, its hash value corresponding to the first level is hash1(D1), hash2(D1), and hash3(D1), and its hash value corresponding to the second level is hash4(D1), hash2(D1), and hash3(D1), respectively. 2 ), hash5(D1) 2 ), hash6(D1) 2 ), whose corresponding hash values at level 3 are hash7(lg(D1)), hash8(lg(D1)), and hash9(lg(D1)).
[0159] (2) Implementation method for increasing the number of levels when vertically expanding the query array:
[0160] Optionally, the increase in the total number of levels is determined based on the total number of levels before the increase in the query array levels.
[0161] Furthermore, the increase in the total number of stages satisfies the following relationship:
[0162]
[0163] Where △L represents the increase in the total number of levels, and L0 represents the total number of levels in the query array before the increase. The calculation result is rounded up.
[0164] (3) Implementation of the query array level for hash mapping of any stored data:
[0165] ①As an optional implementation, for any stored data, the stored data is hash-mapped in the query array at all levels.
[0166] Specifically, such as Figure 8 As shown, the pre-query preparation process can specifically include the following steps:
[0167] S110A. Determine the preset hash function for each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database, and set the newly added query position in the query array as an invalid query value.
[0168] S120A: Select a unique piece of data from the database for storage.
[0169] If step S120A is successfully selected, then step S130A is executed; if all stored data has been selected, then the data pre-query preparation process ends. Afterwards, when data to be queried needs to be pre-queried, step S210A of the query process is executed.
[0170] S130A, let L = 1.
[0171] S140A. For any L-level preset hash function, determine the L-level hash value corresponding to the stored data based on the L-level preset hash function.
[0172] Among them, there are multiple preset hash functions at the Lth level.
[0173] S150A. For any of the Lth level hash values, set the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value.
[0174] S160A, Determine if L is equal to L max .
[0175] If the result of step S160A is negative, proceed to step S170A; if the result of step S160A is positive, proceed to step S120A.
[0176] S170A, Increment L by 1. Return to step S140A.
[0177] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z},L max To query the total number of levels in the array.
[0178] Accordingly, such as Figure 9 As shown, the pre-query process for the data to be queried specifically includes the following steps:
[0179] S210A, let L = 1.
[0180] S220A. For any of the Lth-level preset hash functions, determine the Lth-level hash value corresponding to the data to be queried.
[0181] S230A. Based on the query position corresponding to each Lth-level hash value in the Lth-level query array, determine whether the data to be queried necessarily does not exist in the database.
[0182] Optionally, if each of the Lth-level hash values has an invalid query value in the corresponding query position in the Lth-level query array, it is determined that the data to be queried must not exist in the database; otherwise, it is determined that it is impossible to determine whether the data to be queried must not exist in the database.
[0183] If step S230A cannot determine whether the data to be queried must not exist in the database, proceed to step S240A; if step S230A determines that the data to be queried must not exist in the database, proceed to step S260A.
[0184] S240A, Determine if L is equal to L max That is, in this embodiment, the query termination condition is L = L. max .
[0185] If the result of step S240A is yes, the pre-query process ends. Then, a precise query is performed on the data to be queried in the database.
[0186] If the result of step S240A is negative, proceed to step S250A.
[0187] S250A, increment L by 1. Then return to step S220A.
[0188] S260A, Return the query result if the data to be queried does not exist.
[0189] In summary, in this embodiment, all stored data undergoes hash mapping across all levels of the query array. For the same stored data, by controlling that the L1-level hash value and the L2-level hash value corresponding to the stored data are not identical, the distribution of query positions set as valid query values differs across different levels of the query array. When pre-querying the data to be queried, corresponding hash mapping queries are performed sequentially from the first level to the last level of the query array until it can be directly determined that the data to be queried does not exist in the database. This leverages the characteristic that data not existing in the database is unlikely to cause hash collisions with stored data across multiple levels of the query array, thereby filtering out more data that cannot be misjudged by a single-level Bloom filter.
[0190] ②As another optional implementation, the level of hash mapping for different stored data is not necessarily the same.
[0191] Specifically, such as Figure 10 As shown, the pre-query preparation process can specifically include the following steps:
[0192] S110B: Based on the amount of data in the database, determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array; set newly added query positions in the query array as invalid query values; and determine the corresponding flag arrays for each level based on the query arrays at each level, and set newly added flag positions in the flag arrays as invalid flag values. Each query position in the L-th level query array uniquely corresponds to a flag position in the L-th level flag array.
[0193] For example, such as Figure 11 As shown, according to step S110B, a length of m and a total number of stages L are set. max The query array can be set accordingly to have a length of m and a total number of levels of L. max The query array corresponds to the same position in the flag array.
[0194] In practice, to reduce the storage space occupied by the query array, the flag array can optionally be set as a bit array, where each flag position is represented by a binary bit. Accordingly, for each flag bit, 0 can be used as an invalid flag value and 1 as a valid flag value, and the value of the flag bit can be set in subsequent steps using a union operation.
[0195] S120B: Select a unique piece of data from the database for storage.
[0196] If step S120B is successful, then step S130B is executed; if all stored data has been selected, then the data pre-query preparation process ends. Afterwards, when data to be queried needs to be pre-queried, step S210B of the query process is executed.
[0197] S130B, let L = 1.
[0198] S140B. For any Lth-level preset hash function, determine the Lth-level hash value corresponding to the stored data based on the Lth-level preset hash function.
[0199] Among them, there are multiple preset hash functions at the Lth level.
[0200] S150B: For any of the Lth level hash values, set the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value.
[0201] S160B, Determine whether L is equal to L0. Where L0 is the last level corresponding to the stored data.
[0202] If the result of step S160B is negative, proceed to step S171B; if the result of step S160B is positive, proceed to step S120B.
[0203] S171B. For any of the Lth-level hash values, set the flag position corresponding to the Lth-level hash value in the Lth-level flag array to a valid flag value. That is, for any of the Lth-level hash values, set the flag position corresponding to the query position of the Lth-level hash value to a valid flag value.
[0204] For example, according to step S110B, a length of m and a total number of stages of L are set. max The query array, and a sequence of length m with a total number of levels L. max If the flag array and the query array correspond to the same position in the flag array, for a certain stored data D1, its corresponding hash value at level L is C1, C2, C3. Then, after setting the query positions C1, C2, C3 in the query array at level L as valid query values, and determining that L ≠ L0, the flag positions C1, C2, C3 in the flag array at level L are set as valid flag values.
[0205] S172B, Increment L by 1. Return to step S140B.
[0206] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z}, L0≤L max L max To query the total number of levels in the array.
[0207] Among them, such as Figure 12 As shown, the last level L0 corresponding to any stored data is determined in the following way:
[0208] S161B, Let L = 1, and determine that all stored data in the database are level 1 mapped stored data.
[0209] S162B: Select at least a portion of the storage data from all L-level mapping storage data that participate in determining the corresponding L-level hash value as L-level end mapping storage data.
[0210] For example, half of the Lth-level mapping storage data can be randomly selected from all the Lth-level mapping storage data involved in determining the corresponding Lth-level hash value as the Lth-level ending mapping storage data.
[0211] S163B. For the Lth level ending mapping storage data, determine that the Lth level is the last level L0 corresponding to the ending mapping storage data.
[0212] S164B, Determine if the Lth level is equal to L. max .
[0213] If the result of step S164B is negative, proceed to step S165B; if the result of step S164B is positive, end the process of determining the last level L0 corresponding to any stored data.
[0214] S165B: Use the Lth level mapping storage data (which is not the Lth level ending mapping storage data) as the next level mapping storage data.
[0215] S166B, Increment L by 1. Return to step S162B.
[0216] That is, through the above process, in the process of hash mapping of multi-level query arrays of stored data, a portion of the stored data that has been hash mapped at the current level can be selected for the hash mapping process of the next level query array, while the other portion of the stored data that has been hash mapped at the current level ends the hash mapping process of the next level query array (if the stored data that has been hash mapped at the current level is not considered, the hash mapping process of the next level will not be performed).
[0217] Accordingly, when performing a pre-query on the data to be queried, such as Figure 13 As shown, the pre-query process specifically includes the following steps:
[0218] S210B, let L = 1.
[0219] S220B. For any of the Lth-level preset hash functions, determine the Lth-level hash value corresponding to the data to be queried.
[0220] S230B: Based on the query position corresponding to each Lth-level hash value in the Lth-level query array, determine whether the data to be queried necessarily does not exist in the database.
[0221] Optionally, if each of the Lth-level hash values has an invalid query value at its corresponding query position in the Lth-level query array, it is determined that the data to be queried must not exist in the database; and if each of the Lth-level hash values has both a flag query position and a non-flag query position at its corresponding query position in the Lth-level query array, it is determined that the data to be queried must not exist in the database. Here, the flag query position is the query position in the Lth-level flag array where the corresponding flag position is a valid flag value. Otherwise, it is determined that it is impossible to determine whether the data to be queried must not exist in the database. That is, if the query array is a multi-level query array, then in any level of the query array, not all query positions in the hash mapping of the data to be queried are valid query values; as long as there is a query position with an invalid query value, it can be directly determined that the data to be queried must not exist in the database. Furthermore, if the query array is a multi-level query array, then in the query positions of the hash mapping of the data to be queried on any level of the query array, if some query positions correspond to valid flag values and other query positions correspond to invalid flag values, it can be directly determined that the data to be queried must not exist in the database.
[0222] If step S230B cannot determine whether the data to be queried must not exist in the database, step S240B is executed; if step S230B determines that the data to be queried must not exist in the database, step S260B is executed.
[0223] S240B. Determine that the query positions corresponding to each of the Lth-level hash values in the Lth-level query array are all flag query positions. That is, in this embodiment, the query termination condition is that there are no flag query positions corresponding to each of the Lth-level hash values in the Lth-level query array.
[0224] If the result of step S240B is yes, the pre-query process ends. Then, a precise query is performed on the data to be queried in the database.
[0225] If the result of step S240B is negative, proceed to step S250B.
[0226] S250B, Increment L by 1. Then return to step S220B.
[0227] S260B: Return the query result if the data to be queried does not exist.
[0228] In summary, in this embodiment, different stored data undergo hash mapping at different levels of query arrays. Some stored data undergoes hash mapping at multiple levels of query arrays. Furthermore, for the same stored data, by controlling that the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely identical, the distribution of query positions set as valid query values in different levels of query arrays is differentiated. During pre-querying of the data to be queried, after mapping all the data to be queried to the marked query positions, the hash mapping of the data to be queried continues at the next level of query arrays until it can be directly determined that the data to be queried does not exist in the database. This utilizes the characteristic that data to be queried that does not exist in the database is unlikely to cause hash collisions with stored data in multiple levels of query arrays, thereby filtering out more data to be queried that cannot be misjudged by a single-level Bloom filter alone.
[0229] Based on the same inventive concept, embodiments of the present invention also provide a data pre-query device, such as... Figure 14 As shown, it includes:
[0230] Setting unit U1 is used to determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database, and to set newly added query positions in the query array as invalid query values; for any stored data in the database, starting from the first level to the last level L0 corresponding to the stored data, the following steps are executed level by level to set multiple query positions corresponding to the stored data in each level of the query array: for any L-th level preset hash function, determine the L-th level hash value corresponding to the stored data based on the L-th level preset hash function; there are multiple L-th level preset hash functions; for any L-th level hash value, set the query position corresponding to the L-th level hash value in the L-th level query array as a valid query value;
[0231] The pre-query unit U2 is used to determine the L-th level hash value corresponding to the data to be queried for any of the L-th level preset hash functions; the initial value of L is 1; based on the query position corresponding to each L-th level hash value in the L-th level query array, it is determined whether the data to be queried must not exist in the database; if it cannot be determined that the data to be queried must not exist in the database, and it is determined that the query termination condition is not met, then L is incremented by 1, and the step of determining the L-th level hash value corresponding to the data to be queried for any of the L-th level preset hash functions is returned;
[0232] Wherein, for the same stored data, the L1-level hash value and the L2-level hash value corresponding to the stored data are not completely the same, L1, L2 ∈ {L i |1≤L i ≤L0,L i ∈Z}, L0≤Lmax L max To query the total number of levels in the array.
[0233] Optionally, determining whether the data to be queried necessarily does not exist in the database based on the query position corresponding to each of the Lth-level hash values in the Lth-level query array includes:
[0234] If each of the Lth level hash values has an invalid query value in the corresponding query position of the Lth level query array, it is determined that the data to be queried must not exist in the database.
[0235] Otherwise, it cannot be determined whether the data to be queried necessarily does not exist in the database.
[0236] As an optional implementation, the preset hash functions corresponding to any two levels are different.
[0237] As another optional implementation, for the same target data, the function input values corresponding to the target data are different at different levels; wherein, the L-th level hash value corresponding to the target data is determined according to the L-th level function input value corresponding to the target data;
[0238] The target data is either the data to be queried or the stored data.
[0239] Furthermore, as an optional implementation, the preset hash functions corresponding to any two levels are the same;
[0240] For the same preset hash function, the input value of the first-level function corresponding to the target data is the target data; the input value of the non-first-level function corresponding to the target data is the hash value of the previous level corresponding to the target data.
[0241] Furthermore, as an optional implementation, the preset hash functions corresponding to any two levels are the same;
[0242] For the same preset hash function, the Lth level input value corresponding to the target data is obtained by performing an algebraic transformation on the target data, and the algebraic transformation rules are different for any two levels.
[0243] Optionally, determining the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database specifically includes:
[0244] Determine the first amount of data currently stored in the database, and the second amount of data previously stored in the database;
[0245] If the difference between the first data volume and the second data volume is less than a preset first increment threshold, then the total number of levels is increased;
[0246] If the difference between the first data volume and the second data volume is greater than or equal to the preset first incremental threshold, then the length is increased;
[0247] The method further includes:
[0248] If the length is increased, the number of preset hash functions is increased.
[0249] Optionally, if the total number of levels determined this time is greater than the total number of levels determined last time, the amount of increase in the total number of levels is determined based on the total number of levels before the increase in the query array levels;
[0250] Furthermore, the increase in the total number of stages satisfies the following relationship:
[0251]
[0252] Where △L is the increase in the total number of levels, and L0 is the total number of levels before the increase in the query array levels.
[0253] Optionally, if the length determined this time is greater than the length determined last time, the amount of increase in length is determined based on the length of the query array before the increase, the first data volume, and the number of preset hash functions after the query array length is increased.
[0254] Alternatively, the increase in length may satisfy the following relationship:
[0255]
[0256] Wherein, △m is the length increase, n′ is the first data volume, k′ is the number of preset hash functions after the length of the query array is increased, and m0 is the length of the query array before the length is increased;
[0257] The preset hash function hash(x) added each time is:
[0258] hash(x) = x mod m0
[0259] Where x is the data to be queried or the stored data.
[0260] Optionally, the preset first incremental threshold is determined based on at least one of the device hardware performance and the first data volume.
[0261] Optionally, determining that the query termination condition is not met includes:
[0262] If the query position corresponding to each L-th level hash value in the L-th level query array is determined to be a flag query position, it is determined that the query termination condition is not met; wherein, the flag query position is the query position in the L-th level flag array where the corresponding flag position is a valid flag value;
[0263] The flag arrays at each level are determined as follows:
[0264] Determine the flag arrays at each level, and set the newly added flag positions in the flag arrays to invalid flag values, wherein each query position in the L-th level query array uniquely corresponds to a flag position in the L-th level flag array;
[0265] For any stored data in the database, after setting the query position corresponding to the Lth level hash value in the Lth level query array as a valid query value for any Lth level hash value, if it is determined that the Lth level is not the last level L0 corresponding to the stored data, then the flag position corresponding to the Lth level hash value is set as a valid flag value.
[0266] The last level L0 corresponding to any stored data is determined in the following way:
[0267] At least a portion of the stored data from all L-level mapping stored data that participated in determining the corresponding L-level hash value is selected as the L-level ending mapping stored data; where the starting value of L is 1; and the L-level mapping stored data is all stored data in the database;
[0268] For the Lth level ending the mapping storage data, determine that the Lth level is the last level L0 corresponding to the ending mapping storage data; if the Lth level is not the last level L0... max The step involves selecting at least a portion of the stored data from all L-level mapping data that are not the L-level end mapping data as the next level mapping data, incrementing L by 1, and returning the step of selecting at least a portion of the stored data from all L-level mapping data that participated in determining the corresponding L-level hash value as the L-level end mapping data.
[0269] Optionally, the query array is a bit array.
[0270] Optionally, the flag array is a bit array.
[0271] It should be understood that the embodiments of the data pre-query device described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. The functional units in the embodiments may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The integrated unit described above can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0272] For example, such as Figure 15 As shown, the data pre-query device can be configured to include the aforementioned adaptation proxy module M100, adjustment module M200, and judgment module M300, depending on the actual situation of software program development. The adaptation proxy module M100 can be further divided into a database connection module M110, a parsing module M120, and a data volume determination module M130. The adaptation proxy module is used to implement step S110. The database connection module M110 is used to configure relevant configuration information and interact with the database; the parsing module M120 is used to obtain log files from the database and retrieve all events from the obtained log files; the data volume determination module M130 is used to determine the current stored first data volume, the data increment during the current update of the query array compared to the previous update, etc., based on all events, and provides this information to the adjustment module M200. The adjustment module M200 can be further divided into a horizontal and vertical amplification judgment module M210, a vertical amplification module M220, and a horizontal amplification module M230. The horizontal and vertical amplification judgment module M210 is used to implement steps S111-S113, the vertical amplification module M220 is used to implement step S114, and the horizontal amplification module M230 is used to implement step S115. The judgment module is used to implement steps S120-S170 and steps S210-S260.
[0273] Since the principle of the data pre-query device in solving the problem is basically the same as that of the data pre-query method, the implementation of the pre-query device can be referred to the implementation of the data pre-query method, and will not be repeated here.
[0274] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, such as... Figure 16 As shown, it includes: a processor 110 and a memory 120 for storing instructions executable by the processor 110;
[0275] The processor 110 is configured to execute the instructions to implement the data pre-query method.
[0276] In specific implementations, the device may vary significantly due to differences in configuration or performance. It may include one or more processors 110, memory 120, and computer-readable storage media 130. The memory 120 and / or computer-readable storage media 130 may contain one or more application programs 131 or data 132. The memory 120 and / or computer-readable storage media 130 may also contain one or more operating systems 133, such as Windows, Mac OS, Linux, iOS, Android, Unix, FreeBSD, etc. The memory 120 and computer-readable storage media 130 may be temporary or persistent storage. The application program 131 may include one or more of the aforementioned modules (…). Figure 16 (Not shown in the image), each module may include a series of instruction operations. Furthermore, the processor 110 may be configured to communicate with the computer-readable storage medium 130 and execute a series of instruction operations in the computer-readable storage medium 130 on the device. The device may also include one or more power supplies (…). Figure 16 (not shown in the image); one or more network interfaces 140, including wired network interface 141 and / or wireless network interface 142; one or more input / output interfaces 143.
[0277] Based on the same inventive concept, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when the computer program code is run on a computer, causes the computer to execute the data pre-query method.
[0278] Based on the same inventive concept, embodiments of the present invention also provide a computer program product, the computer program product comprising: computer program code, which, when run on a computer, causes the computer to execute the data pre-query method.
[0279] The data pre-query method, apparatus, device, and computer-readable storage medium provided in this invention, by adaptively adjusting the length of the query array based on the amount of data stored in the current database, can make the hash mapping positions of the stored data in the query array at the current level more sparse compared to a Bloom filter that does not adjust the query array length. This reduces the probability of hash collisions between the data to be queried and the stored data during the query process, thus lowering the pre-query false positive rate for the data to be queried. Furthermore, by adaptively increasing the total number of levels of the query array, the different hash mapping positions of the same stored data in the query arrays at more levels change the distribution of hash mapping positions of each stored data from a macroscopic perspective. This makes it difficult for data to be queried that does not exist in the database to simultaneously cause hash collisions with the stored data in multiple levels of the query array during the pre-query process, thus filtering out data that does not exist in the database and concluding that it definitely does not exist in the database, further reducing the pre-query false positive rate for the data to be queried. Therefore, this effectively improves the pre-query efficiency of the data to be queried.
[0280] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0281] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0282] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0283] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0284] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A data pre-query method, characterized in that, include: For any one L A preset hash function is used to determine the first-order hash function corresponding to the data to be queried. L Level hash value; L The initial value is 1; The first L There are multiple preset hash functions at each level; According to each of the above... L Level hash value at the 1st L The corresponding query position in the query array is used to determine whether the data to be queried must not exist in the database. If it cannot be determined whether the data to be queried necessarily does not exist in the database, and it is determined that the query termination condition is not met, then let L Add 1, and return the value for any of the first... L A preset hash function is used to determine the first hash function corresponding to the data to be queried. L Steps for creating a hash value; The query arrays at each level are determined as follows: Based on the amount of data in the database, determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array, and set any newly added query positions in the query array as invalid query values; wherein, at least one of the total number of preset hash functions corresponding to each level, the total number of levels of the query array, and the length of each level of the query array is positively correlated with the amount of data in the database; For any stored data in the database, starting from the first level up to the last level corresponding to the stored data. The following steps are executed sequentially to set the multiple query positions corresponding to the stored data in the query arrays at each level: For any of the first L The first-level preset hash function, based on the first-level preset hash function. L The first-level preset hash function determines the first-level hash function corresponding to the stored data. L Level hash value; For any of the first L The first-level hash value will be the first-level hash value. L The first level of the query array L Set the query position corresponding to the hash value to a valid query value; Among them, for the same stored data, the first corresponding to the stored data Level hash value and the first The hash values are not completely identical. , , To query the total number of levels in the array.
2. The method as described in claim 1, characterized in that, According to each of the above-described first L The level hash value is in the first L The query position in the level query array is used to determine whether the data to be queried must not exist in the database, including: If each of the above-mentioned first L The level hash value is in the first L If there is a query position in the corresponding query position of the level query array that has an invalid query value, it is determined that the data to be queried must not exist in the database. Otherwise, it cannot be determined whether the data to be queried necessarily does not exist in the database.
3. The method as described in claim 1, characterized in that, The preset hash functions corresponding to any two levels are different; And / or, for the same target data, the function input values corresponding to the target data are different at different levels; wherein, the target data corresponds to the first... L The level hash value is based on the first level corresponding to the target data. L The input values for the level function are determined; The target data is either the data to be queried or the stored data.
4. The method as described in claim 3, characterized in that, The preset hash function is the same for any two levels; For the same preset hash function, the input value of the first-level function corresponding to the target data is the target data; the input value of the non-first-level function corresponding to the target data is the hash value of the previous level corresponding to the target data; or, for the same preset hash function, the input value of the first-level function corresponding to the target data is the hash value of the previous level corresponding to the target data. L The level input values are obtained by performing an algebraic transformation on the target data, and the algebraic transformation rules are different for any two levels.
5. The method as described in claim 1, characterized in that, The step of determining the preset hash function, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database specifically includes: Determine the first amount of data currently stored in the database, and the second amount of data previously stored in the database; If the difference between the first data volume and the second data volume is less than a preset first increment threshold, then the total number of levels is increased; If the difference between the first data volume and the second data volume is greater than or equal to the preset first incremental threshold, then the length is increased; The method further includes: If the length is increased, the number of preset hash functions is increased.
6. The method as described in claim 1, characterized in that, If the total number of levels determined this time is greater than the total number of levels determined last time, the amount of increase in the total number of levels is determined based on the total number of levels before the increase in the query array levels; And / or, if the length determined this time is greater than the length determined last time, the amount of increase in length is determined based on the length of the query array before the increase, the first amount of data currently stored in the database, and the number of preset hash functions after the query array length is increased.
7. The method as described in claim 6, characterized in that, The increase in the total number of levels satisfies the following relationship: in, The increase in the total number of levels, The total number of levels in the query array before adding the number of levels; The increase in length satisfies the following relationship: in, The amount by which the length is increased, This is the first data volume. The number of preset hash functions is the sum of the length of the query array and the number of such functions. The length of the query array before the addition; Each time a preset hash function is added hash ( x )for: in, x The data to be queried or the stored data.
8. The method as described in claim 5, characterized in that, The preset first incremental threshold is determined based on at least one of the device hardware performance and the first data volume.
9. The method as described in claim 1, characterized in that, The determination that the query termination condition is not met includes: Determine each of the aforementioned first L The level hash value is in the first L All query positions in the level query array that correspond to the marked query positions are determined to not meet the query termination condition; wherein, the marked query position is the position at the first level. L The corresponding flag position in the level flag array is the query position for a valid flag value; The flag arrays at each level are determined as follows: Determine the flag arrays at each level, and set the newly added flag positions in the flag arrays to invalid flag values, wherein the first... L Each query position in the level query array uniquely corresponds to the first... L A flag position in the level flag array; For any stored data in the database, in any of the first... L The first-level hash value will be the first-level hash value. L The first level of the query array L After setting the query position corresponding to the first-level hash value as a valid query value, the first-level hash value is determined. L The level is not the last level corresponding to the stored data. Then the first L Set the flag position corresponding to the hash value to a valid flag value; The last level corresponding to any stored data is determined in the following way: : Selecting participants determines the corresponding [number]. L All the first-level hash values L At least a portion of the stored data in the level-mapped storage data is the first L Level ends the mapping storage data; where L The initial value is 1; the first-level mapping stores all the stored data in the database. For the first L Level ends mapping and storage data, determining the first L The level is the last level corresponding to the end of the mapping storage data. If the first L The level is not the last level. , will not L The first level ends the mapping storage data. L Level-mapping storage data is used as the next-level mapping storage data, making L Add 1, and return to the selection panel to confirm the corresponding number. L All the first-level hash values L At least a portion of the stored data in the level-mapped storage data is the first L The step of ending the mapping and storage of data.
10. The method as described in claim 1, characterized in that, The query array is a bit array.
11. The method as described in claim 9, characterized in that, The flag array is a bit array.
12. A data pre-query device, characterized in that, include: The setting unit is used to determine the preset hash function corresponding to each level, the total number of levels of the query array, and the length of each level of the query array based on the amount of data in the database, and to set newly added query positions in the query array as invalid query values; wherein, at least one of the total number of preset hash functions corresponding to each level, the total number of levels of the query array, and the length of each level of the query array is positively correlated with the amount of data in the database; for any stored data in the database, starting from the first level up to the last level corresponding to the stored data. The following steps are executed sequentially to set the multiple query positions corresponding to the stored data in each level of the query array: For any given... L The first-level preset hash function, based on the first-level preset hash function. L The first-level preset hash function determines the first-level hash function corresponding to the stored data. L Level hash value; the first L There are multiple preset hash functions for each level; for any of the first-level functions... L The first-level hash value will be the first-level hash value. L The first level of the query array L Set the query position corresponding to the hash value to a valid query value; Pre-query unit, used for querying any of the first... L A preset hash function is used to determine the first-order hash function corresponding to the data to be queried. L Level hash value; L The initial value is 1; according to each of the described... L The level hash value is in the first L The corresponding query position in the query array is used to determine whether the data to be queried must not exist in the database; if it cannot be determined that the data to be queried must not exist in the database, and it is determined that the query termination condition is not met, then let... L Add 1, and return the value for any of the first... L A preset hash function is used to determine the first hash function corresponding to the data to be queried. L Steps for creating a hash value; Among them, for the same stored data, the first corresponding to the stored data Level hash value and the first The hash values are not completely identical. , , To query the total number of levels in the array.
13. An electronic device, characterized in that, include: A processor and a memory for storing processor-executable instructions; The processor is configured to execute the instructions to implement the data pre-query method as described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program code that, when executed on a computer, causes the computer to perform the data pre-query method as described in any one of claims 1-11.
Citation Information
Patent Citations
Data store and method of allocating data to the data store
CN105378685A
Database data query method and device based on state dictionary, electronic equipment and storage medium
CN116303610A