A method, device and medium for implementing data query based on a Bloom filter
Through the combination of multiple hashing algorithms and Redis cache, the misjudgment problem of traditional Bloom filters when judging the existence of data in the database is solved, more efficient data query and reduce database pressure, and improve user experience.
Patent Information
- Application Number
- CN202211260919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-10-14
AI Technical Summary
Traditional Bloom filters cannot correctly determine whether there is data in the database when processing redundant requests, resulting in misjudgment and increasing database query pressure and reducing query efficiency.
The hash value of the requested data identification is calculated through various hashing algorithms, and the Redis cache is used to determine whether the database contains the corresponding site information, and determine whether the site value is a preset target value in the database array, and directly obtain data from the site.
It avoids misjudgment of the Bloom filter, reduces the database query pressure, and improves query response time and user experience.
Smart Images

Figure CN115544329B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, device, and medium for implementing data query based on a Bloom filter. Background Art
[0002] In web development and design, the business backend is often associated with a relational database. Moreover, when requesting to query certain information in the relational database, there may be many requests for data that do not exist in the database. Then, the requests corresponding to the data that does not exist in the database are redundant. At this time, when the amount of data requests is relatively large, the redundant requests will affect the performance of the system and may even cause the system to crash. Currently, when encountering redundant requests, a Bloom filter is generally used to handle the redundant requests.
[0003] However, when the traditional Bloom filter processes redundant requests, due to the misjudgment of the Bloom filter, it is unable to identify the query requests corresponding to the data that does not exist in the database as redundant requests and continues to send the query requests to the database for data query. At this time, it is impossible to implement the query of the data that does not exist in the database, increasing the query pressure on the database and reducing the data query efficiency. Summary of the Invention
[0004] Embodiments of this application provide a method, device, and medium for implementing data query based on a Bloom filter to solve the technical problem that when the existing technology uses a Bloom filter to process redundant requests, the Bloom filter cannot correctly determine whether the data corresponding to the request exists in the database.
[0005] On the one hand, embodiments of this application provide a method for implementing data query based on a Bloom filter, including:
[0006] Receiving a query request sent by a web end to a database and determining a request data identifier in the query request;
[0007] Based on a Bloom filter and through at least one hashing algorithm, calculating at least one hash value corresponding to the request data identifier, and determining whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value;
[0008] If it is determined that the Redis cache contains the site information corresponding to the request data identifier, sending the query request to the database and determining whether the value corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value; where one hash value of the data corresponds to one site in the bitmap, and the preset target value is used to indicate that the site is not empty;
[0009] When determining that the value corresponding to the at least one locus in the array of the database is a preset target value, query and obtain the data corresponding to the request data identifier in the corresponding locus according to the locus information corresponding to the request data identifier, so as to implement the query of the data corresponding to the request data identifier.
[0010] In an implementation manner of the present application, before receiving the query request sent by the Web end to the database, the method further includes:
[0011] Obtain a plurality of data to be stored that need to be stored in the database, and configure corresponding arrays for the plurality of data to be stored in the database;
[0012] Based on a first-level Bloom filter and through at least one first-level hashing algorithm, calculate at least one first-level hash value corresponding to the data to be stored, and determine at least one first-level locus corresponding to the at least one first-level hash value in the database;
[0013] Obtain at least one value corresponding to the at least one first-level locus in the array of the database, and determine whether the at least one first-level locus is empty according to the at least one value;
[0014] If so, store the data to be stored into the corresponding first-level locus to complete the storage of the data to be stored.
[0015] In an implementation manner of the present application, after determining whether the at least one first-level locus is empty according to the at least one value, the method further includes:
[0016] If all the first-level loci in the at least one first-level locus are not empty, then based on a second-level Bloom filter and through at least one second-level hashing algorithm, calculate at least one second-level hash value corresponding to the data to be stored, and determine at least one second-level locus corresponding to the at least one second-level hash value in the database;
[0017] Obtain at least one value corresponding to the at least one second-level locus in the array of the database to determine whether the at least one second-level locus is empty;
[0018] If so, store the data to be stored into the corresponding second-level locus to complete the storage of the data to be stored.
[0019] In an implementation manner of the present application, after obtaining at least one value corresponding to the at least one second-level locus in the array of the database to determine whether the at least one second-level locus is empty, the method further includes:
[0020] If all the secondary sites in the at least one secondary site are not empty, obtain the deletion times corresponding to all the secondary sites in the at least one secondary site, and compare the deletion times corresponding to all the secondary sites with a preset threshold respectively;
[0021] Determine the secondary sites with deletion times less than the preset threshold, and determine one secondary site from the secondary sites with deletion times less than the preset threshold;
[0022] Delete the data stored in the determined secondary site, and store the data to be stored in the determined secondary site.
[0023] In an implementation manner of the present application, after comparing the deletion times corresponding to all the secondary sites with the preset threshold respectively, the method further includes:
[0024] If it is determined that the deletion times corresponding to all the secondary sites are greater than the preset threshold, calculate at least one tertiary hash value corresponding to the data to be stored based on a tertiary Bloom filter and at least one tertiary hash algorithm, and determine at least one tertiary site corresponding to the at least one tertiary hash value in the database;
[0025] If it is determined that at least one tertiary site corresponding to the at least one numerical value obtained in the array is empty according to the at least one tertiary site, store the data to be stored in the corresponding tertiary site to complete the storage of the data to be stored;
[0026] If it is determined that all the at least one tertiary sites are not empty, continue to calculate at least one quaternary hash value corresponding thereto based on a quaternary Bloom filter and at least one quaternary hash algorithm, and store the data to be stored in the corresponding quaternary site when it is determined that at least one quaternary site corresponding to the at least one numerical value obtained in the array corresponding to the at least one quaternary hash value is empty.
[0027] In an implementation manner of the present application, before determining whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value, the method further includes:
[0028] Based on the data identifier corresponding to the data to be stored, store the site information of the site storing the data to be stored in the Redis cache, so as to determine whether the database contains the data to be stored through the Redis cache.
[0029] In one implementation of the present application, after determining whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value, the method further includes:
[0030] If it is determined that the Redis cache does not contain the site information corresponding to the request data identifier, it is determined that the database does not contain the data corresponding to the request data identifier;
[0031] Filter the query request corresponding to the request data identifier, and return a prompt message to the Web end corresponding to the query request.
[0032] In one implementation of the present application, after querying and obtaining the data corresponding to the request data identifier at the corresponding site according to the site information corresponding to the request data identifier, the method further includes:
[0033] Based on the query request of the Web end, determine the address corresponding to the Web end included in the query request;
[0034] Based on the determined address corresponding to the Web end, return the data corresponding to the request data identifier obtained by querying to the Web end to respond to the query request of the Web end.
[0035] On the other hand, the embodiment of the present application also provides a device for implementing data query based on a Bloom filter, and the device includes:
[0036] At least one processor;
[0037] And a memory communicatively connected to the at least one processor;
[0038] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a method for implementing data query based on a Bloom filter as described above.
[0039] On the other hand, the embodiment of the present application also provides a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are set as:
[0040] A method for implementing data query based on a Bloom filter as described above.
[0041] The embodiment of the present application provides a method, a device and a medium for implementing data query based on a Bloom filter, and at least has the following beneficial effects:
[0042] By obtaining the request data identifier in the query request, calculating the corresponding hash value based on the Bloom filter and through multiple hash algorithms, and then querying in the Redis cache according to the site information corresponding to the hash value, to determine whether the database contains the data corresponding to the request data identifier based on whether the site information corresponding to the request data identifier is included in the Redis cache; in the case where the site information corresponding to the request data identifier is included in the Redis cache, sending the query request to the database, and determining whether the number corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value, and in the case where it is the preset target value, determining that the data corresponding to the request data identifier is stored at this site, and then according to this site information, the data corresponding to the request data identifier can be directly queried and obtained from the site corresponding to the site information, thereby realizing the query of the data corresponding to the request data identifier. According to the above solution, the present application can avoid the misjudgment of the traditional Bloom filter, reduce the data query pressure on the database, speed up the response time of the query request, and improve the user experience. Brief Description of the Drawings
[0043] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0044] Figure 1 It is a schematic flowchart of a method for realizing data query based on a Bloom filter provided by an embodiment of the present application;
[0045] Figure 2 It is a schematic flowchart of a method in an application scenario provided by an embodiment of the present application;
[0046] Figure 3 It is a schematic flowchart of a method in another application scenario provided by an embodiment of the present application;
[0047] Figure 4 It is a schematic internal structure diagram of a device for realizing data query based on a Bloom filter provided by an embodiment of the present application. Detailed Description of the Embodiments
[0048] To make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.
[0049] When dealing with redundant requests, the traditional Bloom filter has the following problems:
[0050] First, there is a possibility of false positives in the Bloom filter. For example, when receiving a query request for a data packet 1 to be queried, the Bloom filter calculates the hash value corresponding to data packet 1 through at least one hash algorithm, and finds the three positions 1, 5, and 7 in the bitmap corresponding to the hash value of data packet 1 from the array in the database, so as to determine whether the values corresponding to the three positions of data packet 1 in the array are the preset target values. If so, the Bloom filter determines that the data queried by the query request exists in the database; for data packet 2, the Bloom filter calculates the hash value corresponding to data packet 2 through at least one hash algorithm, and finds the three positions 2, 5, and 8 in the bitmap corresponding to the hash value of data packet 2 from the array in the database, so as to determine whether the values corresponding to the three positions of data packet 2 in the array are the preset target values. If so, the Bloom filter determines that the data queried by the query request exists in the database; for data packet 3 that has not been stored in the database, the Bloom filter calculates the hash value corresponding to data packet 3 through at least one hash algorithm, and finds the three positions 1, 5, and 8 in the bitmap corresponding to the hash value of data packet 3 from the array in the database. However, since there are overlapping parts between the three positions 1, 5, and 8 corresponding to data packet 3 and the positions corresponding to data packet 1 and data packet 2, when the Bloom filter queries the two positions 1 and 5 corresponding to data packet 3, it will query that the values corresponding to the two positions 1 and 5 of data packet 1 in the array are the preset target values, and when querying the position 5 and 8 corresponding to data packet 3, it will query that the values corresponding to the two positions 5 and 8 of data packet 2 in the array are the preset target values. At this time, the Bloom filter will misjudge that the values corresponding to the three positions 1, 5, and 8 of data packet 3 in the array are the preset target values, and think that data packet 3 exists in the database, thus failing to filter out the redundant query request corresponding to data packet 3 that actually does not exist in the database.
[0051] Second, the Bloom filter cannot delete data. When deleting data packet 1, it is necessary to set the three positions 1, 5, and 7 in the bitmap corresponding to data packet 1 to 0. At this time, if the Web end requests to query data packet 2, it will misjudge that data packet 2 does not exist in the array because the position 5 is set to 0 when deleting data packet 1, increasing the possibility of false positives in the Bloom filter. At the same time, it also increases the query pressure on the database, reduces the query efficiency of data, and affects the query experience of users.
[0052] The embodiments of the present application provide a method, device, and medium for implementing data query based on a Bloom filter. By obtaining the request data identifier in the query request, based on the Bloom filter and through multiple hash algorithms, the corresponding hash value is calculated, and then the query is performed in the Redis cache according to the site information corresponding to the hash value, so as to determine whether the database contains the data corresponding to the request data identifier based on whether the Redis cache contains the site information corresponding to the request data identifier; in the case where the Redis cache contains the site information corresponding to the request data identifier, the query request is sent to the database, and it is determined whether the number corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value, and in the case where it is the preset target value, it is determined that the data corresponding to the request data identifier is stored at this site, and then according to this site information, the data corresponding to the request data identifier can be directly queried and obtained from the site corresponding to the site information, thereby realizing the query of the data corresponding to the request data identifier. According to the above solution, the present application can avoid the misjudgment of the traditional Bloom filter, reduce the data query pressure on the database, speed up the response time of the query request, and improve the user experience. It solves the technical problem that when the existing technology processes redundant requests through a Bloom filter, the Bloom filter cannot correctly determine whether the data corresponding to the request exists in the database.
[0053] Figure 1 It is a schematic flowchart of a method for implementing data query based on a Bloom filter provided by an embodiment of the present application. As Figure 1 shown, a method for implementing data query based on a Bloom filter provided by an embodiment of the present application may mainly include the following steps:
[0054] 101. Receive the query request sent by the Web end to the database, and determine the request data identifier in the query request.
[0055] The server receives the query request sent by the Web end, and determines the corresponding request data identifier from the query request, so as to perform subsequent processing according to the determined request data identifier.
[0056] In an embodiment of the present application, before the server receives a query request from the Web side, it first obtains a number of data to be stored that need to be stored in the database. And in the present application, a one-dimensional array structure is configured in the database for the number of data to be stored, so as to facilitate storing the data to be stored in the one-dimensional array structure provided by the database. Then, based on the first-level Bloom filter and through at least one first-level hashing algorithm, the server calculates at least one first-level hash value corresponding to the data to be stored, and determines at least one first-level site corresponding to the at least one first-level hash value in the database. Furthermore, the server obtains at least one value corresponding to the at least one first-level site from the array in the database, and determines whether the at least one first-level site is empty according to the at least one obtained value.
[0057] It should be noted that in the embodiment of the present application, the value corresponding to each site in the array is set to a default value, and the default value is used to indicate that the site is empty and no data is stored. When data is stored in the site, the value corresponding to the site is modified to a preset target value to indicate that the site is not empty through the preset target value. The default value in the present application is the number 0, and the preset target value is the number 1.
[0058] When the server determines that at least one first-level site is empty according to the value corresponding to the first-level site, it can directly store the data to be stored in the empty first-level site, thereby completing the storage of the data to be stored.
[0059] In an embodiment of the present application, after the server determines whether at least one first-level site is empty according to at least one value, when it determines that all the first-level sites in the at least one first-level site are not empty, based on the second-level Bloom filter and through at least one second-level hashing algorithm, the server calculates at least one second-level hash value corresponding to the data to be stored, and determines at least one second-level site corresponding to the at least one second-level hash value in the database. Then, the server obtains at least one value corresponding to the at least one second-level site from the array in the database, and can further determine whether the at least one second-level site is empty according to the at least one corresponding value.
[0060] When the server determines that at least one second-level site is empty according to the value corresponding to the second-level site, it can directly store the data to be stored in the empty second-level site, thereby completing the storage of the data to be stored.
[0061] In one embodiment of the present application, after the server obtains at least one value corresponding to at least one secondary site in the database array and determines whether at least one secondary site is empty, when all the secondary sites among at least one secondary site are not empty, the server needs to obtain the deletion times corresponding to all the secondary sites in at least one secondary site, and compare the deletion times corresponding to all the secondary sites with a preset threshold respectively. Furthermore, when the deletion times corresponding to some secondary sites among all the secondary sites are less than the preset threshold, the data stored in the corresponding secondary site is deleted, so as to store the data to be stored in this secondary node where the data has been deleted, and complete the storage of the data to be stored.
[0062] At this time, the deleted data will continue to execute the above steps to find an empty site in the corresponding site. If there is no empty site, the currently deleted data will also perform the same occupancy operation, so that the currently deleted data can also find a site to store itself. When storing data in the database, the above operations will be looped.
[0063] It should be noted that the preset threshold is set in the embodiment of the present application to avoid infinite loops. When the number of times the data in the site is deleted exceeds the preset threshold, that is, when the number of consecutive kick-out behaviors exceeds the preset threshold, it indicates that the site in this array is full, and other data to be stored needs to be re-placed.
[0064] As Figure 2 shown, when the server stores the data to be stored in the database, it calculates at least one primary site through at least one primary hashing algorithm, and then obtains at least one value corresponding to at least one primary site in the database array, so as to determine whether at least one primary site is empty according to the corresponding at least one value; if so, the data to be stored is directly stored in the empty primary site for occupancy to complete the data storage.
[0065] If not, it is necessary to calculate at least one secondary site through at least one secondary hashing algorithm, and then obtain at least one value corresponding to at least one secondary site in the database array, so as to determine whether at least one secondary site is empty according to the corresponding at least one value; if so, the data to be stored is stored in the empty secondary site for occupancy.
[0066] If not, it is necessary to obtain the deletion times corresponding to the data in at least one secondary site, compare the deletion times with the preset threshold, and then delete the data in the secondary site where the deletion times are less than the preset threshold, so as to store the data to be stored in the corresponding secondary site. Then, the deleted data will continue to execute the primary hashing algorithm and the secondary hashing algorithm until the deleted data finds a corresponding site for storage.
[0067] In one embodiment of the present application, after the server compares the deletion times corresponding to all secondary sites with a preset threshold respectively, and determines that the deletion times of all secondary sites are greater than the preset threshold, the server needs to calculate at least one tertiary hash value corresponding to the data to be stored based on a tertiary Bloom filter and through at least one tertiary hash algorithm, and determine at least one tertiary site corresponding to the at least one tertiary hash value in the array of the database.
[0068] Then, the server obtains at least one value corresponding to at least one tertiary site from the array of the database, and determines whether at least one tertiary site is empty according to the corresponding at least one value. When at least one tertiary site is empty, the server can directly store the data to be stored into the empty tertiary site to complete the storage of the data to be stored.
[0069] It should be noted that after the sites in the array in the embodiment of the present application are fully stored with data, the probability of data collision can be greatly reduced and the utilization rate of space can be effectively improved by increasing the hash algorithm, so that each data to be stored has more than two alternative sites for storage.
[0070] When it is determined that at least one tertiary site is not empty, the server needs to continue to calculate at least one quaternary hash value corresponding to the data to be stored based on a quaternary Bloom filter and through at least one quaternary hash algorithm, and determine at least one quaternary site corresponding to the at least one quaternary hash value in the array of the database. Then, the server obtains at least one value corresponding to at least one quaternary site from the array of the database, and further determines whether at least one quaternary site is empty according to the at least one value. When it is determined that at least one quaternary site is empty, the server stores the data to be stored into the corresponding quaternary site.
[0071] As Figure 3 shown, when the server determines that the deletion times corresponding to all data in at least one secondary site are greater than the preset threshold, it needs to calculate at least one corresponding tertiary site through at least one tertiary hash algorithm, and then obtain at least one value corresponding to the at least one tertiary site in the array of the database, so as to determine whether at least one tertiary site is empty according to the corresponding at least one value; if so, the data to be stored can be directly stored into the empty tertiary site for occupancy to complete the data storage.
[0072] Otherwise, the server needs to continue to calculate at least one fourth-level site corresponding to the at least one fourth-level hash algorithm, and then obtain at least one value corresponding to the at least one fourth-level site in the database array, so as to determine whether the at least one fourth-level site is empty according to the at least one corresponding value; if so, the data to be stored can be directly stored in the empty fourth-level site for occupancy. By using four hash algorithms, the present application can greatly improve the utilization rate of space.
[0073] 102. Based on the Bloom filter and through at least one hash algorithm, calculate at least one hash value corresponding to the request data identifier, and determine whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value.
[0074] According to the way of storing data in the database, after receiving a query request from the Web side, the server will calculate at least one hash value corresponding to the request data identifier based on the Bloom filter and through at least one hash algorithm. According to the at least one calculated hash value, at least one site information corresponding to the at least one hash value can be determined, and then the site information corresponding to the determined request data identifier can be queried in the Redis cache. Thus, according to whether the Redis cache contains the site information corresponding to the request data identifier, it can be determined whether the database contains the data corresponding to the request data identifier.
[0075] In an embodiment of the present application, before determining whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value, the server determines the site information where the data to be stored is stored in the database, and stores the site information corresponding to the data to be stored in the Redis cache based on the data identifier corresponding to the data to be stored. So that when the Web side sends a query request to the database, according to the site information with the data identifier stored in the Redis cache, it can be judged whether the data corresponding to the request data identifier in the query request exists in the database. In this way, the query requests corresponding to the data that does not exist in the database can be filtered out, reducing the query pressure on the database and improving the response efficiency of data query.
[0076] In one embodiment of the present application, after the server determines whether the site information corresponding to the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to at least one hash value, when it is determined through querying in the Redis cache according to the site information corresponding to the request data identifier that the Redis cache does not contain the site information corresponding to the request data identifier, the server can determine that the data corresponding to the request data identifier does not exist in the database. At this time, the server will filter out the query request corresponding to the request data identifier and, at the same time, return response information to the corresponding Web end. This can not only reduce the query pressure on the database but also improve the response efficiency of the Web end query.
[0077] 103. If it is determined that the Redis cache contains the site information corresponding to the request data identifier, send the query request to the database and determine whether the value corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value.
[0078] When the server determines through querying that the Redis cache contains the site information corresponding to the request data identifier, it sends the query request to the database, finds the corresponding site in the array of the database according to the site corresponding to the site information, and obtains the value corresponding to the corresponding site in the array to determine whether the corresponding value is a preset target value, so as to determine whether the data corresponding to the request data identifier is stored in the corresponding site according to whether the value is a preset target value.
[0079] It should be noted that in the embodiments of the present application, one hash value of the data corresponds to one site in the bitmap.
[0080] 104. When it is determined that the value corresponding to at least one site in the array of the database is a preset target value, query and obtain the data corresponding to the request data identifier in the corresponding site according to the site information corresponding to the request data identifier, so as to implement the query of the data corresponding to the request data identifier.
[0081] When the server determines that the value corresponding to at least one site in the array of the database is the preset target value 1, it queries and obtains the data corresponding to the request data identifier in the corresponding site according to the site information corresponding to the request data identifier, thereby implementing the query of the data corresponding to the request data identifier.
[0082] In one embodiment of the present application, after the server queries and obtains the data corresponding to the request data identifier in the corresponding site according to the site information corresponding to the request data identifier, based on the query request sent by the Web end, the server can determine the request address corresponding to the Web end included in the query request. Then, according to the determined request address, the server returns the response information corresponding to the query request to the corresponding Web end, enabling the Web end to complete the query of the data.
[0083] It should be noted that Figure 2 , Figure 3 the method shown is Figure 1 essentially the same as the method shown. Therefore, Figure 2 , Figure 3 for the parts not described in detail in Figure 1 , the relevant descriptions in
[0084] can be specifically referred to, and details will not be repeated in this application. Figure 4 shown.
[0085] Figure 4 FIG. is a schematic internal structure diagram of a device for implementing data query based on a Bloom filter provided by an embodiment of this application. As Figure 4 shown, the device includes:
[0086] At least one processor;
[0087] And a memory communicatively connected to at least one processor;
[0088] Wherein, the memory stores instructions executable by at least one processor, and the instructions are executed by at least one processor to enable at least one processor to:
[0089] Receive a query request sent by the Web side to the database and determine the request data identifier in the query request;
[0090] Based on the Bloom filter and through at least one hashing algorithm, calculate at least one hash value corresponding to the request data identifier, and determine whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value;
[0091] If it is determined that the Redis cache contains the site information corresponding to the request data identifier, send the query request to the database, and determine whether the value corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value; wherein, one hash value of the data corresponds to one site in the bitmap;
[0092] In the case where the value corresponding to at least one site in the array of the database is determined to be the preset target value, query and obtain the data corresponding to the request data identifier at the corresponding site according to the site information corresponding to the request data identifier, so as to implement the query of the data corresponding to the request data identifier.
[0093] The embodiments of the present application also provide a non-volatile computer storage medium storing computer-executable instructions, and the computer-executable instructions are configured to:
[0094] Receive a query request sent by the Web side to the database and determine the request data identifier in the query request;
[0095] Based on the Bloom filter and through at least one hashing algorithm, calculate at least one hash value corresponding to the request data identifier, and determine whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value;
[0096] If it is determined that the Redis cache contains the site information corresponding to the request data identifier, send the query request to the database, and determine whether the value corresponding to at least one site corresponding to the request data identifier in the array of the database is a preset target value; wherein, one hash value of the data corresponds to one site in the bitmap;
[0097] When it is determined that the values corresponding to at least one site in the array of the database are preset target values, query and obtain the data corresponding to the request data identifier at the corresponding site according to the site information corresponding to the request data identifier, so as to implement the query of the data corresponding to the request data identifier.
[0098] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0099] The devices and media provided by the embodiments of the present application correspond one by one to the methods. Therefore, the devices and media also have beneficial technical effects similar to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be elaborated here.
[0100] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0101] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows and / or blocks Figure 1 one or more flows and / or blocks Figure 1 or multiple blocks.
[0104] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0105] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.
[0106] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0107] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0108] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A method for implementing data query based on Bloom filter, characterized in that, The method includes: Receiving a query request sent by a Web client to a database, and determining a request data identifier in the query request; Based on a Bloom filter and through at least one hashing algorithm, calculating at least one hash value corresponding to the request data identifier, and determining whether the Redis cache contains site information corresponding to the request data identifier according to site information corresponding to the at least one hash value; If it is determined that the Redis cache contains site information corresponding to the request data identifier, sending the query request to the database, and determining whether a value corresponding to at least one site corresponding to the request data identifier in an array of the database is a preset target value; wherein, one hash value of data corresponds to one site in a bitmap, and the preset target value is used to indicate that the site is not empty; When it is determined that a value corresponding to the at least one site in the database array is the preset target value, querying and obtaining data corresponding to the request data identifier at a corresponding site according to the site information corresponding to the request data identifier, so as to implement querying of data corresponding to the request data identifier; Before determining whether the Redis cache contains site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value, the method further includes: Based on a data identifier corresponding to data to be stored, storing site information for storing the data to be stored in the Redis cache, so as to determine whether the database contains the data to be stored through the Redis cache.
2. The method for implementing data query based on a Bloom filter according to claim 1, wherein Before receiving the query request sent by the Web client to the database, the method further includes: Obtaining a plurality of data to be stored in the database, and configuring a corresponding array for the plurality of data to be stored in the database; Based on a first-level Bloom filter and through at least one first-level hashing algorithm, calculating at least one first-level hash value corresponding to the data to be stored, and determining at least one first-level site corresponding to the at least one first-level hash value in the database; Obtaining at least one value corresponding to the at least one first-level site in the database array, and determining whether the at least one first-level site is empty according to the at least one value; If so, storing the data to be stored in a corresponding first-level site, and completing storage of the data to be stored.
3. The method for implementing data query based on Bloom filter according to claim 2, characterized in that, After determining whether the at least one first-level site is empty according to the at least one value, the method further includes: If all the first-level sites in the at least one first-level site are not empty, based on a second-level Bloom filter and through at least one second-level hashing algorithm, calculating at least one second-level hash value corresponding to the data to be stored, and determining at least one second-level site corresponding to the at least one second-level hash value in the database; Obtaining at least one value corresponding to the at least one second-level site in the database array to determine whether the at least one second-level site is empty; If so, storing the data to be stored in a corresponding second-level site, and completing storage of the data to be stored.
4. A method for implementing data query based on a Bloom filter according to claim 3, characterized in that, After obtaining at least one value corresponding to the at least one secondary site in the array of the database to determine whether the at least one secondary site is empty, the method further includes: If all the secondary sites among the at least one secondary site are not empty, obtain the deletion times corresponding to all the secondary sites in the at least one secondary site, and compare the deletion times corresponding to all the secondary sites with a preset threshold respectively; Determine the secondary sites whose deletion times are less than the preset threshold, and determine one secondary site from the secondary sites whose deletion times are less than the preset threshold; Delete the data stored in the determined secondary site, and store the data to be stored in the determined secondary site.
5. A method for implementing data query based on Bloom filter according to claim 4, characterized in that, After comparing the deletion times corresponding to all the secondary sites with the preset threshold respectively, the method further includes: If it is determined that the deletion times corresponding to all the secondary sites are greater than the preset threshold, calculate at least one tertiary hash value corresponding to the data to be stored based on a tertiary Bloom filter and through at least one tertiary hash algorithm, and determine at least one tertiary site corresponding to the at least one tertiary hash value in the database; If it is determined that the at least one tertiary site is empty according to the at least one value corresponding to the at least one tertiary site in the array, store the data to be stored in the corresponding tertiary site to complete the storage of the data to be stored; If it is determined that the at least one tertiary site is not empty, continue to calculate at least one quaternary hash value corresponding thereto based on a quaternary Bloom filter and through at least one quaternary hash algorithm, and store the data to be stored in the corresponding quaternary site when it is determined that the at least one quaternary site is empty according to the at least one value corresponding to the at least one quaternary site corresponding to the at least one quaternary hash value in the array.
6. A method for implementing data query based on a Bloom filter according to claim 1, characterized in that After determining whether the Redis cache contains the site information corresponding to the request data identifier according to the site information corresponding to the at least one hash value, the method further includes: If it is determined that the Redis cache does not contain the site information corresponding to the request data identifier, determine that the database does not contain the data corresponding to the request data identifier; Filter the query request corresponding to the request data identifier, and return a prompt message to the Web end corresponding to the query request.
7. A method for implementing data query based on Bloom filter according to claim 1, characterized in that, After querying and obtaining the data corresponding to the request data identifier in the corresponding site according to the site information corresponding to the request data identifier, the method further includes: Based on the query request of the Web end, determine the address corresponding to the Web end included in the query request; Based on the determined address corresponding to the Web end, return the data corresponding to the request data identifier obtained by querying to the Web end to respond to the query request of the Web end.
8. A device for implementing data query based on a Bloom filter, characterized in that The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute a method for implementing data query based on a Bloom filter as described in any one of claims 1-7.
9. A non-volatile computer storage medium storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: A method for implementing data query based on a Bloom filter as described in any one of claims 1-7.
Citation Information
Patent Citations
Data query method and device, storage medium and electronic equipment
CN112632342A