Data resource management system and method based on Bloom filter
By introducing Bloom filters in the data resource management system for front-end filtering, the problem of inefficiency in traditional methods in massive data processing is solved, rapid data query and audit is realized, system load is reduced, response speed and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202510296343.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-08
AI Technical Summary
Traditional data management methods are inefficient when processing massive data, especially in scenarios such as data query, deduplication and verification, which is difficult to meet real-time requirements.
Bloom filter is used to filter data in the front-end processing stage to reduce the amount of data evaluated in the back-end. Bloom filter is created, initialized and maintained through the Bloom filter module, and combined with data resource query cache, content audit and recommendation modules, it can quickly determine the existence of data and deduplication.
It reduces the processing pressure of the back-end server, improves the stability and response speed of the system, reduces the workload of manual audits, improves the overall efficiency of data resource management and the processing speed of recommended algorithms.
Smart Images

Figure CN120277247A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data resource management, and particularly to a data resource management system and method based on a Bloom filter. Background Art
[0002] With the advent of the big data era, data resource management has become one of the core challenges faced by enterprises, institutions and various organizations. The rapid growth and diversification of data resources have made traditional data management methods show obvious limitations when dealing with massive data, especially in scenarios such as data query, deduplication and verification, where the problems of low efficiency and excessive resource consumption are becoming increasingly prominent. And with the increase in the amount of data, the efficiency of traditional query methods in dealing with massive data decreases significantly, and the latency and resource consumption of verification operations also increase significantly. Summary of the Invention
[0003] In order to solve the above technical problems, the present invention is proposed. Embodiments of the present invention provide a data resource management system and method based on a Bloom filter, which can effectively reduce the amount of data that needs to be transmitted to the backend for detailed evaluation in the front-end processing stage based on the Bloom filter, relieve the processing pressure on the backend server, reduce the overall load of the system, and improve the stability and response speed of the system.
[0004] According to one aspect of the present invention, there is provided a data resource management system based on a Bloom filter, including: a Bloom filter module for creating, initializing and maintaining at least one Bloom filter to determine whether an element may exist in a set; a data resource query cache module communicatively connected to the Bloom filter module, the data resource query cache module being used for determining whether the data resource to be queried exists in the cache during the data resource query process; a data resource content review module communicatively connected to the Bloom filter module, the data resource content review module being used for reviewing whether there is any violation information in the data resource to be reviewed; and a data resource recommendation module communicatively connected to the Bloom filter module, the data resource recommendation module being used for filtering out the data resources that have been recommended when generating a recommendation list.
[0005] In one embodiment, the Bloom filter module includes: a creation unit configured to create a Bloom filter according to preset parameters; an initialization unit configured to add target data resource information to the Bloom filter; a calling unit configured to call the Bloom filter to determine whether an element exists; wherein the element includes at least one of a data resource to be queried, a violation information, and a data resource that has been recommended; an expansion unit configured to expand the Bloom filter when the data volume increases; wherein the expansion includes direct expansion or multi-layer expansion.
[0006] In one embodiment, the data resource content auditing module includes: a violation content set establishment unit configured to collect and organize known violation information and add the known violation information as a set to the Bloom filter; a content screening unit configured to, when auditing a data resource, screen whether there is any violation information in the data resource to be audited through the Bloom filter.
[0007] In one embodiment, the data resource recommendation module includes: a user behavior recording unit configured to record the operation behaviors of a user; wherein the operation behaviors of the user include the data resources applied for by the user; a recommendation list generation unit configured to generate a recommendation list according to the data resources applied for by the user and filter out the data resources that have been recommended through the Bloom filter.
[0008] According to another aspect of the present invention, there is provided a data resource management method based on a Bloom filter, including: creating, initializing, and maintaining at least one Bloom filter through the Bloom filter module; during the data resource query process, calling the Bloom filter through the data resource query cache module to determine whether the data resource to be queried exists in the cache; during the data resource content auditing process, calling the Bloom filter through the data resource content auditing module to audit whether there is any violation information in the data resource to be audited; during the data resource recommendation process, calling the Bloom filter through the data resource recommendation module to filter out the data resources that have been recommended when generating a recommendation list.
[0009] In one embodiment, the data resource management method based on a Bloom filter further includes: obtaining new resource data; when there is the new resource data, writing the new resource data into the Bloom filter to form a new set.
[0010] In one embodiment, when there is the new resource data, writing the new resource data into the Bloom filter includes: determining whether to expand the Bloom filter according to the new resource data and the original resource data; when the Bloom filter meets the preset expansion condition, creating a new Bloom filter, and writing both the original resource data and the new resource data into the new Bloom filter; wherein, the storage capacity of the new Bloom filter is a preset multiple of the storage capacity of the Bloom filter before expansion; or when the Bloom filter meets the preset expansion condition, adding a new Bloom filter, and adding and writing the new resource data into any newly added Bloom filter.
[0011] In one embodiment, creating, initializing, and maintaining at least one Bloom filter by a Bloom filter module includes: creating a Bloom filter according to preset parameters; wherein, the preset parameters include the length of the bit array, the number of hash functions, and the number of elements expected to be stored; the number of elements expected to be stored is inversely proportional to the false positive rate of the Bloom filter, and the number of elements expected to be stored is directly proportional to the memory occupied by the Bloom filter; integrating the target data resource information according to a preset rule and adding it to the Bloom filter.
[0012] In one embodiment, the data resource content review module includes: a violation content set establishment unit and a content screening unit; wherein, the data resource management method based on the Bloom filter includes: collecting and sorting known violation information by the violation content set establishment unit, and adding the known violation information as a set to the Bloom filter; wherein, during the data resource content review process, calling the Bloom filter through the data resource content review module to review whether there is violation information in the data resource to be reviewed, including: calling the Bloom filter through the content screening unit to screen whether there is violation information in the data resource to be reviewed.
[0013] In one embodiment, the data resource recommendation module includes: a user behavior record unit and a recommendation list generation unit; wherein, the data resource management method based on the Bloom filter includes: recording the operation behavior of the user; wherein, the operation behavior of the user includes the data resources applied for by the user; during the data resource recommendation process, calling the Bloom filter through the data resource recommendation module, and filtering out the data resources that have been recommended when generating the recommendation list, including: generating a recommendation list according to the data resources applied for by the user; calling the Bloom filter through the recommendation list generation unit to filter out the data resources that have been recommended in the recommendation list, and generating a filtered recommendation list.
[0014] The data resource management system and method based on Bloom filter provided by the present invention. The Bloom filter is a probabilistic data structure with extremely high space efficiency, which can quickly determine whether an element belongs to a set. By taking advantage of the space efficiency and query speed of the Bloom filter, the system can effectively reduce the amount of data that needs to be transferred to the backend for detailed evaluation during the front-end processing stage, reducing the processing pressure on the backend server. It can determine whether the content of the data resource may contain illegal information in a very short time, reducing the workload of manual review, thus significantly improving the overall review efficiency. Moreover, it can quickly determine whether a certain resource data belongs to the set that the user may be interested in, thereby accelerating the processing speed of the recommendation algorithm. Therefore, by virtue of the advantages of space efficiency and query speed, the Bloom filter can effectively solve the efficiency problems in massive data query, verification, and recommendation of traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] By describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present invention will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification, and are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 FIG. is a schematic structural diagram of a data resource management system based on Bloom filter provided by an exemplary embodiment of the present invention.
[0017] Figure 2 FIG. is a schematic flowchart of a data resource management method based on Bloom filter provided by an exemplary embodiment of the present invention.
[0018] Figure 3 FIG. is a schematic logical diagram of creating a Bloom filter provided by an exemplary embodiment of the present invention.
[0019] Figure 4 FIG. is a schematic logical diagram of Bloom filter data query provided by an exemplary embodiment of the present invention.
[0020] Figure 5 FIG. is a schematic logical diagram of Bloom filter expansion provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] Next, exemplary embodiments of the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments of the present invention. It should be understood that the present invention is not limited by the exemplary embodiments described herein.
[0022] In data resource management, the query operation is one of the most common requirements. However, with the explosive growth of data volume, the efficiency of traditional query methods (such as database-based index queries) significantly decreases when dealing with massive data. For example, in cross-departmental or cross-organizational data sharing scenarios, data may be scattered across multiple systems, and query requests need to be retrieved from multiple data sources, resulting in long query response times. In addition, traditional exact query methods require storing and retrieving complete data records, consuming a large amount of storage space and computing resources.
[0023] In data resource management, data validation is a crucial step to ensure data quality and compliance. For example, when data is entered into a database or shared, it is necessary to quickly verify whether a certain piece of data already exists or meets specific conditions. Traditional validation methods usually require traversing the entire dataset or relying on complex index structures, making it difficult to meet real-time requirements. In addition, as the data volume increases, the latency and resource consumption of validation operations also increase significantly.
[0024] In the big data scenario, storage space is an important resource constraint factor. Traditional exact query and deduplication methods require storing complete data records, which will occupy a large amount of storage space on large-scale datasets. For example, in a distributed system, each node may need to store a complete data copy, further exacerbating the storage pressure. In addition, as the data volume increases, the storage cost also rises significantly.
[0025] A Bloom Filter is a probabilistic data structure with high space efficiency that allows testing whether an element is in a set with a low false positive rate. It is constructed based on a bit array and a series of hash functions, and the element is mapped to different positions in the bit array through multiple hash functions. When an element is inserted, the corresponding positions in the bit array are marked as 1. When determining whether an element exists, the same hash mapping operation is performed on the element. If all the corresponding positions in the bit array are 1, it may exist in the set; if any of the positions in the bit array is 0, it definitely does not exist in the set. Although a Bloom Filter may have false positives (i.e., misjudging non-existent elements as existent), it will never have false negatives (i.e., existent elements will always be judged as existent). This characteristic makes the Bloom Filter very useful in scenarios that require efficient processing of large amounts of data and allow a certain false positive rate.
[0026] Therefore, the present application proposes a data resource management system and method based on a Bloom filter. The application of the Bloom filter in data caching is mainly reflected in its efficient data existence judgment ability and effective mitigation of the cache penetration problem. Cache penetration means that when a user queries a non-existent data, since the data does not exist in the cache, each query needs to access the database, which brings a huge pressure to the database. In a data resource management system, the query of a large amount of data is undoubtedly the top priority of the system function. A large number of read and write operations in the database will be normal. The cache is undoubtedly the first checkpoint. In a project, middleware such as redis and Memcached is generally used as the cache component. Before this layer of middleware, a Bloom filter is used for data filtering. The Bloom filter can effectively solve the efficiency problems of traditional methods in massive data query, deduplication, and verification through its advantages in space efficiency and query speed.
[0027] Figure 1 is a schematic structural diagram of a data resource management system based on a Bloom filter provided by an exemplary embodiment of the present invention, as Figure 1 shown, the data resource management system 1 based on a Bloom filter includes: a Bloom filter module 11, which is used to create, initialize, and maintain at least one Bloom filter to determine whether an element may exist in a set; a data resource query cache module 12, which is communicatively connected to the Bloom filter module 11 and is used to determine whether the data resource to be queried exists in the cache during the data resource query process; a data resource content review module 13, which is communicatively connected to the Bloom filter module 11 and is used to review whether there is any illegal information in the data resource to be reviewed; a data resource recommendation module 14, which is communicatively connected to the Bloom filter module 11 and is used to filter out the data resources that have been recommended when generating a recommendation list.
[0028] The Bloom filter is created through the Bloom filter module. When setting the Bloom filter by parameters, the main considerations are how to balance the false positive rate, space efficiency, and performance. When initializing the Bloom filter, the existing data resource information is added to the created Bloom filter, usually after integrating the query conditions. Maintaining the Bloom filter means performing necessary operations and management on it during its use to ensure that its performance, accuracy, and efficiency can meet the actual requirements. The maintenance process may include dynamically adjusting the Bloom filter. For example, if the number of elements exceeds the expectation, a larger Bloom filter can be created and the original data can be migrated to the new filter. Or if the false positive rate is too high or the bit array is full, the Bloom filter can be re-initialized and elements can be added again. It can also be regularly backing up the bit array of the Bloom filter to prevent data loss. In case of system failure or data corruption, the state of the Bloom filter can be restored from the backup. Therefore, the maintenance work can include adding elements, querying elements, handling false positives, dynamic adjustment, performance monitoring, data backup, etc. Through reasonable maintenance, the advantages of the Bloom filter in terms of space efficiency and query speed can be fully utilized to meet the needs of actual applications.
[0029] The data resource query cache module can be used to solve the problem of cache penetration. It calls the Bloom filter for data filtering to reduce the frequency of accessing the database and relieve the database pressure. It effectively reduces the amount of data that needs to be passed to the backend for detailed evaluation during the front-end processing stage, which greatly reduces the processing pressure on the backend server, reduces the overall load of the system, and improves the stability and response speed of the system.
[0030] In the data resource management system, the created data resources often need to be audited before they can be applied for use by other organizations or units. The audit indicators generally include but are not limited to: data usage purpose, data security, qualification of the creating unit, content of the data resources, etc. Among them, the audit of the content is often time-consuming and laborious. Therefore, a data resource content audit module is set up to introduce the Bloom filter. Its application in content audit is mainly reflected in quickly determining whether a certain content (such as text, picture, video, etc.) belongs to a known set of bad or illegal content, which accelerates the audit process, reduces the workload of manual audit, and thus significantly improves the overall audit efficiency.
[0031] Having excellent resource recommendation capabilities is often a plus for a data resource management system. The application of Bloom filters in data recommendation is mainly reflected in quickly and efficiently determining whether a certain resource data exists in a specific set, thereby optimizing the performance and efficiency of the recommendation algorithm. A data resource recommendation module is set up to introduce a Bloom filter. When users browse a list of a large number of data resource sets, the system records the users' operation behaviors, and then uses the Bloom filter to record the data resources that the users have applied for. When generating a recommendation list, the recommended data resources are first filtered out by the Bloom filter, and the remaining recommendation results are then presented to the users, thereby improving the accuracy and diversity of the recommendation results.
[0032] In one embodiment, the Bloom filter module includes: a creation unit, which is used to create a Bloom filter according to preset parameters; an initialization unit, which is used to add target data resource information to the Bloom filter; a call unit, which is used to call the Bloom filter to determine whether an element exists; wherein, the element includes at least one of the data resource to be queried, violation information, and the recommended data resources; an expansion unit, which is used to expand the Bloom filter when the data volume increases; wherein, the expansion includes direct expansion or multi-layer expansion.
[0033] The key parameters of the Bloom filter include the length (m) of the bit array, the number (k) of hash functions, and the number (n) of elements to be stored. The key parameters can be set according to the actual data volume of the business. For example, the larger the length of the bit array, the lower the false positive rate of the Bloom filter, but the larger the storage space occupied. The more hash functions, the higher the accuracy of the Bloom filter, but too many hash functions will waste space and the calculation cost will also be higher. Too few hash functions will lead to an increase in the false positive rate. Appropriate parameters need to be selected according to the number of elements to be stored and the acceptable false positive rate. If the false positive rate is too high or the bit array is full, the Bloom filter can be re-initialized by the initialization unit and elements can be added again. The call unit is used to call the Bloom filter to determine whether an element exists when needed. The expansion unit is used to expand the Bloom filter when the data volume increases, including direct expansion or multi-layer expansion. For example, if the number of elements exceeds the expectation, a larger Bloom filter can be created and the original data can be migrated to the new filter, which is the method of direct expansion, or multiple filters can be added and searched by traversing multiple filters in a loop, and the new data is added to the new Bloom filter. Based on the fast judgment ability of the Bloom filter, the system can effectively reduce the amount of data that needs to be transferred to the backend for detailed evaluation during the front-end processing stage, which greatly reduces the processing pressure on the backend server, reduces the overall load of the system, and improves the stability and response speed of the system.
[0034] In one embodiment, the data resource content review module includes: a violation content set establishment unit, which is used to collect and organize known violation information and add the known violation information as a set to the Bloom filter; a content screening unit, which is used to screen whether there is violation information in the data resource to be reviewed through the Bloom filter when reviewing the data resource.
[0035] The violation content set establishment unit is used to collect and organize known bad or violation content and add it to the Bloom filter. The content screening unit is used to quickly screen through the Bloom filter when reviewing the content of new data resources. If it is judged to be a violation, the application will be automatically rejected. When new review content is added, the newly added content to be reviewed and screened also needs to be added to the Bloom filter. Based on the set already added in the Bloom filter, with its efficient query performance, the Bloom filter can judge whether the data resource content may contain violation or bad information in an extremely short time. This feature accelerates the review process, reduces the workload of manual review, and thus significantly improves the overall review efficiency. This advantage is particularly obvious in scenarios where a large number of data resources need to be reviewed.
[0036] In one embodiment, the data resource recommendation module includes: a user behavior recording unit, which is used to record the operation behaviors of the user; among them, the operation behaviors of the user include the data resources applied by the user; a recommended list generation unit, which is used to generate a recommended list according to the data resources applied by the user and filter out the data resources that have been recommended through the Bloom filter.
[0037] The user behavior recording unit is used to record the operation behaviors of the user, especially the data resources that the user has applied. The recommended list generation unit is used to filter out the recommended data resources through the Bloom filter when generating the recommended list to improve the diversity and relevance of the recommended results. That is, the data resources that the user has queried or applied to query can be screened for recommended items through the Bloom filter to avoid repeated recommendation of the same content. When generating the recommended list, first filter out the data resources that have been recommended through the Bloom filter, and then display the remaining recommended results to the user, thereby improving the accuracy and diversity of the recommended results.
[0038] The Bloom filter can be introduced through the following methods: In a data resource management system using the Java language, since there is no direct implementation of the Bloom filter provided, but a third-party library such as Google's Guava library can be used, and it can be applied by introducing dependencies in a Maven project; for a system based on the Go language, a third-party library such as google / bloomfilter can also be used.
[0039] Figure 2 This is a schematic structural diagram of a data resource management device based on a Bloom filter provided by an exemplary embodiment of the present invention. As Figure 2 shown, the data resource management method based on a Bloom filter includes:
[0040] S100: Create, initialize, and maintain at least one Bloom filter through a Bloom filter module.
[0041] In some embodiments, S100 may include: creating a Bloom filter according to preset parameters; wherein the preset parameters include the length of the bit array, the number of hash functions, and the number of elements expected to be stored; the number of elements expected to be stored is inversely proportional to the false positive rate of the Bloom filter, and the number of elements expected to be stored is proportional to the memory occupied by the Bloom filter; integrating the target data resource information according to a preset rule and adding it to the Bloom filter.
[0042] Creating a Bloom filter according to preset parameters is the first and crucial step in using a Bloom filter. The performance of a Bloom filter (such as false positive rate, storage space occupancy, etc.) directly depends on the setting of its parameters. Before creating a Bloom filter, it is necessary to clarify the preset parameters, that is, the number of elements expected to be added to the Bloom filter (the number of elements expected to be stored), the acceptable false positive rate, the length of the bit array, and the number of hash functions. The hash function needs to have good distribution to reduce conflicts and should be computationally efficient to improve the performance of the Bloom filter, such as MurmurHash, FNV Hash, etc. If k hash functions are required, different hash functions can be used, or the same hash function can be used, but different processing is performed on the input (such as salting). Initialization can be to initialize all bits of the bit array to 0. Then, to add the target data resource information to the created Bloom filter, calculate the hash values of the elements to be added using k hash functions, and set the corresponding bits in the bit array to 1. Generally, the query conditions are integrated and added, such as: resource name, resource type, open conditions, department, etc.
[0043] In a possible implementation manner, Figure 3 This is a schematic logical diagram of creating a Bloom filter provided by an exemplary embodiment of the present invention. As Figure 3 shown, first create a Bloom filter (see Figure 3 S31 in Figure 3 ), then perform parameter settings: the length of the bit array (m), the number of hash functions (k), and the number of subsequent storage elements (n) (see Figure 3 S32 in Figure 3in S34).
[0044] S200: During the data resource query process, the Bloom filter is called through the data resource query cache module to determine whether the data resource to be queried exists in the cache.
[0045] Figure 4 is a logical schematic diagram of Bloom filter data query provided by an exemplary embodiment of the present invention. As Figure 4 shown, first, the data query process starts (see S41 in Figure 4 ), then the Bloom filter is called (see S42 in Figure 4 ) to determine whether the data resource to be queried exists in the cache. If so, it may exist, and then the storage system is queried in detail (see S43 in Figure 4 ), and finally it ends (see S45 in Figure 4 ). If the Bloom filter determines that the data resource to be queried does not exist in the cache, it definitely does not exist, and an empty set is returned (see S44 in Figure 4 ), and finally it ends (see S45 in Figure 4 ).
[0046] S300: During the data resource content review process, the Bloom filter is called through the data resource content review module to review whether there is any violation information in the data resource to be reviewed.
[0047] In some embodiments, the data resource content review module includes: a violation content set establishment unit and a content screening unit; wherein, the data resource management method based on the Bloom filter may include: collecting and organizing known violation information through the violation content set establishment unit and adding the known violation information as a set to the Bloom filter; therefore, S300 may include: screening whether there is any violation information in the data resource to be reviewed by calling the Bloom filter through the content screening unit.
[0048] In a data resource management system, the created data resources often need to be audited before they can be applied for use by other organizations or units. The audit metrics generally include, but are not limited to: data usage purpose, data security, qualification of the creating unit, content of the data resources, and so on. Among them, the audit of content is often time-consuming and laborious. The application of Bloom Filter in content audit is mainly reflected in quickly judging whether a certain content (such as text, picture, video, etc.) belongs to a known set of bad or illegal content. Although the Bloom Filter itself does not directly handle the complex logic of content audit (such as semantic analysis, image recognition, etc.), it can efficiently assist this process by quickly excluding the items that obviously do not belong to the illegal content, thereby reducing the burden of subsequent detailed audit. For example, first establish a set of illegal content. Collect and organize the known bad or illegal content, such as sensitive words, characteristics of illegal information, etc. And add them to the Bloom Filter. The creation process of the Bloom Filter refers to the above process. Then perform content screening. When there is new data resource content that needs to be audited, first quickly screen it through the Bloom Filter. If the Bloom Filter determines that the content may exist in the set of illegal content (that is, all mapped positions are 1), the application for creating the data resource is automatically rejected. If it does not contain, perform subsequent other audit processes. Finally, if new audit content is added, the newly added content that needs to be audited and screened also needs to be added to the Bloom Filter, and situations such as expansion also need to be considered.
[0049] S400: During the data resource recommendation process, the Bloom Filter is called through the data resource recommendation module, and when generating the recommendation list, the recommended data resources are filtered out.
[0050] In an embodiment, the data resource recommendation module includes: a user behavior record unit and a recommendation list generation unit; wherein, the data resource management method based on the Bloom Filter may include: recording the operation behavior of the user; wherein, the operation behavior of the user includes the data resources applied for by the user; S400 may include: generating a recommendation list according to the data resources applied for by the user; calling the Bloom Filter through the recommendation list generation unit to filter out the recommended data resources in the recommendation list, and generating a filtered recommendation list.
[0051] In a data resource management system, to improve the utilization efficiency of data resources and the user experience, personalized recommendation lists are usually generated based on the user's historical behaviors (such as applications, accesses, downloads, etc.). However, to avoid recommending data resources that have already been applied for, the recommendation list needs to be filtered through a Bloom filter to generate a deduplicated and accurate recommendation list. For example, the system first collects the user's historical behavior data, including the data resources applied by the user, access frequencies, download records, etc., and extracts the user's interest preferences and demand characteristics by analyzing the user's behavior data. For example, if a user frequently applies for a certain type of data resource, the system can infer that the user has a high interest in this type of resource. Based on the user's behavior data, algorithms such as collaborative filtering, content-based recommendation, or deep learning are used to generate an initial recommendation list. The results output by the recommendation algorithm are sorted into an ordered recommendation list, which contains the identifiers (such as IDs, names, etc.) of multiple data resources. When the system starts, the Bloom filter is initialized according to preset parameters (such as the expected number of elements, false positive rate). The size of the bit array and the number of hash functions of the Bloom filter need to be optimized according to actual requirements. The identifiers (such as IDs) of the data resources that the user has already applied for are added to the Bloom filter. The Bloom filter can efficiently store and query large-scale data resource identifiers while occupying very little storage space. For each data resource identifier in the recommendation list, the Bloom filter is called for querying. The data resources with the query result of "not recommended yet" in the Bloom filter are retained in the recommendation list, and the data resources with the query result of "possibly recommended" are removed from the recommendation list. After being filtered by the Bloom filter, the recommendation list only contains the data resources that the user has not applied for, avoiding duplicate recommendations. According to the weights output by the recommendation algorithm or the user interest matching degree, the filtered recommendation list is sorted to ensure that the most relevant and valuable data resources are ranked at the front. The filtered recommendation list is displayed to the user for the user to select and apply. By combining user behavior analysis and Bloom filter technology, a deduplicated recommendation list can be efficiently generated, thereby improving the intelligent level and user experience of the data resource management system.
[0052] In one embodiment, the data resource management method based on the Bloom filter further includes: obtaining new resource data; when there is new resource data, writing the new resource data into the Bloom filter to form a new set. After adding new resource data, the data also needs to be written into the Bloom filter at the same time to ensure subsequent continuous use.
[0053] In one embodiment, when there is new resource data, writing the new resource data into the Bloom filter includes: determining whether to expand the Bloom filter according to the new resource data and the original resource data; when the Bloom filter meets the preset expansion condition, creating a new Bloom filter and writing both the original resource data and the new resource data into the new Bloom filter; wherein, the storage capacity of the new Bloom filter is a preset multiple of the storage capacity of the Bloom filter before expansion; or when the Bloom filter meets the preset expansion condition, adding a new Bloom filter and adding and writing the new resource data into any newly added Bloom filter.
[0054] If the quantity of the new resource data exceeds the expectation, as the data volume increases, the false positive rate of the Bloom filter may gradually increase. Therefore, it is necessary to expand it to maintain performance. The expansion methods of the Bloom filter mainly include direct expansion and multi-layer expansion. Direct expansion means creating a larger Bloom filter on the basis of the original Bloom filter and migrating the original data to the new filter. For example, according to the new expected number of stored elements and the false positive rate, calculating the length of the new bit array and the number of hash functions, traversing all elements in the original Bloom filter, recalculating the hash values, and setting the corresponding bits into the new bit array. If the original Bloom filter supports element query (such as counting Bloom filter), the elements can be directly read and added to the new filter. Replacing the original filter with the new Bloom filter completes the expansion. The implementation logic of direct expansion is clear, and the false positive rate can be precisely controlled by adjusting the size of the bit array and the number of hash functions. Multi-layer expansion means realizing dynamic expansion by constructing multiple Bloom filters (usually called hierarchical Bloom filter or cascaded Bloom filter) and distributing the data to Bloom filters at different levels. For example, creating multiple Bloom filters, the size of the bit array and the number of hash functions of each filter can be the same or different. Usually, the first-layer filter is smaller and used to store recent data; the subsequent-layer filters are larger and used to store historical data. New elements are first inserted into the first-layer filter. When the first-layer filter is close to saturation (such as the proportion of 1 in the bit array exceeds a certain threshold), its data is migrated to the second-layer filter, and the first-layer filter is cleared. Subsequently, the data can be migrated layer by layer to higher-layer filters. When querying the multi-layer expanded Bloom filter, start querying layer by layer from the first-layer filter until finding the matching layer or traversing all layers. If an element may exist is queried in any layer of the filter, return "may exist"; otherwise, return "certainly does not exist". Multi-layer expansion does not require migrating all data at one time, the expansion process is smoother, recent data is usually stored in the first-layer filter, the query speed is faster, and by storing in layers, the size of the bit array of a single filter can be reduced, saving storage space.
[0055] As a possible implementation manner, Figure 5It is a logical schematic diagram of the expansion of the Bloom filter provided by an exemplary embodiment of the present invention. As Figure 5 shown, first, the process of adding new data starts (see S51 in Figure 5 ), then it is written into the Bloom filter (see S52 in Figure 5 ), and then it is checked whether expansion is needed (see S53 in Figure 5 ). If no expansion is needed, it directly ends (see S58 in Figure 5 ). If expansion is needed, direct expansion can be selected (see S51 in Figure 5 ), that is, a new Bloom filter is established, n times the size, and the original data is added (see S55 in Figure 5 ), and finally it ends (see S58 in Figure 5 ). Or multi-layer expansion can be selected (see S56 in Figure 5 ), that is, multiple filters are added, and the new data is added to the new filter (see S57 in Figure 5 ), and finally it ends (see S58 in Figure 5 ).
[0056] An embodiment of the present invention provides a data resource management device based on a Bloom filter. The device embodiment can be implemented by software, or by hardware or a combination of software and hardware. From the hardware level, in addition to the CPU, memory, network interface, and non-volatile memory, the device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, etc. Taking software implementation as an example, as a logically meaningful device, it is formed by the CPU of the device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running.
[0057] According to another aspect of the present invention, a computer-readable storage medium is provided. The storage medium stores a computer program, and the computer program is used to execute the data resource management method based on the Bloom filter in any of the above embodiments.
[0058] In addition to the above methods and devices, an embodiment of the present invention may also be a computer program product, which includes computer program instructions. When the computer program instructions are run by a processor, the processor executes the steps in the data resource management method based on the Bloom filter in various embodiments of the present invention described above.
[0059] According to another aspect of the present invention, an electronic device is provided. The electronic device includes: a processor; a memory for storing processor-executable instructions; and a processor for executing the data resource management method based on the Bloom filter in any of the above embodiments.
[0060] In addition, an embodiment of the present invention may also be a computer-readable storage medium storing computer program instructions, which, when run by a processor, cause the processor to execute the steps in the data resource management method based on a Bloom filter according to various embodiments of the present invention described above.
[0061] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A data resource management system based on a Bloom filter, characterized in that, include: A Bloom filter module, the Bloom filter module is used to create, initialize and maintain at least one Bloom filter to determine whether an element may exist in a set; A data resource query cache module, the data resource query cache module is in communication with the Bloom filter module, and the data resource query cache module is used to determine whether the data resource to be queried exists in the cache during the data resource query process; A data resource content review module, the data resource content review module is in communication connection with the Bloom filter module, and the data resource content review module is used to review whether there is any illegal information in the data resource to be reviewed; A data resource recommendation module, the data resource recommendation module is in communication connection with the Bloom filter module, and the data resource recommendation module is used to filter out the data resources that have been recommended when generating the recommendation list.
2. The data resource management system based on a Bloom filter according to claim 1, characterized in that The Bloom filter module includes: A creation unit, the creation unit is used to create a Bloom filter according to preset parameters; an initialization unit, the initialization unit being used to add target data resource information to the Bloom filter; A calling unit, the calling unit is used to call the Bloom filter to determine whether an element exists; wherein the element includes at least one of a data resource to be queried, violation information, and a recommended data resource; An expansion unit, wherein the expansion unit is used to expand the Bloom filter when the amount of data increases; wherein the expansion includes direct expansion or multi-layer expansion.
3. The data resource management system based on the Bloom filter according to claim 1, characterized in that, The data resource content review module includes: a violation content set establishing unit, the violation content set establishing unit being used to collect and organize known violation information, and add the known violation information as a set to the Bloom filter; The content screening unit is used to screen whether there is any illegal information in the data resources to be reviewed through a Bloom filter when reviewing the data resources.
4. The data resource management system based on the Bloom filter according to claim 1, characterized in that, The data resource recommendation module includes: A user behavior recording unit, the user behavior recording unit is used to record the user's operation behavior; wherein the user's operation behavior includes data resources applied for by the user; A recommendation list generating unit is used to generate a recommendation list according to data resources applied for by the user, and filter out the recommended data resources through a Bloom filter.
5. A data resource management method based on a Bloom filter, characterized in that, include: Creating, initializing and maintaining at least one Bloom filter via a Bloom filter module; During the data resource query process, the Bloom filter is called through the data resource query cache module to determine whether the data resource to be queried exists in the cache; During the data resource content review process, the Bloom filter is called through the data resource content review module to review whether there is any illegal information in the data resource to be reviewed; In the process of data resource recommendation, the Bloom filter is called through the data resource recommendation module to filter out the recommended data resources when generating the recommendation list.
6. The data resource management method based on the Bloom filter according to claim 5, characterized in that, The data resource management method based on Bloom filter also includes: Get new resource data; When the newly added resource data exists, the newly added resource data is written into the Bloom filter to form a new set.
7. The data resource management method based on the Bloom filter according to claim 6, wherein When there is the new resource data, writing the new resource data into the Bloom filter includes: Determining whether to expand the Bloom filter according to the new resource data and the original resource data; When the Bloom filter meets the preset expansion condition, creating a new Bloom filter and writing both the original resource data and the new resource data into the new Bloom filter; wherein, the storage capacity of the new Bloom filter is a preset multiple of the storage capacity of the Bloom filter before expansion; or When the Bloom filter meets the preset expansion condition, adding a new Bloom filter and adding and writing the new resource data into any newly added Bloom filter.
8. The data resource management method based on the Bloom filter according to claim 5, characterized in that, Creating, initializing, and maintaining at least one Bloom filter through a Bloom filter module, including: Creating a Bloom filter according to preset parameters; wherein, the preset parameters include the length of the bit array, the number of hash functions, and the number of elements expected to be stored; the number of elements expected to be stored is inversely proportional to the false positive rate of the Bloom filter, and the number of elements expected to be stored is directly proportional to the memory occupied by the Bloom filter; Integrating the target data resource information according to a preset rule and adding it to the Bloom filter.
9. The data resource management method based on the Bloom filter according to claim 5, characterized in that The data resource content review module includes: a violation content set establishment unit and a content screening unit; wherein, the data resource management method based on the Bloom filter includes: Collecting and organizing known violation information through the violation content set establishment unit and adding the known violation information as a set to the Bloom filter; Wherein, during the data resource content review process, calling the Bloom filter through the data resource content review module to review whether there is violation information in the data resource to be reviewed, including: Calling the Bloom filter through the content screening unit to screen whether there is violation information in the data resource to be reviewed.
10. The data resource management method based on the Bloom filter according to claim 5, characterized in that, The data resource recommendation module includes: a user behavior record unit and a recommendation list generation unit; wherein, the data resource management method based on the Bloom filter includes: Recording the operation behaviors of the user; wherein, the operation behaviors of the user include the data resources applied for by the user; During the data resource recommendation process, calling the Bloom filter through the data resource recommendation module to filter out the data resources that have been recommended when generating the recommendation list, including: Generating a recommendation list according to the data resources applied for by the user; Calling the Bloom filter through the recommendation list generation unit to filter out the data resources that have been recommended in the recommendation list and generating a filtered recommendation list.
Citation Information
Patent Citations
Method and device of filtering candidate results of recommendation videos
CN108133031A
Method and device for recommending multimedia content and calculation equipment
CN111159436A
Content security identification method and device, storage medium and electronic equipment
CN112600834A
Query request filtering method applied to Redis client and Redis client
CN112818019A
Method and system for reconstructing bloom filter based on misjudgment rate
CN116991888A
Cited By
Data filtering, desensitization and real-time storage method and device based on bloom filter and ChatGPT
CN121233817A