An Approximate Membership Query Method and System for Dynamic k-mer Datasets

By adopting approximate member query method and multi-layer Bloom filter index architecture in the dynamic k-mer dataset, the problem of low storage and query efficiency in the existing technology is solved, efficient space utilization and query performance is achieved, and incremental update and deletion operations are supported.

CN119380828BActive Publication Date: 2025-06-13SUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411920384.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-06-13
Estimated Expiration
2044-12-25

AI Technical Summary

Technical Problem

When processing dynamic genomic data sets, the prior art cannot take into account both spatial efficiency and query speed, or does not support incremental updates and deletion, making it difficult to meet the use needs of dynamically stored gene sequences.

Method used

The approximate member query method for dynamic k-mer data sets is adopted to determine the element cardinality of the data set through cardinality estimation calculation method, select the appropriate number of Bloom filters and organizational form, realize the multi-layer architecture of the index, and support incremental updates and deletion.

Benefits of technology

On the basis of ensuring that the space occupied is as small as possible and the number of query times as small, incremental updates and deletion are supported to meet the needs of dynamically storing gene sequences, and improve space utilization and query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119380828B_ABST
    Figure CN119380828B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data storage, and specifically provides an approximate membership query method and system for a dynamic k-mer data set. The method includes: calculating the element cardinality of a target data set to be stored; selecting a target first unit to be stored according to the element cardinality; through the grouping function of the target first unit, generating a second unit index of a target element according to the target element of the target data set, and dividing the target element into a target second unit corresponding to the second unit index; storing the target element in a target third unit in the target second unit, and allocating it to a Bloom filter of the target data set. Furthermore, it solves the technical problem that the storage architecture and query method of gene sequences in the related art cannot balance space efficiency and query speed, or do not support incremental updates and deletions, resulting in difficulty in meeting the usage requirements of dynamically stored gene sequences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene data storage, and particularly to an approximate membership query method and system for a dynamic k-mer data set. Background Art

[0002] In genomic research, data sets change frequently due to the release of new sequences, the generation of new experimental data, and the revision of existing data. Such frequently changing data sets are called dynamic data sets. Facing the increasingly large scale of data sets, the paradigm of storing and indexing data sets as substrings of gene sequences of length k (k-mer) has become increasingly prominent.

[0003] This kind of k-mer data search problem can be abstracted as a multi-set multi-membership query problem: abstracting a database storing various DNA sequences as multiple data sets, and for a DNA sequence of interest, the goal is to detect in which sets this sequence exists.

[0004] Although there have been some methods in the related art that have made remarkable progress in dealing with the sequence search problem of large-scale genomic data. For example, setting up a Bloom filter for each data set separately to form a Bloom filter matrix, but this method wastes a lot of space. Other methods have limitations when dealing with dynamic data sets, unable to balance space efficiency and query speed, or not supporting incremental updates and deletions, resulting in difficulty in meeting the usage requirements of dynamically stored gene sequences.

[0005] Regarding the storage architecture and query method of gene sequences in the related art, in the case of a dynamically stored data set, it is impossible to balance space efficiency and query speed, or does not support incremental updates and deletions, resulting in difficulty in meeting the usage requirements of dynamically stored gene sequences. At present, no effective solution has been proposed. Summary of the Invention

[0006] An approximate membership query method and system for a dynamic k-mer data set provided by an embodiment of the present invention at least solve the problem that in the case of a dynamically stored data set, the storage architecture and query method of gene sequences in the related art cannot balance space efficiency and query speed, or do not support incremental updates and deletions, resulting in difficulty in meeting the usage requirements of dynamically stored gene sequences.

[0007] According to one aspect of the present invention, there is provided a method for generating an approximate membership query index for a dynamic k-mer data set, including: determining the element cardinality of a target data set to be stored according to a cardinality estimation algorithm; selecting a target first unit of the first level to be stored according to the element cardinality, wherein the number of Bloom filters of multiple first units of the first level is different, and the capacity of the Bloom filter of each first unit is the same. The Bloom filter is used to store the corresponding elements in the target data set, and the element cardinality that can be stored in the first unit is not less than the element cardinality of the corresponding data set; through the grouping function of the target first unit, according to the target element of the target data set, a second-level second unit index of the target element is generated, and the target element is divided into the target second unit corresponding to the second unit index, wherein the first unit includes multiple second units, and the number of second units of each data set stored in the first unit is the same as the number of Bloom filters required for the corresponding data set; storing the target element into a target third unit of the third level in the target second unit and assigning it to the Bloom filter of the target data set, wherein the second unit includes multiple third units with a fixed size, and the organizational form of the third unit is an interleaved Bloom filter, and the third unit reads the same bit of different Bloom filters at one time when querying elements.

[0008] As an alternative embodiment, selecting the target first unit of the first level to be stored according to the cardinality includes: comparing the element cardinality of the target data set with the element cardinality that can be stored in the first unit; in the case where the element cardinality of the target data set is less than or equal to the element cardinality that can be stored in the compared first unit and greater than the element cardinality that can be stored in the previous first unit, taking the compared first unit as the target first unit; wherein the compared first unit is one of multiple first units, and the previous first unit is a first unit with a preset difference less than the number of Bloom filters of the compared first unit.

[0009] As an alternative embodiment, storing the target element into a target third unit of the third level in the target second unit and assigning it to the Bloom filter of the target data set includes: storing the target element into a target third unit of the third level in the target second unit; using multiple independent hash functions of the target third unit to perform hash processing on the target element and mapping it to the corresponding position in the corresponding Bloom filter, and setting the corresponding position to 1.

[0010] As an alternative embodiment, the method further includes: receiving a first data set to be inserted; determining a corresponding first insertion unit according to the element cardinality of the first data set; converting the organizational form of the third unit of the first insertion unit from an interleaved Bloom filter to a Bloom filter matrix; finding a corresponding insertion position in the first insertion unit, where the insertion position is found through the Bloom filter; inserting the first data set into the insertion position and allocating a corresponding Bloom filter to the first data set; converting the organizational form of the third unit of the first insertion unit from the Bloom filter matrix to the interleaved Bloom filter.

[0011] As an alternative embodiment, the method further includes: receiving a second data set to be deleted; querying the first unit to which the second data set belongs; converting the organizational form of the third unit of the belonging first unit from an interleaved Bloom filter to a Bloom filter matrix; in the case where the element set of the second data set is a partial element set of the third unit, deleting the Bloom filter corresponding to the second data set and retaining the empty space; in the case where the element set of the second data set is the entire element set of the third unit, deleting the Bloom filter corresponding to the second data set and releasing the space resources of the third unit; converting the organizational form of the third unit of the belonging first unit from the Bloom filter matrix to the interleaved Bloom filter.

[0012] According to another aspect of the present invention, there is provided an approximate membership query method for a dynamic k-mer data set, characterized by including: generating, through a grouping function of a first unit at a first level and a target element to be queried, a second unit index at a second level in the first unit for the target element to identify the target second unit where the target element is located, where the target element is indexed level by level in a multi-level manner, and when stored in the second level, a corresponding second unit index is generated using the grouping function of the belonging first unit and stored in the target second unit corresponding to the second unit index; in the case where the target second unit exists in the first unit, determining that the target element is stored in the target second unit of the first unit; obtaining an interleaved Bloom filter of a third unit at a third level in the target second unit, and obtaining a plurality of hash indexes according to the interleaved Bloom filter, where the third unit is a component of the belonging second unit, and when the target element is stored in the third unit, a Bloom filter of the belonging target data set is allocated to it, the organizational form of the third unit is an interleaved Bloom filter, the number of bits of the hash index is the same as the number of Bloom filters of the target second unit, and the Bloom filters correspond to different data sets; querying the target data set to which the target element belongs according to the hash index.

[0013] As an alternative embodiment, querying the target data set to which the target element belongs according to the hash index includes: performing a bitwise AND operation on a plurality of the hash indexes to obtain a target index, where the hash index is data of the same bit of different Bloom filters, the hash index is a mapping of the same hash function to elements of different data sets, different bits of the hash index represent values of different Bloom filters, and the value characterizes whether the target element exists in the Bloom filter; determining the target data set to which the target element belongs according to the target index.

[0014] As an alternative embodiment, after querying the target data set to which the target element belongs according to the hash index, the method further includes: storing the target data set in an output list; continuing to query whether there is an index of the target element in other first units; in the case where there is an index of the target element in the first unit, obtaining the target data set to which it belongs and storing it in the output list; in the case where all the first units are traversed, outputting the output list.

[0015] As an alternative embodiment, the first unit of the first level is a segment unit, the segment unit includes a plurality of Bloom filters, the capacities of the Bloom filters of each segment unit are the same, one Bloom filter corresponds to recording elements in a corresponding data set, and the cardinality of the elements of the data set that the segment unit can store is greater than or equal to the cardinality of the elements of the data set to be stored; the second unit of the second level is a group unit, the number of group units of each data set stored in the segment unit is the same as the number of Bloom filters required for the corresponding data set; the third unit of the third level is a block unit, the block unit is a data block with a fixed data size, which is a constituent unit of the group unit, and stores different elements in an interleaved Bloom filter organization form. When querying an element, the block unit reads all the data at the positions corresponding to the query element in different Bloom filters at one time according to the interleaved Bloom filter.

[0016] According to another aspect of the present invention, there is provided an approximate membership query system for a dynamic k-mer data set, including: a processor, and a memory storing a program, characterized in that the program includes instructions that, when executed by the processor, cause the processor to execute the above method.

[0017] The index generation method provided by the embodiment of the present invention selects the target first unit of the first level to be stored according to the element cardinality, and uses the first unit including a fixed number of Bloom filters to store data sets corresponding to a fixed number of corresponding element cardinalities, so that sufficient and as small as possible space can be allocated for each data set to avoid space waste.

[0018] Through the grouping function of the target first unit, according to the target elements of the target data set, generate the second unit index of the second level of the target elements, divide the target elements into the corresponding target second units, and the number of second units of each stored data set is the same as the number of Bloom filters required for the corresponding data set. Each second unit stores different element sets of the data set. In this way, when querying, for an element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0019] Store the target elements in the target third unit of the third level in the target second unit, and allocate them to the Bloom filters of the target data set. Using the staggered Bloom filters of the third unit, when querying elements in the third unit, the same bit of different Bloom filters can be read at one time according to the staggered Bloom filters, so as to determine the target data aggregation where the query element is located at one time, avoiding the situation that when using a Bloom filter matrix, only one Bloom filter's data can be read at a time and it is necessary to traverse the Bloom filters to determine the data set to which the query element belongs.

[0020] Furthermore, it solves the problems in the storage architecture and query method of gene sequences in the related art that in the case of dynamically storing data sets, it is impossible to balance space efficiency and query speed, or does not support incremental updates and deletions, resulting in difficulty in meeting the usage requirements of dynamically stored gene sequences. It achieves the technical effect of supporting incremental updates and deletions on the basis of ensuring that the occupied space is as small as possible and the number of queries is as small as possible to meet the usage requirements of the data. Brief Description of the Drawings

[0021] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other embodiments according to these drawings without creative efforts.

[0022] Figure 1 It is a flowchart of an approximate membership query index generation method for a dynamic k-mer data set according to an embodiment of the present invention.

[0023] Figure 2 It is a flowchart of an approximate membership query method for a dynamic k-mer data set according to an embodiment of the present invention.

[0024] Figure 3 It is a schematic diagram of the index hierarchical architecture according to an embodiment of the present invention.

[0025] Figure 4It is a schematic diagram of the index construction process for the data of the embodiment of the present invention.

[0026] Figure 5 It is a schematic diagram of the element query process for the data index of the embodiment of the present invention.

[0027] Figure 6 It is a schematic diagram of the data set addition process for the data index of the embodiment of the present invention.

[0028] Figure 7 It is a schematic diagram of the data set deletion process for the data index of the embodiment of the present invention.

[0029] Figure 8 It is a schematic diagram of an approximate membership query index generation device for a dynamic k-mer data set according to an embodiment of the present invention.

[0030] Figure 9 It is a schematic diagram of an approximate membership query device for a dynamic k-mer data set according to an embodiment of the present invention.

[0031] Figure 10 It is a schematic diagram of the structure of an approximate membership query system for a dynamic k-mer data set according to the present invention. Specific Embodiments

[0032] The embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present invention. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0033] Approximate Membership Queries (AMQ) represent a set in a compact space, answer whether a query element belongs to the set, have high space efficiency and query efficiency, and allow a certain degree of false positives. The Bloom Filter is the most classic approximate membership query data structure. It maintains a bit array of length m, maps elements to several positions in the bit array through multiple hash functions, and sets the bits at these positions to "1" to mark the existence of the element. The Bloom Filter supports fast insertion and lookup but cannot delete elements.

[0034] Approximate member queries have a wide range of applications in multiple fields, including genome sequence searches. With the development of genome sequencing technology, especially the application of next-generation sequencing (NGS) technology, sequencing throughput has been greatly improved, driving the rapid growth of data generated by biotechnology protocols such as WGS and RNA-Seq. These data are stored in public databases such as the European Nucleotide Archive (ENA) and the Sequence Read Archive (SRA). However, the exponential growth of data size makes it extremely challenging to efficiently search for DNA sequences in them. In order to make full use of existing resources, it is crucial to achieve fast and efficient genome sequence searches.

[0035] The following three strategies are mainly used in related technologies for sequence search:

[0036] 1) The classic approach is to assign a Bloom filter to each set and organize it into a Bloom filter matrix. This structure is more flexible, and the addition and deletion of data sets can be achieved by creating or destroying Bloom filters. However, in order to maintain the expected false alarm rate, the Bloom filter size needs to be set according to the maximum set, which causes serious space waste for small sets.

[0037] 2) By using partitioned hash functions to divide the set into multiple regions, each region is assigned an approximate member query structure. Compared with the matrix structure, this method reduces the number of Bloom filters accessed during the query, which can speed up the query when the data scale is small. However, since Bloom filters do not support deletion, the method of storing multiple set elements in a mixed manner does not support the deletion of data sets, otherwise it will cause false negatives.

[0038] 3) Organize the Bloom filter into a hierarchical tree structure. At each level, large sets are split and small sets are merged to deal with the unbalanced set size distribution, thereby alleviating space waste to a certain extent. However, the tree structure complicates the query process and there is a problem of index reconstruction when inserting and deleting data sets.

[0039] In order to solve the problem that the above-mentioned existing approximate member query structures for genome sequence search cannot handle dynamic data sets well, this embodiment aims to design an indexing and query method for dynamically stored data sets, which supports incremental update and deletion of data sets while maintaining good space efficiency and query speed.

[0040] Figure 1 : is a flowchart of a method for generating an approximate member query index for a dynamic k-mer data set according to an embodiment of the present invention, such as Figure 1 As shown, the embodiment of the invention provides a method for generating an approximate member query index for a dynamic k-mer data set. The steps of the method are as follows:

[0041] Step S101, determine the element cardinality of the target data set to be stored according to the cardinality estimation algorithm;

[0042] Step S102, select the target first unit of the first level to be stored according to the element cardinality. Among them, the number of Bloom filters of multiple first units at the first level is different, and the capacity of the Bloom filter of each first unit is the same. The Bloom filter is used to store the corresponding elements in the target data set, and the element cardinality that the first unit can store is not less than the element cardinality of the corresponding data set;

[0043] Step S103, through the grouping function of the target first unit, generate the second-level second unit index of the target element according to the target element of the target data set, and divide the target element into the target second unit corresponding to the second unit index. Among them, the first unit includes multiple second units, and the number of second units of each data set stored in the first unit is the same as the number of Bloom filters required by the corresponding data set;

[0044] Step S104, store the target element in the target third unit of the third level in the target second unit and allocate it to the Bloom filter of the target data set. Among them, the second unit includes multiple third units with fixed sizes. The organization form of the third unit is an interleaved Bloom filter, and the third unit reads the same bit of different Bloom filters at one time according to the interleaved Bloom filter when querying elements.

[0045] The above-mentioned method for generating an index of a dynamically stored data set provided by the embodiment of the present invention creates selects the target first unit of the first level to be stored according to the element cardinality, and uses the first unit containing a fixed number of Bloom filters to store the data set corresponding to the fixed number of element cardinalities, so that sufficient and as small as possible space can be allocated for each data set to avoid space waste.

[0046] Through the grouping function of the target first unit, generate the second-level second unit index of the target element according to the target element of the target data set, and divide the target element into the target second unit corresponding to the second unit index. The number of second units of each data set stored is the same as the number of Bloom filters required by the corresponding data set, and each second unit stores different element sets of the data set. In this way, when querying, for an element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0047] Store the target element in the target third unit at the third level in the target second unit, and allocate it to the Bloom filter of the target data set. By using the staggered Bloom filter of the third unit, when querying elements in the third unit, the same bit of different Bloom filters can be read at one time according to the staggered Bloom filter, so as to determine the target data cluster where the query element is located at one time, avoiding the situation that when using a Bloom filter matrix, only one Bloom filter's data can be read at a time and it is necessary to traverse the Bloom filters to determine the data set to which the query element belongs.

[0048] Furthermore, it achieves the technical effect of supporting incremental updates and deletions on the basis of ensuring that the occupied space is as small as possible and the number of queries is as small as possible, so as to meet the usage requirements of the data.

[0049] The above data set can be a k-mer data set, and the above target data set can be a target k-mer data set.

[0050] The execution entity of the above steps can be a storage device for k-mer data. This storage device can include a processor or a processing chip for data processing to implement different storage, reading, querying and other functions.

[0051] In the above step S101, the k-mer data set can be a set containing multiple k-mer data. The k-mer data can be understood as a gene sequence composed of k bases. For example, CGAT can be simply called 4-mer data. The above radix can be understood as the possible number of classifications after classifying the elements in the k-mer data set.

[0052] The above radix estimation algorithm can be the HyperLogLog algorithm, and the element radix of the target k-mer data set to be stored can be determined according to the radix estimation algorithm.

[0053] Since in the index architecture, the first level includes multiple first units. The difference between different first units is that the number of Bloom filters in each first unit is different, and the capacities of the Bloom filters in the same first unit are the same.

[0054] The first unit can store multiple k-mer data sets. For each stored k-mer data set, the corresponding number of Bloom filters corresponding to the element radix is allocated, and the corresponding number of second units corresponding to the element radix is allocated.

[0055] Adopting such an architecture is mainly to allocate different k-mer data sets to the first units corresponding to their element radixes according to the element radixes of different k-mer data sets, and occupy as little space as possible while ensuring the index records of the k-mer data sets.

[0056] In the above step S102, according to the element cardinality, the target first-level first unit to be stored is selected, that is, the first unit whose element cardinality of the dataset that can be stored is not less than the element cardinality of the corresponding k-mer dataset is selected as the first unit to be stored.

[0057] That is, a Bloom filter of a first unit at the first level corresponds to storing the element index of a k-mer dataset. The element cardinality of the dataset that the first unit can store is not less than the element cardinality of the corresponding k-mer dataset to at least meet the element storage requirements of the k-mer dataset.

[0058] However, the element cardinality of the dataset that the first unit can store should not exceed the element cardinality of the k-mer dataset by too much, otherwise it is easy to cause waste of space.

[0059] Preferably, in this embodiment, the element cardinality of the dataset that the first unit can store is equal to the element cardinality of the corresponding k-mer dataset. This can ensure the effective storage of the element index of the k-mer dataset and minimize the occupation of memory space as much as possible, thereby improving the space utilization rate.

[0060] In the above step S103, the target first unit is also the first unit that meets the element cardinality requirements of the k-mer dataset. Each first unit includes second-level second units; the number of second units of each k-mer dataset stored in the first unit is the same as the number of Bloom filters required by the corresponding k-mer dataset, so that different element sets of different k-mer datasets can be stored separately by the same number of second units.

[0061] Such a storage structure can, when querying, only query one Bloom filter for a query element to determine whether the element is in the second unit and the first unit to which it belongs.

[0062] Specifically, through the grouping function of the target first unit, according to the target element of the target k-mer dataset, the second-level second unit index of the target element is generated; this second unit index can establish a one-to-one mapping between its target element and the second unit.

[0063] When querying an element, only need to directly generate the corresponding second unit index for the query tuple through the grouping function corresponding to the first unit being queried, and determine whether the second unit to which the query element belongs exists in the first unit being queried by querying whether the second unit index exists in the first unit being queried. And when it is determined that the second unit to which the query element belongs exists in the first unit being queried, determine which second unit this second unit is.

[0064] When generating the index, the target element is partitioned into the target second unit corresponding to the second unit index. Specifically, as in step S104 above, the target element is stored in the target third unit at the third level in the target second unit and assigned to the Bloom filter of the target k-mer dataset.

[0065] The second unit includes multiple third units of fixed size, such as 64 bits. Additionally, the third units can be created and deleted according to requirements.

[0066] In the related art, the organization form of data blocks is generally a Bloom filter matrix, that is, different Bloom filters are arranged in a matrix form, and all data of the same Bloom filter can be read in one read operation. However, different positions in the same Bloom filter correspond to the existence of different elements. Thus, if it is queried whether a certain element exists in this data block, it is necessary to traverse multiple Bloom filters corresponding to this data block.

[0067] Moreover, only after the final traversal is completed can it be determined that this data block does not store this element. And the staggered Bloom filter, as the name implies, means that different positions of different Bloom filters can be staggered. The different Bloom filters are arranged according to positions, so that in one read operation, the same position corresponding to the query element in different Bloom filters can be read, and then it can be determined whether the query element exists in this third unit through one read operation, and specifically which Bloom filter this query element belongs to, and its belonging k-mer dataset can be determined according to the belonging Bloom filter.

[0068] Therefore, in this embodiment, the organization form of the third unit is a staggered Bloom filter, and the third unit reads the same bit of different Bloom filters at one time according to the staggered Bloom filter when querying elements.

[0069] In summary, through the multi-layer architecture of the first unit, the second unit, and the third unit and their characteristics, it is possible to ensure a smaller occupied space during storage, higher efficiency and better accuracy during querying. Moreover, the entire architecture can be flexibly adjusted, does not depend on the initial dataset, creates an index, and corresponding datasets can be added and deleted according to requirements. The specific methods of addition and deletion will be described later.

[0070] As an alternative embodiment, selecting a target first unit of the first level to be stored according to the cardinality includes: comparing the cardinality of elements in the target k-mer data set with the cardinality of elements that can be stored in the first unit; when the cardinality of elements in the target k-mer data set is less than or equal to the cardinality of elements that can be stored in the compared first unit and greater than the cardinality of elements that can be stored in the previous first unit, taking the compared first unit as the target first unit; where the compared first unit is one of multiple first units, and the previous first unit is the first unit with a preset difference less in the number of Bloom filters than the compared first unit.

[0071] In the first units of the first level, ideally, for each k-mer data set to be indexed, a first unit with the same cardinality of elements can be found. However, this requires a large number of first units. Between the first units with the closest adjacent numbers of Bloom filters, the difference in the number of Bloom filters is only 1. This may require a very large span in the number of first units, which is not conducive to space allocation.

[0072] Considering the large difference in the number of elements in different k-mer data sets, the difference in the number of Bloom filters between adjacent first units with Bloom filters can be increased. In this way, when selecting the target first unit of the first level to be stored according to the cardinality, usually the cardinality of elements in the k-mer data set satisfied by each first unit is a numerical range.

[0073] In this way, when selecting the target first unit of the first level to be stored according to the cardinality, usually a first unit with exactly the same number of Bloom filters cannot be found. When the cardinality of elements in the target k-mer data set is less than or equal to the cardinality of elements that can be stored in the compared first unit and greater than the cardinality of elements that can be stored in the previous first unit, the compared first unit can be taken as the target first unit.

[0074] Thus, it is possible to take into account minimizing the index storage space while minimizing the number of first units.

[0075] As an alternative embodiment, storing a target element into a target third unit of the third level in a target second unit and allocating a Bloom filter to the target k-mer data set includes: storing the target element into the target third unit of the third level in the target second unit; using multiple independent hash functions of the target third unit to perform hash processing on the target element and map it to the corresponding positions in the corresponding Bloom filter, and setting the corresponding positions to 1.

[0076] Storing the target element into the Bloom filter allocated to the k-mer data set in the target second unit. The specific storage method is to use k independent hash functions h 0 ,…h k-1, each function randomly and uniformly maps the target elements in the k-mer dataset to an integer in {0, ..., m-1}, that is, any integer in the size m of the Bloom filter.

[0077] And set the corresponding k positions mapped in this Bloom filter to 1, that is, B[h 0 (x)], …, B[h k-1 (x)] are set to 1.

[0078] As an optional embodiment, the method further includes: receiving a first k-mer dataset to be inserted; determining a corresponding first insertion unit according to the element cardinality of the first k-mer dataset; converting the organizational form of the third unit of the first insertion unit from an interleaved Bloom filter to a Bloom filter matrix; finding a corresponding insertion position in the first insertion unit, where the insertion position is found through the Bloom filter; inserting the first k-mer dataset into the insertion position and allocating a corresponding Bloom filter for the first k-mer dataset; converting the organizational form of the third unit of the first insertion unit from the Bloom filter matrix to the interleaved Bloom filter.

[0079] When inserting the first k-mer dataset, the corresponding first insertion unit can be determined first according to the element cardinality of the first k-mer dataset. After determining the first unit, convert the organizational form of the third unit in the first unit from the interleaved Bloom filter to the Bloom filter matrix.

[0080] In this way, during the insertion operation, a single insertion operation can be performed on the entire Bloom filter, instead of inserting the Bloom filter to be inserted multiple times for each position.

[0081] After insertion, converting the organizational form of the third unit of the first insertion unit from the Bloom filter matrix to the interleaved Bloom filter can facilitate subsequent element queries.

[0082] It should be noted that when inserting the first k-mer dataset, it can be first checked whether there is a vacant position generated by a deleted Bloom filter in the third unit of the first unit. If so, insert the first k-mer dataset into this vacant position. Since the size m and the element cardinality of the Bloom filters in the same first unit are the same, the vacant position of the deleted Bloom filter can directly insert the newly added Bloom filter.

[0083] If there is no vacant position generated by a deleted Bloom filter in the third unit of the first unit, then insert the first k-mer dataset into the subsequent insertable position in the third unit where there is no space. It should be noted that when inserting, it is adjacent to the inserted Bloom filter to improve space utilization.

[0084] As an alternative embodiment, the method further includes: receiving a second k-mer data set to be deleted; querying the first unit to which the second k-mer data set belongs; converting the organization form of the third unit of the first unit from an interleaved Bloom filter to a Bloom filter matrix; when the element set of the second k-mer data set is a partial element set of the third unit, deleting the Bloom filter corresponding to the second k-mer data set and retaining the empty positions; when the element set of the second k-mer data set is the entire element set of the third unit, deleting the Bloom filter corresponding to the second k-mer data set and releasing the space resources of the third unit; converting the organization form of the third unit of the first unit from a Bloom filter matrix to an interleaved Bloom filter.

[0085] Similarly, when deleting the second k-mer data set, the first unit to which the second k-mer data set belongs can be queried first. After determining the first unit, convert the organization form of the third unit in the first unit from an interleaved Bloom filter to a Bloom filter matrix.

[0086] In this way, during the deletion operation, a single deletion operation can be performed on the entire Bloom filter, instead of deleting the Bloom filter to be deleted multiple times according to the positions.

[0087] After deletion, converting the organization form of the third unit inserted into the first unit from a Bloom filter matrix to an interleaved Bloom filter can facilitate subsequent element queries.

[0088] It should be noted that when the element set of the second k-mer data set is a partial element set of the third unit, after deleting the Bloom filter corresponding to the second k-mer data set, if the third unit still stores the element information of other Bloom filters, only the empty positions are retained. When inserting a Bloom filter subsequently, the inserted Bloom filter can be inserted into the empty positions to improve the space utilization efficiency.

[0089] When the element set of the second k-mer data set is the entire element set of the third unit, it means that after deleting the Bloom filter corresponding to the second k-mer data set, the third unit no longer stores any Bloom filter data, so the space resources of the third unit can be released. For example, deleting the third unit can reduce the memory occupancy of the index.

[0090] As an alternative embodiment, the first unit of the first level is a segment unit. The segment unit includes a plurality of Bloom filters, and the capacities of the Bloom filters of each segment unit are the same. One Bloom filter corresponds to recording an element in a k-mer dataset. The cardinality of the elements of the dataset that the segment unit can store is greater than or equal to the cardinality of the elements of the k-mer dataset to be stored. The second unit of the second level is a group unit. The number of group units of each dataset stored in the segment unit is the same as the number of Bloom filters required for the corresponding dataset. The third unit of the third level is a block unit. The block unit is a data block with a fixed data size and is a constituent unit of the group unit. Different elements are stored in the form of an interleaved Bloom filter. When querying an element, the block unit reads all the data corresponding to the query element at the corresponding positions in different Bloom filters at one time according to the interleaved Bloom filter.

[0091] By adopting the above segmentation method, sufficient and as small as possible space can be allocated for each k-mer dataset, thereby optimizing the space utilization rate. The Bloom filter group of each segment unit is divided into multiple group units, and a grouped hash function is used to determine the group index, which can ensure that each k-mer dataset is only queried one Bloom filter during a k-mer query. This grouping strategy reduces the number of Bloom filters that need to be accessed during the query process, thereby improving the query efficiency.

[0092] The organization form of the block unit does not use the classic Bloom filter matrix, but uses the form of an interleaved Bloom filter. By interleaving the bits of multiple Bloom filters for storage, the target bits obtained by one hash operation are as continuous as possible in the storage space, thereby further improving the query performance.

[0093] Figure 2 It is a flowchart of an approximate membership query method for a dynamic k-mer dataset according to an embodiment of the present invention. As Figure 2 shown, an embodiment of the present invention also provides an approximate membership query method for a dynamic k-mer dataset, and the method includes the following steps:

[0094] Step S201, through the grouping function of the first unit of the first level and the target element to be queried, generate the index of the second unit of the second level in the first unit for the target element to identify the target second unit where the target element is located. Among them, the target element is indexed level by level in multiple levels. When storing into the second level, the corresponding second unit index is generated by using the grouping function of the first unit to which it belongs and stored in the target second unit corresponding to the second unit index.

[0095] Step S202, when there is a target second unit in the first unit, determine that the target element is stored in the target second unit of the first unit.

[0096] Step S203: Obtain the staggered Bloom filter of the third-level third unit in the target second unit. According to the staggered Bloom filter, obtain multiple hash indexes. Here, the third unit is a component of the second unit to which it belongs. When the target element is stored in the third unit, it will be assigned to the Bloom filter of the target data set to which it belongs. The organizational form of the third unit is a staggered Bloom filter. The number of bits of the hash index is the same as the number of Bloom filters of the target second unit. The Bloom filters correspond to different data sets.

[0097] Step S204: Query the target data set to which the target element belongs according to the hash index.

[0098] In the above-mentioned index query method for dynamically stored data sets provided by the embodiments of the present invention, through the grouping function of the first-level first unit and the target element to be queried, the index of the second-level second unit of the target element in the first unit is generated. Using the grouping function of the first unit containing Bloom filters with the same element cardinality, the second unit index is directly generated. By querying whether the second unit indexes are the same, for a target element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0099] Obtain the staggered Bloom filter of the third-level third unit in the target second unit. According to the staggered Bloom filter, obtain multiple hash indexes. Query the target k-mer data set to which the target element belongs. Using the staggered Bloom filter of the third unit, when querying elements in the third unit, the same bit of different Bloom filters can be read at one time according to the staggered Bloom filter, so as to determine the target k-mer number aggregation where the query element is located at one time, avoiding the situation that when using a Bloom filter matrix, only one Bloom filter's data can be read at a time and it is necessary to traverse the Bloom filters to determine the k-mer data set to which the query element belongs.

[0100] Furthermore, the technical effect of supporting incremental update and deletion to meet the usage requirements of k-mer data is achieved on the basis of ensuring that the occupied space is as small as possible and the number of queries is as small as possible.

[0101] The execution subject of the above steps can be a storage device for k-mer data. This storage device can include a processor or a processing chip for data processing to implement different storage, reading, querying and other functions.

[0102] In the above step S201, through the grouping function of the first-level first unit and the target element to be queried, the index of the second-level second unit of the target element in the first unit is generated to identify the target second unit where the target element is located.

[0103] When generating an index, the target elements are indexed level by level in multiple levels. When storing into the second level, the grouping function of the first unit to which they belong is used to generate the corresponding second unit index, which is then stored into the target second unit corresponding to the second unit index.

[0104] In this way, when querying, for a query element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0105] Specifically, through the grouping function of the target first unit, according to the target elements of the target k-mer dataset, the second unit index of the second level of the target elements is generated; this second unit index can establish a one-to-one mapping between its target elements and the second unit.

[0106] When querying an element, only need to directly generate the corresponding second unit index for the query tuple through the grouping function corresponding to the first unit being queried. By querying whether the second unit index exists in the first unit being queried, it can be determined whether the second unit to which the query element belongs exists in the first unit, and when it is determined that the second unit to which the query element belongs exists in the first unit, it can be determined which second unit this second unit is. As in step S202 above, when the target second unit exists in the first unit, it is determined that the target element is stored in the target second unit of the first unit.

[0107] In the above step S203, obtain the staggered Bloom filter of the third unit of the third level in the target second unit, and obtain multiple hash indexes according to the staggered Bloom filter.

[0108] This is because when generating an index, the third unit is a component of the second unit to which it belongs. When the target element is stored into the third unit, it will be assigned to the Bloom filter of the target k-mer dataset to which it belongs. The organizational form of the third unit is a staggered Bloom filter, and the number of bits of the hash index is the same as the number of Bloom filters of the target second unit, and the Bloom filters correspond to different k-mer datasets.

[0109] The second unit includes multiple third units of fixed size, such as 64bit. In addition, the third unit can be created and deleted according to requirements.

[0110] Arrange different Bloom filters in position. In this way, in one read operation, the same position of the corresponding query element in different Bloom filters can be read, and then it can be determined whether the query element exists in the third unit through one read operation, and specifically which Bloom filter the query element belongs to, and its belonging k-mer dataset can be determined according to the Bloom filter to which it belongs.

[0111] Therefore, in this embodiment, the organizational form of the third unit is an interleaved Bloom filter, and the object of the original single read / delete operation is changed from the data of a single Bloom filter to all the data of multiple Bloom filters at multiple positions. When querying an element, the third unit reads all the data of different Bloom filters at multiple positions at one time according to the interleaved Bloom filter to form a hash index.

[0112] On this basis, according to the above step S204, the target k-mer dataset to which the target element belongs can be queried according to the hash index, and the query of the target element is completed.

[0113] Subsequently, the storage location of the target k-mer data can be found according to the index of the target k-mer dataset determined by the query result, and the target k-mer dataset can be extracted. After the target k-mer dataset is obtained, the query target element in the target k-mer dataset can also be obtained.

[0114] As an optional embodiment, querying the target k-mer dataset to which the target element belongs according to the hash index includes: performing a bitwise AND operation on multiple hash indexes to obtain a target index, where the hash indexes are the data of the same bit of different Bloom filters, the hash indexes are the mappings of the elements of different k-mer datasets by the same hash function, different bits of the hash indexes represent the values of different Bloom filters, and the values characterize whether the target element exists in the Bloom filter; determining the target k-mer dataset to which the target element belongs according to the target index.

[0115] For each third unit in the target second unit, when the Bloom filter obtains multiple hash indexes, k hash indexes h 0 (x),…h k-1 (x) corresponding k row bit vectors are obtained, which are the above-mentioned hash indexes.

[0116] Then, a bitwise AND operation is performed on these hash indexes to generate an equal-length bit vector, which is equivalent to the above-mentioned target index. The target indexes obtained by each third unit are concatenated. Assuming that the number of sets stored in the first unit is n, the final query result of the first unit is a bit vector with a length of n, where each bit indicates whether the corresponding set contains the target element q. Specifically, if a certain bit is 1, it means that the corresponding set contains q; if it is 0, it means that q is not contained.

[0117] As an alternative embodiment, after querying the target k-mer dataset to which the target element belongs according to the hash index, the method further includes: storing the target k-mer dataset in an output list; continuing to query whether there is an index of the target element in other first units; in the case that there is an index of the target element in the first unit, obtaining the target k-mer dataset to which it belongs and storing it in the output list; in the case that all the first units have been traversed, outputting the output list.

[0118] The above output list can be synchronously created when the query task is established and initialized as an empty list after creation to store the queried target k-mer dataset.

[0119] As an alternative embodiment, the first unit of the first level is a segment unit. The segment unit includes multiple Bloom filters, and the capacities of the Bloom filters of each segment unit are the same. One Bloom filter corresponds to recording the elements in a k-mer dataset. The element cardinality of the dataset that the segment unit can store is greater than or equal to the element cardinality of the k-mer dataset to be stored; the second unit of the second level is a group unit. The number of group units of each dataset stored in the segment unit is the same as the number of Bloom filters required by the corresponding dataset; the third unit of the third level is a block unit. The block unit is a data block with a fixed data size and is a constituent unit of the group unit. Different elements are stored in the form of an interleaved Bloom filter. When querying an element, the block unit reads all the data corresponding to the query element in different Bloom filters at one time according to the interleaved Bloom filter.

[0120] As Figure 3 shown, by adopting the above segmentation method, sufficient and as small as possible space can be allocated for each k-mer dataset, thereby optimizing the space utilization rate. The Bloom filter group of each segment unit is divided into multiple group units, and a grouped hash function is used to determine the group index, which can ensure that each k-mer dataset is only queried one Bloom filter during a k-mer query. This grouping strategy reduces the number of Bloom filters that need to be accessed during the query process, thereby improving the query efficiency.

[0121] The organization form of the block unit does not use the classical Bloom filter matrix, but uses the form of an interleaved Bloom filter. By storing the bits of multiple Bloom filters interleaved, the target bits obtained by one hash operation are as continuous as possible in the storage space, thereby further improving the query performance.

[0122] It should be noted that this embodiment also provides an alternative implementation manner, which will be described in detail below.

[0123] The approximate membership query method for dynamic k-mer datasets proposed in this embodiment mainly proposes a new storage architecture and a core process dependent on the architecture, including index construction and element query, and supports incremental addition and deletion operations of datasets. For the initial dataset, an index is first constructed to support subsequent query tasks. When a new dataset needs to be added, incremental updates can be achieved without relying on the initial dataset, avoiding redundant backups of the original dataset and thus saving storage space.

[0124] This embodiment proposes an adaptive Bloom filter allocation strategy, that is, setting the relevant parameters of the Bloom filter and then allocating an appropriate number of Bloom filters for each data set. At the same time, combined with the grouped hash function, it is ensured that as few Bloom filters as possible are accessed during the query phase without generating false negative false positives.

[0125] The approximate membership query structure proposed in this embodiment has three levels of basic units, namely segments (units), groups (units), and blocks (units):

[0126] Segment: A segment represents a set group of a logically k-mer dataset. Each segment contains sets that require the same number of Bloom filters. Through this segmentation method, sufficient and as small as possible space can be allocated for each k-mer dataset, thus optimizing the space utilization rate.

[0127] Group: The Bloom filter group of each segment is divided into multiple groups. Using the grouped hash function g(.), it can be ensured that each k-mer dataset is only queried one Bloom filter during a k-mer query. This grouping strategy reduces the number of Bloom filters that need to be accessed during the query process, thus improving the query efficiency.

[0128] Block: In order to maximize the reading efficiency of machine words, each group consists of multiple blocks. The organization form of the block does not use the classic Bloom filter matrix, but uses the form of an interleaved Bloom filter. By interleaving the bits of multiple Bloom filters, the target bits obtained by one hash operation are as continuous as possible in the storage space, thus further improving the query performance.

[0129] Figure 3 It is a schematic diagram of the index hierarchical architecture of the embodiment mode of the present invention, as Figure 3 shown, segment 1 stores the k-mer datasets that require 1 Bloom filter, segment 2 stores the k-mer datasets that require 2 Bloom filters, and segment 3 stores the k-mer datasets that require 3 Bloom filters.

[0130] Segment 1 includes at most 1 group, segment 2 includes at most 2 groups, and segment 3 includes at most 3 groups. That is to say, segment 2 has the ability to store the k-mer data set stored in segment 1, but it will cause waste of space.

[0131] The number of groups in each segment is equal to the number of Bloom filters required for the stored k-mer data set.

[0132] Suppose a k-mer data set requires two Bloom filters. Then it is stored in segment 2. Each group in segment 2 is assigned one Bloom filter for this k-mer data set to store approximately half of the partial elements in this k-mer data set. Suppose there is an element x in this k-mer data set. It is determined through a grouping function whether to store the element x in the Bloom filter assigned to this k-mer data set in group 1 or in the Bloom filter assigned to this k-mer data set in group 2.

[0133] Each group can be divided into multiple blocks. The relative positions of the Bloom filters assigned to the same k-mer data set by different groups in the same segment are consistent.

[0134] The size of each block is fixed and can be 64 bits. Each bit corresponds to at most one Bloom filter. In a 64-bit block, the Bloom filter matrix is as Figure 3 shown. The five bits 00010 and 00000 circled by the ellipse in the figure can be five Bloom filters at different positions for two hash functions and require five read operations to obtain. The same row can be regarded as the same position and corresponds to the same hash function.

[0135] That is to say, the target bit of one hash is also a row in the block, as shown by the rectangular box in Figure 3 . In this way, each read operation of the block can only obtain the data of all positions of the same column of the same Bloom filter. In this way, multiple read operations are required to determine whether a certain query element belongs to this block. It is not until all Bloom filters are read that it can be determined that this block does not have this query element and the specific Bloom filter to which this query element belongs.

[0136] The interleaved Bloom filter strings together the data of the same hash position of different Bloom filters in a head-to-tail connection manner. The data formed in this way is arranged in rows in the block according to the order of the hash positions. In this way, through one read operation, a query can be performed on multiple Bloom filters at once, and it can be determined whether multiple k-mer data sets contain the query element.

[0137] The specific process of index construction is as follows: First, perform data preprocessing, estimate the cardinality of the input n k-mers sets, and the HyperLogLog algorithm is used.

[0138] HyperLogLog (HLL) is a probabilistic algorithm used to estimate the number of distinct elements (cardinality) in a large dataset. Its basic principle is based on logarithmic probability distribution and hash functions. When an element enters the dataset, it is first mapped to a binary hash value through a hash function. Then, observe the number of consecutive zeros starting from the left of this hash value. For example, if the hash value is "00101", the number of consecutive zeros is 2. By performing such operations on a large number of elements, the maximum value of these consecutive zero counts is used to estimate the cardinality of the dataset.

[0139] Specifically, HLL divides the dataset into multiple buckets (usually a power of 2 number of buckets), and each bucket records the maximum value of the number of consecutive zeros in the hash values of the elements therein. Then, based on these maximum values in the buckets, an estimated value of the dataset cardinality is calculated through a complex mathematical formula.

[0140] Then, determine the element cardinality b stored in the base Bloom filter and the size m of the base Bloom filter. According to the number of Bloom filters required by the set, divide the n sets into different segments.

[0141] Segment L stores the sets that require L base Bloom filters. Logically, segment L is divided into L groups, that is, the sequence number L of segment L is the same as the number of Bloom filters it contains and the number of corresponding groups. Each k-mer dataset stored in this segment L has a Bloom filter in each group. For the k-mer dataset S i containing the element x, when constructing, first determine the group index of x stored in segment L, and the specific calculation formula is:

[0142] g x = g(x) mod L

[0143] Store the element x into the Bloom filter allocated to the k-mer dataset S x in group g i Specifically, the storage method is to use k independent hash functions h 0 , … h k-1 , and each hash function randomly and uniformly maps the element x in the k-mer dataset S i to an integer in {0,..., m - 1}. Set the corresponding k positions in this Bloom filter to 1, that is, set B[h 0 (x)], …, B[h k-1 (x)] to 1.

[0144] The specific process of element query is as follows: For the query element q, initialize M q as an empty list, which is used to store the datasets containing the query k-mer q.

[0145] Perform queries segment by segment. Query the current segment L, and use the same grouped hash function as in the construction phase to calculate the group index to determine the group g to be queried for the current segment q . The specific calculation formula is as follows:

[0146] g q = g(q) mod L

[0147] For each block in group g q , obtain the k hash indices h 0 (x), … h k-1 (x) corresponding k row bit vectors, and perform a bitwise AND operation on these bit vectors to generate a bit vector of the same length. Concatenate the bit vectors obtained for each block. Assuming the number of sets stored in segment L is n, the final query result for segment L is a bit vector of length n, where each bit indicates whether the corresponding set contains the element q. Specifically, if a certain bit is 1, it means the corresponding set contains q; if it is 0, it means it does not contain q.

[0148] Query one group for each segment, and add the sets that contain the element q to the query result M q . Continue to query the next segment.

[0149] The specific process for incremental addition of the dataset is as follows: Insert the k-mer dataset s. First, perform cardinality estimation to determine the segment L to be stored s .

[0150] Convert segment L s from the interleaved Bloom filter form to the Bloom filter matrix form to support the insertion operation.

[0151] Find the insertion position for s in segment L s . The insertion position can be an empty position left by a deletion operation or a newly created position for the k-mer dataset s.

[0152] Store the elements in the k-mer dataset s into the L s Bloom filters allocated to it s .

[0153] Convert the involved blocks from the Bloom filter matrix form back to the interleaved Bloom filter form to support subsequent queries.

[0154] The specific process for deletion of the dataset is as follows: Delete the k-mer dataset s and determine that the k-mer dataset s is stored in the index

[0155] If the k-mer dataset s is stored in the index, then obtain the segment L where the k-mer dataset s is stored from the metadata s。

[0156] Convert segment L s from the interleaved Bloom filter form to the Bloom filter matrix form to support query operations.

[0157] Empty or delete all Bloom filters owned by the k-mer dataset s. If the number of deleted sets is small, keep the empty positions. If all sets in an entire block are deleted, release the space resources of the block.

[0158] Convert the involved blocks from the Bloom filter matrix form back to the interleaved Bloom filter form to support subsequent queries.

[0159] Figure 4 is a schematic diagram of the index construction process of k-mer data in the embodiment mode of the present invention, as Figure 4 shown. The specific implementation of index construction includes the following steps:

[0160] S11: Use the HyperLogLog algorithm to estimate the cardinality of the input n k-mers sets.

[0161] S12: Determine the element cardinality b stored in the base Bloom filter and the base Bloom filter size m, and divide the n sets into corresponding segments. The specific division strategy is as follows: For the k-mer dataset S with cardinality c i , if (L - 1)b < c < Lb, then divide the k-mer dataset S i into segment L.

[0162] S13: For the element x in the k-mer dataset S that requires L Bloom filters i , during construction, then determine the group index where x is stored in segment L through the grouped hash function, and then store it in the Bloom filter assigned to the k-mer dataset S i in this group.

[0163] Figure 5 is a schematic diagram of the element query process of the k-mer data index in the embodiment mode of the present invention, as Figure 5 shown. The specific implementation steps of the element query process are as follows:

[0164] S21: For the query element q, initialize M q as an empty list to store the datasets containing k-mer q.

[0165] S22: Perform queries segment by segment.

[0166] S23: Query the current segment L. Segment L has L groups, and the sets stored in L each have a Bloom filter in each group.

[0167] S24: Calculate the group index g in which the element q is stored in L using the same grouping function as in the construction phase q , that is, if q is stored in L, it must be stored in group g q .

[0168] S25: Query all the Bloom filters in g q and add the set containing the element q to the query result M q . Continue to query the next segment

[0169] S26: Return the query result M after all segments have been queried q .

[0170] Figure 6 is a schematic diagram of the process for adding a data set to the k-mer data index in the embodiment of the present invention. As Figure 6 shown, the specific implementation steps for incrementally adding a data set are as follows

[0171] S31: Insert the k-mer data set s

[0172] S32: First, perform cardinality estimation to determine the segment L to be stored in s .

[0173] S33: Convert the segment L s from the interleaved Bloom filter form to the Bloom filter matrix form

[0174] S34: Find the insertion position for the k-mer data set s in the segment L s . The insertion position can be an empty position left by a deletion operation or a newly created position for the k-mer data set s. Each of the L s groups in the segment L s will allocate a Bloom filter for the k-mer data set s, and the relative positions of the k-mer data set s in each group remain consistent

[0175] S35: Store the elements in the k-mer data set s into the L s Bloom filters allocated to it

[0176] S36: Convert the involved blocks from the Bloom filter matrix form back to the interleaved Bloom filter form to support subsequent queries

[0177] Figure 7 is a schematic diagram of the process for deleting a data set from the k-mer data index in the embodiment of the present invention. As Figure 7 shown, the specific implementation steps for deleting a data set are as follows

[0178] S41: Delete the k-mer dataset s.

[0179] S42: Search for the k-mer dataset s in the metadata and determine that the k-mer dataset s is stored in the index.

[0180] S43: If the k-mer dataset s is stored in the index, obtain the segment L where the k-mer dataset s is stored from the metadata. s 。

[0181] S44: Convert the segment L s from the interleaved Bloom filter form to the Bloom filter matrix form.

[0182] S45: Empty or destroy all Bloom filter implementations corresponding to the set to implement deletion.

[0183] S46: Convert the involved blocks back from the Bloom filter matrix form to the interleaved Bloom filter form to support subsequent queries.

[0184] Figure 8 is a schematic diagram of an approximate membership query index generation device for a dynamic k-mer dataset according to an embodiment of the present invention. As Figure 8 shown, based on the above approximate membership query index generation method for a dynamic k-mer dataset provided by the embodiment of the present invention, the embodiment of the present invention further provides an approximate membership query index generation device for a dynamic k-mer dataset. The device includes:

[0185] A cardinality module 801 for determining the element cardinality of the target k-mer dataset to be stored according to a cardinality estimation algorithm;

[0186] A selection module 802 connected to the above cardinality module 801 for selecting the target first unit of the first level to be stored according to the element cardinality. Among them, the number of Bloom filters of multiple first units of the first level is different, and the capacity of the Bloom filter of each first unit is the same. The Bloom filter is used to store the corresponding elements in the target k-mer dataset, and the element cardinality that can be stored in the first unit is not less than the element cardinality of the corresponding k-mer dataset;

[0187] A grouping module 803 connected to the above selection module 802 for generating a second-level second unit index of the target element according to the target element of the target k-mer dataset through the grouping function of the target first unit, and dividing the target element into the target second unit corresponding to the second unit index. Among them, each first unit includes multiple second units, and the number of second units of each k-mer dataset stored in the first unit is the same as the number of Bloom filters required for the corresponding k-mer dataset;

[0188] The allocation module 804, connected to the above-mentioned grouping module 803, is used to store the target element into the target third unit at the third level in the target second unit and allocate it to the Bloom filter of the target k-mer data set. Wherein, the second unit includes a plurality of third units with fixed sizes, the organization form of the third unit is an interleaved Bloom filter, and the third unit reads the same bit of different Bloom filters at one time according to the interleaved Bloom filter when querying elements.

[0189] The above-mentioned approximate membership query index generation device for a dynamic k-mer data set provided by the embodiment of the present invention selects the target first unit at the first level to be stored according to the element cardinality, and uses the first unit containing a fixed number of Bloom filters to store the k-mer data set corresponding to the fixed number of corresponding element cardinalities. In this way, sufficient and as small as possible space can be allocated for each k-mer data set, avoiding space waste.

[0190] Through the grouping function of the target first unit, according to the target element of the target k-mer data set, the second unit index at the second level of the target element is generated, and the target element is divided into the target second unit corresponding to the second unit index. The number of second units of each stored k-mer data set is the same as the number of Bloom filters required by the corresponding k-mer data set. Each second unit stores different element sets of the k-mer data set. In this way, when querying, for an element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0191] The target element is stored into the target third unit at the third level in the target second unit and allocated to the Bloom filter of the target k-mer data set. By using the interleaved Bloom filter of the third unit, when the third unit queries elements, it can read the same bit of different Bloom filters at one time according to the interleaved Bloom filter, so as to determine the target k-mer number aggregation where the query element is located at one time, avoiding the situation that when using a Bloom filter matrix, only one Bloom filter's data can be read at a time and it is necessary to traverse the Bloom filters to determine the k-mer data set to which the query element belongs.

[0192] Furthermore, it achieves the technical effect of supporting incremental updates and deletions on the basis of ensuring that the occupied space is as small as possible and the number of queries is as small as possible to meet the usage requirements of k-mer data.

[0193] Figure 9 It is a schematic diagram of an approximate membership query device for a dynamic k-mer data set according to an embodiment of the present invention, as Figure 9As shown, based on the above approximate membership query method for a dynamic k-mer dataset provided by the embodiments of the present invention, the embodiments of the present invention also provide an approximate membership query device for a dynamic k-mer dataset, and the device includes:

[0194] A generation module 901, configured to generate a second unit index of a second level in a first unit for a target element through a grouping function of a first unit of a first level and the target element to be queried, so as to identify a target second unit where the target element is located. Wherein, the target element is indexed level by level in a multi-level manner. When stored in the second level, the corresponding second unit index is generated by using the grouping function of the first unit to which it belongs, and is stored in the target second unit corresponding to the second unit index.

[0195] A determination module 902, connected to the above generation module 901, is configured to determine that the target element is stored in the target second unit of the first unit when the target second unit exists in the first unit.

[0196] An indexing module 903, connected to the above determination module 902, is configured to obtain an interleaved Bloom filter of a third unit of a third level in the target second unit, and obtain a plurality of hash indexes according to the interleaved Bloom filter. Wherein, the third unit is a component of the second unit to which it belongs. When the target element is stored in the third unit, it is assigned to the Bloom filter of the target k-mer dataset to which it belongs. The organization form of the third unit is an interleaved Bloom filter, and the number of bits of the hash index is the same as the number of Bloom filters of the target second unit. The Bloom filters correspond to different k-mer datasets.

[0197] A query module 904, connected to the above indexing module 903, is configured to query the target k-mer dataset to which the target element belongs according to the hash index.

[0198] For the above approximate membership query device for a dynamic k-mer dataset provided by the embodiments of the present invention, through the grouping function of the first unit of the first level and the target element to be queried, a second unit index of the second level in the first unit for the target element is generated. By using the grouping function of the first unit containing Bloom filters with the same element cardinality, the second unit index is directly generated. By querying whether the second unit indexes are the same, for a target element, only one Bloom filter needs to be queried to determine whether the element is in the second unit and the first unit to which it belongs.

[0199] Obtain the staggered Bloom filter of the third unit at the third level in the target second unit. According to the staggered Bloom filter, obtain multiple hash indexes, and query the target k-mer dataset to which the target element belongs. Using the staggered Bloom filter of the third unit, when querying elements in the third unit, the same bit of different Bloom filters can be read at one time according to the staggered Bloom filter, so as to determine the target k-mer data cluster where the query element is located at one time, avoiding the situation that when using a Bloom filter matrix, only one Bloom filter's data can be read at a time, and it is necessary to traverse the Bloom filter to determine the k-mer dataset to which the query element belongs.

[0200] Furthermore, it achieves the technical effect of supporting incremental updates and deletions on the basis of ensuring that the occupied space is as small as possible and the number of queries is as small as possible, so as to meet the usage requirements of k-mer data.

[0201] An embodiment of the present invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.

[0202] An embodiment of the present invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a processor of a computer, is used to cause the computer to execute the method of the embodiment of the present invention.

[0203] An embodiment of the present invention also provides an approximate membership query system for a dynamic k-mer dataset, including: an electronic device, the electronic device includes at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program capable of being executed by the at least one processor, and the computer program, when executed by the at least one processor, is used to cause the electronic device to execute the method of the embodiment of the present invention.

[0204] Reference Figure 10 , the structural block diagram of the electronic device that can be used as a server or a client in the embodiment of the present invention will now be described. It is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0205] As Figure 10As shown, the electronic device includes a computing unit 1001, which can execute various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0206] Multiple components in the electronic device are connected to the I / O interface 1005, including: an input unit 1006, an output unit 1007, a storage unit 1008, and a communication unit 1009. The input unit 1006 can be any type of device capable of inputting information into the electronic device. The input unit 1006 can receive input digital or character information and generate key signal inputs related to the user settings and / or function controls of the electronic device. The output unit 1007 can be any type of device capable of presenting information and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 1008 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 1009 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks and can include, but is not limited to, a modem, a network card, an infrared communication device, and / or a wireless communication transceiver, such as a Bluetooth device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0207] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a CPU, a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 1002 and / or the communication unit 1009. In some embodiments, the computing unit 1001 can be configured to execute the above methods in any other appropriate manner (e.g., by means of firmware).

[0208] The computer program for implementing the method of the embodiment of the present inventive concept may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0209] In the context of the embodiment of the present inventive concept, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable signal medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, or infrared system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0210] It should be noted that the term "including" and its variations used in the embodiments of the present inventive concept are open-ended, that is, "including but not limited to". The term "based on" is "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "a plurality" mentioned in the embodiments of the present inventive concept are illustrative rather than restrictive. Those skilled in the art should understand that, unless clearly specified otherwise in the context, it should be understood as "one or more".

[0211] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the embodiments of the present inventive concept are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select authorization or rejection.

[0212] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the protection scope. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the appended claims.

Claims

1. A method for generating approximate membership query index for dynamic k-mer datasets, characterized in that: include: Determine the element cardinality of the target data set to be stored according to a cardinality estimation algorithm; Selecting a target first unit of a first level to be stored according to the element cardinality; The first units of the first level have different numbers of Bloom filters, and the capacities of the Bloom filters of the first units are the same. The Bloom filters are used to store corresponding elements in the target data set, and the cardinality of elements that can be stored in the first unit is not less than the cardinality of elements in the corresponding data set. Generate a second unit index of a second level of the target element according to the target element of the target data set through the grouping function of the target first unit, and divide the target element into target second units corresponding to the second unit index; the first unit includes a plurality of second units, and the number of second units of each data set stored in the first unit is the same as the number of Bloom filters required by the corresponding data set; The target element is stored in a target third unit of a third level in the target second unit, and is assigned to a Bloom filter of the target data set; the second unit includes a plurality of third units of fixed size, the third units are organized in an interleaved Bloom filter, and the third units read the same bit of different Bloom filters at one time according to the interleaved Bloom filter when querying an element; When inserting the first k-mer data set, first determine the corresponding first unit to be inserted according to the element cardinality of the first k-mer data set; after determining the first unit, convert the organization form of the third unit in the first unit from the interleaved Bloom filter to the Bloom filter matrix; After the insertion, converting the organization form of the third unit inserted into the first unit from a Bloom filter matrix to an interleaved Bloom filter; When deleting the second k-mer data set, first query the first unit to which the second k-mer data set belongs; after determining the first unit, convert the organization form of the third unit in the first unit from the interleaved Bloom filter to the Bloom filter matrix; After the deletion, the organization of the third unit inserted into the first unit is converted from a Bloom filter matrix to an interleaved Bloom filter.

2. The method according to claim 1, characterized in that Selecting a target first unit of a first level to be stored according to the cardinality includes: The element cardinality of the target data set is compared with the element cardinality that can be stored in the first unit; When the element cardinality of the target data set is less than or equal to the element cardinality that can be stored in the comparison first unit and greater than the element cardinality that can be stored in the previous first unit, the comparison first unit is used as the target first unit; The first comparison unit is one of a plurality of first units, and the previous first unit is a first unit whose number of Bloom filters is less than that of the first comparison unit by a preset difference.

3. The method according to claim 1, characterized in that Storing the target element in a target third unit of a third level in the target second unit and assigning it to a Bloom filter of the target data set includes: storing the target element into a target third unit of a third level in the target second unit; The target element is hashed using multiple independent hash functions of the target third unit and mapped to a corresponding position in the corresponding Bloom filter, and the corresponding position is set to 1.

4. The method according to claim 1, characterized in that: The method further comprises: Receive a first data set to be inserted; Determine, according to the element cardinality of the first data set, a corresponding first insertion unit; Converting the organization form of the third unit inserted into the first unit from an interleaved Bloom filter to a Bloom filter matrix; Searching for a corresponding insertion position in the first insertion unit, wherein the insertion position is found through a Bloom filter; Inserting the first data set into the insertion position, and assigning a corresponding Bloom filter to the first data set; The organization form of the third unit inserted into the first unit is converted from a Bloom filter matrix to an interleaved Bloom filter.

5. The method according to claim 1, characterized in that The method further comprises: receiving a second data set to be deleted; Query the first unit corresponding to the second data set; Converting the organization form of the third unit belonging to the first unit from an interleaved Bloom filter to a Bloom filter matrix; When the element set of the second data set is a partial element set of the third unit, the Bloom filter corresponding to the second data set is deleted, and an empty position is reserved; In a case where the element set of the second data set is the entire element set of the third unit, deleting the Bloom filter corresponding to the second data set, and releasing the space resources of the third unit; The organization form of the third unit belonging to the first unit is converted from a Bloom filter matrix to an interleaved Bloom filter.

6. An approximate membership query method for dynamic k-mer datasets, characterized in that: According to the method for generating an approximate member query index for a dynamic k-mer data set according to any one of claims 1 to 5, querying the generated index comprises: Generate a second unit index of the second level of the target element in the first unit through a grouping function of the first unit of the first level and a target element to be queried, so as to identify the target second unit where the target element is located, wherein the target element is indexed level by level in a multi-level manner, and when being stored in the second level, the corresponding second unit index is generated by using the grouping function of the first unit to which it belongs, and is stored in the target second unit corresponding to the second unit index; In the case where the target second cell exists in the first cell, determining that the target element is stored in the target second cell of the first cell; Obtain an interleaved Bloom filter of a third unit of a third level in the target second unit, and obtain a plurality of hash indexes according to the interleaved Bloom filter, wherein the third unit is a component of the second unit to which it belongs, and when the target element is stored in the third unit, it will be assigned to the Bloom filter of the target data set to which it belongs, and the third unit is organized in an interleaved Bloom filter, and the number of bits of the hash index is the same as the number of Bloom filters of the target second unit, and the Bloom filters correspond to different data sets; The target data set to which the target element belongs is queried according to the hash index.

7. The method according to claim 6, characterized in that Querying a target data set to which the target element belongs according to the hash index includes: Performing a bitwise AND operation on the plurality of hash indexes to obtain a target index, wherein the hash index is data of the same bit of different Bloom filters, the hash index is a mapping of the same hash function to elements of different data sets, different bits of the hash index represent values ​​of different Bloom filters, and the values ​​represent whether the Bloom filter has the target element; The target data set to which the target element belongs is determined according to the target index.

8. The method according to claim 6, characterized in that After querying the target data set to which the target element belongs according to the hash index, the method further includes: storing the target data set in an output list; Continue to query whether the index of the target element exists in other first units; If the index of the target element exists in the first unit, obtain the corresponding target data set and store it in the output list; When all first units have completed traversal, the output list is output.

9. The method according to any one of claims 1 to 8, characterized in that The first unit of the first level is a segment unit, the segment unit includes a plurality of Bloom filters, the capacity of the Bloom filters of each segment unit is the same, one Bloom filter records a corresponding element in a data set, and the element cardinality of the data set that can be stored by the segment unit is greater than or equal to the element cardinality of the data set that needs to be stored; The second unit of the second level is a group unit, and the number of the group units of each data set stored in the segment unit is the same as the number of Bloom filters required by the corresponding data set; The third unit of the third level is a block unit, which is a data block with a fixed data size and is a component unit of the group unit. Different elements are stored in the form of an interleaved Bloom filter. When querying an element, the block unit reads all data at positions corresponding to the query element in different Bloom filters at one time according to the interleaved Bloom filters.

10. An approximate membership query system for dynamic k-mer datasets, comprising: A processor and a memory storing a program, wherein the program comprises instructions, which, when executed by the processor, cause the processor to perform the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Systems, methods, and devices for near data processing

    CN113628647A

  • Efficient lookup in multiple bloom filters

    US20170154099A1