Large-scale protein sequence clustering method, system and product based on minimum hashing

By adopting a minimal hash-based method in protein sequence clustering, the problem of excessive computational complexity in large-scale data processing is solved, and a more efficient clustering process and lower resource requirements are achieved.

CN119851771BActive Publication Date: 2025-05-16SHANDONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510314828.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-05-16
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

When traditional protein clustering algorithms deal with large-scale protein sequences, the computational complexity is too high, resulting in huge time overhead and serious demand for computing resources.

Method used

A large-scale protein sequence clustering method based on minimum hash is used. Using the same hash function, you only need to traverse the protein sequences once, build a minhash set of each sequence, and group the sequences in combination with the set grouping rules, and finally cluster them through the heuristic greedy incremental method.

Benefits of technology

It significantly reduces the time overhead and computational complexity of clustering, reduces the demand for computing resources, maintains the accuracy of clustering results, and improves the efficiency of large-scale protein sequence clustering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119851771B_ABST
    Figure CN119851771B_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of protein sequence clustering, and provides a large-scale protein sequence clustering method, system and product based on minimum hashing. The technical scheme is to construct a unique ID for each sequence, and construct a k-mer set of the sequence; use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the minimum M hash values ​​as the minhash set of the sequence; combine the minhash set of the sequence and the set grouping rule to group the protein sequence; wherein the set grouping rule is that if there is a collision between the minhash sets of two sequences, they are divided into the same group; cluster each group of sequences separately, and integrate the clustered results to obtain the final clustering result. The time overhead of clustering is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of protein sequence clustering, and in particular relates to a large-scale protein sequence clustering method, system and product based on minimum hashing. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] With the development of genomics and bioinformatics, protein clustering, as a key method for analyzing protein sequence and structural similarity, is widely used in protein function prediction, protein family identification and biological evolution research. Through clustering, proteins with similar functions or evolutionary relationships can be classified into the same category, helping scientists to gain a deeper understanding of the biological functions and evolutionary paths of proteins. Protein clustering technology not only improves the efficiency of bioinformatics analysis, but also effectively reduces data redundancy and reduces computing and storage costs.

[0004] However, with the rapid development of high-throughput sequencing technology, the number of protein sequences has increased dramatically, which has posed a severe challenge to traditional protein clustering algorithms. First, with the explosive growth of protein data, the time complexity of clustering tasks has also increased. Especially when faced with massive protein sequences, the computational complexity of many traditional clustering algorithms has increased exponentially. For example, traditional greedy incremental clustering methods usually require each sequence to be compared with other sequences one by one, resulting in a sharp drop in the efficiency of the clustering process, especially when the number of sequences reaches millions or even hundreds of millions, the time overhead increases exponentially.

[0005] In addition, large-scale data not only increases computing time, but also greatly increases the demand for computing resources. When clustering protein sequences, a large number of sequences must be loaded into memory for processing, which not only requires high memory consumption, but also leads to a significant increase in data access and processing time. As the amount of data increases, the clustering task's dependence on computing resources continues to increase, and may require high-performance computing clusters or even distributed computing systems to support it.

[0006] Existing clustering methods for large-scale protein sequences, such as the one with publication number CN119007838A and invention name "Grouping-based Protein Sequence Clustering Method and System", use multiple hash functions during clustering, and each hash function needs to traverse the protein sequence once and select the smallest hash value. In this way, it is necessary to traverse the protein sequence m times. When the scale of the protein sequence is particularly large, the time overhead is also very huge. Summary of the invention

[0007] In order to solve at least one of the technical problems existing in the above-mentioned background technology, the present invention provides a large-scale protein sequence clustering method and system based on minimum hashing, which aims to maintain the accuracy of the clustering results while further reducing the complexity of the calculation. Only one representative hash function is used, and only one protein sequence needs to be traversed, which greatly reduces the time overhead of clustering.

[0008] In order to achieve the above object, the present invention adopts the following technical solution:

[0009] The first aspect of the present invention provides a large-scale protein sequence clustering method based on minimum hashing, comprising the following steps:

[0010] Obtain protein sequence data to be clustered;

[0011] For each sequence, construct a unique ID and build a k-mer set for the sequence;

[0012] Use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence;

[0013] The protein sequences are grouped by combining the minhash set of the sequences and the set grouping rules. The set grouping rule is that if the minhash sets of two sequences collide, they are grouped into the same group.

[0014] Each group of sequences is clustered separately, and the clustering results are integrated to obtain the final clustering result.

[0015] Furthermore, the protein sequences are grouped by combining the minhash set of the sequences and the set grouping rules, including:

[0016] Use a hash table to store the mapping from minhash to sequence ID, traverse each minhash value of all sequences, if the minhash value already exists in the hash table, use union-find to group the current sequence ID and the existing sequence ID into the same group, if the minhash value is not in the hash table, create a new group and insert the mapping in the hash table.

[0017] Furthermore, the sliding window method is used to generate the k-mer set for each sequence.

[0018] Furthermore, the same hash function is used to map the k-mers in the k-mer set into 64-bit unsigned integer hash values ​​one by one.

[0019] Furthermore, for the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

[0020] Furthermore, a heuristic greedy incremental method is used when clustering each group of sequences separately.

[0021] The second aspect of the present invention provides a large-scale protein sequence clustering system based on minimum hashing, comprising:

[0022] A data acquisition module, which is used to acquire protein sequence data to be clustered;

[0023] The minhash set construction module is used to construct a unique ID for each sequence and construct the k-mer set of the sequence; use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence;

[0024] A grouping module is used to group protein sequences by combining the minhash set of the sequences and the set grouping rules; wherein the set grouping rule is that if there is a collision between the minhash sets of two sequences, they are divided into the same group;

[0025] The clustering module is used to cluster each group of sequences separately and integrate the clustering results to obtain the final clustering result.

[0026] In the grouping module, the protein sequences are grouped by combining the minhash set of the sequence and the set grouping rules, including:

[0027] Furthermore, a hash table is used to store the mapping from minhash to sequence ID, and each minhash value of all sequences is traversed. If the minhash value already exists in the hash table, the current sequence ID and the existing sequence ID are grouped into the same group using union-find. If the minhash value is not in the hash table, a new group is created and the mapping is inserted into the hash table.

[0028] Furthermore, for the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

[0029] A third aspect of the present invention provides a program product.

[0030] A program product, which is a computer program product, comprises a computer program, and when the computer program is executed by a processor, the steps in the large-scale protein sequence clustering method based on minimum hashing as described above are implemented.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] The large-scale protein sequence clustering algorithm based on minimum hashing designed by the present invention addresses the clustering needs of large-scale protein sequences and solves the limitations of traditional algorithms in scalability. It only uses one representative hash function and only needs to traverse the protein sequence once, which greatly reduces the time overhead of clustering and significantly reduces the computational complexity. Compared with existing large-scale protein sequence clustering algorithms, the present invention can further reduce time overhead while ensuring accuracy.

[0033] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0035] Figure 1 It is a flow chart of a large-scale protein sequence clustering method based on minimum hashing provided by an embodiment of the present invention;

[0036] Figure 2 This is the minhash value construction process of the protein sequence under a specific hash function provided by the embodiment of the present invention;

[0037] Figure 3 It is a schematic diagram of a grouping process provided by an embodiment of the present invention;

[0038] Figure 4 This is a grouping process visualization process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0039] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0040] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0041] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0042] As mentioned in the background, with the rapid increase in the number of protein sequences, the efficiency of traditional clustering algorithms has dropped significantly due to their high computational complexity, especially when dealing with billions of sequences. The existing large-scale protein sequence clustering methods, although also based on the minimum hashing method, use m hash functions, each of which needs to traverse the protein sequence and select the smallest hash value. In this way, m When the scale of protein sequences is particularly large, the time cost is also very huge. To solve this problem, the present invention designs a large-scale protein sequence clustering method based on minimum hashing, which aims to maintain the accuracy of clustering results while further reducing the complexity of calculation, using only one representative hash function, and only needing to traverse the protein sequence once, which greatly reduces the time cost of clustering.

[0043] Embodiment 1

[0044] like Figure 1 As shown, this embodiment provides a large-scale protein sequence clustering method based on minimum hashing, comprising the following steps:

[0045] Step 1: Obtain the protein sequence data to be clustered. For each sequence, construct a unique ID and construct a k-mer (k-mer refers to a DNA fragment with a length of k) set of the sequence.

[0046] In this embodiment, the initial protein sequence is input in the format of a FASTA file, and the input protein sequence FASTA file is parsed. A unique ID is first created for each protein sequence, and a corresponding k-mer set is generated for each sequence using a sliding window method.

[0047] Step 2: Use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence;

[0048] In this embodiment, the same hash function is used to map the k-mers in the k-mer set into 64-bit unsigned integer hash values ​​one by one. Among the multiple hash values ​​generated, a priority queue is used to select the smallest M hash values ​​as the minhash set of the sequence, thereby obtaining a sketch of the protein sequence.

[0049] like Figure 2 As shown in the figure, taking an actual sequence as an example, the process of constructing the minhash set is explained. Assume that the sequence is MAGHLASDFAFFD, and the k-mer set of this sequence includes MAGHLASD, AGHLASDF, GHLASDFA, HLASDFAF, etc. Use the same hash function to map the k-mers in the k-mer set into hash values ​​corresponding to hash value 1, hash value 2, hash value 3, ..., hash value n one by one. The priority queue selects M minimum hash values ​​to obtain the minhash set.

[0050] Step 3: Group the protein sequences by combining the minhash set of the sequences and the set grouping rules;

[0051] The grouping process is as follows Figure 3 As shown, the set grouping rules specifically include: if there is a collision between the minhash sets of two sequences, they are divided into the same group;

[0052] To simplify the grouping, only the constructed protein sequence ID was retained as the identifier. After the grouping was completed, the corresponding sequence information (sequence name and content, etc.) was extracted according to the sequence ID of each group, and each group of sequences was written into an independent FASTA file for subsequent clustering analysis.

[0053] Specifically, a hash table is used to store the mapping from minhash to sequence ID to support fast search and insertion operations; each minhash value of all sequences is traversed. If the minhash value already exists in the hash table, the current sequence ID and the existing sequence ID are grouped into the same group using union-find. If the minhash value is not in the hash table, a new group is created and the mapping is inserted into the hash table. Through this strategy, hash collisions can be detected quickly and efficiently, and a preliminary grouping structure can be established.

[0054] Visualization process Figure 4As shown, taking an actual sequence grouping as an example, the specific grouping process is explained, including: if the ID of a sequence is 5, the minimum 4 hash values ​​corresponding to the sequence include 185, 201, 246, and 353 as the minhash set of the sequence; traverse each minhash value of all sequences, where hash value 185 and hash value 201 exist in the hash table, the sequence corresponding to hash value 185 in the hash table is ID3, and the sequence corresponding to hash value 201 in the hash table is ID1, then use the union-find set to classify the current sequence ID and the existing sequence ID into the same group, group the grouping results corresponding to hash value 185 as ID5 and ID3 as a group, and group the grouping results corresponding to hash value 201 as ID5, ID3, and ID1 as a group; and hash value 246 and hash value 356 do not exist in the hash table, create a new group and insert the mapping into the hash table as: hash value 246-ID5 and hash value 356-ID5.

[0055] Specific grouping complexity analysis:

[0056] The same hash function is used for each protein sequence. When there are N protein sequences, the length of each protein sequence is L. It is necessary to traverse each sequence and use a priority queue to store M minimum hash values, with a complexity of O(NLlogM).

[0057] When grouping, each sequence has M minimum hash values. Each hash value is traversed to find or insert into the hash table. The complexity of finding and inserting into the hash table is O(1), and each sequence requires M operations.

[0058] For N protein sequences, the complexity is O(NM). After finding the collision in the hash table, use union-find to group them together. The single operation complexity of the union-find operation (with path compression and rank merging) is O(1). For all NM hash values, the complexity is O(NM).

[0059] Therefore, for the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

[0060] In principle, given a minimum number of hash values ​​M, it can be ensured that the minimum Jaccard similarity s between sequences is approximated to 1 / M with a high accuracy.

[0061] This means that when M=10, the grouping strategy of the present invention can correctly group sequences with a Jaccard similarity of more than about 10%. In conventional cluster analysis, more attention is paid to sequences with higher similarity, so the grouping strategy under a smaller M value can meet actual needs. After grouping, the number of sequences in each group is significantly less than the original number of sequences, so that the subsequent cluster analysis for each group greatly reduces the algorithm complexity and improves the overall calculation efficiency.

[0062] By calculating the hash value for the k-mer set of each protein sequence and extracting the minimum M hash values ​​for comparison, the probability that the minimum M hash values ​​of two sets are equal can be directly approximated as their Jaccard similarity. In protein clustering analysis, Jaccard similarity is used to quantify the similarity between two protein sequences. Through the above method, sequences with high similarity can be effectively selected and classified into the same group. Specifically, the higher the similarity of two protein sequences, the more identical k-mers there are, the greater the probability of collision of the minimum M hash values, and the greater the possibility that they are classified into the same group.

[0063] Based on the above principles, whether the sequences have potential similarity can be judged by the collision of the minimum M hash values ​​between two sequences. Specifically, if the minimum M hash values ​​generated by two sequences under any hash function collide, the two sequences can be considered to be similar and can be classified into the same group.

[0064] Assume that the Jaccard similarity s of two sequences is 0.9. For a specific hash function, the probability that the hash values ​​of the two sequences are equal is s=0.9, and the probability that they are not equal is 1-s=0.1. If the smallest M hash values ​​are selected for each protein sequence, as long as any of the M hash values ​​collide, the two sequences will be classified into the same group. The probability that the M smallest hash values ​​of the two sequences do not collide is (1-s) M For example, when s=0.9 and M=15, the probability that the two sequences are divided into different groups is (1−0.9) 15 =0.1 15 , almost zero. This grouping strategy can effectively ensure that the similarity of sequences between groups is lower than the set Jaccard similarity threshold, thereby achieving efficient and accurate grouping. In addition, the number of sequences in the group after grouping is usually much smaller than the number of original sequences. In the subsequent clustering process, the time complexity of the algorithm will be greatly reduced.

[0065] Step 4: Perform clustering operation for each group of sequences separately;

[0066] After the grouping is completed, a clustering operation is performed separately for each group of sequences. In this embodiment, a classic clustering algorithm can be used for analysis. Here, a heuristic greedy incremental method is taken as an example to gradually build clusters for each group, specifically including the following steps:

[0067] First, a subsequence set of length k is constructed for all protein sequences as the basis for measuring similarity in the clustering process;

[0068] Then, the sequences of each group are sorted according to the sequence length, and the longest sequence among the unclustered sequences is selected as the initial clustering center, because in general, longer sequences contain more effective information and can more comprehensively represent the category to which they belong.

[0069] For each new sequence after sorting, it is compared with the existing cluster centers in turn. First, check whether the sequence shares a sufficient number of identical subsequences with the cluster center. If the conditions are met, further global alignment is performed; otherwise, the current sequence is used as a new cluster center. Only when the proportion of identical characters in the global alignment of two sequences exceeds the clustering threshold are they classified into the same category.

[0070] The clustering threshold is a similarity criterion that determines whether two protein sequences should be classified into the same cluster.

[0071] Since there is no overlap in clustering information between the groups after grouping, the clustering task can be naturally parallelized. When each group of sequences is clustered independently on a multi-core platform, there is no need for communication between the processing cores. This design significantly reduces communication overhead and achieves completely independent calculations between groups, thereby maximizing the efficiency of parallel processing. At the same time, this grouping clustering strategy makes full use of the computing power of modern multi-core platforms, avoiding synchronization costs and optimizing overall performance. In this way, the execution efficiency of large-scale protein sequence clustering tasks can be greatly improved while ensuring clustering accuracy.

[0072] Effect analysis

[0073] The existing large-scale protein sequence clustering algorithms, such as the method in CN119007838A, have a time complexity of O(MNL)+O(MNlogN)+O(L 2 N 2 / G), while the time complexity of the algorithm proposed in this invention is O(NLlogM)+O(NM)+O(L 2 N 2 / G), where N represents the number of protein sequences, M represents the minimum M hash values ​​of each protein sequence, L represents the sequence length, and G represents the number of groups. In practical applications, M can be fixed to a constant of about 10, and the size of G is close to N.

[0074] Therefore, the time complexity of the method of the present invention can be approximated as O(NL)+O(N)+O(NL 2 ), the time complexity of existing large-scale protein sequence clustering algorithms can be approximated as O(NL)+O(NlogN)+O(NL 2 ). In the case of a large protein scale, the value of N is particularly large, and O(N) is significantly lower than O(NlogN). Therefore, in this case, the computational cost of this method is significantly lower than that of existing large-scale protein sequence clustering algorithms. Thanks to the optimization of complexity, the software implemented by the present invention has a significantly higher operating efficiency than existing large-scale protein sequence clustering algorithms in actual tests.

[0075] At the same time, according to the grouping method proposed by the present invention, the probability of false negatives occurring in two sequences with a certain similarity s is (1-s) M For example, for two sequences with a similarity of 90%, the probability of a false negative is only 0.110 when M=10. This shows that the method of the present invention has high reliability and accuracy in the grouping process and can better meet the actual needs of large-scale protein sequence clustering.

[0076] Embodiment 2

[0077] This embodiment provides a large-scale protein sequence clustering system based on minimum hashing, including:

[0078] A data acquisition module, which is used to acquire protein sequence data to be clustered;

[0079] The minhash set construction module is used to construct a unique ID for each sequence and construct the k-mer set of the sequence; use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence;

[0080] A grouping module is used to group protein sequences by combining the minhash set of the sequences and the set grouping rules; wherein the set grouping rule is that if there is a collision between the minhash sets of two sequences, they are divided into the same group;

[0081] The clustering module is used to cluster each group of sequences separately and integrate the clustering results to obtain the final clustering result.

[0082] Among them, in the minhash set construction module, the sliding window method is used to generate the k-mer set of each sequence.

[0083] Furthermore, in the minhash set construction module, the same hash function is used to map the k-mers in the k-mer set into 64-bit unsigned integer hash values ​​one by one.

[0084] In the grouping module, the protein sequences are grouped by combining the minhash set of the sequence and the set grouping rules, including:

[0085] Use a hash table to store the mapping from minhash to sequence ID, traverse each minhash value of all sequences, if the minhash value already exists in the hash table, use union-find to group the current sequence ID and the existing sequence ID into the same group, if the minhash value is not in the hash table, create a new group and insert the mapping in the hash table.

[0086] Furthermore, for the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

[0087] Furthermore, in the clustering module, a heuristic greedy incremental method is used when clustering each group of sequences separately.

[0088] It should be noted that the specific implementation of the large-scale protein sequence clustering system based on minimum hashing in the embodiment of the present invention is similar to the specific implementation of the large-scale protein sequence clustering method based on minimum hashing in the embodiment of the present invention. Please refer to the description of the method part for details. In order to reduce redundancy, it will not be repeated here.

[0089] Embodiment 3

[0090] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the large-scale protein sequence clustering method based on minimum hashing as described above are implemented.

[0091] Embodiment 4

[0092] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the large-scale protein sequence clustering method based on minimum hashing as described above are implemented.

[0093] Embodiment 5

[0094] This embodiment provides a program product, which is a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the large-scale protein sequence clustering method based on minimum hashing as described above are implemented.

[0095] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A large-scale protein sequence clustering method based on minimum hashing, characterized in that: The steps include: Obtain protein sequence data to be clustered; For each sequence, construct a unique ID and build a k-mer set for the sequence; Use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence; The protein sequences are grouped by combining the minhash set of the sequences and the set grouping rules. The set grouping rule is that if the minhash sets of two sequences collide, they are grouped into the same group. Each group of sequences is clustered separately, and the clustering results are integrated to obtain the final clustering result.

2. The large-scale protein sequence clustering method based on minimum hashing as claimed in claim 1, characterized in that: Combine the minhash set of the sequence and the set grouping rules to group the protein sequences, including: Use a hash table to store the mapping from minhash to sequence ID, traverse each minhash value of all sequences, if the minhash value already exists in the hash table, use union-find to group the current sequence ID and the existing sequence ID into the same group, if the minhash value is not in the hash table, create a new group and insert the mapping in the hash table.

3. The large-scale protein sequence clustering method based on minimum hashing as claimed in claim 1, characterized in that: The sliding window method is used to generate the k-mer set for each sequence.

4. The large-scale protein sequence clustering method based on minimum hashing according to claim 1, characterized in that: Use the same hash function to map each k-mer in the k-mer set into a 64-bit unsigned integer hash value.

5. The large-scale protein sequence clustering method based on minimum hashing as claimed in claim 1, characterized in that: For the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

6. The large-scale protein sequence clustering method based on minimum hashing as claimed in claim 1, characterized in that: When clustering each group of sequences separately, a heuristic greedy incremental method is used.

7. A large-scale protein sequence clustering system based on minimum hashing, characterized by: include: A data acquisition module, which is used to acquire protein sequence data to be clustered; The minhash set construction module is used to construct a unique ID for each sequence and construct the k-mer set of the sequence; use the same hash function to map the k-mers in the k-mer set into hash values ​​one by one, and select the smallest M hash values ​​as the minhash set of the sequence; A grouping module is used to group protein sequences by combining the minhash set of the sequences and the set grouping rules; wherein the set grouping rule is that if there is a collision between the minhash sets of two sequences, they are divided into the same group; The clustering module is used to cluster each group of sequences separately and integrate the clustering results to obtain the final clustering result.

8. The large-scale protein sequence clustering system based on minimum hashing as claimed in claim 7, characterized in that: In the grouping module, the protein sequences are grouped by combining the minhash set of the sequence and the set grouping rules, including: Use a hash table to store the mapping from minhash to sequence ID, traverse each minhash value of all sequences, if the minhash value already exists in the hash table, use union-find to group the current sequence ID and the existing sequence ID into the same group, if the minhash value is not in the hash table, create a new group and insert the mapping in the hash table.

9. The large-scale protein sequence clustering system based on minimum hashing as claimed in claim 7, characterized in that: For the case of N protein sequences, the total complexity of grouping is O(NLlogM)+O(NM), where M represents the minimum M hash values ​​of each protein sequence and L represents the sequence length.

10. A program product, the program product being a computer program product, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps in the minimum hashing-based large-scale protein sequence clustering method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Grouping-based protein sequence clustering method and system

    CN119007838A

  • Large-scale biological data clustering method and system based on spanning tree

    CN114420215A

  • DNA sequence clustering method and system based on locality sensitive hash function, electronic equipment and readable storage medium

    CN118629513A