Heuristic biological sequence clustering method based on semi-global comparison

Through the heuristic biological sequence clustering method based on semi-global sequence alignment, the problems of low clustering sensitivity and overestimation in the prior art are solved, efficient biological sequence clustering is achieved, and the accuracy and efficiency of clustering are improved.

CN120299521APending Publication Date: 2025-07-11BAOJI UNIV OF ARTS & SCI
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311127149.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-01
Publication Date
2025-07-11

Smart Images

  • Figure CN120299521A_ABST
    Figure CN120299521A_ABST
Patent Text Reader

Abstract

The invention provides a heuristic biological sequence clustering method based on semi-global comparison. The method comprises the following steps: firstly, reading all biological sequences, removing repetitive sequences, and carrying out descending sorting on the biological sequences according to sequence lengths; a first sequence is used as a representative sequence of a first category, then a next sequence is read, the similarity between the current sequence and the representative sequence is calculated by adopting a semi-global sequence comparison algorithm, each sequence can be compared to the most similar area in the representative sequence through semi-global sequence comparison, the similarity between the sequences is found to the maximum extent, and the similarity between the sequences and the representative sequence is calculated. Obtaining a relatively high comparison similarity value between the sequences; if the similarity meets a clustering threshold value, adding the similarity into a category with the same representative sequence, otherwise, taking the similarity as a new representative sequence, and generating a new category; repeating the steps until all the sequences are processed; the number of the final representative sequences is the number of clustering categories, and the categories which are the same as the representative sequences are member sequences of each category.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for biological sequence analysis, mainly a heuristic sequence clustering method based on semi-global sequence alignment. Background Art

[0002] Biological research has made rapid progress in the past few decades. Especially driven by high-throughput sequencing technology, a large amount of biological sequence data has been generated. These data include genomic sequences, protein sequences, RNA sequences, etc., which contain important components of genetic information in organisms. With the continuous accumulation of these data, it is urgent for researchers to develop a method that can effectively process and analyze biological sequence data, and then reveal the functions, structures and evolutionary relationships in organisms. Among them, biological sequence clustering is an important analysis method, aiming to divide similar biological sequences into the same clustering cluster to reveal the similarities and correlations between them. At the same time, sequence clustering is also the basis for subsequent biological information mining. This clustering method helps to discover proteins in the same family, find common gene functions and study the evolutionary relationships between species. However, due to the complexity and diversity of biological sequences, traditional clustering methods face many challenges when dealing with biological sequence data, such as low clustering sensitivity, overestimation of the number of clusters, and other problems. Summary of the Invention

[0003] In order to overcome the deficiencies of existing methods, the present invention provides a method for biological sequence analysis, mainly a heuristic biological sequence clustering method based on semi-global sequence alignment.

[0004] The object of the present invention is to propose a heuristic sequence clustering method based on semi-global sequence alignment for massive biological sequences. Existing heuristic sequence clustering methods have problems such as low clustering sensitivity and overestimation of the number of clusters. This method has high clustering sensitivity and clustering quality, and can obtain fewer clusters, providing an effective technical guarantee for high-throughput sequencing data analysis.

[0005] To achieve the above object, the basic idea of the technical solution of the present invention is as follows: First, read all biological sequences, perform deduplication processing on them, and then sort them in descending order according to the sequence length; take the first sequence (i.e., the longest sequence) as the representative sequence of the first category, and then read the next sequence, and use the semi-global sequence alignment algorithm to calculate the similarity between the current sequence and the representative sequence. Semi-global sequence alignment can align each sequence to the most similar region in the representative sequence to maximize the similarity between the two and obtain a relatively high alignment similarity value between the sequences; if the similarity meets the threshold set for clustering, add it to the category with the same representative sequence, otherwise use it as a new representative sequence and generate a new category; repeat the above steps until all sequences are processed; the final number of representative sequences is the number of clustering categories, and the category with the same representative sequence is the member sequence of each category.

[0006] The sequence clustering method based on semi-global sequence alignment of the present invention includes the following steps:

[0007] Step 1: Sequence deduplication and sorting

[0008] First, perform deduplication processing on all sequences to be sorted, and return the data without duplicate sequences, which can reduce the time calculation complexity. Then, sort these sequences in descending order according to the length. In the case where the sequence lengths vary greatly, it is very important to use the longest sequence as the representative sequence for clustering, which can reduce the number of clusters and improve the accuracy of clustering;

[0009] Step 2: Obtain the similarity between sequences by using semi-global sequence alignment

[0010] First, take the first sequence (i.e., the longest sequence) as the representative sequence of the first clustering unit, and then read the next sequence, and use semi-global sequence alignment to calculate the similarity between the current sequence and the representative sequence; compared with global alignment and local alignment, semi-global sequence alignment can obtain the most similar alignment result between a shorter sequence and a longer sequence, and obtain a relatively high sequence similarity; the specific implementation steps are as follows:

[0011] 1) For the sequence q to be aligned and the representative sequence t, first determine the penalty scores for insertion, deletion, match, and substitution (or mismatch), and then create a scoring (score) matrix score of size [len(q)+1]×[len(t)+1], where len(q) and len(t) represent the lengths of sequences q and t respectively. Then, according to the insertion, deletion, match, and substitution situations, calculate the score of each matrix element by selecting the optimal score, and record the scoring path at the same time; the score calculation formula for each element of the scoring matrix is as follows:

[0012]

[0013] Among them, score(i, j) is the score value at the corresponding element position in the i-th row and j-th column of the score matrix score, γ is the penalty score for insertion and deletion errors (default value is -2), also known as the gap penalty score, and the δ function is the score function for judging the matching position, which is defined as:

[0014]

[0015] Formula (2) indicates that if two bases are the same, that is, a match, the score is α (default value 2), otherwise it is a substitution error and the score is β (default value -2); it can be observed from formula (1) that the score value calculated each time is the maximum value selected from the four cases of match, substitution, insertion, and deletion, that is, the score calculated at each step is the optimal alignment result of the current subsequence. When all the optimal alignment results are calculated, the complete alignment result of the two sequences can be obtained; at the same time, record the calculation path (backtracking direction, that is, the value-taking direction of the current maximum score value) of each element score value. Finally, starting from the position of the maximum value in the last row of the score matrix, find the final sequence alignment result through path backtracking. The formula for path backtracking is:

[0016]

[0017] where q align and t align respectively represent the alignment results (strings) of sequences q and t, '-' represents a gap, the symbol '+' represents string concatenation, and p(i, j) represents the backtracking direction of the element score(i, j) in the score matrix. Formula (3) means that starting from the maximum score in the last row of the matrix, the detailed alignment results of the two sequences are obtained through path backtracking. If the score path comes from the diagonal direction (represented by 0), then according to sequences q and t, one base is taken out in the reverse order (from back to front) and added to q align and t align . If the score path comes from above, extract the base at the corresponding position of sequence q and add it to q align , and add the gap '-' to t align ; if the path comes from the left, extract the base at the corresponding position of sequence q and add it to q align , and add the gap '-' to q align , until backtracking to i = 1, the semi-global alignment results q align and q align of sequences q and t can be obtained;

[0018] Step 3: Generate clustering units

[0019] According to Step 2, the similarity between the sequence to be clustered and the representative sequence can be obtained. If the clustering threshold set by the user is met, the sequence is classified into the category where the representative sequence is located; otherwise, the same operation is performed with the next representative sequence. If it is still not classified after comparing with all the representative sequences, it is used as a new representative sequence and a new category is generated. Then, repeating the above steps can complete the clustering operation of all sequences.

[0020] Preferably, in Step 2, by adopting a semi-global sequence alignment strategy, it is convenient to align the most similar parts between short sequences and long sequences, obtain a higher similarity threshold, overcome the influence of different sequence lengths, and truly reflect the similarity between sequences.

[0021] Preferably, in Step 3, by heuristically clustering sequences, all pairwise sequence alignments can be avoided, and only sequence alignments with representative sequences are performed, reducing the clustering computational complexity and improving the running efficiency of the entire method. The present invention has the following beneficial effects:

[0022] By using a semi-global sequence alignment strategy to perform sequence alignment between the sequence to be clustered and the representative sequence, the most similar sequence alignment result between the two can be obtained, obtaining a higher sequence similarity. Then, it is judged whether to classify according to the similarity threshold, improving the clustering quality and clustering sensitivity and reducing the number of clusters. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of a heuristic sequence clustering method based on semi-global sequence alignment;

[0024] Figure 2 is a schematic diagram of sequence clustering units generated in sorted and unsorted cases. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] On a server with an Intel Xeon Gold 5218R CPU @ 2.10GHz and 256GB of running memory, based on the Linux platform of CentOS 7.5 version, a set of common biological sequence data is selected for comparison and simulation experiments.

[0026] The present invention will be further described below in conjunction with the drawings and embodiments. Among them, Figure 1 is a flowchart of the heuristic sequence clustering method based on semi-global alignment of the present invention; Figure 1 (A) When the similarity between the sequence to be clustered and the representative sequence is greater than (or equal to) the given threshold, it is assigned to the corresponding clustering unit of the representative sequence. Figure 1 (B) When the similarity between the sequence to be clustered and all representative sequences is less than the given threshold, it is used as a new representative sequence and a new category is generated, and it is the distribution of matching k-mer positions. Figure 1(C) The finally generated clustering results, each cluster (dashed circle) is represented by a representative sequence (the central node of each cluster); Figure 2 It is a schematic diagram of the number of clustering units and their class members generated in the sorted and unsorted cases, which contains a total of 3 sequences of different lengths.

[0027] If there are 3 sequences s1, s2, s3, s4 to be clustered, where s4 is the same as s1 sequence, now cluster them and elaborate on the specific clustering implementation steps.

[0028] Step 1: Sequence deduplication and sorting

[0029] Read all sequences. By judgment and comparison, it is found that s4 is the same as s1 sequence, that is, a duplicate sequence. Then remove sequence s4. Finally, keep 3 sequences s1, s2, s3, and sort them according to the sequence length. The obtained order is: s2, s1, s3;

[0030] Step 2: Calculate sequence similarity using semi-global alignment

[0031] Take sequence s2 as the representative sequence of the first class, then read the next sequence s1. Through semi-global sequence alignment, the sequence similarity between s1 and s2 is 87.25% (for example), which is less than the set similarity threshold of 90% (for example). Then take sequence s1 as the representative sequence of the second class; then read the next sequence s3. First, perform semi-global sequence alignment with the first representative sequence s2, and the sequence similarity between s3 and s2 is 80.86% (for example), which is less than the set similarity threshold of 90%. Then perform semi-global sequence alignment with the second representative sequence s1, and the sequence similarity between s3 and s1 is 91.35% (for example), which is greater than the set similarity threshold of 90%, that is: it meets the threshold condition, so classify s3 into the class where the representative sequence s1 is located; since all sequences have been processed, the current clustering process ends.

[0032] Step 3: Calculate and generate clustering units

[0033] Finally, output the clustering units. Since s4 sequence is the same as s1 sequence, classify s4 into the class where s1 is located. The final clustering output is: a total of two classes. The first class is s2, and the second class is: s1, s3, and s4.

[0034] Table 1 shows the running results of the number of clusters for different methods on real biological sequence data. Table 2 shows the comparison results of the sensitivity of representative sequences. The number of clusters can reflect whether there is an overestimation problem in the number of clusters, and the sensitivity of representative sequences can reflect the quality of representative sequences. As can be seen from Table 1, the heuristic biological sequence clustering method based on semi-global sequence alignment has fewer clusters than other methods, effectively reducing the overestimation phenomenon of the number of clusters. As can be seen from Table 2, the heuristic biological sequence clustering method based on semi-global sequence alignment has a higher sensitivity of representative sequences, indicating that the quality of its representative sequences is better.

[0035] Table 1 Comparison results of the number of clusters of the heuristic sequence clustering method based on semi-global alignment and other methods

[0036]

[0037] Table 2 Comparison of the sensitivity of representative sequences of the heuristic sequence clustering method based on semi-global alignment and other methods

[0038]

[0039] The above results show that the heuristic sequence clustering method based on semi-global alignment can obtain fewer clusters, and the quality of its representative sequences is better than that of other methods, improving the clustering sensitivity. It is suitable for the clustering analysis of high-throughput biological sequence data and has great potential application value.

Claims

1. A heuristic biological sequence clustering method based on semi-global alignment, characterized in that, It includes the following steps: Step 1: Sequence deduplication and sorting First, perform deduplication on all sequences and return the data without duplicate sequences, which can reduce the time calculation complexity. Then, sort these sequences in descending order according to their lengths. In the case where the sequence lengths vary greatly, it is very important to use the longest sequence as the representative sequence for clustering, which can reduce the number of clusters and improve the accuracy of clustering. Step 2: Obtain sequence similarity using semi-global sequence alignment First, use the first sequence (i.e., the longest sequence) as the representative sequence of the first clustering unit, and then read the next sequence and calculate the similarity between the current sequence and the representative sequence using semi-global sequence alignment. Compared with global alignment and local alignment, semi-global sequence alignment can obtain the most similar alignment result between a shorter sequence and a longer sequence, resulting in a higher sequence similarity. The specific implementation steps are as follows: 1) For the sequence q to be aligned and the representative sequence t, first determine the penalty scores for insertion, deletion, match, and substitution (or mismatch). Then, create a scoring (score) matrix score of size [len(q)+1]×[len(t)+1], where len(q) and len(t) represent the lengths of sequences q and t respectively. Then, according to the insertion, deletion, match, and substitution situations, calculate the score of each matrix element by selecting the optimal score, and record the scoring path at the same time. The score calculation formula for each element of the scoring matrix is as follows: where score(i,j) is the score value at the corresponding element position in the i-th row and j-th column of the scoring matrix score, γ is the penalty score for insertion and deletion errors (default value is -2), also known as the gap penalty, and the δ function is the scoring function for judging the match position, defined as: Formula (2) means that if two bases are the same, i.e., a match, the score is α (default value 2), otherwise it is a substitution error and the score is β (default value -2). It can be observed from formula (1) that the score value calculated each time is the maximum value selected from the four cases of match, substitution, insertion, and deletion, that is, the score calculated at each step is the optimal alignment result of the current subsequence. After all the optimal alignment results are calculated, the complete alignment result of the two sequences can be obtained. At the same time, record the calculation path (backtracking direction, that is, the value-taking direction of the current maximum score value) of each element score value. Finally, start from the position of the maximum value in the last row of the scoring matrix and find the final sequence alignment result through path backtracking. The formula for path backtracking is: where q align and t align respectively represent the alignment results (strings) of sequences q and t, '-' represents a gap, the symbol '+' represents string concatenation, and p(i, j) represents the backtracking direction of the element score(i, j) in the scoring matrix; formula (3) means that starting from the maximum score in the last row of the matrix, the detailed alignment results of the two sequences are obtained through path backtracking. If the scoring path comes from the diagonal direction (represented by 0), then according to sequences q and t, one base is taken out in the reverse direction (reverse order) from the back to the front and added to q align and t align . If the scoring path comes from above, the base at the corresponding position of sequence q is extracted and added to q align , and the gap '-' is added to t align ; if the path comes from the left, the base at the corresponding position of sequence q is extracted and added to q align , and the gap '-' is added to q align . Until backtracking to i = 1, the semi-global alignment results of sequences q and t, q align and q align ; can be obtained. Step 3: Generate clustering units According to Step 2, the similarity between the sequence to be clustered and the representative sequence can be obtained. If it meets the clustering threshold set by the user, the sequence is classified into the category where the representative sequence is located; otherwise, perform the same operation with the next representative sequence. If it still cannot be classified after comparing with all the representative sequences, it serves as a new representative sequence and a new category is generated. Then, repeat the above steps to complete the clustering operation of all sequences.

2. The heuristic biological sequence clustering method based on semi-global alignment according to claim 1, wherein: In step 2, by adopting the semi-global sequence alignment strategy, it is convenient to align the most similar parts between the short sequence and the long sequence, obtain a relatively high similarity threshold, overcome the influence of different sequence lengths, and truly reflect the similarity between sequences.

3. The heuristic biological sequence clustering method based on semi-global alignment according to claim 1, characterized in that: In step 3, by clustering sequences heuristically, all pairwise sequence alignments can be avoided, and only sequence alignments with representative sequences are performed, reducing the clustering computational complexity and improving the running efficiency of the entire method.

Citation Information

Cited By

  • Sequence recognition method, sequence recognition model training method, sequence recognition model training system, sequence recognition equipment and storage medium

    CN121884946A

  • Sequence recognition methods and their model training methods, systems, devices, and storage media

    CN121884946B