A pan-genome long read alignment method based on weighted syncmers

By using frequency-division and weighted sampling based on weighted syncmer and back-calculation of topological information, the problem of low alignment efficiency of complex pan-genome maps is solved, and efficient and accurate long-read alignment is achieved, which can meet the alignment needs of highly repetitive and mutated regions in complex genomes.

CN121528309BActive Publication Date: 2026-03-20YANTAI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610055671.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-16
Publication Date
2026-03-20
Estimated Expiration
2046-01-16

AI Technical Summary

Technical Problem

Existing genome long read alignment methods are inefficient when faced with complex pan-genome maps, and cannot effectively handle highly repetitive regions and gene mutations. Traditional methods have high sampling redundancy in highly repetitive regions and are not sensitive enough to mutations.

Method used

A weighted syncmer-based approach is used to perform frequency-divided and weighted sampling on the reference pan-genome graph, construct a hash index, obtain the best linear chain through backtracking calculation using topological information, improve matching efficiency by combining the xxHash3 hash function, filter high-frequency redundant k-mers, retain high-specificity k-mers, and generate a candidate anchor array and graph chain set.

Benefits of technology

It significantly improves the efficiency and accuracy of pan-genome sequence alignment, effectively handles complex topological structures, reduces false matches, enhances the sensitivity and accuracy of alignment results, and adapts to gene mutations that are widespread in complex genomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121528309B_ABST
    Figure CN121528309B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of gene sequence alignment, in particular to a pan-genome long read alignment method based on weighted syncmer; the method of the present application firstly processes the segment of the reference pan-genome graph based on the weighted syncmer sampling of frequency division weight reduction, filters high-frequency redundant k-mers, retains high-specificity k-mers, then constructs a hash index to realize the rapid matching of the k-mers of the query sequence, and generates a candidate anchor point array, which greatly reduces the false matching and improves the alignment efficiency; then, based on the topological information of the pan-genome graph, the best linear chain is selected by backtracking calculation and further expanded into the best graph chain, and the hierarchical screening strategy reduces the search complexity of the complex graph structure; finally, the candidate pan-genome graph chain set is integrated to output the alignment result, which guarantees the high precision and integrity of the result, and realizes the efficient and accurate alignment of the complex topological structure pan-genome graph.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of gene sequence alignment, and particularly relates to a pan-genome long read alignment method based on weighted syncmer. BACKGROUND

[0002] As a key breakthrough of new generation sequencing, long read sequencing technology can effectively identify complex genomic structural variations, including large-scale inversion, deletion, translocation and high repetitive and variable regions, which cannot be detected by traditional short read technology. With the progress of sequencing technology, the assembly quality of individual genomes has approached the level of reference genomes, and a large number of high-quality human genome assemblies have been accumulated in GenBank, which further promotes the development of sequence-to-graph alignment methods. As a basic operation in bioinformatics, genomic sequence alignment is often the first step in the genome analysis process, and its sensitivity and accuracy directly determine the reliability of subsequent research.

[0003] Among existing tools, Minigraph, as a classic pan-genome alignment method, uses a minimizer-based subsequence sampling strategy combined with a "seed-chaining-expansion" process to improve the sensitivity and accuracy of alignment. However, the minimizer method has context dependence and is easily disturbed by sequence mutations, with high redundancy in high repetitive regions and often ignoring high-frequency k-mers, resulting in decreased sensitivity and accuracy. In contrast, syncmer, as a context-independent sampling method, has better robustness to mutations, but has the problem of over-sampling repetitive subsequences, affecting the alignment efficiency. Existing research has proposed a genome long sequence alignment method based on weighted syncmer, but this method mainly faces the sequence-to-sequence alignment scenario. When it comes to pan-genome graphs composed of multiple sequences, this method needs to align the sequencing sequence to each linear sequence in the graph one by one, resulting in a large amount of repeated calculations and low efficiency, which cannot be directly applied to the alignment of complex pan-genome graphs. SUMMARY

[0004] The purpose of the present application is to provide a pan-genome long read alignment method based on weighted syncmer.

[0005] The technical scheme of the present application is as follows:

[0006] A pan-genome long read alignment method based on weighted syncmer includes the following operations:

[0007] S1, all segments in the reference pan-genome graph are subjected to weighted syncmer sampling processing based on frequency division weight reduction to obtain a reference weighted syncmer group, a hash index is constructed to obtain a reference pan-genome index; a query weighted syncmer group of a query sequence is obtained, each k-mer is taken as a key value, and matching is performed in the reference pan-genome index to obtain each k-mer that can be successfully matched, the end position of the corresponding segment in the reference pan-genome graph, the end position on the query sequence and the sequence length, respectively, constitute an anchor triplet corresponding to each k-mer, and the anchor triplets are sorted in ascending order at the end position of the corresponding segment to form a candidate anchor array;

[0008] S2, each anchor in the candidate anchor array is traversed, and based on the topological information in the reference pan-genome graph, a linear chain corresponding to the maximum score ending with each anchor is calculated and obtained as a corresponding best linear chain to form a candidate linear chain array; each best linear chain in the candidate linear chain array is traversed to obtain a best graph chain ending with each best linear chain to form a candidate pan-genome graph chain set;

[0009] S3, based on the candidate pan-genome graph chain set, a long read sequence alignment result is obtained.

[0010] The operation of the weighted syncmer sampling processing based on frequency division weight reduction in S1 is as follows: within each segment, a k-mer is selected at each position, an s-mer is selected at each position in the k-mer, the s-mer is converted into a positive number less than 1 using an xxHash3 hash function, and different weights are given to s-mers in different repeat frequency intervals according to the repeat frequency interval of the s-mer; if the s-mer with the smallest weight in the current k-mer appears at a pre-set position, the current k-mer is considered to be a weighted syncmer; all weighted syncmers are counted to obtain a reference weighted syncmer group.

[0011] In the operation of frequency division weight reduction, the greater the repeat frequency of the s-mer in the repeat frequency interval, the greater the weight reduction degree.

[0012] The method for obtaining the repeat frequency interval is as follows: the s-mer sequences with a frequency of 0.02% in the reference pan-genome graph are sorted in descending order of frequency to obtain a sorted s-mer set; the sorted s-mer set is divided into four different repeat frequency intervals, namely, a super-high frequency interval, a high frequency interval, a secondary high frequency interval and a medium frequency interval, and the repeat frequencies decrease in turn, and the frequency interval lengths are in the ratio of 0.02%:(0.02%~30%):(30%~70%):(70%~100%).

[0013] S2 is the maximum score of the linear chain ending with the anchor point i S2 is the maximum score of the linear chain ending with the anchor point The calculation formula of S2 is as follows:

[0014]

[0015]

[0016] S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point i S2 is the score of the linear chain ending with the anchor point j S2 is the score of the linear chain ending with the anchor point S2 is the score of the linear chain ending with the anchor point

[0017] The operation of S3 is specifically: converting the matching interval coordinates in the segment in the reference pan-genome graph into overall reference coordinates to construct a reference path unit; traversing each graph chain in the candidate pan-genome graph chain set, connecting the path units corresponding to each anchor point in the graph chain in the reference path unit in order, merging if adjacent, and summarizing the query sequence name and the total length of the query sequence to output to obtain a long read sequence alignment result.

[0018] A pan-genome long read alignment system based on weighted syncmer is used to implement the pan-genome long read alignment method based on weighted syncmer, and comprises:

[0019] ​​​The candidate anchor array generation module is configured to perform weighted syncmer sampling processing on all segments in the reference pan-genome graph based on frequency division weighting, to obtain a reference weighted syncmer group, to construct a hash index, and to obtain a reference pan-genome index; obtain a query weighted syncmer group of a query sequence, take each k-mer as a key value, and perform matching in the reference pan-genome index to obtain each k-mer that can be successfully matched, the end position of the corresponding segment in the reference pan-genome graph, the end position on the query sequence, and the sequence length, which respectively constitute an anchor triplet corresponding to each k-mer, and the anchor triplets are sorted in ascending order according to the end position of the corresponding segment to form a candidate anchor array;

[0020] The candidate pan-genome graph chain set generation module is configured to traverse each anchor point in the candidate anchor array, to obtain a maximum score corresponding to a linear chain ending with each anchor point based on topological information in the reference pan-genome graph, to obtain a corresponding best linear chain, to form a candidate linear chain array, and to traverse each best linear chain in the candidate linear chain array to obtain a best graph chain ending with each best linear chain to form a candidate pan-genome graph chain set.

[0021] The sequence alignment result generation module is configured to obtain a long read sequence alignment result based on the candidate pan-genome graph chain set.

[0022] A pan-genome long read alignment device based on weighted syncmer, comprising a processor and a memory, wherein the processor implements the pan-genome long read alignment method based on weighted syncmer described above when executing the computer program stored in the memory.

[0023] A computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the pan-genome long read alignment method based on weighted syncmer described above.

[0024] The beneficial effects of the present application are as follows:

[0025] The application provides a long read alignment method based on a weighted syncmer, which comprises the following steps: firstly, a segment of a reference pan-genome graph is processed based on a weighted syncmer sampling with frequency division and weight reduction, high-frequency redundant k-mers are filtered, high-specificity k-mers are reserved, a hash index is constructed to realize fast matching of k-mers of a query sequence, and a candidate anchor array is generated, so that the number of false matches is greatly reduced and the alignment efficiency is improved; secondly, based on the topological information of the pan-genome graph, the best linear chain is screened out through backtracking calculation, and is further expanded into the best graph chain, and the search complexity of the complex graph structure is reduced through a hierarchical screening strategy; and finally, the candidate pan-genome graph chain set is integrated to output an alignment result, so that the high precision and integrity of the result are ensured, and efficient and accurate alignment of the complex topological structure pan-genome graph is realized.

[0026] The application provides a long read alignment method based on a weighted syncmer, which comprises the following steps: firstly, a segment of a reference pan-genome graph is processed based on a weighted syncmer sampling with frequency division and weight reduction, high-frequency redundant k-mers are filtered, high-specificity k-mers are reserved, a hash index is constructed to realize fast matching of k-mers of a query sequence, and a candidate anchor array is generated, so that the number of false matches is greatly reduced and the alignment efficiency is improved; secondly, based on the topological information of the pan-genome graph, the best linear chain is screened out through backtracking calculation, and is further expanded into the best graph chain, and the search complexity of the complex graph structure is reduced through a hierarchical screening strategy; and finally, the candidate pan-genome graph chain set is integrated to output an alignment result, so that the high precision and integrity of the result are ensured, and efficient and accurate alignment of the complex topological structure pan-genome graph is realized.

[0027] The application provides a long read alignment method based on a weighted syncmer, which comprises the following steps: firstly, a segment of a reference pan-genome graph is processed based on a weighted syncmer sampling with frequency division and weight reduction, high-frequency redundant k-mers are filtered, high-specificity k-mers are reserved, a hash index is constructed to realize fast matching of k-mers of a query sequence, and a candidate anchor array is generated, so that the number of false matches is greatly reduced and the alignment efficiency is improved; secondly, based on the topological information of the pan-genome graph, the best linear chain is screened out through backtracking calculation, and is further expanded into the best graph chain, and the search complexity of the complex graph structure is reduced through a hierarchical screening strategy; and finally, the candidate pan-genome graph chain set is integrated to output an alignment result, so that the high precision and integrity of the result are ensured, and efficient and accurate alignment of the complex topological structure pan-genome graph is realized.

[0028] The application provides a long read alignment method based on a weighted syncmer, which comprises the following steps: firstly, a segment of a reference pan-genome graph is processed based on a weighted syncmer sampling with frequency division and weight reduction, high-frequency redundant k-mers are filtered, high-specificity k-mers are reserved, a hash index is constructed to realize fast matching of k-mers of a query sequence, and a candidate anchor array is generated, so that the number of false matches is greatly reduced and the alignment efficiency is improved; secondly, based on the topological information of the pan-genome graph, the best linear chain is screened out through backtracking calculation, and is further expanded into the best graph chain, and the search complexity of the complex graph structure is reduced through a hierarchical screening strategy; and finally, the candidate pan-genome graph chain set is integrated to output an alignment result, so that the high precision and integrity of the result are ensured, and efficient and accurate alignment of the complex topological structure pan-genome graph is realized. BRIEF DESCRIPTION OF DRAWINGS

[0029] The schemes and advantages of the present application will become clear to those skilled in the art through reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not considered as limiting the present application.

[0030] In the drawings:

[0031] Figure 1 For the embodiment, the flowchart of the method of the embodiment is shown in the figure;

[0032] Figure 2 For the embodiment, the partial topological structure diagram of the reference pan-genome graph is shown in the figure;

[0033] Figure 3 Figure 6 is a subsequence sampling number plot for an example simulation comparison experiment based on the human X chromosome pan-genome map;

[0034] Figure 4 Figure 7 is a subsequence sampling repeat rate plot for an example simulation comparison experiment based on the human X chromosome pan-genome map. DETAILED DESCRIPTION

[0035] In order to make the objects, technical solutions and advantages of the exemplary embodiments of the present application clearer, the technical solutions in the exemplary embodiments of the present application will be described clearly and completely below with reference to the drawings. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application.

[0036] Based on the exemplary embodiments shown in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work under the premise that the present application falls within the scope of protection of the present application. In addition, although the disclosure in the present application is introduced according to one or more examples, it should be understood that each aspect of these disclosures can also constitute a complete technical solution independently.

[0037] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclusive inclusion, for example, a product or device containing a series of components is not necessarily limited to those components clearly listed, but includes other components not clearly listed or inherent to these products or devices.

[0038] The term "module" used in the present application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic or combination of hardware or / and software code capable of performing functions related to the element.

[0039] Embodiment 1

[0040] The present embodiment provides a pan-genome long read alignment method based on weighted syncmer, see Figure 1 , comprising the following operations:

[0041] S1, all segments in the reference pan-genome graph are subjected to weighted syncmer sampling processing based on frequency-based weight reduction to obtain a reference weighted syncmer group, a hash index is constructed to obtain a reference pan-genome index; a query weighted syncmer group of a query sequence is obtained, each k-mer is taken as a key value, matching is performed in the reference pan-genome index, each k-mer that can be successfully matched is obtained, the end position of the corresponding segment in the reference pan-genome graph, the end position on the query sequence and the sequence length respectively constitute an anchor triplet corresponding to each k-mer, and the anchor triplets are sorted in ascending order according to the end position of the corresponding segment to form a candidate anchor array;

[0042] S2, each anchor in the candidate anchor array is traversed, and based on the topological information in the reference pan-genome graph, a linear chain corresponding to the maximum score ending with each anchor is calculated and obtained as a corresponding best linear chain to form a candidate linear chain array; each best linear chain in the candidate linear chain array is traversed to obtain a best graph chain ending with each best linear chain to form a candidate pan-genome graph chain set;

[0043] S3, based on the candidate pan-genome graph chain set, a long read sequence alignment result is obtained.

[0044] The specific steps are as follows.

[0045] S1, all segments in the reference pan-genome graph are subjected to weighted syncmer sampling processing based on frequency-based weight reduction to obtain a reference weighted syncmer group, a hash index is constructed to obtain a reference pan-genome index; a query weighted syncmer group of a query sequence is obtained, each k-mer is taken as a key value, matching is performed in the reference pan-genome index, each k-mer that can be successfully matched is obtained, the end position of the corresponding segment in the reference pan-genome graph, the end position on the query sequence and the sequence length respectively constitute an anchor triplet corresponding to each k-mer, and the anchor triplets are sorted in ascending order according to the end position of the corresponding segment to form a candidate anchor array.

[0046] Firstly, all segments in the reference pan-genome graph are traversed to obtain a reference weighted syncmer set by weighted syncmer sampling based on frequency-based weight reduction. Specifically, the Gfa format reference pan-genome graph Ref_graph is read from the hard disk, all vertices (i.e. segments) in the reference pan-genome graph are traversed for sampling, k-mers are selected at each position within each segment, s-mers are selected at each position in the k-mers, the s-mers are converted into positive numbers less than 1 using the xxHash3 hash function, and different weights are assigned to the s-mers according to the repeat frequency interval in which the s-mers are located. If the s-mer with the smallest weight in the current k-mer appears at a pre-set position, the current k-mer is considered to be a weighted syncmer. All weighted syncmers are counted to obtain the reference weighted syncmer set. The xxHash3 is specially optimized for short key values, has a lower collision rate and better distribution characteristics, thereby ensuring more uniform weight distribution and improving sampling quality.

[0047] In the frequency-based weight reduction operation, the greater the repeat frequency of the s-mer in the repeat frequency interval, the greater the weight reduction degree. Differentiated weight reduction according to the frequency of the s-mer in the genome can effectively improve the stability and diversity of the sampling.

[0048] The method for obtaining the repeat frequency interval is: through the meryl tool, the sequence and frequency of the highly repetitive s-mer with the appearance frequency in the top 0.02% (including 0.02%) in the reference pan-genome graph Ref_graph are saved in the repetitive.txt file, through the self-defined script, the s-mer in the repetitive.txt file is sorted in descending order of frequency and saved as a repetitive_sorted.txt file, the sequence of the s-mer with the appearance frequency in the top 0.02% in the reference pan-genome graph is sorted in descending order of frequency, and the sorted s-mer set is obtained and saved as a repetitive_sorted.txt file; the sorted s-mer set in the repetitive_sorted.txt file is divided into four different repeat frequency intervals according to the frequency of the s-mer in the repetitive_sorted.txt file, which are: super-high frequency interval, high frequency interval, secondary high frequency interval, and medium frequency interval, the repeat frequency decreases in turn, and the frequency interval length ratio is 0.02%: (0.02%~30%): (30%~70%): (70%~100%), that is, the super-high frequency interval with the repeat frequency of the top 0.02% (<0.02%), the high frequency interval with the repeat frequency of the top 0.02%~30% ([0.02%, 30%]), the secondary high frequency interval with the repeat frequency of the top 30%~70% ((30%, 70%)), and the medium frequency interval with the repeat frequency of the top 70%~100% ([70%, 100%]); the information of the four intervals is read from the repetitive_interval.txt, and the s-mer in the repetitive_sorted.txt is read in order, and finally saved as four sets: R 1 、 R 2 、 R 3 、 R 4 , respectively corresponding to the super-high frequency interval, the high frequency interval, the secondary high frequency interval, and the medium frequency interval.

[0049] The above process can refer to the sampling logic in Table 1, wherein the getOrder function is a weighting function combined with the frequency division and weight reduction strategy, which efficiently converts the s-mer into a positive number less than 1 through the xxHash3 hash function, and identifies which high frequency interval the highly repetitive s-mer belongs to, and accurately reduces the weight. For the super-high frequency interval, the s-mer hash value is reduced by 16 times, for the high frequency interval, the weight will be reduced by 8 times, which effectively reduces the possibility of different degrees of high repeat s-mer interfering with the stability of sampling.

[0050] Table 1: segments internal sub-sequence sampling algorithm

[0051]

[0052] Then, based on the reference weighted syncmer group, a hash index is constructed to obtain a reference pan-genome index. Specifically, the sampling results in the array are stored in the weighted syncmer hash index Hash, the key value of which is the base sequence of the weighted syncmer, and the value value is the position set of the corresponding weighted syncmer in the reference pan-genome graph . After the reference pan-genome index is constructed, it includes the weighted syncmer hash index Hash and the high-repetition s-mer set .

[0053] Next, the Fasta format query sequence Read_seq is read from the hard disk, and the algorithm logic in Table 1 is used, i.e. the method of obtaining the query weighted syncmer group of the query sequence is used (the query sequence is regarded as an array with a length of 1, and the only element is , i.e. , and the query weighted syncmer group data is stored in the array , which saves the sequence, position information (segment to which it belongs and offset position) of all k-mers obtained by sampling in , and each k-mer of the query weighted syncmer group is used as the key value to match in the reference pan-genome index.

[0054] Each k-mer that can be successfully matched is obtained, and the end position of the corresponding segment in the reference pan-genome graph x , the end position on the query sequence y and the sequence length w (also denoted as k value) form the anchor triplet corresponding to each k-mer, which is sorted in ascending order according to the end position of the segment to form the candidate anchor array, denoted as the array , and the technical logic is shown in Table 2.

[0055] Table 2: Candidate anchor array acquisition algorithm

[0056]

[0057] S2, traverse each anchor point in the candidate anchor point array, based on the topological information in the reference pan-genome graph, backtrack to calculate the linear chain corresponding to the maximum score ending with each anchor point as the corresponding best linear chain, forming a candidate linear chain array; traverse each best linear chain in the candidate linear chain array, obtain the best graph chain ending with each best linear chain, forming a candidate pan-genome graph chain set.

[0058] First, traverse each anchor point in the candidate anchor point array Each anchor point is a triple , indicating the matching of the query sequence interval and the interval of a segment in the reference pan-genome graph During the traversal process, based on the topological information in the reference pan-genome graph (see the connection relationship between the base information corresponding to the anchor points in the reference pan-genome graph in Figure 2 Each box in Figure 2 can be regarded as an anchor point), obtain the linear chain ending with the traversed anchor point, each linear chain is a set of anchor points sorted by the end position x value of the segment , where , i.e. the end position of the previous anchor point in the linear chain is less than the end position of the current anchor point in the segment, backtrack to calculate the linear chain corresponding to the maximum score ending with each anchor point, as the corresponding best linear chain, forming a candidate linear chain array .

[0059] The calculation formula of the maximum score of the linear chain ending with the anchor point i is as follows:

[0060] ,

[0061] ,

[0062] is the score of the linear chain ending with the anchor point j , the end position of the anchor point j is less than the end position of the anchor point i , is the number of matching bases of the anchor point i and the anchor point j , , , are the matching base information of the anchor point i and the anchor point​j The end position in the query sequence, , Anchor points i With anchor point j At the end of the segment, anchor point i The sequence length, Anchor point i With anchor point j Open shot penalty, anchor point i With anchor point j The length of the empty space, when hour, ;when hour, ; anchor point i With anchor point j The minimum distance, in graph construction mode. , , These are the first coefficient and the second coefficient, respectively. , , =0.01, backtracking calculation can greatly improve computational efficiency.

[0063] Then, each best linear chain in the candidate linear chain array is traversed to obtain the best graph chain ending with each best linear chain, thus forming a candidate pan-genome graph chain set. The method for obtaining the best linear chain is the same as that for obtaining the linear chain corresponding to the maximum score ending at each anchor point, except that the best linear chain is treated as a larger anchor point (e.g., Figure 2 Two connected boxes can be considered as a single linear chain. This information, combined with topological information from the reference pan-genome map, is used to optimize the connections between these linear chains. (Array) Each graph chain in the graph is an ordered set of linear chains, such as... It can also be viewed as a sequence of anchor points (sorted in ascending order by the y-value of the end position on the query sequence), with each anchor point consisting of a quadruple. constitute( , , , (representing the segment, end position, end position on the query sequence, and sequence length of the m-th graph chain, respectively), meaning that the m-th graph chain... m Intervals on each segment interval of the query sequence Matching is essentially an extension of linear chains.

[0064] S3, obtaining the long read sequence alignment result based on the candidate pan-genome graph chain set. Specifically: according to the reference pan-genome graph, the SN and S0 tags in the Ref_graph file are used to convert the segment internal matching interval coordinates into overall reference coordinates to construct a reference path unit; for traversing the candidate pan-genome graph chain set G[], each graph chain is processed in order, the path unit corresponding to each anchor point in the graph chain in the reference path unit is connected in order, if adjacent, merge, and summarize the query sequence name and total length of the query sequence, and output to obtain the long read sequence alignment result.

[0065] To verify the effect of the method of the embodiment, the following experiments were performed.

[0066] Experiment 1: Simulation data comparison experiment based on human X chromosome pan-genome graph.

[0067] The human X chromosome has rich repetitive regions, and the centromere region is a highly repetitive region. After constructing several different human X chromosomes into a pan-genome graph, the gene mutations, especially the repetitive regions, are more abundant, which further puts higher requirements on the sensitivity and accuracy of the alignment method. In the experiment, the method of the embodiment was compared with the existing advanced technology Minigraph to verify the advantages of the alignment result of the method. The experiment was run on a T640 server, the processor of which was Intel Xeon Gold 6226r (main frequency 2.9 GHZ, 16 cores), the memory was 128 GB, and the operating system was Ubuntu 20.04.

[0068] The specific experimental steps of the experiment are as follows.

[0069] (1) The first data in the data set of Table 3, i.e. the X chromosome of CHM13, was simulated at a depth of 10 times using the simulation sequence tool PBSIM under the condition of default parameters, and then the.fastq format file generated was used as the query sequence input file for alignment. All the data in the data set of Table 3 are from human haplotypes or haplotype folding assemblies.

[0070] Table 3: Data set of experiment 1

[0071]

[0072] (2) Using Minigraph, construct a human X chromosome pan-genome map of all X chromosomes in the above dataset. Align the query sequence with the human X chromosome pan-genome map using both Minigraph and the method of this embodiment. If the position mapped to a read overlaps with the position of the answer by 80% or more, the read is considered to have been correctly aligned. If 80% of the reads in the simulated sequence are aligned to the pan-genome map, the sensitivity of the alignment result is considered to be 80%. The alignment results are shown in Table 4.

[0073] Table 4: Comparison results of Minigraph and the method of this embodiment in Experiment 1

[0074]

[0075] (3) During subsequence sampling in index construction, each megabase read is considered a super window. Output the number of subsequence samples and their repetition rate within each super window. If a super window samples... The number of subsequences and the total number of subsequence types appeared. Then the repetition rate of sampling within this super window is... The number of subsequence samples and their repetition rate are compared between the Minigraph and the method of this embodiment, as shown below. Figure 3 , Figure 4 As shown, the method of this embodiment achieves stable subsequence sampling in regions where the number of Minigraph samples decreases sharply. In most regions, the repetition rate of subsequence sampling is significantly lower than that of the latter, proving that the method of this embodiment can achieve more diverse sampling.

[0076] Experiment 2: Comparison of real data based on the human X chromosome pangenome map.

[0077] In the experiment, the method of this embodiment was compared with the existing advanced technology Minigraph to verify the advantages of the comparison results of this method. The experiment was run on a T640 server, which has an Intel Xeon Gold 6226r processor (2.9GHz, 16 cores), 128GB of memory, and Ubuntu 20.04 operating system.

[0078] The specific experimental steps are as follows.

[0079] (1) Use Minigraph to construct a human X chromosome pan-genome map of all X chromosomes except HG00099 in the datasets in Table 5 below. The datasets in Table 5 are from human haplotypes or haplotype fold assemblies, where the human individual HG00099 is the original sequencing data (SRR741411).

[0080] Table 5: Dataset of Experiment 2

[0081]

[0082] (2) Using HG00099 as the query sequence, the query sequence is aligned with the human X chromosome pan-genome map constructed in the previous step using Minigraph and the method of the present embodiment respectively, and the alignment results are shown in Table 6. Since there is no answer position in Experiment 2, only the number of reads aligned in the query sequence is evaluated. As can be seen from Table 6, the present embodiment obtains more reads aligned and has higher sensitivity.

[0083] Table 6: Alignment results of Minigraph and the method of the present embodiment in Experiment 2

[0084]

[0085] Experiment 3: Simulation data comparison experiment based on chimpanzee X chromosome pan-genome map.

[0086] The X chromosome of chimpanzee has a similar structure to that of human, and also has rich repetitive regions and centromere regions. In order to further prove the applicability of the method of the present embodiment to closely related species, an experiment based on chimpanzee X chromosome pan-genome map is designed. In the experiment, the method of the present embodiment is compared with the existing advanced technology Minigraph to verify the advantages of the alignment results of the present method. The experiment is run on a T640 server, the processor of which is Intel Xeon Gold 6226R (main frequency 2.9 GHz, 16 cores), the memory of which is 128 GB, and the operating system of which is Ubuntu 20.04.

[0087] The specific experimental steps are as follows.

[0088] (1) Using Minigraph, all X chromosomes in the following Table 7 dataset are constructed into a chimpanzee X chromosome pan-genome map. The last piece of human data comes from CHM13.

[0089] Table 7: Dataset of Experiment 3

[0090]

[0091] (2) Using the simulation sequence tool PBSIM under the condition of default parameters, the last data in the data set of Table 5, i.e., the human X chromosome of CHM13, is simulated at a depth of 10 times, and then the fastq format file generated is used as the input file of the query sequence for alignment. The query sequence is aligned with the pan-genome map of the great ape X chromosome constructed in the previous step using Minigraph and the method of the present embodiment, respectively. If the position to which a read is mapped coincides with the answer position by 80% or more, it is considered that the read is correctly aligned. If 80% of the reads in the simulation sequence are aligned to the pan-genome map, the sensitivity of the alignment result is considered to be 80%. The alignment results are shown in Table 8 below. It can be seen from the table that when studying closely related species, the method of the present embodiment can still achieve higher sensitivity and accuracy in alignment.

[0092] Table 8: Alignment results of Experiment 3 of Minigraph and the method of the present embodiment

[0093]

[0094] Experiment 4: Real data comparison experiment based on the pan-genome map of the great ape X chromosome.

[0095] In the experiment, the method of the present embodiment is compared with the existing advanced technology Minigraph to verify the advantages of the alignment results of the present method. The experiment is run on a T640 server, the processor of which is Intel Xeon Gold 6226R (2.9 GHz, 16 cores), the memory is 128 GB, and the operating system is Ubuntu 20.04. The specific experimental steps are as follows.

[0096] (1) The original sequencing data (from chimpanzee Clinton, produced by Oxford Nanopore sequencing technology at a coverage rate of 99) of SRR24707269 is downloaded using SRAToolkit and used as the query sequence.

[0097] (2) The query sequence is aligned with the pan-genome map of the great ape X chromosome constructed in Example 3 using Minigraph and the method of the present embodiment, respectively. Since there is no answer position in the present experiment, only the number of reads aligned is evaluated. The alignment results are shown in Table 9 below, which shows that the method of the present embodiment still completes more alignment and has higher sensitivity.

[0098] Table 9: Alignment results of Experiment 4 of Minigraph and the method of the present embodiment

[0099]

[0100] From the above experiments, it can be seen that the method of the embodiment can effectively improve the stability and diversity of subsequence sampling, and further improve the accuracy and sensitivity of the alignment result.

[0101] The embodiment also provides a weighted syncmer-based pan-genome long read alignment system for implementing the weighted syncmer-based pan-genome long read alignment method.

[0102] The candidate anchor array generation module is configured to perform weighted syncmer sampling processing on all segments in the reference pan-genome graph based on frequency division and weight reduction to obtain a reference weighted syncmer group, construct a hash index to obtain a reference pan-genome index, obtain a query weighted syncmer group of a query sequence, match each k-mer as a key value in the reference pan-genome index, obtain each k-mer that can be successfully matched, and obtain an ending position of a corresponding segment in the reference pan-genome graph, an ending position on the query sequence, and a sequence length, which respectively constitute an anchor triple corresponding to each k-mer, and form a candidate anchor array in ascending order of the ending position of the corresponding segment.

[0103] The candidate pan-genome graph chain set generation module is configured to traverse each anchor point in the candidate anchor array, backtrack and calculate a linear chain corresponding to a maximum score ending with each anchor point based on topological information in the reference pan-genome graph, as a corresponding best linear chain, form a candidate linear chain array, and traverse each best linear chain in the candidate linear chain array to obtain a best graph chain ending with each best linear chain, and form a candidate pan-genome graph chain set.

[0104] The sequence alignment result generation module is configured to obtain a long read sequence alignment result based on the candidate pan-genome graph chain set.

[0105] The embodiment also provides a weighted syncmer-based pan-genome long read alignment device, including a processor and a memory, wherein the processor implements the weighted syncmer-based pan-genome long read alignment method when executing a computer program stored in the memory.

[0106] The embodiment also provides a computer readable storage medium for storing a computer program, wherein the computer program is executed by a processor to implement the weighted syncmer-based pan-genome long read alignment method.

[0107] This embodiment provides a pan-genome long read alignment method based on weighted syncmers. First, weighted syncmer sampling based on frequency division and weight reduction processes the segments of the reference pan-genome map, filtering out high-frequency redundant k-mers and retaining high-specificity k-mers. Then, a hash index is constructed to achieve fast matching of query sequence k-mers and generate a candidate anchor array, significantly reducing false matches and improving alignment efficiency. Next, based on the topological information of the pan-genome map, the best linear chain is screened through backtracking calculation and further expanded into the best graph chain. The hierarchical screening strategy reduces the search complexity of complex graph structures. Finally, the alignment results are output by integrating the candidate pan-genome graph chain set, ensuring high accuracy and completeness of the results, and enabling efficient and accurate alignment of pan-genome maps with complex topological structures.

[0108] This embodiment provides a pan-genome long read alignment method based on weighted syncmers. It linearizes the complex graph structure, treats the pan-genome graph as a single sequence composed of different segments, and applies a weighted syncmer-based sampling method to construct a pan-genome graph index, which greatly improves the efficiency of pan-genome sequence alignment.

[0109] This embodiment provides a pan-genome long read alignment method based on weighted syncmers. By combining a frequency-based weighting strategy, it reduces the weights according to the repetition frequency of s-mers, accurately suppressing the interference of repetitive regions with different repetition levels on sampling. This improves the stability and sampling diversity of traditional weighted sampling strategies, enabling them to better cope with the complex and widespread gene mutations in the pan-genome and improve the sensitivity and accuracy of long read alignment results.

[0110] This embodiment provides a pan-genome long read alignment method based on weighted syncmer. The weighted default murmurhash hash function is replaced with the higher-performance xxHash3 hash function, which is specifically optimized for small key values, thereby improving the efficiency of sampling and alignment.

[0111] While exemplary embodiments of the invention have been described herein, many other variations or modifications conforming to the principles of the invention can be directly determined or derived from the disclosure of this invention without departing from its spirit and scope. Therefore, the scope of the invention should be understood and recognized to cover all such other variations or modifications.

Claims

1. A pan-genome long read alignment method based on weighted syncmer, characterized in that, This includes the following operations: S1. Perform frequency-based weighted syncmer sampling on all segments in the reference pan-genome map to obtain the reference weighted syncmer group, construct a hash index, and obtain the reference pan-genome index. The weighted syncmer sampling process based on frequency division and weight reduction is as follows: Within each segment, a k-mer is selected at each position. For each position within the k-mer, an s-mer is selected, and the s-mer is converted to a positive number less than 1 using the xxHash3 hash function. Then, based on the repetition frequency interval of the s-mer, frequency division and weight reduction are applied to s-mers with different repetition frequencies, assigning them different weights. In the current k-mer, if the s-mer with the smallest weight appears at a pre-defined position, the current k-mer is considered a weighted syncmer. All weighted syncmers are counted to obtain a reference weighted syncmer group. In the frequency division and weight reduction operation, the greater the repetition frequency of the repetition frequency interval of the s-mer, the greater the corresponding weight reduction degree. Obtain the query-weighted syncmer group of the query sequence, use each k-mer as a key value, match it in the reference pan-genome index, obtain each k-mer that can be successfully matched, and the end position of the corresponding segment in the reference pan-genome map, the end position on the query sequence, and the sequence length constitute the anchor triplet corresponding to each k-mer. Sort them in ascending order according to the end position of the corresponding segment to form a candidate anchor array. S2. Traverse each anchor point in the candidate anchor point array. Based on the topological information in the reference pangenome graph, backtrack to calculate and obtain the linear chain corresponding to the maximum score ending at each anchor point. This is the corresponding best linear chain, forming a candidate linear chain array. By traversing each best linear chain in the candidate linear chain array, the best graph chain ending with each best linear chain is obtained, forming a candidate pangenome graph chain set; S3. Based on the candidate pan-genome graph chain set, the long read sequence alignment results are obtained.

2. The method for obtaining the repetition frequency interval in the pan-genome long read alignment method based on weighted syncmer according to claim 1 is as follows: The s-mer sequences that appear in the top 0.02% of the reference pan-genome map are sorted in descending order of frequency to obtain the sorted s-mer set. The sorted s-mer set is then divided into four different repeat frequency intervals: ultra-high frequency interval, high frequency interval, second-highest frequency interval, and intermediate frequency interval, with the repeat frequency decreasing in that order. The ratio of the frequency interval lengths is 0.02%:(0.02%~30%):(30%~70%):(70%~100%).

3. The pan-genome long read alignment method based on weighted syncmer according to claim 1, characterized in that, In S2, anchor points i Maximum score of the final linear chain The calculation formula is as follows: , , For anchor points j Score of the final linear chain, anchor point j The end position of the segment is less than the anchor point. i At the end of the segment, anchor point i With anchor point j The number of matching bases, Anchor point i With anchor point j Open shot penalty, anchor point i The sequence length, anchor point i With anchor point j The length of the empty space, anchor point i With anchor point j The minimum distance, , These are the first coefficient and the second coefficient, respectively.

4. The pan-genome long read alignment method based on weighted syncmer according to claim 1, characterized in that, The specific operation of S3 is as follows: The coordinates of the matching intervals within the segment in the reference pan-genome graph are converted into global reference coordinates to construct reference path units. Each graph chain in the candidate pan-genome graph chain set is traversed, and the path units corresponding to each anchor point in the graph chain in the reference path unit are connected in order. If they are adjacent, they are merged. The query sequence name and the total length of the query sequence are summarized, and the long read sequence alignment results are output.

5. A pan-genome long read alignment system based on weighted syncmer, used to implement the pan-genome long read alignment method based on weighted syncmer as described in claim 1, characterized in that, include: The candidate anchor array generation module is used to perform frequency-based weighted syncmer sampling on all segments in the reference pan-genome graph to obtain a reference weighted syncmer group, construct a hash index, and obtain the reference pan-genome index. The query weighted syncmer group of the query sequence is obtained, and each k-mer is used as a key value to match in the reference pan-genome index. For each k-mer that can be successfully matched, the end position of the corresponding segment in the reference pan-genome graph, the end position on the query sequence, and the sequence length constitute the anchor triplet corresponding to each k-mer. The triplet is sorted in ascending order according to the end position of the corresponding segment to form a candidate anchor array. The candidate pangenome graph chain set generation module is used to traverse each anchor point in the candidate anchor point array, and based on the topological information in the reference pangenome graph, backtrack to calculate and obtain the linear chain corresponding to the maximum score ending at each anchor point, which is taken as the corresponding best linear chain, forming a candidate linear chain array; traverse each best linear chain in the candidate linear chain array to obtain the best graph chain ending at each best linear chain, forming a candidate pangenome graph chain set; The sequence alignment result generation module is used to obtain long-read sequence alignment results based on the candidate pan-genome graph chain set.

6. A pan-genome long read alignment device based on weighted syncmer, characterized in that, It includes a processor and a memory, wherein the processor executes a computer program stored in the memory to implement the pan-genome long read alignment method based on weighted syncmer as described in any one of claims 1-4.

7. A computer-readable storage medium, characterized in that, Used to store a computer program, wherein the computer program, when executed by a processor, implements the pan-genome long read alignment method based on weighted syncmer as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Sequence alignment method and device based on local window, equipment and storage medium

    CN118351937A

  • Alignment method, system and equipment of genome long sequence and storage medium

    CN118866104A