Methods and systems for phasing, mapping, and assembling sequence reads and detecting structural variants
By employing signature phasing k-mers and linkage information, the methods and systems enhance the phasing and mapping of sequence reads, facilitating the detection of smaller structural variants with improved accuracy and efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-04-02
AI Technical Summary
Traditional nucleic acid sequencing methods lose information about the original position and connectivity of sequence fragments, making it difficult to accurately detect structural variants in genomic DNA samples.
The methods and systems utilize signature phasing k-mers to phase sequence reads into haplotypes, leverage linkage information for accurate mapping, and employ a k-mer signature-based filter to recruit unmapped reads, enabling robust phasing and mapping of sequence reads, particularly in regions with heterozygous variants.
This approach allows for the detection of smaller structural variants and improves computational efficiency, enabling faster and more accurate detection of structural variants, especially those less than 10k base pairs long.
Smart Images

Figure US2025044659_02042026_PF_FP_ABST
Abstract
Description
ILLINC.848WO / IP-2820-PCT PATENTMETHODS AND SYSTEMS FOR PHASING, MAPPING, AND ASSEMBLING SEQUENCE READS AND DETECTING STRUCTURAL VARIANTSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 700,569, filed September 27, 2024, the content of which is incorporated by reference in its entirety.BACKGROUNDField
[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, this disclosure relates to systems and methods for phasing sequence reads, mapping sequence reads, assembling sequence reads, and detecting structural variants in a genomic DNA sample.Description
[0003] Structural variants (SVs) of a genome are relatively large variants in the genome with respect to a known reference genome. For example, a structural variant may be an insertion of 1000 nucleotides into a new position of the genome as compared to the reference genome. One widely-used threshold of variant-size for classifying variants as “small” or “large” has been 50 base pairs (bp). However, some methods may use 35 bp as the threshold to denote the difference between a small or large variant. SVs include DNA losses, gains, and rearrangements relative to the reference genome. Simple SVs are typically classified as deletions (DNA loss), insertions and duplications (DNA gain), inversions and translocations (DNA rearrangements). Simple SVs can combine to cause more complex SVs in a genomic locus. For example, one genomic region may have a segment of DNA deleted, and a new segment of DNA inserted so that the single position is missing a nucleotide, and also includes new nucleotides as compared to a reference genome. The importance of SVs has already been well established due to their role in gene regulation, various diseases, and ethnic diversity. Others have commented on the importance of SVs and their applications in medicine and molecular biology. Mahmoud, M. et al. (2019) ‘Structural variant calling: the long and theshort of it’, Genome Biology, 20(1), p. 246. Available at: doi.org / 10.1186 / sl3059-019-1828- 7.
[0004] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods, including Sequencing by Synthesis (SBS), use a shotgun approach to sequence large genomic DNA fragments, sometimes called template genomic sequences. During SBS sequencing, the template genomic sequences are first fragmented into smaller pieces that are amenable to next-generation sequencing methods on a flow cell. One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been sequenced, knowledge of their original position in the genome, and their connectivity and proximity to each other in the original template genomic sequence is lost.SUMMARY
[0005] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand how the features of the devices and methods disclosed herein provide advantages over other known devices and methods.
[0006] Disclosed herein are methods for phasing sequence reads in a target region. In some embodiments, the method includes receiving an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyzing the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assigning unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0007] In some embodiments, the signature phasing k-mer is included in sequence reads phased to the first haplotype and is not included in sequence reads phased to a second haplotype of the two or more haplotypes. In some embodiments, the method comprises generating sequence reads from nucleic acids. In some embodiments, generating sequence reads comprises fragmenting the genomic DNA sample on a flow cell into a plurality of sequence fragments. In some embodiments, the phased sequence reads are initially phased into two or more haplotypes based on their geographic location on the flow cell. In some embodiments, the phased sequence reads are initially phased into two or more haplotypes based on read phasing, small variants-based phasing, statistical phasing, or trio phasing. In some embodiments, the target region is a putative structural variant region.
[0008] Disclosed herein are methods of detecting a structural variant. In some embodiments, the method includes phasing sequence reads according to the methods described herein; generating one or more assemblies from the first sequence reads; and detecting a structural variant based on the one or more assemblies.
[0009] Further disclosed herein are systems for phasing sequence reads in a target region. In some embodiments, the system includes one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyzing the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assigning unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0010] Further disclosed herein are non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyze the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequenceread is associated with a first haplotype of the two or more haplotypes; and assign unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0011] Further disclosed herein are methods for mapping sequence reads in a target region. In some embodiments, the method comprises receiving an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identifying anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identifying unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyzing unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and storing the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
[0012] In some embodiments, the method further includes assigning a haplotype to the unmapped sequence reads having a link to an anchor sequence read and which include the common k-mer signature, based on the phasing of linked anchor sequence reads. In some embodiments, storing unmapped sequence reads having a link to an anchor sequence read and which include the common k-mer signature comprises identifying unmapped sequence reads having a link to an anchor sequence read and that share a common k-mer signature with a greater than a threshold number of unmapped sequence reads having a link to an anchor sequence read. In some embodiments, k-mer signatures are weighted based on the frequency of the k-mer signature in the unmapped sequence reads or reference genome. In some embodiments, the method comprises indexing the unmapped sequence reads in a sequence index. In some embodiments, analyzing unmapped sequence reads and anchor sequence reads comprises querying the sequence index based on the geographic distance between a sequence cluster comprising the anchor sequence reads and a sequence cluster comprising unmappedsequence reads on a flow cell. In some embodiments, the one or more flanking regions comprise an upstream flanking region and a downstream flanking region.
[0013] Further disclosed herein are methods of detecting a structural variant. In some embodiments, the method includes mapping sequence reads according to the methods described herein; generating one or more sequence assemblies from the first sequence reads; and detecting a structural variant based on the one or more assemblies.
[0014] Further disclosed herein are systems for mapping sequence reads in a target region. In some embodiments, the system comprises one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identifying anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identifying unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyzing unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and storing the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
[0015] Further disclosed herein are non-transitory computer-readable media. In some embodiments the non-transitory computer readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identify anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identify unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyze unmappedsequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and store the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
[0016] Further disclosed herein are methods for generating a phased haplotype assembly from sequence reads. In some embodiments, the method includes receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generating an assembly graph de novo from the partitioned phased sequence reads; selecting a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolving paths in the assembly graph based on the selected guide sequence.
[0017] In some embodiments, the method further comprises randomly assigning unphased sequence reads to a haplotype of the two or more haplotypes. In some embodiments, the method comprises selecting a guide sequence from a database of SV-containing haplotypes based on sequence similarity between the guide sequence and the phased sequence reads.
[0018] Further disclosed herein are methods of detecting a structural variant, comprising: generating one or more phased haplotype assemblies from sequence reads according to the methods disclosed herein; aligning contigs from each phased haplotype assembly to a reference sequence; and detecting a structural variant based on the alignment to the reference sequence.
[0019] Further disclosed herein are systems for generating a phased haplotype assembly from sequence reads. In some embodiments, the system comprises one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for eachhaplotype group: generating an assembly graph de novo from the partitioned phased sequence reads; selecting a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolving paths in the assembly graph based on the selected guide sequence.
[0020] Further disclosed herein are non-transitory computer-readable media comprising plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partition phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generate an assembly graph de novo from the partitioned phased sequence reads; select a guide sequence from a database of SV- containing haplotypes based on the phased sequence reads; and resolve paths in the assembly graph based on the selected guide sequence.BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity, reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.
[0022] FIG. 1 is a flow diagram that schematically illustrates an exemplary method for phasing sequence reads in a target region.
[0023] FIG. 2 is a flow diagram that schematically illustrates an exemplary method for mapping sequence reads in a target region.
[0024] FIG. 3 is a flow diagram that schematically illustrates an exemplary method for assembling a phased haplotype from sequence reads.
[0025] FIG. 4A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.
[0026] FIG. 4B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 4A.
[0027] FIG. 5 schematically illustrates an alignment of sequence reads to a reference genome.
[0028] FIG. 6 schematically illustrates an alignment of sequence reads to a reference genome and methods for mapping sequence reads.
[0029] FIG. 7 is an Integrative Genomics Viewer (IGV) screenshot of one example based on real data showing results of the signature phasing k-mer propagation.
[0030] FIG. 8 is a heatmap that depicts the differences between SVs in various categories in the subset vs those in the whole genome for HG002 reference genome.
[0031] FIG. 9 is a bar plot of Recall scores for various variant calls using 3 different stringency settings.
[0032] FIG. 10 is a bar plot of fl scores for various variant calls using 3 different stringency settings.
[0033] FIG. 11 is a set of bar plots of recall scores stratified by INS / DEL (insertion / deletion) and TR / NTR (inside / outside Tandem Repeats) for the seq len settings. Panel a) shows recall for an insertion (INS) in a region inside a tandem repeat (TR) for the most strict (seq len) settings. Panel b) shows recall for a deletion (DEL) in a region inside a tandem repeat (TR) for the most strict (seq_len) settings. Panel c) shows recall for an insertion (INS) in a region outside a tandem repeat (NTR) for the most strict (seq_len) settings. Panel d) shows recall for a deletion (DEL) in a region outside a tandem repeat (NTR) for the most strict (seq_len) settings.
[0034] FIG. 12 is a set of bar plots of recall scores stratified by INS / DEL (insertion / deletion) and TR / NTR (inside / outside Tandem Repeats) for the NOseq_NOlen settings. Panel a) shows recall for an insertion (INS) in a region inside a tandem repeat (TR) for the most lenient (NOseq NOlen) settings. Panel b) shows recall for a deletion (DEL) in a region inside a tandem repeat (TR) for the most lenient (NOseq_NOlen) settings. Panel c) shows recall for an insertion (INS) in a region outside a tandem repeat (NTR) for the mostlenient (N0seq_N01en) settings. Panel d) shows recall for a deletion (DEL) in a region outside a tandem repeat (NTR) for the most lenient (NOseq_NOlen) settings.DETAILED DESCRIPTION
[0035] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.
[0036] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entirety for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.
[0037] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures,can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0038] Structural variant (SV) detection is an active area of research with new tools and methods being developed for different sequencing technologies using different bioinformatics techniques. The computational and scientific challenges of identifying SVs and benchmarking different methods and tools reflects in the fact that, unlike small variants, there are no well-established SV calling and benchmarking tools / pipelines. The SV calling methods’ landscape has been surveyed by others where they compared different methods calling SVs from long / short reads using mappings and assemblies. See Mahmoud, M. et al. (2019) ‘Structural variant calling: the long and the short of it’, Genome Biology, 20(1), p. 246.
[0039] The present disclosure relates to methods and systems for phasing sequence reads, mapping sequence reads, and assembling sequence reads. Each of these processes may be particularly useful in, for example, methods for detecting structural variants. For example, a method for detecting structural variants may use the phasing, mapping, and assembly techniques described herein or a combination thereof.
[0040] In some embodiments, sequence reads comprise a barcode. As used herein “barcode” refers to short, unique nucleic acid sequence used to tag or label different nucleic acid samples or different nucleic acid molecules. In some embodiments, sequence reads in the flanking regions of a candidate locus are used to recruit barcoded reads. In some embodiments, recruited reads are used for gap-filling of the flanking regions, for example as a form of sequence assembly.
[0041] The methods and systems described herein are able to advantageously detect structural variants at smaller candidate loci (for example, less than 10k base pairs long), in contrast to other methods and systems that only detect structural variants at larger loci that are about 50k base pairs or more in length.. The inventive methods and systems described herein are faster and less compute demanding than currently available methods and systems.Phasing
[0042] In diploid organisms such as humans, two haplotypes exist for each chromosome: one for the paternal homolog and one for the maternal homolog. Phasing, also referred to as “haplotagging”, is the process of assigning sequence reads into one of two ormore haplotypes. In humans, this would include assigning each sequence read to either the maternal or the paternal homolog. Phasing information can be useful for improving downstream sequence assembly and for detecting variants. For example, it is clinically relevant for some variants whether two or more variants are inherited in the same haplotype (cis) or different haplotypes (trans).
[0043] Phasing can traditionally be achieved using a variety, or combination, of methods. A first method is statistical inferential phasing, which includes comparing the sequence read to a catalog or database of variants of multiple individuals and selecting the correct haplotype based on a best match of the variant to either the maternal or paternal homolog. Another method is pedigree-based phasing wherein the sequence reads and variants are compared against the known sequences from parents or ancestors to assign the read to a particular homolog. A final method can include read-backed phasing where sequence information from overlapping sequence reads is used to identify a relationship between sequence reads which are known to be assigned to a particular homolog, and adjacent sequence reads that can be assigned to the same homolog. Read-backed phasing can make use of several sequencing technologies including short / long reads, strand-seq, HiC, 10X long-range reads, etc. to link the known, phased sequence read to the unknown, unphased sequence read. Across these methods, one valuable step in phasing is leveraging the fact that many homologs include heterozygous (HET) variants, where each homolog differs in sequence from the other homolog.
[0044] The above methods have been shown to improve the accuracy of assembly, particularly when there are no reference sequences available and the sequencing system needs to perform a de novo assembly of the sequence reads without reference to another sequence. Even though relatively short phased sequences obtained from the above methods may be made larger by using long-range information to infer a relationship between genomically distant sequence reads, when such relationships between sequence reads are difficult to determine, or sparse throughout the sequence space, and particularly when phasing information is not available, there may be areas of an alignment where few or no sequence reads can be assigned to a particular haplotype. When there are too many areas with insufficient amounts of phased reads, it may not be possible to leverage phasing in sequence read assembly or variant calling.
[0045] To address these issues, for example in systems which perform sequencing with relatively short reads, the present disclosure analyzes the presence of k-mers within each sequence read to help determine phasing for the read. Embodiments first discover “signature” k-mers within the sequence reads, and use these signature k-mers to determine the phase of sequence reads. In this embodiment, the systems and methods identify signature phasing k- mers within each sequence read that have already been phased to a particular homolog. The system may then use the different signature phasing k-mers which are included in the sequence reads found for the two haplotypes to assign unphased reads to one of the haplotypes.
[0046] For example, in some embodiments, the methods and systems identify a k- mer that appears in sequence reads of a first haplotype and not in a second haplotype. The methods and systems can then look for the occurrence of the k-mer in an unphased read. Because the k-mer appears in only a single haplotype, the sequence read which contains that k-mer is very likely to also be part of the same haplotype. Thus, the methods and systems can assign the unphased read to that the same haplotype as the phased sequence reads that include that particular k-mer.
[0047] This method can advantageously allow for robust phasing in regions without heterozygous variants and information connecting various sequence reads (for example, “links” as further described herein), where other methods relying on heterozygous variants and information connecting phased reads would fail.Signature k-mers
[0048] Sequence reads may comprise common signature k-mers. A “k-mer” represents a nucleic acid sequence of length k, that is contained within a sequence read. “K- mers” as used herein also includes derivatives of k-mers such as minimizers, strobemers, and syncmers. K-mers are representative subsequences of a larger sequence.
[0049] In the context of phasing sequence reads in the methods described herein, a “signature phasing k-mer” may be a k-mer that appears in reads that are phased to a first haplotype and does not appear in reads that are phased to a second haplotype. A “signature phasing k-mer” may also be a k-mer that appears more frequently in reads phased to a first haplotype than reads phased to a second haplotype. In an embodiment, a signature phasing k- mer is a k-mer that appears at least n times more frequently in the sequence reads phased tothe first haplotype than in the sequence reads phased to the second haplotype. In this embodiment, n is any integer for example 2, 3, 4 or 5. Optionally, a signature phasing k-mer may be a k-mer that appears at least two times, at least three times, at least four times, at least five times, or at least ten times in the sequence reads phased to the first haplotype. Thus, the methods and systems may determine whether sequence reads phased to a first haplotype comprise common signature phasing k-mers by partitioning the sequence reads phased to the first haplotype into k-mers and partitioning the sequence reads phased to the second haplotype into k-mers. The methods and systems may then compare the first haplotype k-mers and the second haplotype k-mers, and determine which k-mers appear in the first haplotype k-mers and not in the second haplotype k-mers, or which k-mers appear more frequently in the first haplotype k-mers than in the second haplotype k-mers. The methods and systems may then assess the k-mers which appear in the first haplotype k-mers and not (or less frequently) in the second haplotype k-mers and count them. Any k-mers which appear at least twice, at least three times, at least four times, at least five times, or at least ten times in the first haplotype k- mers and not in the second haplotype k-mers may be designated as signature phasing k-mers. Any k-mers that appear less than n, less than 5, less than 4, less than 3, or once in the first haplotype k-mers and not (or less frequently) in the second haplotype k-mers may be a result of a sequencing error and may be disregarded.
[0050] In the context of mapping sequence reads as described herein, a “signature mapping k-mer” (also referred to herein as a “k-mer signature”) may be a k-mer that appears in more than a threshold number of unmapped reads that are linked to mapped reads in a flanking region. In an embodiment, a signature mapping k-mer is a k-mer that appears at least n times in the unmapped reads that have a link to a flanking region, wherein n is any integer for example 2, 3, 4, 5, 10, or more. Optionally, a signature mapping k-mer is a k-mer that appears at least two times, at least three times, at least four times, at least five times, or at least ten times in the unmapped sequence reads that are linked to a sequence read in a flanking region. Thus, the methods and systems may determine whether unmapped sequence reads comprise common signature mapping k-mers by partitioning the unmapped sequence reads into k-mers. The methods and systems may then assess which k-mers appear more than a threshold number of times, for example at least twice, at least three times, at least four times, at least five times, or at least ten times, in the unmapped sequence reads that have a link to asequence read from a flanking region. Any k-mers that appear less than n, less than 5, less than 4, less than 3, or once in the first haplotype k-mers and not (or less frequently) may be a result of an error and could be disregarded. The value of k can be selected by the methods and systems, and can be any value. Optionally, the value of k is at least 5, at least 10, at least 15, less than 100, less than 50, less than 25, between 5 and 100, between 10 and 50, or between 15 and 25. Generally, the user will select a value of k which is as long as possible, while ensuring that the fraction of k-mers in a read that contains one or more sequencing errors is low. In one embodiment, the proportion of Aimers in a read that contains sequencing errors is less than 50%, less than 40%, less than 30%, between 0% and 50%, between 0% and 40%, or between 0% and 30%.Mapping (Read ‘Recruitment”)
[0051] During massively parallel sequencing processes, such as those found in Next Generation Sequencing (NGS) systems, relatively short reads of 100-500 nucleotides are typically generated. During this process, sample nucleic acids may be fragmented by tagmentation of the sample directly on the flow cell, in some embodiments. In an example process, transposons are linked to the flow cell, and when contacted by sample nucleic acids may fragment the sample nucleic acids resulting in many fragments from the same nucleic acid molecule being generated. This example process may result in fragments from the same nucleic acid molecule being bound to the flow cell in geographically adjacent or nearby locations. These fragments then give rise to nucleic acid clusters at nearby locations on the flow cell, and these nucleic acid clusters will participate in sequencing reactions to generate reads or read pairs. Therefore, reads or read pairs that are generated by sequencing these fragments which are located near each other on the flow cell may be determined by the disclosed systems or methods to be “connected” or “linked”, or that there are “links” among the reads or read pairs since they may have derived from the same originating nucleic acid molecule. This type of linking information or connectivity information may enable reconstruction of a nucleic acid molecule by way of grouping together linked read pairs which are bound to the flow cell at geographically nearby locations on the flow cell.
[0052] In read-alignment-based SV calling, sequence reads determined from a sample genome are aligned to a reference genome. If the sample genome has an insertion in alocus with respect to the reference genome, the reads originating from the inserted sequence may be unmapped, for example if the insertion is not represented in the reference genome, or incorrectly mapped to the original position in the genome, particularly if the insertion is a mobile element.
[0053] The present disclosure relates to methods and systems that can “recruit” these unmapped / incorrectly mapped sequence reads using linking information between two reads. As used herein, the term “recruit,” “recruiting,” and “recruitment,” refers to identifying and storing reads which can be used for mapping to a reference genome.
[0054] As further described herein, “links” between reads originating from the same molecule / template can be deduced by their location on the flow cell. For example, long (greater than 500 bp) DNA fragments may be flowed across a sequencing flow cell and may attach to transposome complexes that are embedded on the surface of the flow cell. The transposome complexes can fragment the long DNA fragments into shorter fragments of less than 500 bp, and attach sequencing adapters. Each shorter fragment can be amplified to form clusters, and the sequence of each shorter fragment can be determined using the clusters. The physical location of each cluster can be recorded in addition to sequence reads. An initial linking model can be developed based on the physical flow cell distance and genomic distance between an initial subset of the sequence reads. Then, given the physical location of two clusters on a flow cell and the genomic distance when mapped to a reference genome, sequence reads may be linked in that a probability can be determined related to whether sequence reads derive from the same original long DNA fragment.
[0055] Linkage information may improve the accuracy and efficiency of the alignment of sequence reads to a reference sequence. By improving the alignment of sequence reads to the reference genome, linkage information can help with methods of detecting structural variants. Linkage information can be leveraged to align sequence reads to regions of the reference genome that are difficult to map, and leads to improved accuracy in variant detection capabilities over such regions.
[0056] In some embodiments of the present disclosure, the links of the reads in the flanking regions of a potential SV candidate are leveraged to retrieve the “missing” reads of an insertion. A previous patent application (App. No. 63 / 600,492, incorporated by reference in its entirety) describes aspects of a method to retrieve such reads using flow cell information.
[0057] However, simply using proximity on a flow cell to recruit reads may bring in too many reads generated from other regions of the genome and which should not be mapped to the candidate SV region. This is because the linking process may generate spurious links between sequence reads which are not truly genomically proximate to one another. Sequence reads which have such spurious links are referred to herein as “noise” reads. The larger the flanking regions used to recruit reads, the more “noise” reads may be recruited because there are more sequence reads and thus more possible spurious links. Including too many noise reads leads to noisy assembled contigs in the assembly next stage, leading to missed or wrong SV identification, and also makes the assembly step itself too computationally demanding to be used for a real-world application.
[0058] Embodiments relate to systems and methods of employing a filtering strategy while recruiting linked reads to filter out noise reads. For example, because not all nearby reads on a flow cell come from the same original molecule, not all nearby reads on the flow cell are genomically close to each other. Thus, some links based on flow cell proximity are not useful for mapping. To reduce the number of reads which are not useful for mapping, in some embodiments, the methods and systems include a k-mer signature-based filter where only unmapped reads that both a) have a link to a sequence read in a recruiter region and b) include a k-mer that is shared with more than a threshold number of other unmapped linked reads, are recruited for mapping. For example, sequence reads originating from an insertion may be linked to the reads from genomic regions flanking the insertion based on flow cell proximity, and would have enough coverage to share their k-mer signature with other reads from the insertion. In contrast, spuriously linked sequence reads would not have enough coverage to share a k-mer signature with other reads. Additionally, in some embodiments, the phasing information of the “recruiter” read from the flanking regions is propagated to the read being “recruited” and thus the recruited reads are assigned to a haplotype as well. The use of a k-mer signature filter for noise control allows more accurate mapping / assembly and processes fewer “noisy” reads, which results in more computationally efficient SV calling and allows for the detection of smaller SVs than other SV detection methods (for example, SVs that are about 10 kbp in size instead of 50 or more kbp in size).Read Assembly
[0059] Genome assemblers can be classified as de novo and reference-guided. De novo assemblers make use of assembly graphs, for example a de Bruijn graph, or an overlap layout consensus (OLC) graphs, resulting from all reads and traverse paths in the graph to produce larger assembled sequences known as contigs. Others have described various de novo sequence assembly approaches. See Ayling, M., Clark, M.D. and Leggett, R.M. (2020) ‘New approaches for metagenome assembly with short reads’, Briefings in Bioinformatics, 21(2), pp. 584-594. Reference-guided approaches make use of a reference sequence, either during assembly, or after assembling together different contigs generated in a de novo process.
[0060] The use of a reference runs the risk of introducing biases and misassemblies in the diverging regions, but de novo assemblers struggle in resolving paths in difficult regions such as repeats longer than the read lengths. However, a combination of the two can make use of the complementary strengths of the two approaches.
[0061] Moreover, most of the assemblers collapse the haplotypes, losing the context for heterozygous variants. For example, in a heterozygous region where there are differences in maternal and paternal haplotypes while the flanks of this region have the exact same maternal and paternal haplotypes, the heterozygous region would appear as a “bubble” in an assembly graph because there are two possible paths from the left flank to the right flank. Many prior art assemblers would traverse only one of these two paths to produce a contig and ignore the alternate path, thus “collapsing” of maternal and paternal haplotypes to produce only one contiguous sequence (“contig”). The single contig might be completely from one of the parents, or might have some maternal sequence and some paternal sequence if the contig includes two or more “bubbles.” Previous assembly approaches can thus result in a contig that only reflects one haplotype or is an artificial combination of two haplotypes. Although some long-read assemblers use haplotype-aware approaches to produce diploid assemblies, currently short-read assemblers do not produce separate assemblies for haplotypes.
[0062] The present disclosure relates to methods and systems that segregate the sequence reads into haplotypes and perform separate local assemblies of the potential SV candidate region. Unlike the short reads SV detection methods that try to assemble breakends (endpoints of an SV) in haplotype-unaware manner, producing separate haplotype assemblies simplifies the assembly graph by removing paths, which may present on the assembly graphas bubbles / branches, originating from reads of other haplotypes. This improves computational efficiency by simplifying the assembly process. In addition, in some embodiments, the methods and systems use a guide sequence, such as a sequence from a database of SV- containing haplotypes of multiple individuals inferred to be the most similar to the sample, to resolve paths in the assembly graphs. This technique of using a guide sequence to determine ambiguous paths in the assembly graph further improves sequence assembly.Definitions
[0063] Although the following terms are believed to be well understood by one of skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.
[0064] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of skill in the art.
[0065] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.
[0066] As used herein, the term “about,” when used in reference to a measurable value such as an amount of mass, dose, time, temperature, and the like, is meant to encompass variations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.
[0067] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).
[0068] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.
[0069] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that thecompositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.
[0070] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0071] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.
[0072] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. Anucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series.
[0073] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. Examples of reference genomes include GRCh38 from the Genome Reference Consortium.
[0074] As used herein, the term “sample” refers to a specimen, culture, or the like that is suspected of including a target nucleic acid. In some embodiments, the sample comprises DNA, ribonucleic acid (RNA), peptide nucleic acid (PNA), locked nucleic acid (LNA), chimeric or hybrid forms of nucleic acids as targets. The sample can likewise include any biological, clinical, surgical, agricultural-atmospheric, or aquatic-based specimen containing one or more nucleic acids. A sample also includes any isolated or extracted nucleicacid sample from an organism, such a genomic DNA, fresh-frozen, or formalin-fixed paraffin- embedded nucleic acid specimen. In some cases, accordingly, a sample can include a full genome or partial genome that is isolated or extracted (e.g., in whole or in part by a kit) from an organism and that is prepared to undergo sequencing or an assay in a sequencing device. A sample can be from a single individual, a collection of nucleic acid samples from genetically related members, nucleic acid samples from genetically unrelated members, nucleic acid samples (matched) from a single individual such as a tumor sample and normal tissue sample, or sample from a single source that contains two distinct forms of genetic material, such as maternal and fetal DNA obtained from a maternal subject, or the presence of contaminating bacterial DNA in a sample that contains plant or animal DNA. In some embodiments, the source of nucleic acid material can include nucleic acids obtained from a newborn, for example as typically used for newborn screening.
[0075] The sample can include high molecular weight material, such as genomic DNA (gDNA). The sample can include low molecular weight material such as nucleic acid molecules obtained from FFPE or archived DNA samples. In another implementation, low molecular weight material includes enzymatically or mechanically fragmented DNA. The sample can include cell-free circulating DNA. In some implementations, the sample can include nucleic acid molecules obtained from biopsies, tumors, scrapings, swabs, blood, mucus, urine, plasma, semen, hair, laser capture micro-dissections, surgical resections, and other clinical or laboratory obtained samples. In some implementations, the sample can be an epidemiological, agricultural, forensic, or pathogenic sample. In some implementations, the sample can include nucleic acid molecules obtained from an animal such as a human or mammalian source. In another implementation, the sample can include nucleic acid molecules obtained from a non-mammalian source such as a plant, bacteria, virus, or fungus. In some implementations, the source of the nucleic acid molecules may be an archived or extinct sample or species.
[0076] As further used herein, the term “sequencing run” refers to an iterative process on a sequencing device to determine a primary structure of nucleotide sequences from a sample (e.g., genomic sample). In particular, a sequencing run includes cycles of sequencing chemistry and imaging performed by a sequencing device (including an imaging device, such as a CCD or CMOS) that incorporate nucleobases into growing oligonucleotides to determinenucleotide reads from nucleotide sequences extracted from a sample (or other sequences within a library fragment) and seeded throughout a flow cell or other nucleotide-sample slide. In some cases, a sequencing run includes replicating oligonucleotides derived or extracted from one or more genomic samples seeded in clusters throughout a flow cell. Upon completing a sequencing run, a sequencing device can generate base-call data in a file, such as a binary base call (BCL) sequence file or a fast-all quality (FASTQ) file.
[0077] Relatedly, the term “sequencing cycle” (or “cycle”) refers to an iteration of adding or incorporating one or more nucleobases to one or more oligonucleotides representing or corresponding to a sample’s sequence (e.g., a genomic or transcriptomic sequence from a sample) or a corresponding adapter sequence. In some cases, a sequencing cycle includes an iteration of both incorporating nucleobases into clusters of oligonucleotides using sequencing chemistry and capturing images of such clusters attached to a nucleotide-sample slide (e.g., a flow cell). Accordingly, cycles can be repeated as part of sequencing a nucleic-acid polymer (e.g., a sample genomic sequence). For example, in one or more embodiments, each sequencing cycle involves incorporating nucleobases into either a single nucleotide read in which DNA or RNA strands are read in only a single direction or paired-end reads in which DNA or RNA strands are read from both ends but in different cycles. Further, in certain cases, each sequencing cycle involves a camera taking an image of the nucleotide-sample slide or multiple sections of the nucleotide-sample slide to generate image data for determining a particular nucleobase added or incorporated into particular oligonucleotides. Following the image capture stage, a sequencing system can remove certain fluorescent labels from incorporated nucleobases and perform another sequencing cycle until the nucleic-acid polymer has been completely sequenced. In one or more embodiments, a sequencing cycle includes a cycle within an SBS run. A sequencing cycle can include one or both of an indexing cycle and a genomic sequencing cycle. For instance, one cluster of oligonucleotides or a set of clusters of oligonucleotides may be undergoing a genomic sequencing cycle in which nucleobases corresponding to a sample genomic sequence are incorporated and another cluster of oligonucleotides or another set of clusters of oligonucleotides may be concurrently undergoing an indexing cycle in which nucleobases corresponding to an indexing sequence for a nucleotide read are incorporated.
[0078] Further, as used herein, the term “nucleotide-sample slide” (or “nucleotide- sample substrate”) refers to a plate or substrate, such as a flow cell, comprising oligonucleotides for sequencing nucleotide sequences from genomic samples or other sample nucleic-acid polymers. In particular, a nucleotide-sample slide can refer to a substrate containing fluidic channels through which reagents and buffers can travel as part of sequencing. For example, in one or more embodiments, a flow cell (e.g., a patterned flow cell or non-patterned flow cell) may comprise small fluidic channels and oligonucleotide samples that can be bound to adapter sequences on the substrate. In other implementations, a nucleotide-sample slide can be an open substrate with one or more regions for oligonucleotide samples to be analyzed and the oligonucleotide samples may be positioned using charged pads or other means. In yet another implementation, the nucleotide-sample slide can be a membrane having a nanopore through which one or more oligonucleotide samples may pass.
[0079] Relatedly, as used herein, the term “region of a nucleotide-sample slide” (or “nucleotide-sample slide region”) refers to an area that is part of a nucleotide-sample slide. In particular, a region of a nucleotide-sample slide can refer to a discrete portion of a nucleotide- sample slide that differs from other portions of the nucleotide-sample slide. For instance, a region of a nucleotide-sample slide can include a subsection of patterned flow cell comprising one or more wells (e.g., a nano- wells) or a discrete subsection of a non-pattered flow cell (e.g., a subsection corresponding to one or more clusters). In some cases, a region (e.g., section) of a nucleotide-sample slide includes a tile or a sub-tile of a flow cell having clusters of oligonucleotides growing in parallel.
[0080] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, thatcan be aligned and specifically assigned to a chromosome or genomic region or gene. For example, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0081] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood of the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is on a particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.
[0082] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period forimplementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0083] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.
[0084] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.
[0085] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.
[0086] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragment thereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith- Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.
[0087] Relatedly, as used herein, the term “mapping quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In someembodiments, for example, a mapping quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong}, rounded to the nearest integer. In the alternative to a mean or median mapping quality, in some implementations, a mapping quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.
[0088] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 300 bp, and includes paired end sequence reads.
[0089] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more. A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed by assembling multiple short sequence reads.
[0090] As used herein, a “file” includes electronic files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.
[0091] As used herein, “haplotag,” “haplotagging,” and “haplotagged” may refer to sequence reads that have been assigned to a haplotype and have been tagged with haplotype information. For example, a sequence read may be phased to a haplotype as described herein, and haplotype information (e.g., “Hl,” “H2”) may be stored with the sequence read in addition to the nucleic acid sequence information.
[0092] A “k-mer” is a subsequence of k contiguous letters of a nucleotide sequence, for example, of a sequence read or assembled contig. “K-mers” as used herein also encompasses derivatives of k-mers (e.g., minimizers, strobemers, syncmers etc.). [Moeckel, C. et al. (2024) ‘A survey of k-mer methods and applications in bioinformatics’, Computational and Structural Biotechnology Journal, 23, pp. 2289-2303. Available at: https: / / doi.Org / 10.1016 / j.csbj.2024.05.025] presents a comprehensive review of the methods, applications, and significance of k-mers and its derivatives in genomic and proteomic data analyses, and is hereby incorporated by reference in its entirety.
[0093] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquid. The substrate can be non-porous or porous. The substrate canoptionally be capable of taking up a liquid (e.g. due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solid supports for some embodiments are located within a flow cell apparatus. Exemplary flow cells are set forth in further detail below.
[0094] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.
[0095] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations, wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1mm2, 500 pm2, 100 pm2, 25 pm2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2. Alternatively, oradditionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 gm2, 2.5 gm2, 5 gm2, 10 gm2, 100 gm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.
[0096] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantially flat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.
[0097] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complementary DNA (cDNA), and derivatives or analogs of these nucleic acids.
[0098] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element- containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleicacid. Typically, the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.
[0099] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is a double stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-type or mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, nonnatural bases, modified backbone, and can comprise nicks in one or both strands.
[0100] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from a solid support surface) can be leveraged to identify fragments that are generated from a single long input fragment and subsequentially be used to improve mapping reads in ambiguous positions.
[0101] In one or more embodiments, the system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing datafile” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.
[0102] Moreover, in one or more embodiments, one or more sequencing data files in which the system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignment data file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.
[0103] Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a nonpatterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.
[0104] Embodiments of the present disclosure relate to methods and systems which use “links” or “linkage information” between sequence reads. The “link” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long nucleic acids, such as genomic DNA, from a sample are sheared to create shorter fragments which can be sequenced in a single read. The shearing process can create these shorter fragments which land on theflow cell and the spatial location of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments which come from the same nucleic acid molecule have been found to bind closer together on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are close together spatially and also close together on the genome, the clusters are more likely to have come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability that two reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.
[0105] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). As described above, according to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in spatially nearby positions as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surface of the flow cell. In regular library preps with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell spatial information is retained. This spatial information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.
[0106] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used tobind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina, Inc. In some embodiments, the complexed strand and transposome is in solution, and is then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.
[0107] Once the fragments have been bound to substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. The location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster. Once the nucleotide sequence of each cluster has been determined, the method can start to map those reads to determine the original target polynucleotide from which the read originated. In some embodiments, the mapping process takes into account the spatial location of each cluster, such that clusters which are closer to each other on the flow cell are more likely to have originated from the same target polynucleotide. In some embodiments, the library preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment associated with the systems. Furthermore, by mapping the sequenced fragments to target polynucleotides using the spatial information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the spatial location of each cluster into account during the mapping process. Therefore, spatial information that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the read quality of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. Increasing the read quality of these previously discarded reads, by improving the confidence of read pair’s alignment based on linking information with a high link quality score, may improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling. Processing DNA samplessuitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments.
[0108] In some embodiments, linking information is determined by analyzing, for example, statistically analyzing with a model, the genomic distance between two reads and spatial distance (also known as flow cell proximity), between the two reads (or read clusters) on the flow cell. In some embodiments, the methods and systems determine whether the genomic distance and / or spatial distance between reads is below a threshold. In some embodiments, the methods and systems determine the presence or absence of a link between the two sequence reads. In some embodiments, the methods and systems determine a linking quality score between the two sequence reads. In some embodiments, the methods and systems analyze genomic distance and spatial distance for a plurality of pairs of two sequence reads (for example, each possible pair) of two sequence reads in a dataset. Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT7US2024 / 035447 and PCT7US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U.S. Provisional Patent Application Nos. 63 / 600460, 63 / 614066, 63 / 700,049, and 63 / 700,262, the disclosure of each of which is incorporated herein by reference in its entirety.Methods for Phasing Sequence Reads in a Target Region
[0109] In one aspect, disclosed herein are methods for phasing sequence reads in a target region. In some embodiments, the sequence reads are from genomic DNA. In some embodiments, the method is a method for propagating phasing information from phased sequence reads to unphased sequence reads in a target region. In some embodiments, the target region is a putative structural variant region.
[0110] In some embodiments, the method allows for the propagation of phasing from a set of phased sequence reads to unphased sequence reads. This may allow for phasing information to be extended to areas of the genome that traditionally are difficult to phase, including regions without heterozygous variants, and regions where there are not many links to other phased sequence reads, where other callers relying on heterozygous variants and / or density of links for phased reads would be unable to phase the unphased sequence reads.
[0111] The ability to extend phasing information to additional sequence reads, as described herein, may allow for improved sequence assembly as further described herein. For example, phasing information may be used to generate phased sequence assemblies. By providing phasing information before the assembly stage, efficiency may be improved during the computationally heavy assembly step, and sequence information related to haplotype may be retained.
[0112] In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. In some embodiments, generating sequence reads from fragments of the genomic DNA sample bound to a flow cell comprises: providing transposome complexes, wherein the transposome complexes comprise a transposase and a first polynucleotide comprising an end sequence and a first tag; contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides; and amplifying the fragmented target polynucleotides to form a plurality of nucleic acid clusters on a substrate. In some embodiments, the transposome complexes are attached to a substrate of the flow cell. In some embodiments, the sequence reads comprise paired end sequence reads.
[0113] In some embodiments, the method includes identifying a target region based on the first sequence reads. In some embodiments, the target region is a putative structural variant region. Various methods may be used to identify a putative structural variant based on sequence reads. For example, in some embodiments, a putative structural variant region is identified by scanning an alignment of sequence reads to a reference genome to identify regions with putative structural variants. For example, regions of the reference genome that include more than a threshold number of “abnormal” alignments such as split reads, improper pairs, or abnormal template length may be identified as a putative structural variant. A priori population-based information about regions of the genome that potentially have structural variants may also be used to identify a putative structural variant region, or any similar methods known to those skilled in the art. The step of identifying a target region, such as by identifying a putative structural variant region, may be performed in parallel across different subsections of the reference genome, thus improving computational efficiency.
[0114] FIG. 1 is a flow diagram that schematically illustrates an exemplary method 100 for phasing sequence reads in a target region. In some embodiments, the method 100 isimplemented on a processor. The method 100 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system, for example the computing system 4000. When the method 100 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device.Receiving an alignment with phased and unphased sequence reads
[0115] As shown in FIG. 1, the method 100 for phasing sequence reads in a target region may start from start block 110. The method 100 may proceed to block 120, wherein an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome is received.
[0116] The alignment may be, for example, retrieved from a storage, for example the data store 490. The sequence reads may be generated as further described herein. The sequence reads may be in any electronic file format. It will be apparent that one of ordinary skill in the art may use a variety of approaches to align the sequence reads to the reference genome.
[0117] In some embodiments, the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads.
[0118] The set of phased sequence reads may have been already phased using a variety of methods. For example, in some embodiments, the method comprises fragmenting the genomic DNA sample on a flow cell into plurality of sequence fragments as and wherein the phased sequence reads are initially phased into two or more haplotypes. In some embodiments, the initial phasing is based at least in part on the geographic location of sequence reads on the flow cell. As further described herein, the proximity of two clusters on a flow cell may be used to map sequence reads to a reference genome and to phase sequence reads. For example, a sequence read in a target region may be assigned to one haplotype if it has a link based on flow cell cluster proximity to another sequence read that is assigned to the haplotype.
[0119] It will also be apparent to one of ordinary skill in the art that other techniques may be used for initial phasing of sequence reads. For example, in some embodiments, the phased sequence reads are initially phased into two or more haplotypes based on read phasing, small variants-based phasing, statistical phasing, or trio phasing.
[0120] In some embodiments, the unphased sequence reads are sequence reads that initial phasing techniques as described above were not able to assign to a haplotype.
[0121] The method 100 may also be performed in parallel on multiple target regions. For example, the method may include a step (not shown) of binning (e.g., dividing) the alignment of first sequence reads to the reference genome into bins (e.g., subsections of a similar size to one another), and later steps in the workflow (such as the steps of block 130 and block 140) may be performed independently within each bin. This parallelization can further improve computational efficiency.Analyzing the subset of phased sequence reads to determine a signature phasing k-mer
[0122] The method 100 may proceed to block 130, wherein the phased sequence reads are analyzed to determine a signature phasing k-mer included in a plurality of phased sequence reads.
[0123] In some embodiments, the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes. For example, the phased sequence reads may be separated by haplotype and analyzed to determine a signature phasing k-mer that is included in sequence reads that are phased to a first haplotype and is not included or is included less frequently in sequence reads that are phased to a second haplotype.
[0124] In some embodiments, the signature phasing k-mer is included in sequence reads phased to the first haplotype and is not included in sequence reads phased to a second haplotype of the two or more haplotypes.
[0125] In some embodiments, the signature phasing k-mer is included more frequently in sequence reads phased to the first haplotype than in sequence reads phased to a second haplotype of the two or more haplotype. For example, if there are 5 sequence reads phased to a first haplotype and 5 sequence reads phased to a second haplotype, and a k-mer is included 5 out of 5 times in sequence reads phased to the first haplotype and only 1 out of 5 time in sequence reads phased to the second haplotype, the k-mer may be used as a signature phasing k-mer.Assigning unphased sequence reads that include the signature phasing k-mer
[0126] The method 100 may proceed to block 140, wherein unphased sequence reads that include the signature phasing k-mer are assigned to the first haplotype, thereby phasing unphased sequence reads.
[0127] For example, the set of unphased sequence reads may be analyzed to identify sequence reads which include the signature phasing k-mer identified in block 130, which is associated with a first haplotype of the two or more haplotypes. Unphased sequence reads which are identified as having the signature phasing k-mer may be assigned (for example, haplotagged) to the first haplotype. Thus, unphased sequence reads may be phased based in inclusion of a signature phasing k-mer which is included in sequence reads phased to a certain haplotype.
[0128] The method 100 may end at an end block 150.Methods for Mapping Sequence Reads in a Target Region
[0129] In one aspect, disclosed herein are methods for mapping sequence reads in a target region. In some embodiments, the sequence reads are from genomic DNA. In some embodiments, the target region is a putative structural variant region. In some embodiments, the target region includes a putative insertion.
[0130] In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. This may be accomplished as described above with respect to the methods of phasing sequence reads.
[0131] In some embodiments, an initial mapping step is performed. This initial mapping step may be performed, for example using flow cell cluster information as further described herein, or using other mapping methods known to those skilled in the art. In some embodiments, the initial mapping step results in some sequence reads which are mapped (for example, with a quality score (e.g., MAPQ) above a predetermined threshold) to a reference genome, and some unmapped sequence reads which are not able to be mapped to the reference genome or are mapped with a quality score below the threshold. In some embodiments, the method includes “recruiting” these unmapped sequence reads, for example, identifying previously unmapped sequence reads based on having a link based on flow cell proximity tosequence reads which were previously able to be mapped to a region flanking the target region, and storing such reads for mapping.
[0132] The ability to store and map additional sequence reads, as described herein, may allow for improved sequence assembly as further described herein. For example, by increasing the number of sequence reads that can be mapped, the methods and systems can more accurately determine a nucleic acid sequence in regions that previously were difficult to do so.
[0133] Moreover, in some embodiments, the presently disclosed methods include a filtering step where only sequence reads that share a signature mapping k-mer with a sufficient number of other sequence reads are retrieved, thereby reducing the number of noise sequence reads that are inaccurately retrieved based on spurious links. If too many sequence reads are inaccurately identified and stored for mapping, downstream sequence assembly may be computationally inefficient and may produce assemblies that falsely suggest a structural variant or conceal a true structural variant. Thus, the k-mer signature filtering step described below improves recall, specificity, and computational efficiency.
[0134] In some embodiments, the method includes identifying a target region based on the first sequence reads. This may be accomplished as described above with respect to the methods of phasing sequence reads.
[0135] FIG. 2 is a flow diagram that schematically illustrates an exemplary method 200 for mapping sequence reads in a target region. In some embodiments, the method 200 may be implemented on a processor. The method 200 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system, for example the computing system 4000. When the method 200 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device.Receiving an alignment to a target region, upstream flanking region, and downstream flanking region
[0136] As shown in FIG. 2, the method 200 for phasing sequence reads in a target region may start from start block 210. The method 200 may proceed to block 220, wherein analignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, is received.
[0137] In some embodiments, the alignment comprises mapped sequence reads and unmapped sequence reads. For example, the alignment may have been generated through an initial mapping step as described above, where some sequence reads are able to be mapped to the reference genome (mapped sequence reads) and some sequence reads are not able to be mapped to the reference genome (unmapped sequence reads).
[0138] In some embodiments, the one or more flanking regions comprise an upstream flanking region (for example, a region which is upstream of the target region and which flanks the target region) or a downstream flanking region (for example, a region which is downstream of the target region and which flanks the target region). In some embodiments, the one or more flanking regions comprise an upstream flanking region and a downstream flanking region. In some embodiments, the one or more flanking regions are each about 1 kbp, about 5 kbp, about 10 kbp, about 20 kbp, about 50 kbp, about 100 kbp, or a range created from any of the aforementioned values, in size. In some embodiments, the method 200 includes a step (not shown) of identifying the one or more flanking regions.
[0139] The method 200 may also be performed in parallel on multiple target regions. For example, the method may include a step (not shown) of binning (e.g., dividing) the alignment of first sequence reads to the reference genome into bins (e.g., subsections of a similar size to one another), and later steps in the workflow may be performed independently within each bin. This parallelization can further improve computational efficiency.Identifying anchor sequence reads that are mapped to the one or more flanking regions
[0140] The method 200 may proceed to block 230, wherein anchor sequence reads are identified. As used herein, “anchor sequence reads” refer to first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold. Thus, sequence reads which are mapped to one of the one or more flanking regions (for example, an upstream flanking region or a downstream flanking region) with a confidence score above a first predetermined threshold are identified at block 230. These mapped sequence reads may be used in downstream steps in order to “recruit” unmapped sequence reads, as will be further described.
[0141] The first predetermined threshold can be a MAPQ score of 5, 10, 20, 30, 40, 50, 60, 70, or more, or any value therebetween, or a range constructed from any of the aforementioned values. It will be understood that one of ordinary skill in the art could select a suitable predetermined threshold in order to capture a sufficient amount of sequence reads with a sufficient statistical probability that they are accurately mapped to a location on the reference genome.Identifying unmapped sequence reads
[0142] The method 200 may proceed to block 240, wherein unmapped sequence reads are identified.
[0143] In some embodiments, unmapped sequence reads comprise unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold are identified. Thus, “unmapped” sequence reads may refer to both sequence reads that the methods and systems cannot initially map to a reference genome, as well as sequence reads that are mapped but with a quality score that is below a predetermined threshold.
[0144] In some embodiments, unmapped sequence reads and sequence reads that are mapped with a quality score below the second predetermined threshold are sequence reads that could be coming from a potential structural variant, such as an insertion in the sample genome. For example, a structural variant such as an insertion can cause sequence reads to map unreliably.
[0145] The second predetermined threshold can be a MAPQ score of 5, 10, 20, 30, 40, 50, 60, 70, or more, or any value therebetween, or a range constructed from any of the aforementioned values. The second predetermined can be the same value as the first predetermined threshold described above, or it can be a different value. It will be understood that one of ordinary skill in the art could select a suitable predetermined threshold in order to capture sequence reads that are mapped with a lower than desirable confidence to a location of the reference genome.
[0146] In some embodiments, the method 200 includes a step (not shown) of indexing the unmapped sequence reads in a sequence index. In some embodiments, thesequence index stores sequence information for the unmapped sequence reads and additional information such as flow cell cluster location information (for example, coordinates).
[0147] In some embodiments, the method 200 includes a step (not shown) of storing k-mer signatures of the unmapped sequence reads in a data structure, for example, a hash-map, for example in the data store 490. The k-mer signatures may be used in later steps in the method 200, as further described herein. In some embodiments, k-mer signatures are filtered, for example, k-mers that are solely homopolymers may be excluded from the data structure that holds k-mer signatures.
[0148] While block 240 is described as taking place after block 230 in the method 200, it will be apparent to one of ordinary skill in the art that the order of these steps could be reversed in some embodiments.Analyzing unmapped sequence reads and anchor sequence reads to determine unmapped sequence reads having a link to an anchor sequence read
[0149] The method 200 may proceed to block 250, wherein unmapped sequence reads and anchor sequence reads are analyzed based on the geographic distance between sequence clusters on a flow cell. For example, at block 250, unmapped sequence reads and anchor sequence reads are analyzed based on the geographic distance between a) sequence clusters comprising the anchor sequence reads, and b) sequence clusters comprising the unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read. U.S. App. No. 63 / 600,492, incorporated by reference in its entirety, describes aspects of a method to link unmapped sequence reads to anchor sequence reads such reads using flow cell information.
[0150] In some embodiments, the method includes analyzing the geographic or flow cell distance between a given cluster on a flow cell corresponding to an anchor sequence read and each of the clusters corresponding to the unmapped sequence reads. In some embodiments, clusters corresponding to the unmapped sequence reads that are within a predetermined geographic distance of the cluster corresponding to an anchor sequence read are identified as unmapped sequence reads having a link based on flow cell proximity to the anchor sequence read.
[0151] In some embodiments, the predetermined geographic distance is a threshold distance of, for example and not by way of limitation, 10,000 nm, 5,000 nm, 3,000 nm, 2,000 nm, or 1,000 nm. In some embodiments, the predetermined geographic distance is a threshold distance measured in flow cell units, for example, within 10 flow cell units, 50 flow cell units, 100 flow cell units, 500 flow cell units, or any value therebetween, where one flow cell unit may correspond to about 1 micron, 20 micron, 40 micron, 60 micron, 80 micron, 100 micron, 200 micron, or any value therebetween, in different sequencing instruments. In some embodiments, the predetermined geographic distance may mean within a certain number of proximate wells. For example, on a substrate which includes wells for each read cluster, the number of wells between clusters may be much greater than 50, than 100, or than 200 wells. In some embodiments, the predetermined geographic distance may depend on x / y direction as the diffusion pattern may not be uniform after fragmentation. For example, the links may form an oval-like pattern on the flow cell, or any suitable shape. For example, in some embodiments, the methods and systems construct a two-dimensional histogram of the count of the relative displacement between two clusters on the flow cell associated with two read pairs in the first set of read pairs. In some cases, the two read pairs were derived from two regions on the same nucleic acid fragment, wherein the two regions are separated by a given genomic distance. In some embodiments, the contour lines of the histogram may comprise an oval-like shape. In some embodiments, the major axis of the oval-like shape may align with the direction of flow of sequencing reaction reagents on the flow cell. In some embodiments, the oval-like shape may center at the origin. In some embodiments, the contour lines of the histogram may comprise a circle. In some embodiments, the circle may center at the origin. In some embodiments, the contour lines of the histogram may comprise a pair of blobs. In some embodiments, the pair of blobs may be symmetric with respect to the x-axis and the y-axis of the histogram. For example, the y-axis of the histogram may align with the direction of flow of sequencing reaction reagents on the flow cell. In some embodiments, each blob of the pair of blobs may be a distance away from the x-axis of the histogram.
[0152] In some embodiments, analyzing unmapped sequence reads and anchor sequence reads to determine unmapped sequence reads having a link to an anchor sequence read comprises querying the sequence index (created as described above) based on thegeographic, or flow cell, distance between a sequence cluster comprising the anchor sequence reads and a sequence cluster comprising unmapped sequence reads on a flow cell.Storing the unmapped sequence reads having a link to an anchor sequence read which include a common k-mer signature
[0153] The method 200 may proceed to block 260, wherein unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature are stored for mapping to the reference genome at the target region. For example, the sequence reads may be stored in the data store 490.
[0154] For example, in some embodiments, the method includes partitioning sequence reads (including unmapped sequence reads having a link to an anchor sequence read as determined in block 250) into k-mers. The partitioning into k-mers may take place earlier in the workflow and may have been stored in a data structure as described above. In some embodiments, k-mer signatures are analyzed to identify common k-mer signatures (such as k- mers that are found in multiple unmapped sequence reads having a link to an anchor sequence read).
[0155] For example, in some embodiments, a common k-mer signature is a k-mer signature that appears in a greater than a threshold number of unmapped sequence reads having a link to an anchor sequence read. It will be apparent to one of ordinary skill in the art that the threshold number may be predetermined and may be a number of occurrences, for example occurring at least 2, 3, 4, 5, 6, 7, 10, 20 or more times in unmapped sequence reads having a link to an anchor sequence read, or may be a fraction of total unmapped sequence reads having a link to an anchor sequence read (for example, occurrence of the signature mapping k-mer in greater than 5%, 10%, 20%, etc. of unmapped sequence reads having a link to an anchor sequence read).
[0156] In some embodiments, k-mer signatures are weighted based on the frequency of the k-mer signature in the unmapped sequence reads or reference genome. Thus, in some embodiments, a common k-mer signature is a k-mer signature which has greater than a threshold weight value.
[0157] In some embodiments, the stored unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature may be used formapping to the reference genome at the target region. For example, in some embodiments, the method 200 includes a step (not shown) of mapping the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature to the reference genome at the target region.
[0158] In some embodiments, the method 200 includes a step (not shown) of assigning a haplotype to the unmapped sequence reads having a link to an anchor sequence read and which include the common k-mer signature, based on the phasing of linked anchor sequence reads. For example, in some embodiments, if an unmapped sequence reads which includes the common k-mer signature has a link to an anchor sequence read that is phased to a first haplotype, then the linked unmapped sequence read is assigned to that first haplotype.
[0159] The method 200 may end at end block 270.Methods for Generating a Phased Haplotype Assembly from Sequence Reads
[0160] In one aspect, disclosed herein are methods for generating a phased haplotype assembly from sequence reads. In some embodiments, the sequence reads are from genomic DNA. In some embodiments, the phased haplotype assembly includes a target region. In some embodiments, the target region is a putative structural variant region.
[0161] In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. In some embodiments, sequence reads are generated from fragments of the genomic DNA sample bound to a flow cell as described above with respect to the methods of phasing sequence reads.
[0162] In some embodiments, the method includes mapping and phasing sequence reads. In some embodiments, the mapping and phasing steps are performed according to any method as described herein.
[0163] In some embodiments, the method includes identifying a target region based on the first sequence reads. This may be accomplished as described above with respect to the methods of phasing sequence reads.
[0164] FIG. 3 is a flow diagram that schematically illustrates an exemplary method 300 for generating a phased haplotype assembly from sequence reads. In some embodiments, the method 300 is implemented on a processor. The method 300 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or moredisk drives, of a computing system, for example the computing system 4000. When the method 300 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device.Receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome
[0165] As shown in FIG. 3, the method 300 for generating a phased haplotype assembly from sequence reads may start from start block 310. The method 300 may proceed to block 320, wherein an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome is received.
[0166] In some embodiments, at least a subset of the first sequence reads are phased into two or more haplotypes. In some embodiments, phasing of the sequence reads to generate a subset of phased sequence reads is accomplished according to any method as described herein, for example by using signature phasing k-mers.
[0167] In some embodiments, the first sequence reads are sequence reads which have been mapped to the reference genome at a target region
[0168] The method 300 may also be performed in parallel on multiple target regions. For example, the method may include a step (not shown) of binning (e.g., dividing) the alignment of first sequence reads to the reference genome into bins (e.g., subsections of a similar size to one another), and later steps in the workflow may be performed independently within each bin. This parallelization can further improve computational efficiency.Partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads
[0169] The method 300 may proceed to block 330, wherein phased sequence reads are partitioned into two or more haplotype groups based on the phasing of the phased sequence reads.
[0170] For example, in some embodiments, all sequence reads that are haplotagged with (assigned to) a first haplotype are put into a first group, and all sequence reads that are assigned to a second haplotype are put into a separate second group. If there are sequence readsassigned to additional (third, fourth, etc.) haplotypes, then additional (third, fourth, etc.) groups of sequence reads may be created.
[0171] In some embodiments, the step of partitioning sequence reads according to haplotype takes place on a target region-by-target region basis. In some embodiments, the step of partitioning sequence reads takes place across all of the first sequence reads.
[0172] In some embodiments, the method 300 includes a step (not shown) of randomly assigning unphased sequence reads from the first sequence reads to a haplotype group. Thus, in some embodiments, unphased sequence reads are randomly assigned to a haplotype of the two or more haplotypes. For example, if unphased sequence reads remain unphased, unphased sequence reads may be randomly assigned to either haplotype because their sequences may be sufficiently similar such that the two haplotypes are identical at the region that the unphased sequence reads cover. For example, in some embodiments, these randomly assigned unphased sequence reads are added to a randomly assigned haplotype group alongside the phased sequence reads.Generating an assembly graph de novo from the partitioned phased sequence reads
[0173] The method 300 may proceed to block 340, wherein an assembly graph is generated de novo from the partitioned phased sequence reads. In some embodiments, the assembly graph is generated separately for each haplotype group created at block 330 above.
[0174] In the context of block 340, generating an assembly graph de novo includes generating an assembly graph without use of a reference sequence, and based only on the phased sequence reads that are assigned to a haplotype group.
[0175] It will be readily apparent to those of skill in the art that a variety of assembly graphs could be used, for example, De Bruijn graphs or overlap layout consensus (OLC) graphs, for the assembly graph of block 340. Exemplary assembly graphs are described in [Li et al., Comparison of the two major classes of assembly algorithms: overlap-layout- consensus and de-bruijn-graph, Briefings in Functional Genomics, Volume 11, Issue 1, January 2012, Pages 25-37, doi.org / 10.1093 / bfgp / elr035],
[0176] Briefly, in some embodiments, an assembly graph represents graph-based representation of the relationships between different sequence fragments (reads) used to reconstruct the original genome or DNA sequence. An assembly graph includes nodes, whichtypically represent sequence reads, contigs, or k-mers (short subsequences of length k extracted from the reads). The nodes are connected by edges, which represent overlaps or shared sequences between the nodes. For example, when an edge between two nodes indicates that the sequences they represent have a significant overlap, this suggests the sequences of the two nodes are adjacent in the original genome. In an overlap graph, nodes represent reads, and edges indicate overlaps between these reads. In a De Bruijn graph, nodes represent k-mers, and edges connect nodes that correspond to consecutive k-mers in the reads
[0177] It will also be readily apparent to those of skill in the art that a variety of methods could be used to generate an assembly graph based on sequence reads, including de novo assemblers that output an assembly graph and are known in the art, for example, SPAdes [Bankevich et al., SPAdes: A New Genome Assembly Algorithm and Its Applications to Single-Cell Sequencing, J Comput. Biol. 2012 May; 19(5): 455-477] or String Graph Assembler (SGA) [Simpson et al., Efficient de novo assembly of large genomes using compressed data structures, Genome Res. 2012 Mar;22(3): 549-56. doi:10.1101 / gr.126953.111. Epub 2011 Dec 7],Select a guide sequence from a database of SV-containing haplotypes
[0178] The method 300 may proceed to block 350, wherein a guide sequence is selected from a database of structural variant (SV)-containing haplotypes based on the phased sequence reads. In some embodiments, the guide sequence is selected separately for each haplotype group.
[0179] In some embodiments, the database of structural variant (SV)-containing haplotypes includes reference sequences with different structural variants in the target region. In some embodiments, the database of structural variant (SV)-containing haplotypes includes reference sequences from a plurality of individuals.
[0180] In some embodiments, a guide sequence is selected from a database of SV- containing haplotypes based on sequence similarity between the guide sequence and the phased sequence reads. In some embodiments, sequence reads are a guide sequence is selected by analyzing a tag on one or more sequence reads in the initial de novo assembly. In some embodiments, the tag includes flow cell location information. In some embodiments, the tag includes information related to a reference sequence or to a haplotype containing an SV that isidentified within a population. Population-based, multigenome graphs (references constructed from haplotypes associated with a sample as known to those skilled in the art) allow for customizable references for mapping and alignment. These multigenome graphs enable more accurate mapping, alignment, and variant identification as compared with standard, single reference genomes, such as HG002. US20230420082A1, hereby incorporated in its entirety, describes methods for selecting reference structural variant haplotypes.Resolve paths in the assembly graph based on the selected guide sequence
[0181] The method 300 may proceed to block 350, wherein paths in the assembly graph are resolved based on the selected guide sequence.
[0182] In some embodiments, the method includes aligning the assembly graph to the selected guide sequence. In some embodiments, the method includes orienting and / or linking unitigs (sequences without branches) in the assembly graph based on the selected guide sequence.
[0183] For example, a path connecting some nodes of an assembly graph may be ambiguous due to a complex region. For example, in the case of a segmental duplication or a region which includes repetitive sequences, the assembly graph may have ambiguities about which nodes are connected because of sequence similarity between the two segmental duplication areas. In some embodiments, these ambiguous paths in the assembly graph may be resolved based on the guide sequence, for example, a path may be selected which resembles the guide sequence most closely.
[0184] The method 300 may proceed to decision state 370, wherein the system queries whether there are additional haplotype groups. If yes, the method 300 may return to block 340 and the workflow may proceed as described above. If no, the method 300 may proceed to decision state 280, wherein the system queries whether there are additional target regions. If yes, the method 300 may return to block 330, and the workflow may proceed as described above. If no, the method 300 may proceed and may end at end block 390.Methods of Detecting Structural Variants
[0185] According to a further aspect, disclosed herein are methods of detecting structural variants (such as insertions, deletions, inversions, etc.) from a genomic sample.
[0186] In some embodiments, the method includes phasing sequence reads according to the methods described herein, generating one or more assemblies from the first sequence reads, and detecting a structural variant based on the assembly of first sequence reads. For example, the phasing methods disclosed herein improve structural variant detection because they enable a greater number of sequence reads to be assigned to a haplotype. One haplotype may have a structural variant while the other may not. By increasing the ability to separate sequence reads into haplotypes, structural variants can be determined separately for each haplotype. This helps avoid situations where the sequence reads of both haplotypes are analyzed together to detect structural variants, and the presence / absence of a structural variant in one haplotype masks the presence / absence of the structural variant in the other haplotype. In some embodiments, the method includes mapping sequence reads according to the methods described herein, generating one or more assemblies from the first sequence reads, and detecting a structural variant based on the assembly of first sequence reads. For example, the mapping methods disclosed herein improve structural variant detection because linkage information allows a greater number of sequence reads to be identified and stored for mapping, increasing the number of sequence reads the methods and systems can use to determine whether there is a structural variant at a candidate site. At the same time, the k-mer signature filter step ensures that fewer noise reads based on spurious links are identified and stored for mapping, which greatly improves computational efficiency.
[0187] In some embodiments, the method includes generating one or more phased haplotype assemblies from sequence reads according to the methods described herein, aligning contigs from each phased haplotype assembly to a reference sequence; and detecting a structural variant based on the alignment to the reference sequence. For example, the sequence assembly methods disclosed herein improve structural variant detection because sequence reads are assembled separately by haplotype, which can allow the methods and systems to more accurately detect structural variants because the presence / absence of a structural variant in one haplotype masks the presence / absence of the structural variant in the other haplotype. Moreover, the approach of assembling first without a reference sequence allows for assembly that is not biased by a reference sequence, which can be useful in detecting structural variants that are not included in the reference sequence. Then, a reference sequence can be selected andused to further resolve the initial de novo assembly, which can polish the assembly to provide a further boost to the accuracy of structural variant detection.
[0188] For example, in some embodiments, contigs generated using any of the phasing, mapping, or phased haplotype assembly methods herein can be aligned to a reference sequence, for example, at one or more candidate structural variant regions. Then, the methods and systems can compare the aligned contigs to the reference sequence to detect structural variants, including insertions, deletions, inversions. In some embodiments, information about structural variants at two or more candidate structural variant regions (e.g., variant calls) is merged together and stored together in a single file.
[0189] In some embodiments, the method further includes storing information related to structural variant detection at a plurality of target regions in a digital file. In some embodiments, the method further includes storing information related to structural variant detection for two or more haplotypes in a digital file, for example, whether a variant call is heterozygous or homozygous.Systems
[0190] Further disclosed herein are electronic systems. In some embodiments, the system includes a processor configured to perform any of the computer-implemented methods described herein.
[0191] Disclosed herein are systems for phasing sequence reads in a target region. In some embodiments, the system includes a processor configured to perform a method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyzing the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assigning unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0192] Disclosed herein are systems for mapping sequence reads in a target region. In some embodiments, the system includes a processor configured to perform a methodcomprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identifying anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identifying unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyzing unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and storing the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region
[0193] Disclosed herein are systems for generating a phased haplotype assembly from sequence reads. In some embodiments, the system includes a processor configured to perform a method comprising: receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generating an assembly graph de novo from the partitioned phased sequence reads; selecting a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolving paths in the assembly graph based on the selected guide sequence.
[0194] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to perform any of the computer-implemented methods described herein, for example the method 100, the method 200, or the method 300.
[0195] In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprisephased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyze the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assign unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0196] In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identify anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identify unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyze unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and store the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region .
[0197] In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partition phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generate an assembly graph de novo from the partitioned phased sequence reads; select a guide sequence from a database of SV- containing haplotypes based on the phased sequence reads; and resolve paths in the assembly graph based on the selected guide sequence.
[0198] FIG. 4A illustrates a diagram of an environment in which a system for detecting structural variants from a genomic DNA sample can operate in accordance with one or more implementations. The following paragraphs describe the structural variant detection system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 4A illustrates a schematic diagram of a computing system 4000 in which a structural variant application 4106 operates in accordance with one or more implementations. As illustrated, the computing system 4000 includes one or more server device(s) 4102 connected to a user client device 4108, a local device 4118, and a sequencing device 4114 via a network 4112. The network 4112 can comprise any suitable network over which computing devices can communicate.
[0199] As shown in FIG. 4A, the computing system 4000 includes the server device(s) 4102. In various implementations, the server device(s) 4102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 4102 receive various data from the sequencing device 4114, such as data from a sample genome and / or sequence reads. The server device(s) 4102 may also communicate with the user client device 4108. In particular, the server device(s) 4102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 4108.
[0200] As shown, the server device(s) 4102 includes a sequencing application 4110. In general, the sequencing application 4110 analyzes the data (such as call data) received from the sequencing device 4114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 4110 can receive raw data from the sequencing device 4114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 4110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.
[0201] As also shown, the sequencing application 4110 includes the structural variant application 4106. In some embodiments, the structural variant application 4106 can phase sequence reads in a target region. For example, in some embodiments, the structural variant application can receive an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequencereads; analyze the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assign unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
[0202] In some embodiments, the structural variant application 4106 can map sequence reads in a target region. For example, in some embodiments, the structural variant application can receive an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identify anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identify unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyze unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and store the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
[0203] In some embodiments, the structural variant application 4106 can generate a phased haplotype assembly from sequence reads. For example, in some embodiments, the structural variant application can receive an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partition phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generate an assembly graph de novo from the partitioned phased sequence reads; select a guide sequence from a database of SV- containing haplotypes based on the phased sequence reads; and resolve paths in the assembly graph based on the selected guide sequence
[0204] While the sequencing application 4110 has been described as including the structural variant application 4106, other systems or methods may be included within thesequencing application 4110, such as an application to determine single nucleotide polymorphisms (not illustrated).
[0205] Moreover, while the structural variant application 4106 is described being implemented on the server device(s) 4102, as part of the sequencing application 4110, in some implementations, the structural variant application 4106 is implemented by (such as located entirely or in part) on the user client device 4108, the sequencing device 4114, and / or the local device 4118. As mentioned, in some implementations, structural variant application 4106 is implemented by one or more other components of the computing system 4000, such as the sequencing device 4114. In particular, the structural variant application 4106 can be implemented in a variety of different ways across the server device(s) 4102, the network 4112, the user client device 4108, the local device 4118, and the sequencing device 4114.
[0206] As further shown in FIG. 4A, the computing system 4000 includes the user client device 4108. In various implementations, the user client device 4108 can generate, store, receive, and send digital data. In particular, the user client device 4108 can receive the data from the sequencing device 4114. As further illustrated, the user client device 4108 includes a sequencing application 4110. The sequencing application 4110 may be a web application or a native application stored and executed on the user client device 4108 (for example, a mobile application, desktop application, or web application). The sequencing application 4110 can receive data from the sequencing application 4110 and / or structural variant application 4106. For example, the user client device 4108 can receive variant call files and / or alignment files from the sequencing application 4110.
[0207] The sequencing application 4110 can also include instructions that (when executed) cause the user client device 4108 to receive data from the structural variant application 4106 and present data from the sequencing device 4114 and / or the server device(s) 4102. Furthermore, the sequencing application 4110 can instruct the user client device 4108 to display data for variant calls, such as nucleobase calls or an indication of a structural variant. Indeed, the user client device 4108 can display nucleobase call results for a genome sample and / or an indication of a predicted structural variant.
[0208] As further shown in FIG. 4A, the computing system 4000 includes the sequencing device 4114. In various implementations, the sequencing device 4114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 4114analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 4114. More particularly, the sequencing device 4114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 4114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 4112, in some implementations, the sequencing device 4114 bypasses the network 4112 and communicates directly with the user client device 4108.
[0209] As further depicted in FIG. 4A, in some implementations, the server device(s) 4102 includes a distributed collection of servers, where the server device(s) 4102 include several server devices distributed across the network 4112 and located in the same or different physical locations. For instance, the server device(s) 4102 can be implemented, in whole or in part, on the local device 4118. To illustrate, the local device 4118 may implement the sequencing application 4110 and / or the structural variant application 4106. Further, the server device(s) 4102 and / or the local device 4118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0210] The user client device 4108 illustrated in FIG. 4A can include various types of client devices. For example, in some implementations, the user client device 4108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. In various implementations, the user client device 4108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.
[0211] Though FIG. 4A illustrates the components of the computing system 4000 communicating via the network 4112, in certain implementations, the components of computing system 4000 can also communicate directly with each other, bypassing the network 4112. For instance, in some implementations, the user client device 4108 communicates directly with the sequencing device 4114. Additionally, in some implementations, the user client device 4108 communicates directly with the structural variant application 4106 and / or the server device(s) 4102. In some implementations, the user client device 4108 communicates directly with the local device 4118. Moreover, the structural variant application 4106 can access one or more databases housed on or accessed by the server device(s) 4102 or elsewhere in the computing system 4000.
[0212] FIG. 4B is a block diagram of an exemplary server device 4102 that may be used in connection with the computing system 4000 of FIG. 4 A. The server device 4102 may be configured to detect one or more structural variants from a genomic DNA sample. The general architecture of the server device 4102 depicted in FIG. 4B includes an arrangement of computer hardware and software components. The server device 4102 may include many more (or fewer) elements than those shown in FIG. 4B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 4102 includes a processing unit 410, a network interface 420, a computer readable medium drive 430, an input / output device interface 440, a display 450, and an input device 460, all of which may communicate with one another by way of a communication bus. The network interface 420 may provide connectivity to one or more networks or computing systems. The processing unit 410 may thus receive information and instructions from other computing systems or services via a network. The processing unit 410 may also communicate to and from memory 470 and further provide output information for an optional display 450 via the input / output device interface 440. The input / output device interface 440 may also accept input from the optional input device 460, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0213] The memory 470 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 410 executes in order to implement one or more embodiments. The memory 470 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 470 may store an operating system 472 that provides computer program instructions for use by the processing unit 410 in the general administration and operation of the server device 4102. The memory 470 may store a reference genome 473, such as for use by the sequencing application 4110. The memory 470 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0214] For example, in one embodiment, the memory 470 includes a sequencing application 4110, which may include a structural variant application 4106. The structural variant application 4106 can perform the methods disclosed herein. In addition, memory 470 may include or communicate with the data store 490 and / or one or more other data stores thatstore one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference genomes.
[0215] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.
[0216] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.
[0217] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++, Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.
[0218] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintainedin another location relative to where the data is being produced, such as that provided by a third party service provider.
[0219] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.
[0220] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In some embodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal locationto the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.
[0221] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0222] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.
[0223] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.
[0224] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Electronic Files
[0225] The methods and systems described herein can include a step of storing information generated by the methods and systems, in an electronic file, or retrieving information from an electronic file. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.
[0226] In some embodiments, the electronic file includes information related to a co-proximity property, e.g., links as described herein, where reads that are nearby within an individual’s genome are also nearby on the flow cell surface. This property enables improved mapping, variant calling, and phasing, among other enhancements. The following is a list of terms that are used to describe data in Tables 1 and 2 below.
[0227] Template molecule: A template molecule is a long DNA molecule (e.g., at least 500 bp) from either a standard or high molecular weight extraction that is loaded into the sequencing cartridge for on-flow cell tagmentation.
[0228] Proximal reads: Proximal reads are reads from the same template molecule that are from clusters located near one another on the flow cell surface.
[0229] Proximity group: A proximity group is a set of reads that are likely to have have been derived from the same template molecule.
[0230] Flow cell proximity: The flow cell proximity is the distance between DNA nanowells of reads within a proximity group, measured either in common units of measurementincluding nanometers, millimeters, inches, or in flow cell units. Genomic proximity: Genomic proximity is the genomic distance between reads within a proximity group.
[0231] Template length: Genomic distance (e.g., distance on a reference genome, measured in bp) spanned by a proximity group from the same original template.
[0232] Co-proximity rate: Percentage of all reads that are co-proximal with at least one other read.
[0233] Co-proximity quality: A Phred-scaled quality score that estimates the probability that two reads derive from the same original template molecule.
[0234] In some embodiments, the method includes a step of storing sequence reads or other sequence information in a BAM file or a VCF file. In some embodiments, the BAM file has tags and / or VCF file has fields as described below. In some embodiments, metrics for these tags are stored in a separate CSV file with the definitions provided.
[0235] Table 1 describes new tags for BAM files. In some embodiments, the BAM file also includes other tags, for example debug tags (xn, xp, xf, xa, xb, Id, ad, bd, ws, wm, and pq)-Table 1 : New BAM tags
[0236] In some embodiments, the VCF file includes fields related to variant calling and haplotype information (for example, Whatshap phase).
[0237] In some embodiments, the method includes creating an electronic file that includes one or more of the following metrics:Table 2: Description of MetricsEXAMPLES
[0238] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1
[0239] The following example relates to an exemplary software implementation of a structural variant detection method incorporating the methods described herein.
[0240] In this example, haplotagged alignments of the reads to a reference (e.g. in bam format) is taken as an input and large (genotyped) variant calls are produced (in vcf format). The alignments could be initially haplotagged using multiple approaches, for example, by explicitly calling and phasing small variants and using tools like Whatshap, Illumina ® DRAGEN, or implicitly using heterozygous small variants and a catalog / database of multiple related / unrelated samples. The phasing blocks and haplotags are also extended using the linking information based on the geographic location of clusters on a flow cell.
[0241] Broadly, the method has the following steps:Index all “recruitable ” reads
[0242] First, sequence reads that could be coming from a potential insertion in the sample genome are indexed. These sequence reads can be identified by mapping the reads to the reference genome. Any read that is either unmapped or mapped with a low mapping quality (e.g. a MAPQ score less than a threshold; e,g, 60) is marked as “recruitable”.
[0243] A sequence index storing sequence reads is created. The sequence index includes flow cell location information and is able to answer queries for reads in the flow cell neighborhood of a given read. To accomplish this, the sequence index includes K-D trees storing X,Y coordinates on the flow cell as the key.
[0244] An additional data-structure, for example in the data store 490, holds the k- mer signatures of the “recruitable” reads so that the k-mer signature filter can later be applied while recruiting reads from this index. The additional data structure is a hash-map with readnames based keys storing such signatures. Any kmer-derivative (such as minimizers, syncmers, strobmers), optionally more than one, can serve as the k-mer signatures. Additionally, some of the k-mers are excluded from being a signature like those consisting of only homopolymers.
[0245] Since the decision of whether a read is recruitable and should be indexed depends only on its own alignment record, this step can be parallelized.Idenlify candidate SV regions
[0246] Candidate SV regions may be identified with putative SVs based on “abnormal” (split reads, improper pairs, abnormal template length etc.) alignments.
[0247] This step can be performed in parallel on different subsections of the alignment.
[0248] For each SV candidate region, in parallel:
[0249] Identify recruiter regions.
[0250] These regions are the upstream and downstream flanking regions of the candidate locus. In some embodiments, the size of these flanking regions is in the range of 20k base pairs.
[0251] Propagate phasing information based on signature phasing k-mers in the candidate and its recruiter regions.
[0252] K-mers of the already haplotagged reads in a bin, which appear in only one haplotype serve as the “differentiating signature”. An unphased read containing one of the differentiating signatures are assigned the same phasing information and haplotag as the “source” haplotagged read. This step can also operate in parallel on independent bins.
[0253] FIG. 5 illustrates how this step would be able to divide the reads into haplotypes, especially those around SVs. FIG. 5 shows reads 501 mapped to region around an SV 502. There are no HET variants in the vicinity of the SV that could phase the reads involved in the SV, reads 503 (shown here as with slanted bar in the beginning to indicate that the SV- related sequence might be soft-clipped). A far-away heterozygous variant induced phasing of reads at that site into two haplotypes. Reads 504 (including read Rl) are phased to a first haplotype, and reads 505 (including read R2) are phased to a second haplotype. Some of these reads might be linked to the reads in the SV based on flow cell cluster location. Here, for example, read Rl has a link to read R4 and read R2 has a link to read R5. This causes read R4 to have same phase / haplotag as read Rl and read R5 to have same phase / haplotag as read R2. Analysis of the kmers of R4 and R5 shows that a k-mer is different between them, k-mer 520, which can be used as a differentiating k-mer. The remaining unphased reads R3, R6, R7, R8 (shown with dashed lines) are haplotagged depending on whether they have the differentiatingk-mer 520. Note that the reads in the middle, reads 507, do not have any differentiating k-mer. In this case, reads 507 are randomly assigned to one of the haplotypes.
[0254] Recruit reads for mapping.
[0255] The reads mapped “reliably” (e.g. with mapping quality greater than a MAPQ threshold, e g., MAPQ 50) in the recruiter regions (“recruiter reads”) then pull / recruit reads from the recruitable indexed reads. The systems and methods can query the sequence index to identify unmapped sequence reads that are within a predetermined distance of a recruiter read.
[0256] Amongst these neighbor reads that the index returns, a k-mer signature filter is applied. With this kmer-based filter, an umapped read must share its k-mer signature with more than a threshold number of other unmapped sequence reads returned by the query, for it to be “recruited,” e.g., identified and stored for mapping. Additionally, k-mer signatures may be assigned scores or weights, for example, depending on the frequency of the k-mer signatures in the index and / or sample or reference genome. K-mer signature scoring can be used to ensure that sequence reads that are not proximate on the flow cell to a “recruiter read” just by chance (e.g. sequence reads from a mobile element not found on the reference genome) are identified and stored for mapping.
[0257] A “recruited” read will get assigned the phasing / haplotag information of its recruiter read.
[0258] FIG. 6 below depicts the working of this step. For a candidate region 601 with structural variants 602, 603, there is an upstream recruiter region 610 and downstream recruiter region 611. Reads reliably mapped in a recruiter region (recruiter reads 620, only upstream shown; downstream works in the same way) will pull / fetch their flow cell neighbors from the index 630 of recruitable reads 635. Of the pulled neighbors 640, those that share their k-mer signatures (shown with ovals) with at least N (here, 3) others from this collection will be recruited (recruited reads 650.
[0259] Prepare reads for assemblies.
[0260] The reads (directly) mapped in the candidate locus as well as the recruited reads for that locus, along with their mates, are partitioned in two haplotypes depending on their haplotags.
[0261] Unphased (direct) reads can be randomly assigned to one haplotype since the k-mer propagation indicates that they are possibly from a region which has no difference between two haplotypes. Unphased recruited reads could optionally undergo a round of phasing based on signature phasing k-mers as described above.
[0262] Perform haplotype-resolved assemblies of the phased reads.
[0263] Each of the two sets of reads is assembled separately. Any short reads de novo assembler that outputs an assembly graph (e.g. Spades) can be used for this step.
[0264] A guide sequence is then chosen from a database of SV containing haplotypes that is closest to the reads being assembled in terms of the sequence.
[0265] The guide can then be used to resolve paths in the assembly graph. It could either be done by aligning the guide to the assembly graph or by orienting and linking unitigs.
[0266] Detect variants from the assembled contigs.
[0267] Each haplotypes’ contigs can be aligned to the local reference sequence and the alignments can be parsed to detect insertions, deletions, inversions, etc.Merge the variants from multiple candidate regions
[0268] Variants from all candidate regions for each haplotype are merged together. The merged variants from both haplotypes can then be collapsed, encoding the similar variants as homozygous variants.Example 2
[0269] In the following example, sequence reads from real data were phased using different techniques and compared to truth sets.
[0270] FIG. 7 is a screenshot from an Integrative Genomics Viewer (I GV) of one example based on real data showing results of the signature phasing k-mer propagation. Panel 701 and panel 702 are truth haplotypes (derived from the T2T HG002 genome) which have 2 different sized deletions, deletion 711 and deletion 712. Panel 720 is the original haplotagged reads, including Hl -tagged reads 723 and H2-tagged reads 724, which were phased based on flow cell geographic location information and other techniques. Panel 730 is the reads after k- mer based phasing method described in Example 1 in the section entitled “Propagate phasing information based on signature phasing k-mers in the candidate and its recruiter regions,” Hl- tagged reads 733 and H2 -tagged reads 734. Of the Hl -tagged reads 733 and H2 -tagged reads734, sequence reads that were haplotagged by the phasing method described in Example 1 are shown in color while sequence reads that were phased in earlier steps are shown in gray.
[0271] FIG. 7 shows that many of the unphased sequence reads 725 from panel 720 were correctly assigned to the right haplotype in panel 730, while the remaining unphased sequence reads 735 are truly homozygous reads that can be randomly assigned to either haplotype because the sequence does not differ between haplotypes.
[0272] Without the additional phasing shown in panel 730, downstream sequence assembly that segregates sequence reads by phasing produced deletions of same length in both haplotypes, instead of showing the deletion 711 in one haplotype and deletion 712 in the second haplotype.Example 3
[0273] In the following example, the method described in Example 1 was implemented in a prototype and performance was evaluated on real data. Structural variants determined from the methods of Example 1 (using Illumina paired-end short reads) was compared with structural variants determined from each of the following methods were compared: native long reads (two native long reads public VCFs, “Long Reads Method 1” and “Long Reads Method 2”), synthetic long reads, and other methods using Illumina paired-end short reads that does not make use of the linking or phasing information. The method described in Example 1 was also compared with other prototype variations of the method of Example 1, with one or more steps described in Example 1 missing, as shown in Table 3 below. In Table 3 below, kmerpulled j)hased_recruitFiltered PopAlts refers to the method of Example 1.Table 3. LabelsEvaluation Method
[0274] 1000 SV-containing regions were randomly selected for HG002 and NIST1.0 truthset was used for benchmarking using Truvari 4.0. Truvari was run with different stringency settings. In each setting, a match was found for SVs (>50 bp while some gave a chance for SVs >35bp to participate in a match without being counted towards FNs;) within a distance of 2000bp on reference. Sequence and size similarity thresholds were changed for different stringencies:•
[0275] seq_len: Sequence similarity and size similarity over 70% to be a match.•
[0276] NOseq_len: Sequence similarity was NOT checked; only size similarity should be over 70% to be a match.•
[0277] NOseq_NOlen: Both sequence similarity and size similarity were not checked.
[0278] The subset of 1000 regions had roughly the same SV profile as that of the whole genome. FIG. 8 is a heatmap that depicts the differences between SVs in various categories in the subset vs those in the whole genome for HG002. In FIG. 8, the Y axis marks the size of SVs and X axis marks the type of SVs (ins: Insertion, del: Deletion, tr: tandem repeats, ntr: outside tandem-repeats).Results
[0279] FIG. 9 is a bar plot of Recall scores for various VCFs in 3 different stringency settings: NOseq_NOlen (Breakend is correct); NOseq_len (In addition, SV lengthis > 70% correct); seq_len (In addition, the sequence of a SV is >70% correct). FIG. 10 is a bar plot of fl scores (measuring precision and recall) for various VCFs in 3 different stringency settings: NOseq_NOlen (Breakend is correct); NOseq_len (In addition, SV length is > 70% correct); seq_len (In addition, sequence of an SV is >70% correct).
[0280] FIG. 9 shows that kmerpulled j)hased recruitFiltered PopAlts (hybrid haplotype assemblies from k-mer propagated phased reads with reads recruited incorporating noise reduction) approach resulted in the best recall with respect to all short reads methods as well as synthetic long reads. In FIGS. 9-10, there is still a gap between native long reads and the method of Example 1 , which uses Illumina® short reads and flow cell cluster location linking. In FIG. 10, the lower margin in the Fl score is due to a gap in precision, which should also be managed with the incorporation of a validation module.
[0281] Next, recall was stratified by INS / DEL (insertion / deletion) and TR / NTR (inside / outside Tandem Repeats) for the most strict (seq len) and most lenient (NOseq_NOlen) settings. This data is shown in the bar plots of FIGS. 11-12.
[0282] FIGS. 11-12 show that there is better sequence resolution of all categories using the method of Example 1 in comparison to all short-read-based methods (and also synthetic long reads), kmerpulled j)hased recruitFiltered PopAlts (the method of Example 1) is on-par or better in DELS with respect to native long reads. In breakends resolution, with respect to native long reads, kmerpulled j)hased_recruitFiltered PopAlts is on-par or better for DELs and NTR INS. The only remaining gap to close is -10% TR INS.
[0283] From the data in FIGS. 9-12, the steps making the largest impact in improving recall and Fl are using phased assemblies and introducing guides (from the population SV-containing haplotypes). These two ideas have caused a jump in recall in the vicinity of 10 percentage points each. Using flow cell-based recruitment and kmer-based phasing propagation have resulted in relatively lower bumps in recall (-1 and 2 percentage points, respectively) but kmer-based phasing propagation has a larger impact on retrieving DELs. Although the kmer-based filter is not very impactful for the overall recall (except sequence resolution of NTR INS), the kmer-based filter is the most important factor in improving computational efficiency and speeding up the assembly step (and to keep the running time similar to that of widely-used short reads SV-callers).Other Considerations
[0284] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.
[0285] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.
[0286] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (such as a measurement value) is close to a targeted value, where close can mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.
[0287] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.
[0288] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that variousomissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
[0289] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.
[0290] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.
Claims
WHAT IS CLAIMED IS:
1. A method for phasing sequence reads in a target region, the method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyzing the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assigning unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
2. The method of claim 1, wherein the signature phasing k-mer is included in sequence reads phased to the first haplotype and is not included in sequence reads phased to a second haplotype of the two or more haplotypes.
3. The method of claim 1 or claim 2, wherein the method comprises generating sequence reads from nucleic acids.
4. The method of claim 3, wherein generating sequence reads comprises fragmenting the genomic DNA sample on a flow cell into a plurality of sequence fragments.
5. The method of claim 4, wherein the phased sequence reads are initially phased into two or more haplotypes based on their geographic location on the flow cell.
6. The method of any of claims 1-5, wherein the phased sequence reads are initially phased into two or more haplotypes based on read phasing, small variants-based phasing, statistical phasing, or trio phasing.
7. The method of any of claims 1-6, wherein the target region is a putative structural variant region.
8. A method of detecting a structural variant, comprising: phasing sequence reads according to the method of any of claims 1-7; generating one or more assemblies from the first sequence reads; and detecting a structural variant based on the one or more assemblies.
9. A system for phasing sequence reads in a target region, comprising one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyzing the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assigning unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
10. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive an alignment of first sequence reads from a genomic DNA sample to a target region of a reference genome, wherein the first sequence reads comprise phased sequence reads that are phased into two or more haplotypes, and unphased sequence reads; analyze the phased sequence reads to determine a signature phasing k-mer included in a plurality of phased sequence reads, wherein the signature phasing k-mer indicates a sequence read is associated with a first haplotype of the two or more haplotypes; and assign unphased sequence reads in the genomic DNA sample which include the signature phasing k-mer to the first haplotype, thereby phasing unphased sequence reads.
11. A method for mapping sequence reads in a target region, the method comprising:receiving an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identifying anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identifying unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyzing unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and storing the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
12. The method of claim 11, wherein the method further includes assigning a haplotype to the unmapped sequence reads having a link to an anchor sequence read and which include the common k-mer signature, based on the phasing of linked anchor sequence reads.
13. The method of claim 11, wherein storing unmapped sequence reads having a link to an anchor sequence read and which include the common k-mer signature comprises identifying unmapped sequence reads having a link to an anchor sequence read and that share a common k-mer signature with a greater than a threshold number of unmapped sequence reads having a link to an anchor sequence read.
14. The method of claim 11 or claim 13, wherein k-mer signatures are weighted based on the frequency of the k-mer signature in the unmapped sequence reads or reference genome.
15. The method of any of claims 11-14, wherein the method comprises indexing the unmapped sequence reads in a sequence index.
16. The method of claim 13, wherein analyzing unmapped sequence reads and anchor sequence reads comprises querying the sequence index based on the geographicdistance between a sequence cluster comprising the anchor sequence reads and a sequence cluster comprising unmapped sequence reads on a flow cell.
17. The method of any of claims 7-16, wherein the one or more flanking regions comprise an upstream flanking region and a downstream flanking region.
18. A method of detecting a structural variant, comprising: mapping sequence reads according to the method of any one of claims 11-17; generating one or more sequence assemblies from the first sequence reads; and detecting a structural variant based on the one or more assemblies.
19. A system for mapping sequence reads in a target region, comprising one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identifying anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identifying unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyzing unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and storing the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
20. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to:receive an alignment of first sequence reads from a genomic DNA sample to a target region and one or more flanking regions of a reference genome, wherein the alignment comprises mapped sequence reads and unmapped sequence reads; identify anchor sequence reads comprising first sequence reads that are mapped to the one or more flanking regions with a quality score above a first predetermined threshold; identify unmapped sequence reads comprising unmapped first sequence reads and first sequence reads that map to the target region with a quality score below a second predetermined threshold; analyze unmapped sequence reads and anchor sequence reads based on the geographic distance between sequence clusters comprising the anchor sequence reads and unmapped sequence reads on a flow cell, to determine unmapped sequence reads having a link to an anchor sequence read; and store the unmapped sequence reads having a link to an anchor sequence read and which include a common k-mer signature, for mapping to the reference genome at the target region.
21. A method for generating a phased haplotype assembly from sequence reads, the method comprising: receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generating an assembly graph de novo from the partitioned phased sequence reads; selecting a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolving paths in the assembly graph based on the selected guide sequence.
22. The method of claim 21, wherein thee method further comprises randomly assigning unphased sequence reads to a haplotype of the two or more haplotypes.
23. The method of any of claims 21-22, wherein the method comprises selecting a guide sequence from a database of SV-containing haplotypes based on sequence similarity between the guide sequence and the phased sequence reads.
24. A method of detecting a structural variant, comprising: generating one or more phased haplotype assemblies from sequence reads according to the method of any of claims 21-23; aligning contigs from each phased haplotype assembly to a reference sequence; and detecting a structural variant based on the alignment to the reference sequence.
25. A system for generating a phased haplotype assembly from sequence reads, comprising one or more processors having instructions that when executed perform a method comprising: receiving an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partitioning phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generating an assembly graph de novo from the partitioned phased sequence reads; selecting a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolving paths in the assembly graph based on the selected guide sequence.
26. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to:receive an alignment of first sequence reads from the genomic DNA sample to a plurality of target regions of a reference genome, wherein at least a subset of the first sequence reads are phased into two or more haplotypes, and for each target region: partition phased sequence reads into two or more haplotype groups based on the phasing of the phased sequence reads, and for each haplotype group: generate an assembly graph de novo from the partitioned phased sequence reads; select a guide sequence from a database of SV-containing haplotypes based on the phased sequence reads; and resolve paths in the assembly graph based on the selected guide sequence.
Citation Information
Patent Citations
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1
Transposon end compositions and methods for modifying nucleic acids
US20100120098A1
Generating and implementing a structural variation graph genome
US20230420082A1
Determining a mudweight of drilling fluids for drilling through naturally fractured formations
US62636004P0
Interactive techniques for speeding up homomorphic linear equation solving on encrypted data
US62637000P0