Methods and systems for detecting copy number variants
By leveraging linkage information and a correction factor, the method addresses the challenge of accurately determining copy number variants in nucleic acid sequencing, enhancing detection accuracy and reliability.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2026-04-02
AI Technical Summary
Traditional nucleic acid sequencing methods face challenges in accurately determining copy number variants due to loss of connectivity and proximity of sequence fragments, compounded by factors like GC content, sequencing bias, and repetitive sequences, leading to inaccurate estimation of copy numbers.
The method involves generating sequence reads from genomic DNA fragments, determining linkage information based on geographic location on a flow cell, aligning reads to a reference genome, and applying a correction factor to sequencing depth to estimate copy numbers, using metrics such as linking rates and mapping quality scores to improve accuracy.
This approach enhances the detection of copy number variants by correcting for false positives and negatives, particularly in difficult-to-map regions, improving specificity and sensitivity in variant detection.
Smart Images

Figure US2025044665_02042026_PF_FP_ABST
Abstract
Description
ILLINC.846WO / IP-2818-PCT PATENTMETHODS AND SYSTEMS FOR DETECTING COPY NUMBER VARIANTSCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 700486, filed September 27, 2024, the content of which is incorporated by reference in its entirety.BACKGROUNDField
[0002] The present disclosure relates to DNA sequencing systems and methods. In particular, this disclosure relates to systems and methods for detecting copy number variants in a genomic sequence based on sequencing depth. A correction factor based on links between pairs of reads on a sequencing flow cell may be applied to the sequencing depth to determine the copy number variants.Description
[0003] Genomic DNA is often too long to be directly sequenced using modern sequencing technologies. Library preparation is a step performed before genome sequencing to facilitate the sequencing process and ensure accurate and efficient analysis of the genomic DNA. Library preparation involves fragmenting the DNA into smaller, more manageable nucleic acid segments. This fragmentation can be achieved through physical or enzymatic methods. Fragmented DNA allows for more efficient sequencing and enables the reconstruction of the original genome during data analysis. Library preparation also involves attaching adapter sequences to the fragmented DNA. Adapters contain specific sequences that are recognized by the particular sequencing platform. These adapters provide priming sites and identification tags which may be used to prime the strand being sequenced and identify each sequence read to later alignment and identification processes.
[0004] Traditional nucleic acid sequencing methods, and several types of nextgeneration sequencing methods such as Sequencing by Synthesis (SBS), use a shotgun approach to sequence large genomic DNA fragments, called template genomic sequences. Specifically, template genomic sequences are first fragmented in solution into smaller piecesthat are amenable to next-generation sequencing methods on a flow cell. In many cases, the library preparation process results in many copies of the same genomic sequences from multiple cells are fragmented and bound to a flow cell. This results in the same sequences being read multiple times during the sequencing process which results in the SBS process having a particular sequencing depth. The sequencing depth, which is also known as “read depth” or “depth of coverage”, is the number of times the same region of the DNA is read during the sequencing process. For example, in an SBS sequencing run, the same region may have a sequencing depth of 30, meaning that this region was sequenced 30 times during the sequencing run. Increasing the sequencing depth can increase the confidence that a particular region was sequenced accurately and variants determined in the region were not an artifact of the sequencing process.
[0005] One of the difficulties of this approach is that by the time the smaller sequence fragments from the template genomic sequences have been read, knowledge of their connectivity and proximity to each other in the original template genomic sequence is lost.
[0006] One way to estimate copy number at a locus is to sequence genomic DNA, align the sequence reads to a reference, and analyze the sequencing depth at the locus by comparing it to the sequencing depth at other loci. Generally, a lower sequencing depth at the locus compared to other loci can indicate a lower copy number at the locus (for example, due to a deletion) and a higher sequencing depth can indicate a higher copy number at the locus (for example, due to a duplication). However, factors such as differing GC content between reads, sequencing bias, repetitive sequences, and sequence similarity between segmental duplications can complicate the estimation of copy number based on sequencing depth.SUMMARY
[0007] The methods disclosed herein each have several aspects, no single one of which is solely responsible for their desirable attributes. Without limiting the scope of the claims, some prominent features will now be discussed briefly. Numerous other embodiments are also contemplated, including embodiments that have fewer, additional, and / or different components, steps, features, objects, benefits, and advantages. The components, aspects, and steps may also be arranged and ordered differently. After considering this discussion, and particularly after reading the section entitled “Detailed Description”, one will understand howthe features of the devices and methods disclosed herein provide advantages over other known devices and methods.
[0008] Disclosed herein are methods for detecting copy number variants in a nucleic acid segment from a genomic DNA sample. In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detecting a copy number variant in the genomic DNA based on the estimated copy numbers.
[0009] In some embodiments, a higher corrected sequencing depth in a first region as compared to other regions corresponds to a higher estimated copy number, and a lower corrected sequencing depth in the first region as compared to other regions corresponds to a lower estimated copy number.
[0010] In some embodiments, estimating the one or more copy numbers for the nucleic acid segment based on the correction factor comprises: binning the alignment of the sequence reads to the reference genome; determining a depth-per-bin of sequence reads with a mapping quality score above a threshold; adjusting the depth-per-bin for GC content; normalizing the depth-per-bin by a median depth-per-bin; and applying the correction factor to the normalized depth-per-bin to generate a corrected depth-per-bin signal.
[0011] In some embodiments, determining depth-per-bin of sequence reads with a mapping quality score above a threshold comprises, for each bin, counting the number of sequence reads with a MAPQ score of 3 or greater which align to the bin. In some embodiments, the method further comprises segmenting the corrected depth-per-bin signal. In some embodiments, the method further includes converting segmented depth-per-bin signals to integer copy numbers to determine a copy number genotype.
[0012] In some embodiments, the correction factor is determined based on one or more of the following: a fraction of sequence reads that do not have a link to another sequence read and are aligned within a first bin of the reference genome with confidence above apredetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold.
[0013] In some embodiments, the fraction of sequence reads that do not have a link to another sequence read and that are aligned within the first bin of the reference genome with confidence above a predetermined threshold, is estimated based on a fraction of sequence reads that have a mapping quality score above a predetermined threshold, over all sequence reads within the first bin that do not have a link to another sequence read. In some embodiments, sequence reads that have a mapping quality score above a predetermined threshold are sequence reads that have a MAPQ score of 3 or greater.
[0014] In some embodiments, the linking rate is estimated based on all of the sequence reads, or a random subset thereof. In some embodiments, the linking rate is estimated per bin.
[0015] In some embodiments, the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold comprises a fraction of sequence reads which have a MAPQ score below the predetermined threshold, over all linked sequence reads within the first bin. In some embodiments, the predetermined threshold for the MAPQ score is 3.
[0016] In some embodiments, the correction factor is determined using primary and secondary alignments of each sequence read to the reference genome. In some embodiments, one or more of the linking rate and the fraction of linked sequence reads that are mapped to the reference genome with confidence below a predetermined threshold, are determined using primary and secondary alignments of each sequence read to the reference genome.
[0017] In some embodiments, the correction factor is determined based on each of: a fraction of sequence reads that do not have a link to another sequence read and are aligned within a first bin of the reference genome with confidence above a predetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads thatare mapped to the reference genome within the first bin with confidence below a predetermined threshold.
[0018] In some embodiments, the linkage information is determined based on location information of clusters of the fragments on the flow cell. In some embodiments, the linkage information comprises information linking a first sequence read from a first cluster to a second sequence read from a second cluster based on a distance or proximity between the first and second clusters in the flow cell. In some embodiments, the location information comprises a first spatial coordinate and a second spatial coordinate in a cartesian coordinate system. In some embodiments, determining linkage information between the sequence reads comprises: obtaining location information for each cluster on a substrate of the flow cell, and assigning sequence reads to target polynucleotides using the obtained location information. In some embodiments, assigning sequence reads to the target polynucleotides using the obtained location information comprises: determining distances between each cluster, and using the determined distances to assign sequence reads to a specific target polynucleotide. In some embodiments, the sequence reads comprise paired end sequence reads.
[0019] Further disclosed herein are methods for determining one or more copy numbers for a nucleic acid segment from a genomic DNA sample. In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information.
[0020] Further disclosed herein are methods for detecting a pathogenic copy number variant in a nucleic acid segment from a genomic DNA sample. In some embodiments, the methods include determining one or more copy numbers for a nucleic acid segment from a genomic DNA sample, and detecting a pathogenic copy number variant based on the one or more copy numbers. In some embodiments, the plurality of locations in the reference genome comprise one or more of an IKBKG gene, an IKBKGP1 pseudogene, a RCCX locus, a STK19 gene, a STK19B pseudogene, a C4A gene, a C4B gene, a CYP21A1P pseudogene, a CYP21A2 gene, a TNXA pseudogene, a TNXB gene, a STRC gene, a STRCP1 pseudogene, a PMS2 gene,a PMS2CL pseudogene, a OTOA gene, a OTOAP1 pseudogene, a GBA gene, a GBAP1 pseudogene, a HYDIN gene, or a HYDIN2 pseudogene.
[0021] In some embodiments, the methods further comprise creating an electronic file comprising one or more copy numbers, information related to a copy number variant, or information related to a copy number genotype.
[0022] Further disclosed herein are systems for detecting copy number variants in a nucleic acid segment from a genomic DNA sample. In some embodiments, the systems comprise one or more processors having instructions that when executed perform a method comprising: receiving sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detecting a copy number variant in the genomic DNA based on the estimated copy numbers.
[0023] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable media comprise a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimate one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detect a copy number variant in the genomic DNA based on the estimated copy numbers.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Features of examples of the present disclosure will become apparent by reference to the following detailed description and drawings, in which like reference numerals correspond to similar, though perhaps not identical, components. For the sake of brevity,reference numerals or features having a previously described function may or may not be described in connection with other drawings in which they appear. In addition to the features described herein, additional features and variations will be readily apparent from the following descriptions of the drawings and exemplary embodiments. It is to be understood that these drawings depict typical embodiments, and are not intended to be limiting in scope.
[0025] FIG. 1A is a flow diagram that schematically illustrates an exemplary method for detecting copy number variants in a nucleic acid segment from a genomic DNA sample.
[0026] FIG. IB is a flow diagram that schematically illustrates a process of estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor determined based on linkage information, that may take place within the method of FIG. 1A.
[0027] FIG. 2 is a schematic representation of how estimates of correction factors may be skewed when copy number variants are present.
[0028] FIG. 3 A is a block diagram of an exemplary sequencing system that may be used to perform the disclosed methods.
[0029] FIG. 3B is a block diagram of an exemplary computing device that may be used in connection with the exemplary sequencing system of FIG. 3 A.
[0030] FIG. 4A is a graphic showing a baseline alignment of sequence reads to a reference genome and an alignment that uses linkage information, for the IKBKG segmental duplication.
[0031] FIG. 4B is a graphic showing a baseline alignment of sequence reads to a reference genome and an alignment that uses linkage information, for the RCCX locus which covers the CYP21A2 gene.
[0032] FIG. 4C is a graphic showing a baseline alignment of sequence reads to a reference genome and an alignment that uses linkage information, for the STRC segmental duplication.
[0033] FIG. 4D is a graphic showing a baseline alignment of sequence reads to a reference genome and an alignment that uses linkage information, for the PMS2 segmental duplication.
[0034] FIG. 5 A is a line graph showing normalized sequencing depth at the IKBKG locus for an alignment that used linkage information.
[0035] FIG. 5B is a line graph showing normalized sequencing depth at the STRC locus for an alignment that used linkage information.
[0036] FIG. 5C is a line graph showing normalized sequencing depth at the RCCX locus for an alignment that used linkage information.
[0037] FIG. 6A is a line graph showing normalized, GC-corrected sequencing depth at the STRC and STRCP1 loci for a baseline alignment that did not use linkage information.
[0038] FIG. 6B is a line graph showing normalized, GC-corrected sequencing depth at the STRC and STRCP1 loci locus for an alignment that used linkage information but not a correction factor.
[0039] FIG. 6C is a line graph showing normalized, GC-corrected sequencing depth at the STRC and STRCP1 loci for an alignment that used linkage information and a correction factor.
[0040] FIGs. 7A and 7B include line graphs showing normalized, GC-corrected sequencing depth at the PMS2 (Fig. 7A) and PMS2CL (Fig. 7B) loci for a baseline alignment that did not use linkage information.
[0041] FIGs. 7C and 7D include line graphs showing normalized, GC-corrected sequencing depth at the PMS2 (Fig. 7C) and PMS2CL (Fig. 7D) loci for an alignment that used linkage information but not a correction factor.
[0042] FIG. 7E and 7F include line graphs showing normalized, GC-corrected sequencing depth at the PMS2 (Fig. 7E) and PMS2CL (Fig. 7F) loci for an alignment that used linkage information and a correction factor.
[0043] FIGs. 8A and 8B include graphs showing normalized, GC-corrected sequencing depth at the OTOA (Fig. 8A) and OTOAP1 (Fig. 8B) loci for a baseline alignment that did not use linkage information.
[0044] FIGs. 8C and 8D include line graphs showing normalized, GC-corrected sequencing depth at the OTOA (Fig. 8C) and OTOAP1 (Fig. 8D) loci for an alignment that used linkage information but not a correction factor.
[0045] FIGs. 8E and 8F include line graphs showing normalized, GC-corrected sequencing depth at the OTOA (Fig. 8E) and OTOAP1 (Fig. 8F) loci for an alignment that used linkage information and a correction factor.
[0046] FIG. 9A is a line graph showing normalized, GC-corrected sequencing depth at the IKBKG locus for an alignment that used linkage information and a correction factor based on primary alignments only.
[0047] FIG. 9B is a line graph showing normalized, GC-corrected sequencing depth at the IKBKG locus for an alignment that used linkage information and a correction factor based on primary alignments and secondary alignments.DETAILED DESCRIPTION
[0048] The foregoing and other aspects of the present disclosure will now be described in more detail with respect to the description and methodologies provided herein. This description is not intended to be a detailed catalogue of all the ways in which the embodiments of the present disclosure may be implemented, or of all the features that may be added to the present disclosure. For example, features illustrated with respect to one embodiment may be incorporated into other embodiments, and features illustrated with respect to a particular embodiment may be deleted from that embodiment. In addition, numerous variations and additions to the various embodiments suggested herein, which do not depart from the instant disclosure, will be apparent to those skilled in the art in light of the instant detailed description, figures and claims. Hence, the following specification is intended to illustrate some particular embodiments, and not to exhaustively specify all permutations, combinations and variations thereof.
[0049] All patents, patent applications, and other publications, including all sequences disclosed within these references, referred to herein are expressly incorporated herein by reference, to the same extent as if each individual publication, patent or patent application was specifically and individually indicated to be incorporated by reference. All documents cited are, in relevant part, incorporated herein by reference in their entireties for the purposes indicated by the context of their citation herein. However, the citation of any document is not to be construed as an admission that it is prior art with respect to the present disclosure.
[0050] In the following detailed description, reference is made to the accompanying drawings, which form a part hereof. In the drawings, similar symbols typically identify similar components, unless context dictates otherwise. The illustrative embodiments described in the detailed description, drawings, and claims are not meant to be limiting. Other embodiments may be utilized, and other changes may be made, without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein, and illustrated in the figures, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.
[0051] Embodiments relate to systems and methods for determining the copy number of repeated segments of a genomic sequence. In many cases, the number of repeated segments of a particular nucleic acid sequence, hereinafter “Copy Number Variants” or “CNVs” in a genome is correlated with specific disease states. CNVs can include repeated segments of all, or part, of a particular gene. In some cases, the copied segment may include several genes. The variation in copy number can be due to increased numbers of copies of the repeated segment, or deletions resulting in fewer numbers of copies of the repeated segment as compared to a reference sequence. These CNVs can result in disease phenotypes in humans. For example, disease states associated with variations in copy number include autosomal dominant diseases such as Potocki-Lupski syndrome, CMT1A, Adult-onset leukodystrophy, 7ql 1.23 duplication syndrome, Spinocerebellar ataxia type 20 and Microduplication 22ql 1.2. Copy number variation has also been associated with autosomal recessive disease states, such as Familial juvenile nephronophthisis, Gaucher disease, beta-thalassemia, and Spinal muscular atrophy.
[0052] During massively parallel sequencing processes, such as those found in Next Generation Sequencing (NGS) systems, relatively short reads of 100-500 nucleotides are typically generated. During this process, sample nucleic acids may be fragmented by tagmentation of the sample directly on the flow cell, in some embodiments. In an example process, transposons are linked to the flow cell, and when contacted by sample nucleic acids may fragment the sample nucleic acids resulting in many fragments from the same nucleic acid molecule being generated. This example process may result in fragments from the same nucleic acid molecule being bound to the flow cell in geographically adjacent or nearbylocations. These fragments then give rise to nucleic acid clusters at nearby locations in flow cell proximity to one another on the flow cell, and these nucleic acid clusters will participate in sequencing reactions to generate reads or read pairs. Therefore, reads or read pairs that are generated by sequencing these fragments which are located nearby, in flow cell proximity, to each other on the flow cell may be determined by the disclosed systems or methods to be “connected” or “linked”, or that there are “links” among the reads or read pairs since they may have derived from the same originating nucleic acid molecule. This type of linking information or connectivity information may enable reconstruction of a nucleic acid molecule by way of grouping together linked read pairs which are bound to the flow cell at geographically nearby locations on the flow cell.
[0053] As discussed above, NGS systems can provide sequences which have been determined at varying sequence depths. Some embodiments relate to methods of determining CNVs based on the depth of sequencing sequence reads and their linkage information. Processes for depth-based copy number variant (CNV) detection generally involve three major steps: binning (e.g., dividing) an alignment into similarly-sized bins (e.g., subsections) and determining the high-MAPQ depth-per-bin; GC-correcting and normalizing the depth in each bin by the median depth-per-bin across all bins; and segmenting the normalized depth-per-bin signal to identify consecutive bins with consistent increases (copy number gain) or decreases (copy number loss) in sequencing depth. In a normal whole genome sequencing (WGS) dataset, difficult-to-map regions are excluded from consideration in the CNV detection process due to lack of reliably mapped reads and consequently unreliable depth signal.
[0054] To improve the accuracy of the depth-based assignments of each sequence read, some embodiments of the present disclosure incorporate linkage information when determining the depth of difficult-to-map regions. In some embodiments, linkage information includes the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule based on the physical location of clusters for each sequence read in a flow cell and on the mapping location on a reference genome. For example, long (greater than 500 base pairs (bp)) DNA fragments may be flowed across a sequencing flow cell and may attach to transposome complexes that are embedded on the surface of the flow cell. The transposome complexes can fragment the long DNA fragments into shorter fragments of less than 500 bp, and attach sequencing adapters. Each shorter fragment can beamplified to form clusters, and the sequence of each shorter fragment can be determined by sequencing the clusters. The physical location of each cluster can be recorded in addition to sequence reads. Then, given the physical location of two clusters on a flow cell and genomic distance when mapped to a reference genome, sequence reads may be linked in that a probability can be determined related to whether the sequence reads derive from the same original long DNA fragment.
[0055] Linkage information may improve the accuracy and efficiency of alignment of sequence reads to a reference sequence, which is useful in determining sequencing depth at various locations of the reference sequence, and particularly where sequences are repeated in the genome so that assignment of a particular sequence read to a specific region of the genome may be difficult. As mentioned above, sequencing depth can be used to estimate copy numbers. For example, if a sample has a significantly higher sequencing depth at one location compared to other locations, it may indicate that there is a higher copy number at that location. Conversely, a lower sequencing depth at one location compared to others may indicate a lower copy number. By improving the alignment of sequence reads to the reference genome, linkage information can help with methods of estimating copy numbers and detecting copy number variants. Linkage information can be leveraged to align sequence reads to regions of the reference genome that are difficult to map, and leads to improved accuracy in variant detection capabilities over such regions.
[0056] However, in some cases, even after aligning based on linkage information, sequencing depth may not accurately reflect whether the sample has a copy number variant, leading to false negatives or false positives. This may be particularly true through segmental duplication areas such as a gene and pseudogene. For example, it may not be possible to determine a link from every read pair (in the case of paired-end sequencing) to another read pair. Furthermore, some links may be between two sequence reads that are both in difficult-to- map regions (such as reads linked to other reads that also cannot be uniquely placed), thus in this case, the alignment of the two linked reads would still have low confidence. The rate of read mapping in difficult-to-map regions is dependent both on the linking rate (percentage of reads with links) of the specific dataset, which is influenced by multiple assay factors, and by the local genomic context. Therefore, while difficult-to-map regions in a dataset may have significantly higher coverage by sequences reads which are mapped with a high degree ofcertainty (high MAPQ), the high-MAPQ depth in such regions may be significantly lower than what is observed in uniquely mappable regions with relatively low MAPQ. Further, the relative drop in high-MAPQ depth relative to unique regions will be dependent on the local genomic context.
[0057] Given that CNV calling has the underlying assumption that regions without CNVs will have similar sequencing depth, the lower high-MAPQ depth in difficult-to-map regions in a dataset may make it difficult to detect copy number variants in such regions, even though linkage information allows for better alignment of sequence reads to a reference genome. Importantly, many pathogenic CNVs occur in difficult-to-map regions of the genome and currently either cannot be detected, or need specialized callers for detection in regular short read (less than 500 bp) WGS. For example, if sequence reads are incorrectly aligned to one copy of a segmental duplication instead of another, the sequencing depth may appear artificially high at one copy and artificially low at another copy. The sequencing depth at one location may lead to an inaccurate determination that there is a duplication at one location, as an example of a false positive. As another example, if sequence reads do not align with sufficient confidence to a region, the resulting sequencing depth may appear as a deletion at the region (another false positive CNV). As an example of a false negative, a sequencing depth that appears consistent with a copy number of 2 may mask the presence of a deletion at a location.
[0058] Additional techniques are presented herein which improve the sensitivity and specificity of methods of detecting copy number variants. Embodiments of the present disclosure relate to methods for detecting copy number variants in a nucleic acid segment from a genomic DNA sample. Sequence reads may be generated from fragments of the genomic DNA sample bound to a flow cell. Linkage information between the sequence reads may be determined based on the geographic location of each fragment on the flow cell, and the sequence reads may be aligned to a reference genome using the linkage information.
[0059] Next, the copy number for the nucleic acid segment may be estimated to determine a CNV for the segment, based on a sequencing depth and by applying a correction factor to the sequencing depth to result in corrected sequencing depth for the particular segment. With the corrected sequencing depth, bins (e.g., subsections of the alignment) that cover difficult-to-map regions may advantageously provide CNV detection data, allowing fordetection of CNVs over many difficult-to-map regions without introducing false discoveries due to assay-related shifts in depth. Thus, a copy number variant may be detected in the genomic DNA based on the estimated copy numbers. The correction factor can correct both for artificially low sequencing depth at a locus due to sequence reads failing to map or mapping elsewhere, and for artificially high sequencing depth due to sequence reads incorrectly mapping to the locus. Thus, the correction factor can improve both specificity and recall for CNV detection.
[0060] In some embodiments, the correction factor is determined based on one or more metrics related to linkage information, for example, related to an analysis of the geographic proximity of clusters of sequence reads on a flow cell. Some or all of the components are determined per bin (e.g., subsection of the alignment). First, in some embodiments, the correction factor is based on the fraction of reads in the bin that are aligned uniquely even without having a link to another sequence read based on flow cell proximity. For example, this can be determined as the fraction of “unlinked” sequence reads (sequence reads that do not have a link to another sequence read, or do not have a link with a quality score above a threshold) that are aligned within a first bin of the reference genome with confidence above a predetermined threshold. Second, in some embodiments, the correction factor is based on the proportion of sequence reads which do have a link to another sequence read based on flow cell proximity, e.g., a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold. Third, in some embodiments, the correction factor is based on the fraction of sequence reads which do have a link to another sequence read based on flow cell proximity but nevertheless do not have a sufficiently high mapping quality score, e.g., the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold. In some embodiments, the correction factor is based on a combination of these three metrics, including all three metrics.
[0061] In some embodiments, the correction factor is determined according to following equation:where H is the fraction of unlinked sequence reads that are aligned within a first bin of the reference genome with MAPQ that is higher than a first predetermined threshold (e.g., 3 or greater), L is the linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold (determined per bin), and P is the fraction of linked sequence reads that are mapped to the reference genome within the first bin with a MAPQ below a predetermined threshold (e.g., below 3).Definitions
[0062] Although the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate understanding of the presently disclosed subject matter.
[0063] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques or substitutions of equivalent techniques that would be apparent to one of ordinary skill in the art.
[0064] As used herein, the terms “a” or “an” or “the” may refer to one or more than one. For example, “a” marker can mean one marker or a plurality of markers.
[0065] As used herein, the term “about,” when used in reference to a measurable value such as an amount of mass, dose, time, temperature, and the like, is meant to encompass variations of 20%, 10%, 5%, 1%, 0.5%, or even 0.1% of the specified amount.
[0066] As used herein, the term “and / or” refers to and encompasses any and all possible combinations of one or more of the associated listed items, as well as the lack of combinations when interpreted in the alternative (“or”).
[0067] Throughout this specification, unless the context requires otherwise, the words “comprise,” “comprises,” and “comprising” will be understood to imply the inclusion of a stated step or element or group of steps or elements but not the exclusion of any other step or element or group of steps or elements.
[0068] As used herein, the term “consists essentially of’ (and grammatical variants thereof), as applied to the compositions and methods of the present disclosure, means that the compositions / methods may contain additional components so long as the additional components do not materially alter the composition / method.
[0069] The term “nucleic acid” or “polynucleotide” refers to a deoxyribonucleotide or ribonucleotide polymer in either single- or double-stranded form, and unless otherwise limited, encompasses known analogs of natural nucleotides that hybridize to nucleic acids in manner similar to naturally occurring nucleotides, such as peptide nucleic acids (PNAs) and phosphorothioate DNA. Unless otherwise indicated, a particular nucleic acid sequence includes the complementary sequence thereof. Nucleotides include, but are not limited to, ATP, dATP, CTP, dCTP, GTP, dGTP, UTP, TTP, dUTP, 5-methyl-CTP, 5-methyl-dCTP, ITP, diTP, 2-amino-adenosine-TP, 2-amino-deoxyadenosine-TP, 2-thiothymidine triphosphate, pyrrolo-pyrimidine triphosphate, and 2-thiocytidine, as well as the alphathiotriphosphates for all of the above, and 2'-O-methyl-ribonucleotide triphosphates for all the above bases. Modified bases include, but are not limited to, 5-Br-UTP, 5-Br-dUTP, 5-F-UTP, 5-F-dUTP, 5-propynyl dCTP, and 5-propynyl-dUTP.
[0070] As used herein, the term "fragment," when used in reference to a first nucleic acid, is intended to mean a second nucleic acid having a part or portion of the sequence of the first nucleic acid. Generally, the fragment and the first nucleic acid are separate molecules. The fragment can be derived, for example, by physical removal from the larger nucleic acid, by replication or amplification of a region of the larger nucleic acid, by degradation of other portions of the larger nucleic acid, a combination thereof or the like. The term can be used analogously to describe sequence data or other representations of nucleic acids. As used herein, the term "haplotype" refers to a set of alleles at more than one locus inherited by an individual from one of its parents. A haplotype can include two or more loci from all or part of a chromosome. Alleles include, for example, single nucleotide polymorphisms (SNPs), short tandem repeats (STRs), gene sequences, chromosomal insertions, chromosomal deletions etc. The term "phased alleles" refers to the distribution of the particular alleles from a particular chromosome, or portion thereof. Accordingly, the "phase" of two alleles can refer to a characterization or representation of the relative location of two or more alleles on one or more chromosomes.
[0071] As used herein, the term "nucleotide sequence" or simply “sequence” is intended to refer to the order and type of nucleotide monomers in a nucleic acid polymer. A nucleotide sequence is a characteristic of a nucleic acid molecule and can be represented in any of a variety of formats including, for example, a depiction, image, electronic medium, series of symbols, series of numbers, series of letters, series of colors, etc. The information can be represented, for example, at single nucleotide resolution, at higher resolution (e.g. indicating molecular structure for nucleotide subunits) or at lower resolution (e.g. indicating chromosomal regions, such as haplotype blocks). A series of "A," "T," "G," and "C" letters is a well-known sequence representation for DNA that can be correlated, at single nucleotide resolution, with the actual sequence of a DNA molecule. A similar representation is used for RNA except that "T" is replaced with "U" in the series.
[0072] As used herein, the term “reference genome” or “reference sequence” refers to any particular known genome sequence, whether partial or complete, of any organism or virus which may be used to reference identified sequences from a subject. For example, a reference genome used for human subjects as well as many other organisms is found at the National Center for Biotechnology Information at ncbi.nlm.nih.gov. In various embodiments, the reference sequence is significantly larger than the reads that are aligned to it. For example, it may be at least about 100 times larger, or at least about 1000 times larger, or at least about 10,000 times larger, or at least about 105times larger, or at least about 106times larger, or at least about 107times larger. In one example, the reference sequence is that of a full-length genome. Such sequences may be referred to as genomic reference sequences. Other examples of reference sequences include genomes of other species, such as of control organisms as disclosed herein, as well as chromosomes, sub-chromosomal regions (such as strands), etc., of any species. In various embodiments, the reference sequence is a consensus sequence or other combination derived from multiple individuals. However, in certain applications, the reference sequence may be taken from a particular individual. Examples of reference genomes include GRCh38 from the Genome Reference Consortium.
[0073] The term “nucleic acid sample” herein may refer to a sample, typically derived from one or more biological fluids, cells, tissues, organs, or organisms, comprising a nucleic acid or a mixture of nucleic acids comprising at least one nucleic acid sequence that is to be screened for copy number variation. In certain embodiments the nucleic acid samplecomprises at least one nucleic acid sequence whose copy number is suspected of having undergone variation. Such samples may include, but are not limited to sputum / oral fluid, amniotic fluid, blood, a blood fraction, or fine needle biopsy samples (such as surgical biopsy, fine needle biopsy, etc.), urine, peritoneal fluid, pleural fluid, and the like. Although the sample is often taken from a human subject (such as a patient), the sample may be from any mammal, including, but not limited to dogs, cats, horses, goats, sheep, cattle, pigs, etc. The sample may be used directly as obtained from the biological source or following a pretreatment to modify the character of the sample. For example, such pretreatment may include preparing plasma from blood, diluting viscous fluids and so forth. Methods of pretreatment may also involve, but are not limited to, filtration, precipitation, dilution, distillation, mixing, centrifugation, freezing, lyophilization, concentration, amplification, nucleic acid fragmentation, inactivation of interfering components, the addition of reagents, lysing, etc. If such methods of pretreatment are employed with respect to the sample, such pretreatment methods are typically such that the nucleic acid(s) of interest remain in the test sample, sometimes at a concentration proportional to that in an untreated test sample (such as namely, a sample that is not subjected to any such pretreatment method(s)). Such “treated” or “processed” samples are still considered to be biological “test” samples with respect to the methods described herein. A “nucleic acid sample” may also include nucleic acid sequence information stored in a memory, and which was originally obtained from a source such as one or more biological fluids, cells, tissues, organs, or organisms.
[0074] The term “read” or “sequence read” (or sequencing reads) refers to a sequence obtained from a portion of a nucleic acid sample. A read may be represented by a string of nucleotides sequenced from any part or all of a nucleic acid molecule. Typically, though not necessarily, a read represents a short sequence of contiguous base pairs in the sample. The read may be represented symbolically by the base pair sequence (in A, T, C, or G) of the sample portion. It may be stored in a memory device and processed as appropriate to determine whether it matches a reference sequence or meets other criteria. A read may be obtained directly from a sequencing apparatus or indirectly from stored sequence information concerning the sample. In some cases, a read is a DNA sequence of sufficient length (such as at least about 25 bp) that can be used to identify a larger sequence or region, for example, that can be aligned and specifically assigned to a chromosome or genomic region or gene. Forexample, a sequence read may be a short string of nucleotides (such as 20-150 bases) sequenced from a nucleic acid fragment, a short string of nucleotides at one or both ends of a nucleic acid fragment, or the sequencing of the entire nucleic acid fragment that exists in the biological sample. Sequence reads may be obtained by any method known in the art. For example, a sequence read may be obtained in a variety of ways, such as using sequencing techniques or using probes, such as in hybridization arrays or capture probes, or amplification techniques, such as the polymerase chain reaction (PCR) or linear amplification using a single primer or isothermal amplification. Sequence reads can be generated by techniques such as sequencing by synthesis, sequencing by binding, or sequencing by ligation. Sequence reads can be generated using instruments such as MINISEQ, MISEQ, NEXTSEQ, HISEQ, and NOVASEQ sequencing instruments from Illumina, Inc. (San Diego, CA).
[0075] As used herein, a “short sequence read” refers to a sequence read of between 50-500 bp, for example, about 50 - 300 bp, and includes paired end sequence reads.
[0076] As used herein, a “long sequence read” refers to a sequence read of more than about 500 bp, for example 500 - 250,000 bp or more. A long sequence read may be obtained from a long-read sequencing technology, or may be synthetically constructed by assembling multiple short sequence reads.
[0077] As used herein, a “file” includes electronic files. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.
[0078] As used herein, the terms “aligned,” “alignment,” or “aligning” refer to the process of comparing a read or tag to a reference sequence and thereby determining the likelihood of the reference sequence contains the read sequence. If the reference sequence contains the read, the read may be mapped to the reference sequence or, in certain embodiments, to a particular location in the reference sequence. For example, the alignment of a read to the reference sequence for human chromosome 13 will tell the likelihood that the read is present in the reference sequence for chromosome 13. In some cases, an alignment additionally indicates a location where the read or tag maps to in the reference sequence. For example, if the reference sequence is the whole human genome sequence, an alignment may indicate that a read is present on chromosome 13, and may further indicate that the read is ona particular strand and / or site of chromosome 13. A “site” may be a unique position on a polynucleotide sequence or a reference genome (i.e. chromosome ID, chromosome position and orientation). In some embodiments, a site may provide a position for a residue, a sequence tag, or a segment on a sequence.
[0079] Aligned reads or tags are one or more sequences that are identified as a match in terms of the order of their nucleic acid molecules to a known sequence from a reference genome. Alignment can be done manually, although it is typically implemented by a computer algorithm, as it would be impossible to align reads in a reasonable time period for implementing the methods disclosed herein. The matching of a sequence read in aligning can be a 100% sequence match or less than 100% (non-perfect match).
[0080] Alignment may be performed by modifications and / or combinations of methods such as Burrows-Wheeler Aligner (BWA), iSAAC, BarraCUDA, BFAST, BLASTN, BLAT, Bowtie, CASHX, Cloudburst, CUDA-EC, CUSHAW, CUSHAW2, CUSHAW2-GPU, drFAST, ELAND, ERNE, GNUMAP, GEM, GensearchNGS, GMAP and GSNAP, Geneious Assembler, LAST, MAQ, mrFAST and mrsFAST, MOM, MOSAIK, MPscan, Novoaligh & NovoalignCS, NextGENe, Omixon, PALMapper, Partek, PASS, PerM, PRIMEX, QPalma, RazerS, REAL, cREAL, RMAP, rNA, RT Investigator, Segemehl, SeqMap, Shrec, SHRiMP, SLIDER, SOAP, SOAP2, SOAP3 and SOAP3-dp, SOCS, SSAHA and SSAHA2, Stampy, SToRM, Subread and Subjunc, Taipan, UGENE, VelociMapper, XpressAlign, and ZOOM.
[0081] The term “mapping” used herein refers to specifically assigning a sequence read to a larger sequence, e.g., a reference genome, by alignment.
[0082] As used herein, the term “paired-end reads” or “paired end reads” refers to paired reads generated from sequencing the forward and reverse ends of a larger nucleic acid fragment. In some examples, the forward and reverse ends of a larger nucleic acid fragment may share the same name. The paired-end reads may be generated from paired end sequencing that obtains one read from each end of a nucleic acid fragment.
[0083] Moreover, as used herein, the term “alignment score” refers to a numeric score, metric, or other quantitative measurement evaluating an accuracy of an alignment between one or more nucleotide reads or a fragment of a nucleotide read and another nucleotide sequence from a reference genome. In particular, an alignment score includes a metric indicating a degree to which the nucleobases of one or more nucleotide reads (or a fragmentthereof) match or are similar to a reference sequence or an alternate contiguous sequence from a reference genome. In certain implementations, an alignment score takes the form of a Smith- Waterman score or a variation or version of a Smith- Waterman score for local alignment, such as various settings or configurations used by DRAGEN by Illumina, Inc. for Smith-Waterman scoring.
[0084] Relatedly, as used herein, the term “mapping quality score” refers to a metric or other measurement quantifying a quality or certainty of an alignment of nucleotide reads (or other nucleotide sequences or subsequences) with a reference genome. In some embodiments, for example, a mapping quality score includes mapping quality (MAPQ) scores for nucleobase calls at genomic coordinates, where a MAPQ score represents -10 loglO Pr {mapping position is wrong}, rounded to the nearest integer. In the alternative to a mean or median mapping quality, in some implementations, a mapping quality score includes a full distribution of mapping qualities for all nucleotide reads aligning with a reference genome at a genomic coordinate.
[0085] As used herein, a “primary alignment” refers to an alignment of a sequence read to a reference genome at a location that has a highest determined probability (e.g., mapping quality).
[0086] As used herein, a “secondary alignment” refers to an alignment of a sequence read to a reference genome at with mapping quality above a desired threshold but less than that of the primary alignment.
[0087] As used herein, the term "solid support" refers to a rigid substrate that is insoluble in aqueous liquid. The substrate can be non-porous or porous. The substrate can optionally be capable of taking up a liquid (e.g. due to porosity) but will typically be sufficiently rigid that the substrate does not swell substantially when taking up the liquid and does not contract substantially when the liquid is removed by drying. A nonporous solid support is generally impermeable to liquids or gases. Exemplary solid supports include, but are not limited to, glass and modified or functionalized glass, plastics (including acrylics, polystyrene and copolymers of styrene and other materials, polypropylene, polyethylene, polybutylene, polyurethanes, Teflon™, cyclic olefins, polyimides etc.), nylon, ceramics, resins, Zeonor, silica or silica-based materials including silicon and modified silicon, carbon, metals, inorganic glasses, optical fiber bundles, and polymers. Particularly useful solidsupports for some embodiments are located within a flow cell apparatus. Exemplary flow cells are set forth in further detail below.
[0088] As used herein, the term "flow cell" is intended to mean a chamber having a surface across which one or more fluid reagents can be flowed. Generally, a flow cell will have an ingress opening and an egress opening to facilitate the flow of fluid. A flow cell can have multiple surfaces. Examples of flow cells and related fluidic systems and detection platforms that can be readily used in the methods of the present disclosure are described, for example, in Bentley et al, Nature 456:53-59 (2008), WO 04 / 018497; US 7,057,026; WO 91 / 06678; WO 07 / 123744; US 7,329,492; US 7,211,414; US 7,315,019; US 7,405,281, and US 2008 / 0108082, each of which is incorporated herein by reference.
[0089] In many embodiments, a solid support to which nucleic acids are attached in a method set forth herein will have a continuous or monolithic surface. Thus, fragments can attach at spatially random locations, wherein the distance between nearest neighbor fragments (or nearest neighbor clusters derived from the fragments) will be variable. The resulting arrays will have a variable or random spatial pattern of features. Alternatively, a solid support used in a method set forth herein can include an array of features that are present in a repeating pattern. In such embodiments, the features provide the locations to which modified nucleic acid polymers, or fragments thereof, can attach. Particularly useful repeating patterns are hexagonal patterns, rectilinear patterns, grid patterns, patterns having reflective symmetry, patterns having rotational symmetry, or the like. The features to which a modified nucleic acid polymer, or fragment thereof, attach can each have an area that is smaller than about 1mm2, 500 pm2, 100 pm2, 25 pm2, 10 pm2, 5 pm2, 1 pm2, 500 nm2, or 100 nm2. Alternatively, or additionally, each feature can have an area that is larger than about 100 nm2, 250 nm2, 500 nm2, 1 pm2, 2.5 pm2, 5 pm2, 10 pm2, 100 pm2, or 500 pm2. A cluster or colony of nucleic acids that result from amplification of fragments on an array (whether patterned or spatially random) can similarly have an area that is in a range above or between an upper and lower limit selected from those exemplified above.
[0090] As used herein, the term "surface," when used in reference to a material, is intended to mean an external part or external layer of the material. The surface can be in contact with another material such as a gas, liquid, gel, polymer, organic polymer, second surface of a similar or different material, metal, or coat. The surface, or regions thereof, can be substantiallyflat. The surface can have surface features such as wells, pits, channels, ridges, raised regions, pegs, posts or the like. The material can be, for example, a solid support, gel, or the like.
[0091] As used herein, the term "target," when used in reference to a nucleic acid polymer, is intended to linguistically distinguish the nucleic acid, for example, from other nucleic acids, modified forms of the nucleic acid, fragments of the nucleic acid, and the like. Any of a variety of nucleic acids set forth herein can be identified as target nucleic acids, examples of which include genomic DNA (gDNA), messenger RNA (mRNA), copy or complimentary DNA (cDNA), and derivatives or analogs of these nucleic acids.
[0092] As used herein, the term "transposase" is intended to mean an enzyme that is capable of forming a functional complex with a transposon element- containing composition (e.g., transposons, transposon ends, transposon end compositions) and catalyzing insertion or transposition of the transposon element-containing composition into a target DNA with which it is incubated, for example, in an in vitro transposition reaction. The term can also include integrases from retrotransposons and retroviruses. Transposases, transposomes and transposome complexes are generally known to those of skill in the art, as exemplified by the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Although many embodiments described herein refer to Tn5 transposase and / or hyperactive Tn5 transposase, it will be appreciated that any transposition system that is capable of inserting a transposon element with sufficient efficiency to tag a target nucleic acid can be used. In particular embodiments, a preferred transposition system is capable of inserting the transposon element in a random or in an almost random manner to tag the target nucleic acid. As used herein, the term "transposome" is intended to mean a transposase enzyme bound to a nucleic acid. Typically the nucleic acid is double stranded. For example, the complex can be the product of incubating a transposase enzyme with double-stranded transposon DNA under conditions that support non-covalent complex formation. Transposon DNA can include, without limitation, Tn5 DNA, a portion of Tn5 DNA, a transposon element composition, a mixture of transposon element compositions or other nucleic acids capable of interacting with a transposase such as the hyperactive Tn5 transposase.
[0093] As used herein, the term "transposon element" is intended to mean a nucleic acid molecule, or portion thereof, that includes the nucleotide sequences that form a transposome with a transposase or integrase enzyme. Typically, the nucleic acid molecule is adouble stranded DNA molecule. In some embodiments, a transposon element is capable of forming a functional complex with the transposase in a transposition reaction. As non-limiting examples, transposon elements can include the 19-bp outer end ("OE") transposon end, inner end ("IE") transposon end, or "mosaic end" ("ME") transposon end recognized by a wild-type or mutant Tn5 transposase, or the R1 and R2 transposon end as set forth in the disclosure of US Pat. App. Pub. No. 2010 / 0120098, which is incorporated herein by reference. Transposon elements can comprise any nucleic acid or nucleic acid analogue suitable for forming a functional complex with the transposase or integrase enzyme in an in vitro transposition reaction. For example, the transposon end can comprise DNA, RNA, modified bases, nonnatural bases, modified backbone, and can comprise nicks in one or both strands.
[0094] A standard NGS sequencing run yields millions of short sequences that are eventually mapped on a reference genome. A percentage of good-quality reads (1-5%) are discarded because of ambiguous genomic location. Increasing read length (2x500 or long-read sequencing), designing a specialized algorithm to map reads on specific regions of the genome (targeted callers), using expensive and time-consuming library preparation (Illumina CLR), or a combination thereof may be implemented to address the need for disambiguating such reads that would normally be discarded. However, such approaches are costly, laborious, and time intensive. Spatial information (X and Y coordinates) obtained from a solid support surface can be leveraged to identify fragments that are generated from a single long input fragment and subsequently be used to improve mapping reads in ambiguous positions.
[0095] In one or more embodiments, the system identifies and / or stores sequencing metrics within one or more sequencing data files. As used herein, the term “sequencing data file” refers to a digital file that includes genetic sequencing information concerning genotype calls or nucleotide reads generated by one or more genomic sequencing procedures. Such sequencing information may include, for example, nucleotide reads, alignment and mapping information, nucleotide reads at one or more genomic coordinates, and so forth.
[0096] Moreover, in one or more embodiments, one or more sequencing data files in which the system identifies or stores sequencing metrics include an alignment data file containing information from a read processing and mapping procedure. As used herein, the term “alignment data file” refers to a digital file that indicates mapping and alignment information for nucleotide reads of a sample nucleotide sequence. For example, an alignmentdata file can include a binary alignment map (BAM) file, a compressed reference-oriented alignment map (CRAM) file, or another file indicating nucleotide reads of a sample nucleotide sequence.
[0097] Moreover, as used herein, the term “cluster of oligonucleotides” (or “cluster” or “oligonucleotide cluster”) refers to a localized group or collection of DNA or RNA molecules on a nucleotide-sample slide, such as a flow cell, or other solid surface. In particular, a cluster includes tens, hundreds, thousands, or more copies of a cloned or the same DNA or RNA segment. For example, in one or more embodiments, a cluster includes a grouping of oligonucleotides immobilized in a section of a flow cell or other nucleotide-sample slide. In some embodiments, clusters are evenly spaced or organized in a systematic structure within a patterned flow cell. By contrast, in some cases, clusters are randomly organized within a nonpatterned flow cell. A cluster of oligonucleotides can be imaged utilizing one or more light signals. For instance, an oligonucleotide-cluster image may be captured by a camera during a sequencing cycle of light emitted by irradiated fluorescent tags incorporated into oligonucleotides from one or more clusters on a flow cell.Methods for Detecting Copy Number Variants
[0098] In an aspect, disclosed herein are methods for detecting copy number variants in a nucleic acid segment from a genomic DNA sample.
[0099] In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell. In some embodiments, generating sequence reads from fragments of the genomic DNA sample bound to a flow cell includes providing transposome complexes bound to the flow cell. The transposome complexes includes a transposase and a first polynucleotide comprising an end sequence and a first tag. A library preparation of the genomic sequences is then contacted with the flow cell and transposomes in order to contact the transposome complexes with the target genomic DNA sample under conditions to cause they transposome to fragment the genomic DNA sample. Because the transposomes are bound to the flow cell, following cleavage of the genomic DNA sample, the resulting fragments become bound to the flow cell. The process may then include amplifying the fragmented genomic DNA to form a plurality of nucleic acid clusters on a flow cell. A sequencing by synthesis process may then be started to sequence the nucleic acids ineach cluster on the flow cell to generate sequence reads. In some embodiments, the sequence reads comprise paired end sequence reads where each nucleic acid is sequenced by two primers bound in opposite directions to one another on the bound fragment.
[0100] FIG. 1A is a flow diagram that schematically illustrates an exemplary method 100 for detecting copy number variants. In some embodiments, the method 100 is implemented on a computer. The method 100 may be embodied in a set of executable program instructions stored on a computer-readable medium, such as one or more disk drives, of a computing system. When the method 100 is initiated, the executable program instructions can be loaded into a memory and executed by one or more processors of a server device or other electronic system, such as a Next Generation Sequencing system.
[0101] As shown in FIG. 1A, the method 100 for detecting copy number variants may start from a start block 110. The method 100 may proceed to block 120, wherein sequence reads sequenced from fragments of a genomic DNA sample bound to a flow cell are obtained, for example, by retrieving the sequence reads from a storage. The sequence reads may be generated as further described above. The sequence reads may be stored and retrieved in any electronic file format, for example such as a BAM file format.Determining Linkage Information and Aligning to a Reference Sequence
[0102] The method 100 may proceed to block 130, wherein linkage information between the sequence reads based on the geographic location of each fragment on the flow cell is determined, and the sequence reads are aligned to a reference genome using the linkage information.
[0103] In some embodiments, the linkage information is determined based on the geographic location information of clusters of the fragments on the flow cell and their putative genetic location on a reference genome. For example, in some embodiments, the linkage information comprises information linking a first sequence read from a first cluster to a second sequence read from a second cluster based on a distance between the first and second clusters on the flow cell. In some embodiments, the location information comprises a first spatial coordinate and a second spatial coordinate in a Cartesian coordinate system. For example, the dimensions of the flow cell may be mapped with a Cartesian coordinate system, e.g., with xand y dimensions, and locations of clusters on a flow cell may be assigned to a spatial coordinate using this system.
[0104] Embodiments of the present disclosure relate to methods and systems which use “links” or “linkage information” between sequence reads. The “link” or “linkage information” as discussed herein refers, in some embodiments, to the probability that two pairs of reads on a sequencing flow cell are derived from the same original nucleic acid molecule. In some next generation sequencing (NGS) systems, fragments of long DNA, such as genomic DNA, from a biological source are sheared to create shorter fragments which can be sequenced in a single read. The shearing process can create these shorter fragments which land on the flow cell and the spatial location of each fragment may be related to the original nucleic acid molecule from which the fragment was derived. For example, fragments which come from the same nucleic acid molecule have been found to bind closer together on the flow cell as compared to fragments which come from different original nucleic acid molecules. Accordingly, if two clusters of reads on a flow cell are in close physical proximity and also close together on the genome, the clusters are more likely to have come from the same nucleic acid molecule. However, unrelated fragments may also bind to the flow cell near one another, which leads to an uncertainty in the probability that adjacent clusters originate from the same molecule. A number of factors could affect the probability that unrelated clusters would land in a similar area, and these factors may change based on a variety of experimental conditions. Embodiments of the disclosure provide a statistical method for calculating the probability that two reads are linked, such that on a flow cell the two reads were derived from the same nucleic acid molecule.
[0105] Embodiments of the disclosure relate to systems and methods for sequencing target nucleic acids by fragmenting the target nucleic acid and distributing the fragments onto a flow cell. As the fragments are distributed along the flow cell, they bind capture primers and are then used to create clusters by well-known technologies, such as those provided by Illumina Inc. (San Diego, CA). As described above, according to the methods of this disclosure, fragments which were derived from the same template genomic sequence are more likely to bind to the flow cell in close physical proximity as compared to fragments that are from different template genomic sequences, particularly when the fragmentation is performed directly on the flow cell using immobilized transposome complexes on the surfaceof the flow cell. In some embodiments, the library preparation steps are performed on the flow cell, which may reduce the complexity and the amount of equipment associated with the systems. In regular library preparations with fragmentation happening prior to loading, fragments can land anywhere in the flow cell independently of whether they came from the same molecule. However, when fragmentation is performed directly on the flow cell spatial information is retained. This spatial information can be used to help guide assembly and variant calling of the original template genomic sequence, as will be described in more detail below.
[0106] For example, transposome complexes may be provided as part of the sequencing process. In some embodiments, the transposome complexes include a transposase and a first polynucleotide having end sequences which can be used to fragment the target polynucleotides and insert into each fragment an end sequence or tag which can be used to bind to capture probes located on the substrate. The method can include contacting the transposome complexes with the target polynucleotides under conditions to fragment the target polynucleotides and add capture sequences to the ends of each fragment. In some embodiments, the capture sequences include P5 or P7 sequences as provided by Illumina®, Inc. (San Diego, CA). In some embodiments, the complexed strand and transposome is in solution, and is then brought towards a substrate and immobilized thereon. In some embodiments, prior to immobilization of the transposome complexes on the substrate, one or more of the transposome complexes bind the target polynucleotides in solution. In this embodiment, the transposome complexes in solution become immobilized to the substrate.
[0107] Once the fragments have been bound to the substrate, the bound fragments can be amplified to form a plurality of nucleic acid clusters on the substrate. The location of each cluster on the flow cell can then be determined before, during or after performing sequencing by synthesis reactions (SBS) to obtain the nucleotide sequence of each fragment located in each cluster.
[0108] Once the nucleotide sequence of each cluster has been determined, the method can start to map those reads to determine the original target polynucleotide from which the read originated. In some embodiments, the mapping process takes into account the spatial location of each cluster, such that clusters which are closer to each other on the flow cell are more likely to have originated from the same target polynucleotide.
[0109] For example, in some embodiments, the methods and systems align an initial set of reads to the reference genome (e.g., a subset of all of the sequence reads) in order to define linking model parameters based on the spatial location and mapping location. In some embodiments, once linking model parameters for the current sample are defined, the full link- informed alignment process begins. For example, when reads are being aligned to a reference sequence, various alignment candidates are evaluated for each read, and potential links for each candidate alignment position of a read are evaluated and a boost to the mapping quality is calculated from the linking information in each possible candidate position. For example, in some embodiments, all candidate alignments are obtained and a pairwise comparison between potentially linked reads is performed. In some embodiments, a link bonus for each candidate alignment is determined based on the linking model parameters.
[0110] Furthermore, by mapping the sequenced fragments to target polynucleotides using the spatial information accompanying each cluster, the method performs more accurate mapping operations as compared to methods that do not take the spatial location of each cluster into account during the mapping process. Therefore, spatial information that includes relative distances between various clusters on a flow cell may be leveraged to adjust mapping information, thereby increasing the read quality and alignment accuracy of previously identified multi-mapped reads. In the past, identified multi-mapped reads may have been discarded. Increasing the read quality of these previously discarded reads, by improving the confidence of read pair’s alignment based on linking information with a high link quality score, may improve the alignment information and quality of information used in certain genomic analysis applications including, but not limited to, variant calling. Processing DNA samples suitable for high-throughput sequencing that retain information on the original configuration of the DNA samples provides useful information on co-located fragments.
[0111] Further details regarding sequencing conditions that result in links or downstream analyses utilizing linking information can be found in International Patent Application Nos. PCT / US2024 / 035447 and PCT / US2024 / 045996, International Patent Application Publication Nos. WO2015 / 189636, WO2015 / 095226 and WO2023 / 122755, and U.S. Provisional Patent Application Nos. 63 / 600,460, 63 / 614,066, 63 / 700,049, and 63 / 700,262, the disclosure of each of which is incorporated herein by reference in its entirety.Estimating One or More Copy Numbers
[0112] The method 100 may proceed to process block 140, wherein a copy number of the nucleic acid segment is estimated based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information. The process block 140 is explained in more detail with reference to FIG. IB, which is a block diagram that illustrates additional details on the process taking place within the process block 140.
[0113] The method 100 may proceed to decision state 150, wherein the system queries whether there are additional bins. If yes, the workflow returns to process block 140 and proceeds as described above. If no, the method 100 may proceed to block 160 below.Detecting a Copy Number Variant
[0114] The method 100 may proceed to block 160, wherein a copy number variant is detected in the genomic DNA based on the estimated copy numbers.
[0115] In some embodiments, a higher corrected sequencing depth in a region than other regions corresponds to a higher estimated copy number, and a lower corrected sequencing depth in a region than other regions corresponds to a lower estimated copy number. For example, if a first region has a higher sequencing depth (after the correction factor has been applied) than other regions of the genome, the method may determine that the first region has a higher copy number than the other regions of the genome.
[0116] A copy number variant may be detected based on variations in the estimated copy numbers. For example, if a nucleic acid segment expected to have a copy number of 2 but has an estimated copy number of 3 or 1, a copy number variant may be detected. A copy number variant may be detected if different regions of the nucleic acid segment have different copy numbers. For example, a first region may have an estimated copy number of 2, a second region may have an estimated copy number of 2, and then a third region may have an estimated copy number of 3. In this example, a copy number variant may be detected in the third region.
[0117] The method 100 may end at block 170.
[0118] In some embodiments, the method 100 further includes a step (not illustrated) of storing one or more copy numbers, information related to a copy number variant, or information related to a copy number genotype in an electronic file, for example, a VCFfile. For example, an electronic file may be created which includes one or more estimated copy numbers for the nucleic acid segment. The electronic file may also include an indication of whether a copy number variant was detected. The electronic file may also include information related to a copy number genotype that has been determined for the genomic DNA sample based on the one or more estimated copy numbers.Process of Estimating One or More Copy Numbers
[0119] FIG. IB is a block diagram that illustrates additional details on the specific methods taking place within the process block 140, wherein a copy number is estimated. As shown in FIG. IB, the process 140 for estimating one or more copy numbers for the nucleic acid segment in the genomic DNA sample based on the correction factor may include the following steps.
[0120] The process 140 may begin from block 210, wherein the alignment of the sequence reads to the reference genome is binned. The bins may be a window of the reference genome of a predetermined size, for example between about 1 kbp to about 100 kbp in size, for example, about 1 kbp, about 5 kbp, about 10 kbp, about 20 kbp, about 50 kbp, about 100 kbp, or a range constructed from any of the aforementioned values. In some embodiments, each bin is of a similar size to other bins. Bin sizes can be adjusted based on a balance of computational efficiency and accuracy. For example, smaller bin sizes may improve accuracy by providing a more granular determination of sequencing depth along the reference genome, but may be computationally more intense to process. A desired bin size can be determined based on these considerations.
[0121] The process 140 may continue to block 220, wherein a high-MAPQ depth- per-bin is determined. In some embodiments, determining a high-MAPQ depth-per-bin comprises for each bin, counting the number of sequence reads with a MAPQ score above a predetermined threshold which align to the bin. In other embodiments, another predetermined threshold may be used, such as a MAPQ score of greater than or equal to any of the following: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range constructed from any of the aforementioned values. In some embodiments, the predetermined threshold is 3 or greater.
[0122] The process 140 may continue to block 230, wherein the high-MAPQ depth-per-bin is GC-corrected. GC-correction may be accomplished, for example, according to any method known to those of skill in the art. For example, sequencing depth signal may be adjusted based on GC content of the nucleic acid sequence at that region of the genome, and based on the GC content of other regions of the genome that are expected to be diploid.
[0123] The process 140 may continue to block 240, wherein the GC-corrected depth-per-bin is normalized by a median depth-per-bin. For example, the GC-corrected depth- per-bin for an individual bin may be adjusted based on a median depth-per-bin that has been determined based on sequencing depths across the genome.
[0124] The process 140 may continue to block 250, wherein a correction factor is applied to the normalized depth-per-bin to generate a corrected depth-per-bin signal. In some embodiments, the correction factor is determined based on one or more of the following: a fraction of unlinked sequence reads that are aligned within a first bin of the reference genome with confidence above a predetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold.
[0125] In some embodiments, the correction factor is determined based on each of: a fraction of unlinked sequence reads that are aligned within a first bin of the reference genome with confidence above a predetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold. In some embodiments, the correction factor is determined based on the following expression:wherein H is the fraction of unlinked sequence reads that are aligned within a first bin of the reference genome with confidence above a predetermined threshold, L is the linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold, and P is the fraction of linked sequence reads that aremapped to the reference genome within the first bin with confidence below a predetermined threshold.
[0126] In some embodiments, the fraction of unlinked sequence reads that are aligned within the first bin of the reference genome with confidence above a predetermined threshold, is estimated based on a fraction of sequence reads that are high-MAPQ, over all sequence reads within the first bin that do not have a link to another sequence read. In some embodiments, sequence reads that are high-MAPQ are sequence reads that have a MAPQ score of 3 or greater. In other embodiments, another predetermined threshold may be used, such as a MAPQ score of greater than or equal to any of the following: 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range constructed from any of the aforementioned values.
[0127] In some embodiments, the linking rate is estimated based on all of the sequence reads, or a random subset thereof. For example, the proportion of sequence reads which have a link to another sequence read is determined with respect to all of the sequence reads, or a random subsample thereof. In other embodiments, the linking rate is estimated per bin, and thus may be determined for each bin individually. In some embodiments, sequence reads which have a link to another sequence read refers to sequence reads which have a link based on linkage information that is above a predetermined threshold. It will be apparent that various predetermined thresholds can be used (such as a quality score, confidence, etc.) when determining a probability that, for example, a sequence read was derived from the same nucleic acid fragment as another sequence from a different cluster on a flow cell.
[0128] In some embodiments, the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold comprises a fraction of sequence reads which have a MAPQ score below the predetermined threshold, over all linked sequence reads within the first bin. In some embodiments, the predetermined threshold for the MAPQ score is 3. In some embodiments, the predetermined threshold is 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, or a range constructed from any of the aforementioned values.
[0129] In some embodiments, the correction factor is determined using primary and secondary alignments of each sequence read to the reference genome. In particular, in some embodiments, one or more of the linking rate and the fraction of linked sequence readsthat are mapped to the reference genome with confidence below a predetermined threshold, are determined using primary and secondary alignments of each sequence read to the reference genome. For example, in some embodiments, copy number variants may skew a correction factor. For example, estimates of linking rates and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold may be skewed. FIG. 2 illustrates an embodiment where a genomic DNA sample includes Paralog 1 and Paralog 2. MAPQO reads are mapped equally between Paralog 1 and Paralog 2, but linked and high MAPQ reads have all been mapped to Paralog 1. This causes the percentage of linked reads (linking rate) to be overestimated in Paralog 1 and underestimated in Paralog 2, and the fraction of linked sequence reads that are mapped with a confidence below a predetermined threshold to be underestimated in Paralog 1 and overestimated in Paralog 2. This situation would result erroneous copy number estimation, with Paralog 1 determined to have a copy number of 1 and Paralog 2 determined to have a copy number of 0. In order to avoid situations such as the one illustrated in FIG. 2, the linking rate and the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold may be determined based on both primary and secondary alignments of sequence reads to the reference genome.
[0130] After the correction factor is applied, the process 140 may then proceed to block 260, wherein the corrected depth-per-bin signal is segmented by identifying consecutive bins with a similar corrected depth-per-bin signal. Segmenting may be accomplished, for example, using any method known to those of skill in the art. For example, segmenting may be accomplished using Shifting Levels Models, Circular Binary Segmentation, or Hidden Markov Models.
[0131] The process 140 may then proceed to block 270, wherein segmented depth- per-bin signals are converted to integer copy numbers with an associated probability score, to determine a copy number genotype. Conversion to an integer copy number (such as 0, 1, 2, 3, 4, and so on) and determination of an associated probability score may be accomplished using, for example, any method known to those of skill in the art. For example, a statistical model such as a Gaussian Mixture Model, or a Student’s T distributions parametrized for specific set of possible copy numbers, may be used to convert the depth-per-segment signal to an integer copy number and to determine an associated probability score that the integer copy number isaccurate to the genomic sample, based on the discrepancy between the depth-per-bin signal and the rounded integer conversion. In some embodiments, the associated probability score is a measure of confidence in the integer copy number. For example, if a depth-per bin signal is determined to be 2.34, the systems and methods can determine an integer copy number of 2 and a probability score associated with the integer copy number based on the difference between the “raw” depth-per-bin signal (2.34) and the integer it has been rounded to (2). For example, if the segmented depth-per-bin signal for a first segment is 2.34 and for a second segment is 2.12, the integer copy number for both may be 2 but the second segment may have a higher associated probability score than the first segment because 2.12 is closer to 2 than 2.34. A copy number genotype may be determined based on the integer copy numbers. For example, a deviation from having two copies of a gene or locus, such as a deletion or duplication, may be determined based on integer copy numbers.Methods for Determining the Copy Number of a Nucleic Acid Segment from a Genomic DNA Sample
[0132] In another aspect, disclosed herein are methods for determining a copy number of a nucleic acid segment from a genomic DNA sample. In some embodiments, the method includes generating sequence reads from fragments of the genomic DNA sample bound to a flow cell, which may be accomplished as further described above. In some embodiments, the method includes determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information, which may be accomplished as further described above. In some embodiments, the method includes estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information, which may be accomplished as further described above.
[0133] In some embodiments, the copy number is stored in an electronic file, for example, a VCF file.Methods for Detecting a Pathogenic Copy Number Variant
[0134] In another aspect, disclosed herein are methods for detecting a pathogenic copy number variant in a nucleic acid segment from a genomic DNA sample. In some embodiments, the method includes detecting copy number variants in a nucleic acid segment from a genomic DNA sample or determining a copy number of a nucleic acid segment from a genomic DNA sample as further described above.
[0135] In some embodiments, the plurality of locations in the reference genome comprise one or more of an IKBKG gene, an IKBKGP1 pseudogene, a RCCX locus, a STK19 gene, a STK19B pseudogene, a C4A gene, a C4B gene, a CYP21A1P pseudogene, a CYP21A2 gene, a TNXA pseudogene, a TNXB gene, a STRC gene, a STRCP1 pseudogene, a PMS2 gene, a PMS2CL pseudogene, a OTOA gene, a OTOAP1 pseudogene, a GBA gene, a GBAP1 pseudogene, a HYDIN gene, or a HYDIN2 pseudogene.Systems for Detecting Copy Number Variants
[0136] Further disclosed herein are electronic systems for detecting copy number variants in a nucleic acid segment from a genomic DNA sample. In some embodiments, the system includes a processor configured to perform a method comprising: receiving sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detecting a copy number variant in the genomic DNA based on the estimated copy numbers.
[0137] Further disclosed herein are non-transitory computer-readable media. In some embodiments, the non-transitory computer-readable medium includes a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimate one or more copy numbers for the nucleic acidsegment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detect a copy number variant in the genomic DNA based on the estimated copy numbers..
[0138] FIG. 3A illustrates a diagram of an environment in which a system for detecting copy number variants in a nucleic acid segment from a genomic DNA sample can operate in accordance with one or more implementations. The following paragraphs describe the copy number variant detection system with respect to illustrative figures that portray example implementations and embodiments. For example, FIG. 3A illustrates a schematic diagram of a computing system 3000 in which a copy number application 3106 operates in accordance with one or more implementations. As illustrated, the computing system 3000 includes one or more server device(s) 3102 connected to a user client device 3108, a local device 3118, and a sequencing device 3114 via a network 3112. The network 3112 can comprise any suitable network over which computing devices can communicate.
[0139] As shown in FIG. 3 A, the computing system 3000 includes the server device(s) 3102. In various implementations, the server device(s) 3102 may generate, receive, analyze, store, and transmit digital data, such as data for nucleobase calls or sequenced nucleic- acid polymers. In some implementations, the server device(s) 3102 receive various data from the sequencing device 3114, such as data from a sample genome and / or sequence reads. The server device(s) 3102 may also communicate with the user client device 3108. In particular, the server device(s) 3102 can send data for sequence reads, direct nucleobase calls, nucleobase calls, and / or sequencing metrics to the user client device 3108.
[0140] As shown, the server device(s) 3102 includes a sequencing application 3110. In general, the sequencing application 3110 analyzes the data (such as call data) received from the sequencing device 3114 or elsewhere to determine nucleobase sequences for nucleic- acid polymers. For example, the sequencing application 3110 can receive raw data from the sequencing device 3114 and determine a nucleobase sequence for a sample genome or a nucleic-acid segment. In some implementations, the sequencing application 3110 determines the sequences of nucleobases in DNA and / or RNA segments or oligonucleotides.
[0141] As also shown, the sequencing application 3110 includes the copy number application 3106. As described below, in some embodiments, the copy number application 3106 can determine one or more copy numbers for a nucleic acid segment from a genomicDNA sample. For example, in some embodiments, the copy number application receives sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determines linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimates one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detects a copy number variant in the genomic DNA based on the estimated copy numbers. For example, the copy number application 3106 can include the methods shown in FIG. 1 A and FIG. IB and described above.
[0142] While the sequencing application 3110 has been described as including the copy number application 3106, other systems or methods may be included within the sequencing application 3110, such as an application to determine single nucleotide polymorphisms or to assemble sequence reads (not illustrated).
[0143] Moreover, while the copy number application 3106 is described being implemented on the server device(s) 3102, as part of the sequencing application 3110, in some implementations, the copy number application 3106 is implemented by (such as located entirely or in part) on the user client device 3108, the sequencing device 3114, and / or the local device 3118. As mentioned, in some implementations, copy number application 3106 is implemented by one or more other components of the computing system 3000, such as the sequencing device 3114. In particular, the copy number application 3106 can be implemented in a variety of different ways across the server device(s) 3102, the network 3112, the user client device 3108, the local device 3118, and the sequencing device 3114.
[0144] As further shown in FIG. 3 A, the computing system 3000 includes the user client device 3108. In various implementations, the user client device 3108 can generate, store, receive, and send digital data. In particular, the user client device 3108 can receive the data from the sequencing device 3114. As further illustrated, the user client device 3108 includes a sequencing application 3110. The sequencing application 3110 may be a web application or a native application stored and executed on the user client device 3108 (for example, a mobile application, desktop application, or web application). The sequencing application 3110 can receive data from the sequencing application 3110 and / or copy number application 3106. Forexample, the user client device 3108 can receive variant call files and / or alignment files from the sequencing application 3110.
[0145] The sequencing application 3110 can also include instructions that (when executed) cause the user client device 3108 to receive data from the copy number application 3106 and present data from the sequencing device 3114 and / or the server device(s) 3102. Furthermore, the sequencing application 3110 can instruct the user client device 3108 to display data for variant calls, such as nucleobase calls or an indication of a copy number variant. Indeed, the user client device 3108 can display nucleobase call results for a genome sample and / or an indication of a predicted copy number variant.
[0146] As further shown in FIG. 3A, the computing system 3000 includes the sequencing device 3114. In various implementations, the sequencing device 3114 can sequence a genomic sample or other nucleic-acid polymer. For example, the sequencing device 3114 analyzes nucleic-acid segments or oligonucleotides extracted from genomic samples to generate data either directly or indirectly on the sequencing device 3114. More particularly, the sequencing device 3114 receives and analyzes, within nucleotide-sample slides (such as flow cells), nucleic-acid sequences extracted from genomic samples. In one or more implementations, the sequencing device 3114 utilizes sequencing by synthesis (SBS) to sequence a genomic sample or other nucleic-acid polymers. In addition to, or in the alternative to communicating across the network 3112, in some implementations, the sequencing device 3114 bypasses the network 3112 and communicates directly with the user client device 3108.
[0147] As further depicted in FIG. 3A, in some implementations, the server device(s) 3102 includes a distributed collection of servers, where the server device(s) 3102 include several server devices distributed across the network 3112 and located in the same or different physical locations. For instance, the server device(s) 3102 can be implemented, in whole or in part, on the local device 3118. To illustrate, the local device 3118 may implement the sequencing application 3110 and / or the copy number application 3106. Further, the server device(s) 3102 and / or the local device 3118 can include a content server, an application server, a communication server, a web-hosting server, or another type of server.
[0148] The user client device 3108 illustrated in FIG. 3 A can include various types of client devices. For example, in some implementations, the user client device 3108 includes non-mobile devices, such as desktop computers or servers, or other types of client devices. Invarious implementations, the user client device 3108 includes mobile devices, such as laptops, tablets, mobile telephones, or smartphones.
[0149] Though FIG. 3A illustrates the components of the computing system 3000 communicating via the network 3112, in certain implementations, the components of computing system 3000 can also communicate directly with each other, bypassing the network 3112. For instance, in some implementations, the user client device 3108 communicates directly with the sequencing device 3114. Additionally, in some implementations, the user client device 3108 communicates directly with the copy number application 3106 and / or the server device(s) 3102. In some implementations, the user client device 3108 communicates directly with the local device 3118. Moreover, the copy number application 3106 can access one or more databases housed on or accessed by the server device(s) 3102 or elsewhere in the computing system 3000.
[0150] FIG. 3B is a block diagram of an exemplary server device 3102 that may be used in connection with the computing system 3000 of FIG. 3 A. The server device 3102 may be configured to determine one or more copy numbers for a nucleic acid segment from a genomic DNA sample. The general architecture of the server device 3102 depicted in FIG. 3B includes an arrangement of computer hardware and software components. The server device 3102 may include many more (or fewer) elements than those shown in FIG. 3B. It is not necessary, however, that all of these generally conventional elements be shown in order to provide an enabling disclosure. As illustrated, the server device 3102 includes a processing unit 310, a network interface 320, a computer readable medium drive 330, an input / output device interface 340, a display 350, and an input device 360, all of which may communicate with one another by way of a communication bus. The network interface 320 may provide connectivity to one or more networks or computing systems. The processing unit 310 may thus receive information and instructions from other computing systems or services via a network. The processing unit 310 may also communicate to and from memory 370 and further provide output information for an optional display 350 via the input / output device interface 340. The input / output device interface 340 may also accept input from the optional input device 360, such as a keyboard, mouse, digital pen, microphone, touch screen, gesture recognition system, voice recognition system, gamepad, accelerometer, gyroscope, or other input device.
[0151] The memory 370 may contain computer program instructions (grouped as modules or components in some embodiments) that the processing unit 310 executes in order to implement one or more embodiments. The memory 370 generally includes RAM, ROM and / or other persistent, auxiliary or non-transitory computer readable media. The memory 370 may store an operating system 372 that provides computer program instructions for use by the processing unit 310 in the general administration and operation of the server device 3102. The memory 370 may store a reference genome 373, such as for use by the sequencing application 3110. The memory 370 may further include computer program instructions and other information for implementing aspects of the present disclosure.
[0152] For example, in one embodiment, the memory 370 includes a sequencing application 3110, which may include a copy number application 3106. The copy number application 3106 can perform the methods disclosed herein. In addition, memory 370 may include or communicate with the data store 390 and / or one or more other data stores that store one or more inputs, one or more outputs, and / or one or more results (including intermediate results) of aligning sequence reads, and / or one or more reference genomes.
[0153] In some embodiments, the disclosed systems and methods may involve approaches for shifting or distributing certain sequence data analysis features and sequence data storage to a cloud computing environment or cloud-based network. User interaction with sequencing data, genome data, or other types of biological data may be mediated via a central hub that stores and controls access to various interactions with the data. In some embodiments, the cloud computing environment may also provide sharing of protocols, analysis methods, libraries, sequence data as well as distributed processing for sequencing, analysis, and reporting. In some embodiments, the cloud computing environment facilitates modification or annotation of sequence data by users. In some embodiments, the systems and methods may be implemented in a computer browser, on-demand or on-line.
[0154] In some embodiments, software written to perform the methods as described herein is stored in some form of computer readable medium, such as memory, CD- ROM, DVD-ROM, memory stick, flash drive, hard drive, SSD hard drive, server, mainframe storage system and the like.
[0155] In some embodiments, the methods may be written in any of various suitable programming languages, for example compiled languages such as C, C#, C++,Fortran, and Java. Other programming languages could be script languages, such as Perl, MatLab, SAS, SPSS, Python, Ruby, Pascal, Delphi, R and PHP. In some embodiments, the methods are written in C, C#, C++, Fortran, Java, Perl, R, Java or Python. In some embodiments, the method may be an independent application with data input and data display modules. Alternatively, the method may be a computer software product and may include classes wherein distributed objects comprise applications including computational methods as described herein.
[0156] In some embodiments, the methods may be incorporated into pre-existing data analysis software, such as that found on sequencing instruments. Software comprising computer implemented methods as described herein are installed either onto a computer system directly, or are indirectly held on a computer readable medium and loaded as needed onto a computer system. Further, the methods may be located on computers that are remote to where the data is being produced, such as software found on servers and the like that are maintained in another location relative to where the data is being produced, such as that provided by a third party service provider.
[0157] An assay instrument, desktop computer, laptop computer, or server which may contain a processor in operational communication with accessible memory comprising instructions for implementation of systems and methods. In some embodiments, a desktop computer or a laptop computer is in operational communication with one or more computer readable storage media or devices and / or outputting devices. An assay instrument, desktop computer and a laptop computer may operate under a number of different computer based operational languages, such as those utilized by Apple based computer systems or PC based computer systems. An assay instrument, desktop and / or laptop computers and / or server system may further provide a computer interface for creating or modifying experimental definitions and / or conditions, viewing data results and monitoring experimental progress. In some embodiments, an outputting device may be a graphic user interface such as a computer monitor or a computer screen, a printer, a hand-held device such as a personal digital assistant (such as PDA, Blackberry, iPhone), a tablet computer (such as iPAD), a hard drive, a server, a memory stick, a flash drive and the like.
[0158] A computer readable storage device or medium may be any device such as a server, a mainframe, a supercomputer, a magnetic tape system and the like. In someembodiments, a storage device may be located onsite in a location proximate to the assay instrument, for example adjacent to or in close proximity to, an assay instrument. For example, a storage device may be located in the same room, in the same building, in an adjacent building, on the same floor in a building, on different floors in a building, etc. in relation to the assay instrument. In some embodiments, a storage device may be located off-site, or distal, to the assay instrument. For example, a storage device may be located in a different part of a city, in a different city, in a different state, in a different country, etc. relative to the assay instrument. In embodiments where a storage device is located distal to the assay instrument, communication between the assay instrument and one or more of a desktop, laptop, or server is typically via Internet connection, either wireless or by a network cable through an access point. In some embodiments, a storage device may be maintained and managed by the individual or entity directly associated with an assay instrument, whereas in other embodiments a storage device may be maintained and managed by a third party, typically at a distal location to the individual or entity associated with an assay instrument. In embodiments as described herein, an outputting device may be any device for visualizing data.
[0159] An assay instrument, desktop, laptop and / or server system may be used itself to store and / or retrieve computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. One or more of an assay instrument, desktop, laptop and / or server may comprise one or more computer readable storage media for storing and / or retrieving software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like. Computer readable storage media may include, but is not limited to, one or more of a hard drive, a SSD hard drive, a CD-ROM drive, a DVD-ROM drive, a floppy disk, a tape, a flash memory stick or card, and the like. Further, a network including the Internet may be the computer readable storage media. In some embodiments, computer readable storage media refers to computational resource storage accessible by a computer network via the Internet or a company network offered by a service provider rather than, for example, from a local desktop or laptop computer at a distal location to the assay instrument.
[0160] In some embodiments, computer readable storage media for storing and / or retrieving computer implemented software programs incorporating computer code for performing and implementing computational methods as described herein, data for use in the implementation of the computational methods, and the like, is operated and maintained by a service provider in operational communication with an assay instrument, desktop, laptop and / or server system via an Internet connection or network connection.
[0161] In some embodiments, a hardware platform for providing a computational environment comprises a processor (such as CPU) wherein processor time and memory layout such as random access memory (such as RAM) are systems considerations. For example, smaller computer systems offer inexpensive, fast processors and large memory and storage capabilities. In some embodiments, graphics processing units (GPUs) can be used. In some embodiments, hardware platforms for performing computational methods as described herein comprise one or more computer systems with one or more processors. In some embodiments, smaller computer are clustered together to yield a supercomputer network.
[0162] In some embodiments, computational methods as described herein are carried out on a collection of inter- or intra-connected computer systems (such as grid technology) which may run a variety of operating systems in a coordinated manner. For example, the CONDOR framework (University of Wisconsin-Madison) and systems available through United Devices are exemplary of the coordination of multiple stand-alone computer systems for the purpose dealing with large amounts of data. These systems may offer Perl interfaces to submit, monitor and manage large sequence analysis jobs on a cluster in serial or parallel configurations.Electronic Files
[0163] The methods and systems described herein can include a step of storing information generated by the methods and systems, in an electronic file, or retrieving information from an electronic file. In some embodiments, a file is on a computer storage medium (such as a computer hard drive, for example a spinning magnetic disk drive or a solid state drive). In some embodiments, the electronic file is stored in the format of a BAM, FASTQ, SAM, CRAM, JSON, CIGAR, or VCF file.
[0164] In some embodiments, the electronic file includes information related to a co-proximity property, where reads that are nearby within an individual’s genome are also nearby on the flow cell surface. This property enables improved mapping, variant calling, and phasing, among other enhancements. The following is a list of terms that are used to describe data in Tables 1 and 2 below.
[0165] Template molecule: A template molecule is a long DNA molecule (e.g., at least 500 bp) from either a standard or high molecular weight extraction that is loaded into the sequencing cartridge for on-flow cell tagmentation.
[0166] Proximal reads: Proximal reads are reads from the same template molecule that are from clusters located near one another on the flow cell surface.
[0167] Proximity group: A proximity group is a set of reads that are likely to have have been derived from the same template molecule.
[0168] Flow cell proximity: The flow cell proximity is the distance between DNA nanowells of reads within a proximity group, measured either in common units of measurement including nanometers, millimeters, inches, or in flow cell units.
[0166] Genomic proximity: Genomic proximity is the genomic distance between reads within a proximity group.
[0169] Template length: Genomic distance (e.g., distance on a reference genome, measured in bp) spanned by a proximity group from the same original template.
[0170] Co-proximity rate: Percentage of all reads that are co-proximal with at least one other read.
[0171] Co-proximity quality: A Phred-scaled quality score that estimates the probability that two reads derive from the same original template molecule.
[0172] In some embodiments, the method includes a step of storing sequence reads or other sequence information in a BAM file or a VCF file. In some embodiments, the BAM file has tags and and / or VCF file has fields as described below. In some embodiments, metrics for these tags are stored in a separate CSV file with the definitions provided.
[0173] Table 1 describes new tags for BAM files. In some embodiments, the BAM file also includes other tags, for example debug tags (xn, xp, xf, xa, xb, Id, ad, bd, ws, wm, and pq)-Table 1 : New BAM tags
[0174] In some embodiments, the VCF file includes fields related to variant calling and haplotype information (for example, Whatshap phase).
[0175] In some embodiments, the method includes creating an electronic file that includes one or more of the following metrics:Table 2: Description of MetricsEXAMPLES
[0176] Some aspects of the embodiments discussed above are disclosed in further detail in the following examples, which are not in any way intended to limit the scope of the present disclosure. Those in the art will appreciate that many other embodiments also fall within the scope of the disclosure, as it is described herein above and in the claims.Example 1
[0177] In the following example, Illumina® paired-end sequence reads were generated from genomic DNA sample extracted from the HG002 / NA24385 epithelial cell line and sequenced on a NovaSeqX® Plus platform with transposase bound to the cellist surface, allowing for on-flow cell fragmentation and adapter ligation. Various techniques were tested to estimate copy number at different genomic regions. The genomic DNA sample was from cell line HG002.
[0178] First, sequence reads were mapped to a reference genome (GRCh38 from the Genome Reference Consortium) at four different regions: the IKBKG segmental duplication (FIG. 4A), the RCCX locus (FIG. 4B), the STRC segmental duplication (FIG. 4C), and the PMS2 segmental duplication (FIG. 4D), without using linkage information or a correction factor (panels labeled “Baseline” and “Baseline Coverage” in FIGS. 4A-4D). Asseen in the panels labeled “Baseline” and “Baseline Coverage” in FIGS. 4A-4D, there are some areas with gaps in the alignment depth.
[0179] Next, alignment to the reference genome was performed again for each region, but using linkage information between the sequence reads, which was determined based on the geographic location of each DNA fragment on the flow cell (panels labeled “Linkage Mapping” and “Linkage Mapping Coverage” in FIGS. 4A-4D).
[0180] As seen in the panels labeled “Linkage Mapping” and “Linkage Mapping Coverage” in FIGS. 4A-4D, the gaps in the alignment depth have been improved compared to the baseline, but there is still some inconsistency in the depth of coverage over the various regions.
[0181] FIGS. 5A-5C show normalized sequencing depth from linkage mapping at three regions: IKBKG, STRC, and RCCX. Linkage information was used to align the sequence reads to the reference genome. The alignment of the sequence reads to the reference genome was binned. A high-MAPQ depth-per-bin was determined, and was normalized by a median depth-per-bin. As seen in FIGS. 5A-5C, there are some variations in the alignment depth.Example 2
[0182] In the following example, a correction factor determined based on linkage information was applied to the linkage mapping sequencing depth. The correction factor was determined according to the following equation:where H is the fraction of unlinked sequence reads that are aligned within a first bin of the reference genome with MAPQ of 3 or greater, L is the linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold (per bin), and P is the fraction of linked sequence reads that are mapped to the reference genome within the first bin with a MAPQ below 3. Primary and secondary alignments were used to determine both the linking rate and the fraction of linked sequence reads that are mapped to the reference genome with confidence below MAPQ = 3.
[0183] The following example used a genomic DNA sample from the HG002 cell line to test for copy number variation at several loci.
[0184] FIGS. 6A-6C show a comparison of sequencing depth for baseline, linkage mapping, and after application of the correction factor in a region that includes STRC and pseudogene STRCP1. As shown in FIG. 6 A, without linkage information or the correction factor, there are large variations in sequencing depth which may appear to be deletions (copy number variants). FIG. 6B shows that when linkage information is used to align the sequence reads, the sequencing depth has changed at the locations of STRC and STRCP1 to appear closer to the standard copy number of 2. FIG. 6C shows that when the correction factor is applied, the sequencing depth has now been corrected to more accurately show that the copy number is 2 at STRC and STRCP1, and would likely not result in the false detection of a copy number variant at these loci.
[0185] FIGS. 7A-7F show a comparison of sequencing depth for baseline, linkage mapping, and after application of the correction factor for PMS2 and its pseudogene PMS2CL. As shown in FIGs. 7A and 7B, without linkage information or the correction factor, there are large variations in sequencing depth which may appear to be deletions or duplications (copy number variants). FIGs. 7C and 7D show that when linkage information is used to align the sequence reads, the sequencing depth has changed at the locations of PMS2 and PMS2CL to appear closer to the standard copy number of 2. FIGs. 7E and 7F show the depth signal after application of the correction factor, which shows a potential copy number variant in PMS2 / PMS2CL.
[0186] FIGs. 8A-8F show a comparison of sequencing depth for baseline, linkage mapping, and after application of the correction factor for OTO A and its pseudogene OTOAP1. FIGs. 8A and 8B show the depth signal without linkage information or the correction factor. FIGs. 8C and 8D show the depth signal when linkage information is used to align the sequence reads. FIGs. 8E and 8F show that when the correction factor is applied, it can be seen that there are some duplications in OTOAP1 and in some of the regions surrounding OTOA.Example 3
[0187] In the following example, correction factors with and without the use of primary and secondary alignments were compared. FIGS. 9A-9B illustrate an example where copy number was estimated based on sequencing depth and a correction factor as described herein for the IKBKG locus. In FIG. 9A, only primary alignments were used to determine thecorrection factor. In FIG. 9B, the linking rate and the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold were determined based on both primary and secondary alignments. As shown in FIG. 9B, leveraging secondary alignments improved copy number variant detection for the IKBKG locus and allowed for better visualization of a deletion at the IKBKGP1 locus.Other Considerations
[0188] Conditional language used herein, such as, among others, “can,” “might,” “may,” “e.g.,” and the like, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or states. Thus, such conditional language is not generally intended to imply that features, elements and / or states are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without author input or prompting, whether these features, elements and / or states are included or are to be performed in any particular embodiment. The terms “comprising,” “including,” “having,” “involving,” and the like are synonymous and are used inclusively, in an open-ended fashion, and do not exclude additional elements, features, acts, operations, and so forth. Also, the term “or” is used in its inclusive sense (and not in its exclusive sense) so that when used, for example, to connect a list of elements, the term “or” means one, some, or all of the elements in the list.
[0189] Disjunctive language such as the phrase “at least one of X, Y or Z,” unless specifically stated otherwise, is otherwise understood with the context as used in general to present that an item, term, etc., may be either X, Y or Z, or any combination thereof (such as X, Y and / or Z). Thus, such disjunctive language is not generally intended to, and should not, imply that certain embodiments require at least one of X, at least one of Y or at least one of Z to each be present.
[0190] The terms “about” or “approximate” and the like are synonymous and are used to indicate that the value modified by the term has an understood range associated with it, where the range can be ±20%, ±15%, ±10%, ±5%, or ±1%. The term “substantially” is used to indicate that a result (such as a measurement value) is close to a targeted value, where closecan mean, for example, the result is within 80% of the value, within 90% of the value, within 95% of the value, or within 99% of the value.
[0191] Unless otherwise explicitly stated, articles such as “a” or “an” should generally be interpreted to include one or more described items.
[0192] While the above detailed description has shown, described, and pointed out novel features as applied to illustrative embodiments, it will be understood that various omissions, substitutions, and changes in the form and details of the devices or algorithms illustrated can be made without departing from the spirit of the disclosure. As will be recognized, certain embodiments described herein can be embodied within a form that does not provide all of the features and benefits set forth herein, as some features can be used or practiced separately from others. All changes which come within the meaning and range of equivalency of the claims are to be embraced within their scope.
[0193] It should be appreciated that all combinations of the foregoing concepts (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.
[0194] The scope of the present disclosure is not intended to be limited by the specific disclosures of examples in this section or elsewhere in this specification, and may be defined by claims as presented in this section or elsewhere in this specification or as presented in the future. The language of the claims is to be interpreted broadly based on the language employed in the claims and not limited to the examples described in the present specification or during the prosecution of the application, which examples are to be construed as nonexclusive.
Claims
WHAT IS CLAIMED IS:
1. A method for detecting copy number variants in a nucleic acid segment from a genomic DNA sample, the method comprising: generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detecting a copy number variant in the genomic DNA based on the estimated copy numbers.
2. The method of claim 1, wherein a higher corrected sequencing depth in a first region as compared to other regions corresponds to a higher estimated copy number, and a lower corrected sequencing depth in the first region as compared to other regions corresponds to a lower estimated copy number.
3. The method of any of claims 1-2, wherein estimating the one or more copy numbers for the nucleic acid segment based on the correction factor comprises: binning the alignment of the sequence reads to the reference genome; determining a depth-per-bin of sequence reads with a mapping quality score above a threshold; adjusting the depth-per-bin for GC content; normalizing the depth-per-bin by a median depth-per-bin; and applying the correction factor to the normalized depth-per-bin to generate a corrected depth-per-bin signal.
4. The method of claim 3, wherein determining depth-per-bin of sequence reads with a mapping quality score above a threshold comprises, for each bin, counting the number of sequence reads with a MAPQ score of 3 or greater which align to the bin.
5. The method of claim 3 or claim 4, further comprising segmenting the corrected depth-per-bin signal.
6. The method of claim 5, further comprising converting segmented depth-per-bin signals to integer copy numbers, to determine a copy number genotype.
7. The method of any of claims 1-6, wherein the correction factor is determined based on one or more of the following: a fraction of sequence reads that do not have a link to another sequence read and are aligned within a first bin of the reference genome with confidence above a predetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold.
8. The method of claim 7, wherein the fraction of sequence reads that do not have a link to another sequence read and that are aligned within the first bin of the reference genome with confidence above a predetermined threshold, is estimated based on a fraction of sequence reads that have a mapping quality score above a predetermined threshold, over all sequence reads within the first bin that do not have a link to another sequence read.
9. The method of claim 8, wherein sequence reads that have a mapping quality score above a predetermined threshold are sequence reads that have a MAPQ score of 3 or greater.
10. The method of any of claims 7-9, wherein the linking rate is estimated based on all of the sequence reads, or a random subset thereof.
11. The method of claim 7-9, wherein the linking rate is estimated per bin.
12. The method of any of claims 7-11, wherein the fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold comprises a fraction of sequence reads which have a MAPQ score below the predetermined threshold, over all linked sequence reads within the first bin.
13. The method of claim 12, wherein the predetermined threshold for the MAPQ score is 3.
14. The method of any of claims 1-13, wherein the correction factor is determined using primary and secondary alignments of each sequence read to the reference genome.
15. The method of claim 14, wherein one or more of the linking rate and the fraction of linked sequence reads that are mapped to the reference genome with confidence below a predetermined threshold, are determined using primary and secondary alignments of each sequence read to the reference genome.
16. The method of any of claims 7-15, wherein the correction factor is determined based on each of: a fraction of sequence reads that do not have a link to another sequence read and are aligned within a first bin of the reference genome with confidence above a predetermined threshold; a linking rate comprising a proportion of sequence reads which have a link to another sequence read with a score over a predetermined threshold; and a fraction of linked sequence reads that are mapped to the reference genome within the first bin with confidence below a predetermined threshold.
17. The method of any of claims 1-16, wherein the linkage information is determined based on location information of clusters of the fragments on the flow cell.
18. The method of claim 17, wherein the linkage information comprises information linking a first sequence read from a first cluster to a second sequence read from a second cluster based on a distance between the first and second clusters in the flow cell.
19. The method of claim 17 or claim 18, wherein the location information comprises a first spatial coordinate and a second spatial coordinate in a cartesian coordinate system.
20. The method of any of claims 1-19, wherein determining linkage information between the sequence reads comprises: obtaining location information for each cluster on a substrate of the flow cell, and assigning sequence reads to target polynucleotides using the obtained location information.
21. The method of claim 20, wherein assigning sequence reads to the target polynucleotides using the obtained location information comprises: determining distances between each cluster, andusing the determined distances to assign sequence reads to a specific target polynucleotide.
22. The method of any of claims 1-21, wherein the sequence reads comprise paired end sequence reads.
23. A method for determining one or more copy numbers for a nucleic acid segment from a genomic DNA sample, the method comprising: generating sequence reads from fragments of the genomic DNA sample bound to a flow cell; determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information.
24. A method for detecting a pathogenic copy number variant in a nucleic acid segment from a genomic DNA sample, comprising the method of any of claims 1-23.
25. The method of any of claims 1-24, wherein the plurality of locations in the reference genome comprise one or more of an IKBKG gene, an IKBKGP1 pseudogene, a RCCX locus, a STK19 gene, a STK19B pseudogene, a C4A gene, a C4B gene, a CYP21A1P pseudogene, a CYP21A2 gene, a TNXA pseudogene, a TNXB gene, a STRC gene, a STRCP1 pseudogene, a PMS2 gene, a PMS2CL pseudogene, a OTOA gene, a OTOAP1 pseudogene, a GBA gene, a GBAP1 pseudogene, a HYDIN gene, or a HYDIN2 pseudogene.
26. The method of any of claims 1-25, further comprising creating an electronic file comprising one or more copy numbers, information related to a copy number variant, or information related to a copy number genotype.
27. A system for detecting copy number variants in a nucleic acid segment from a genomic DNA sample, comprising one or more processors having instructions that when executed perform a method comprising: receiving sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell;determining linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimating one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detecting a copy number variant in the genomic DNA based on the estimated copy numbers.
28. A non-transitory computer-readable medium comprising a plurality of instructions, which when executed by at least one processor, cause the at least one processor to: receive sequence reads sequenced from fragments of the genomic DNA sample bound to a flow cell; determine linkage information between the sequence reads based on the geographic location of each fragment on the flow cell and aligning the sequence reads to a reference genome using the linkage information; estimate one or more copy numbers for the nucleic acid segment based on a sequencing depth by applying a correction factor to the sequencing depth, wherein the correction factor is determined based on linkage information; and detect a copy number variant in the genomic DNA based on the estimated copy numbers.
Citation Information
Patent Citations
Polymerase enzymes and reagents for enhanced nucleic acid sequencing
US20080108082A1
Transposon end compositions and methods for modifying nucleic acids
US20100120098A1
Determining a mudweight of drilling fluids for drilling through naturally fractured formations
US62636004P0
Interactive techniques for speeding up homomorphic linear equation solving on encrypted data
US62637000P0
Method And System For Touchless Gesture Detection And Hover And Touch Detection
US62637002P0