Fast-Forward Sequencing by Synthetic Methods
By hybridizing the polynucleotide with the primer and extending the primer with the labeled nucleotide and specific flow sequence, sequencing data for the first and third regions of the polynucleotide is solved, and efficient and accurate sequencing and variant detection are achieved.
Patent Information
- Application Number
- CN202080048933.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-02-07
- Filing Date
- 2020-05-01
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-05-01
AI Technical Summary
Traditional paired-end sequencing methods are difficult to obtain information about the region between the 3' and 5' ends of polynucleotides, and cannot effectively detect specific variants in unsequencing regions. Long-range sequencing is relatively slow and sequencing errors are prone to occur.
By hybridizing the polynucleotide to the primer, extending the primer with the labeled nucleotide and detecting the integrated labeled nucleotide, sequencing data related to the sequence of the first region of the polynucleotide is generated. The primer is then further extended through the second region using nucleotides provided in the second region flow sequence to generate sequencing data related to the sequence of the third region of the polynucleotide.
Efficient sequencing of polynucleotide unsequencing regions is achieved, and structural variants and single nucleotide polymorphisms can be detected, improving the accuracy and efficiency of sequencing data.
Smart Images

Figure CN114096682B_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of U.S. Provisional Patent Application Serial No. 62 / 842,534, filed on May 3, 2019; U.S. Provisional Patent Application Serial No. 62 / 904,274, filed on September 23, 2019; and U.S. Provisional Patent Application Serial No. 62 / 971,530, filed on February 7, 2020; the respective contents of which are incorporated herein by reference in their entireties.
[0003] Submission of the Sequence Listing as an ASCII Text File
[0004] The content of the following submitted ASCII text file is incorporated herein by reference in its entirety: a sequence listing in computer - readable form (CRF) (filename: 165272000440SEQLIST.TXT, date of record: April 27, 2020, size: 5KB). Field of the Invention
[0005] Methods for sequencing polynucleotides are described herein, including methods for generating paired - end sequencing read pairs, and methods for analyzing sequencing data obtained from such sequencing methods.
[0006] Background
[0007] Paired - end sequencing has been used to obtain sequencing data from the 3' and 5' ends of polynucleotide molecules. Typically, a sequencing primer is hybridized to a DNA polynucleotide to be sequenced, and several bases are sequenced to obtain sequencing data from the first end of the polynucleotide. Then a second sequencing primer is hybridized to the complementary strand near the other end of the polynucleotide, and it is sequenced to determine the sequencing data from the other end of the polynucleotide. Based on the fact that the sequencing data is obtained from the same sequencing cluster, the sequencing data from the 3' and 5' ends of the polynucleotide is coupled. Paired - end sequencing is often used in next - generation sequencing (NGS) protocols.
[0008] However, with traditional paired - end sequencing, little (or no) information is obtained for the region between the 3’ and 5’ ends of the polynucleotide. Although paired - end sequencing data can be used for certain analytical purposes, it cannot be used to detect specific variants in the unsequenced regions of the polynucleotide. Certain long - range sequencing techniques have been developed to sequence polynucleotide regions that are typically missing when using traditional paired - end sequencing methods. However, long - range sequencing is relatively slow and prone to significant sequencing errors.
[0009] Summary of the Invention
[0010] Methods for sequencing polynucleotides are described herein, including methods for generating paired - end sequencing read pairs, and methods for analyzing sequencing data obtained from such sequencing methods.
[0011] Methods for generating paired coupled sequencing reads from a polynucleotide include: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order, or (iii) the extension of the primer through the second region proceeds faster than the extension of the primer in step (b); and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. In some embodiments, the extension of the primer through the second region proceeds faster than the extension of the primer in step (b). In some embodiments, the methods for generating paired coupled sequencing reads include correlating the sequencing data of the first region with the sequencing data of the third region.
[0012] In some embodiments, methods for generating paired coupled sequencing reads from a polynucleotide include (a) hybridizing a primer with a first region of the polynucleotide to form a hybridization template; (b) extending the primer through a second region using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; and (c) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (b) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. In some embodiments, the first region includes a naturally occurring sequence targeted by the primer.
[0013] In some embodiments, the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer. In some embodiments, at least a portion of the nucleotides used to extend the primer through the second region are unlabeled nucleotides. In some embodiments, the nucleotides used to extend the primer through the second region are unlabeled nucleotides.
[0014] In some embodiments, a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order.
[0015] In some embodiments of the method of generating paired sequencing reads, the second region flow order includes five or more nucleotide flows. In some embodiments, each nucleotide flow includes a single nucleotide base. In some embodiments, for 50% or more of the possible SNP arrangements at 5% or more of the random sequencing start positions, the second region flow order induces a signal change at more than two flow positions. In some embodiments, the induced signal change is a change in signal intensity, or a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the induced signal change is a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the second region flow order has an efficiency of integrating 0.6 or more bases per flow.
[0016] In some embodiments, the method of generating paired sequencing reads includes determining the expected sequencing data for the second region using a reference sequence and the second region flow order. In some embodiments, a primer is extended through the third region using nucleotides provided in a third region flow order, and the method further includes determining the expected sequencing data for the third region using the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region. In some embodiments, the third region flow order includes five or more nucleotide flows. In some embodiments, each nucleotide flow includes a single nucleotide base. In some embodiments, for 50% or more of the possible SNP arrangements at 5% or more of the random sequencing start positions, the third region flow order induces a signal change at more than two flow positions. In some embodiments, the induced signal change is a change in signal intensity, or a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the induced signal change is a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the third region flow order has an efficiency of integrating 0.6 or more bases per flow.
[0017] In some embodiments of a method of generating paired sequencing reads, a primer is extended through a third region using nucleotides provided in a third region flow order, and the method further includes determining expected sequencing data for the third region using a reference sequence of a second region, a second region flow order, a third region flow order, and sequencing data associated with the sequence of the third region, where the sequencing data associated with the sequence of the third region is the same or different sequencing data generated for the third region. In some embodiments, the expected reference data for the second region or the third region includes a binary or non-binary flowgram. In some embodiments, the method further includes determining expected test variant sequencing data for the second region using a second region flow order and a second reference sequence of the second region, where the second reference sequence includes a test variant. In some embodiments, a primer is extended through a third region using nucleotides provided in a third region flow order, and the method further includes determining expected test variant sequencing data for the third region using a second reference sequence of the second region, a second region flow order, a third region flow order, and a reference sequence of the third region. In some embodiments, a primer is extended through a third region using nucleotides provided in a third region flow order, and the method further includes determining expected test variant sequencing data for the third region using a second reference sequence of the second region, a second region flow order, a third region flow order, and sequencing data associated with the sequence of the third region, where the sequencing data associated with the sequence of the third region is the same or different sequencing data generated for the third region. In some embodiments, the expected reference sequencing data for the second region or the third region includes a binary or non-binary flowgram.
[0018] In some embodiments, methods of generating paired sequencing reads include determining expected sequencing data for a second region using a reference sequence and a second region flow order. In some embodiments, a primer extended in nucleotide extension step (d) provided in a third region flow order is used, and the method further includes determining expected sequencing data for a third region using a reference sequence of the second region, the second region flow order, the third region flow order, and a reference sequence of the third region. In some embodiments, a primer extended in nucleotide extension step (d) provided in a third region flow order is used, and the method further includes determining expected sequencing data for a third region using a reference sequence of the second region, the second region flow order, the third region flow order, and sequencing data associated with the sequence of the third region, wherein the sequencing data associated with the sequence of the third region is sequencing data generated in the same or a different step (d). In some embodiments, the expected reference data for the second region or the third region includes a binary or non-binary flowgram. In some embodiments, the method includes determining expected test variant sequencing data for the second region using the second region flow order and a second reference sequence of the second region, wherein the second reference sequence includes a test variant. In some embodiments, a primer extended in nucleotide extension step (d) provided in a third region flow order is used, and the method further includes determining expected test variant sequencing data for the third region using the second reference sequence of the second region, the second region flow order, the third region flow order, and a reference sequence of the third region. In some embodiments, a primer extended in nucleotide extension step (d) provided in a third region flow order is used, and the method further includes determining expected test variant sequencing data for the third region using the second reference sequence of the second region, the second region flow order, the third region flow order, and sequencing data associated with the sequence of the third region, wherein the sequencing data associated with the sequence of the third region is sequencing data generated in the same or a different step (d). In some embodiments, the expected reference sequencing data for the second region or the third region includes a binary or non-binary flowgram.
[0019] In some embodiments, generating paired sequencing reads further comprises: (e) further extending the primer extended in step (d) through a fourth region using nucleotides provided in a fourth region flow order, wherein (i) the primer is extended through the fourth region without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer, (ii) a mixture of at least two different types of nucleotide bases is used in at least one step of the fourth region flow order, or (iii) the extension of the primer through the fourth region proceeds faster than the primer extension in step (b) or step (d); and (f) generating sequencing data related to the sequence of a fifth region of the polynucleotide by further extending the primer extended in step (e) using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides. In some embodiments, the method further comprises correlating the sequencing data of the fifth region with the sequencing data of the first region or the sequencing data of the third region.
[0020] Also described herein is a method of mapping paired sequencing reads to a reference sequence, comprising: mapping a first region or a portion thereof or a third region or a portion thereof of the paired sequencing reads to the reference sequence; and using distance information indicative of the length of the second region to map an unmapped first region or a portion thereof or an unmapped third region or a portion thereof to the reference sequence.
[0021] Also provided is a method of detecting a structural variant, comprising: mapping a first region or a portion thereof or a third region or a portion thereof of the paired sequencing reads to the reference sequence; using distance information indicative of the length of the second region to determine an expected locus within the reference sequence of an unmapped first region or a portion thereof or an unmapped third region or a portion thereof; determining expected sequencing data of the sequence at the expected locus based on the reference sequence; and detecting the structural variant by comparing the sequencing data of the unmapped first region or a portion thereof or the unmapped third region or a portion thereof with the expected sequencing data, wherein a difference between the sequencing data of the unmapped first region or a portion thereof or the unmapped third region or a portion thereof and the expected sequencing data indicates the structural variant.
[0022] Also provided herein is a method of detecting a structural variant, comprising mapping a first region or a portion thereof or a third region or a portion thereof of the paired sequencing reads to the reference sequence, wherein the unmapped first region or the unmapped third region is unmappable within the reference sequence. In some embodiments, the method further comprises determining the locus of the structural variant within the reference sequence based on expected distance information indicative of the length of the second region.
[0023] In some embodiments, the unmapped first region or a portion thereof, or the unmapped third region or a portion thereof, is located within an insertion relative to the reference sequence. In some embodiments, the unmapped first region or a portion thereof, or the unmapped third region or a portion thereof, bridges the start or end of an insertion relative to the reference sequence.
[0024] Also provided herein is a method for detecting a structural variant, comprising mapping a first region or a portion thereof and a third region or a portion thereof of a paired sequencing read to a reference sequence; determining mapped distance information between the mapped first region and the mapped third region; and detecting a structural variant by comparing the mapped distance information with expected distance information of a second region, wherein a difference between the mapped distance information and the expected distance information indicates a structural variant. In some embodiments, the structural variant is a chromosome fusion, inversion, insertion, or deletion. In some embodiments, the variant is an insertion or deletion within the second region.
[0025] In some embodiments of the methods described herein, information related to the flow order of the second region and the probability distribution of bases in the second region are used to determine the distance information. In some embodiments, the information related to the flow order of the second region is the number of different types of nucleotide bases simultaneously used to extend a primer in step (c). In some embodiments, the probability distribution of bases in the second region is determined by the distribution of bases within the genome.
[0026] In some embodiments of the methods described herein, the distance information is derived from expected sequencing data of the second region determined using the reference sequence and the flow order of the second region. In some embodiments, the expected sequencing data includes a binary or non-binary flowgram.
[0027] The present disclosure further describes a method for mapping paired - end sequencing reads to a reference sequence, comprising: mapping a first region or a portion thereof and a third region or a portion thereof of the paired - end sequencing reads to the reference sequence at two or more different position pairs comprising a first position and a second position; and for two or more position pairs, using first distance information indicating the length of a second region and second distance information indicating the distance between the first position and the second position to select the correct position pair. In some embodiments, the first distance information is determined using information related to the flow order of the second region and the probability distribution of bases in the second region. In some embodiments, the information related to the flow order of the second region is the number of different types of nucleotide bases simultaneously used to extend the primer in step (c). In some embodiments, the probability distribution of bases in the second region is determined by the distribution of bases within the genome. In some embodiments, the first distance information is derived from the expected sequencing data of the second region determined using the reference sequence and the flow order of the second region. In some embodiments, the expected reference sequencing data includes a binary or non - binary flowgram.
[0028] The present disclosure also describes a method for detecting variants between two sequencing regions of paired - end sequencing reads generated according to any of the above - described methods, wherein a nucleotide provided in the flow order of the third region is used to extend the primer extended in step (d), comprising: mapping the first region or a portion thereof to the reference sequence; using (1) the reference sequence of the second region, the flow order of the second region, the flow order of the third region, and the reference sequence of the third region, or (2) the reference sequence of the second region, the flow order of the second region, the flow order of the third region, and the generated sequencing data related to the sequence of the third region, to determine the expected sequencing data of the third region or a portion thereof, wherein the generated sequencing data related to the sequence of the third region is the same or different from the sequencing data generated in step (d); and detecting the presence of a variant by comparing the expected sequencing data of the third region with the generated sequencing data related to the sequence of the third region. In some embodiments, the variant is a structural variant. In some embodiments, the structural variant is a chromosomal fusion, inversion, insertion, or deletion. In some embodiments, the variant is a single nucleotide polymorphism (SNP). In some embodiments, the method is used to detect a test variant and the reference sequence contains the test variant. In some embodiments, the test variant is selected by identifying the test variant within a second polynucleotide. In some embodiments, the method further comprises associating the detected test variant with the allele sequenced in the first region or the third region of the polynucleotide.
[0029] The present disclosure also describes a method of detecting variants between two sequencing regions of a paired sequencing read generated according to any of the above methods, wherein the extended primer is extended through a third region using nucleotides provided in a third region flow order, the method comprising: mapping the first region or a portion thereof to a reference sequence; using (1) the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region, or (2) the reference sequence of the second region, the second region flow order, the third region flow order, and the generated sequencing data related to the sequence of the third region, to determine the expected sequencing data of the third region or a portion thereof, wherein the generated sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region; and detecting the presence of a variant by comparing the expected sequencing data of the third region with the generated sequencing data related to the sequence of the third region. In some embodiments, the variant is a structural variant. In some embodiments, the structural variant is a chromosomal fusion, inversion, insertion, or deletion. In some embodiments, the variant is a single nucleotide polymorphism (SNP). In some embodiments, the method is used to detect a test variant, and the reference sequence includes the test variant. In some embodiments, the test variant is selected by identifying the test variant within a second polynucleotide. In some embodiments, the method includes associating the detected test variant with an allele sequenced in the first region or the third region of the polynucleotide.
[0030] The present disclosure further describes a method of generating a paired sequencing read for detecting the presence of a base transversion in an unsequenced region of a polynucleotide, the method comprising: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides. In some embodiments, the primer is extended through the second region without detecting the presence of the label of the nucleotides incorporated into the extended primer.
[0031] The present disclosure also describes a method for generating paired sequencing reads from a polynucleotide, comprising: (a) hybridizing a primer to a first region of the polynucleotide to form a hybridization template; (b) extending the primer through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and (c) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (b) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. In some embodiments, the first region comprises a naturally occurring sequence targeted by the primer. In some embodiments, the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer.
[0032] In some embodiments, a method for detecting the presence of base transversions in an unsequenced region of a polynucleotide comprises: mapping a first region or a portion thereof and a third region or a portion thereof of paired sequencing reads generated according to the method described above to a reference sequence, wherein the primer extended in step (d) is extended using nucleotides provided in the flow order of the third region; using the flow order of the second region, the flow order of the third region, and the reference sequence to determine the expected sequencing data for the third region; and detecting the presence of base transversions based on the difference between the expected sequencing data for the third region and the generated sequencing data for the third region. In some embodiments, the flow order of the second region, the flow order of the third region, the reference sequence of the second region, and the reference sequence of the third region are used to determine the expected sequencing data for the third region. In some embodiments, the flow order of the second region, the flow order of the third region, the reference sequence of the second region, and the generated sequence data related to the sequence of the third region are used to determine the expected sequencing data for the third region, wherein the generated sequence data related to the sequence of the third region is the same as or different from the sequence data generated in step (d). In some embodiments, the expected sequencing data for the third region comprises a binary or non-binary flowgram.
[0033] The present disclosure further describes a method of generating one or more consensus sequences, including assembling a plurality of paired-end sequencing reads. In some embodiments, one or more consensus sequences are assembled using distance information indicative of the length of a second region of the plurality of paired-end sequencing reads. In some embodiments, the distance information is determined using the probability distribution of bases in the second region and information related to the flow order of the second region. In some embodiments, the information related to the flow order of the second region is the number of different types of nucleotide bases used to extend the primer in step (c). In some embodiments, the probability distribution of bases in the second region is determined by the distribution of bases within the genome. In some embodiments, the distance information is derived from the expected reference sequencing data of the second region determined using a reference sequence and the flow order of the second region. In some embodiments, the expected reference sequencing data includes a binary or non-binary flowgram.
[0034] In some embodiments, the method of generating one or more consensus sequences further includes verifying a portion of a consensus sequence selected from the one or more consensus sequences, the method using a selected paired-end sequencing read related to the portion of the selected consensus sequence, wherein the primer is extended using nucleotides provided in a third region flow order, the primer being extended in step (d) when generating the selected paired-end sequencing read, the verification including: determining the expected sequencing data of the third region of the selected paired-end sequencing read using the second region flow order, the third region flow order, and the portion of the selected consensus sequence; and verifying the portion of the selected consensus sequence by comparing the expected sequencing data of the third region of the selected paired-end sequencing read with the generated sequencing data of the third region.
[0035] Also described is a method of verifying a test variant status, including: comparing the variant status across a plurality of overlapping paired-end sequencing reads that include a locus corresponding to the locus of the test variant; and verifying the variant status based on the comparison. In some embodiments, the first region or the third region of the selected paired-end sequencing read overlaps at least a portion of the second region of other paired-end sequencing reads among the plurality of overlapping paired-end sequencing reads. In some embodiments, the variant status of the selected paired-end sequencing read indicates a variant in the first region or the third region of the selected paired-end sequencing read. In some embodiments, the second region of the selected paired-end sequencing read overlaps at least a portion of the second region of other paired-end sequencing reads among the plurality of overlapping paired-end sequencing reads. In some embodiments, the variant status of the selected paired-end sequencing read indicates a variant in the second region of the selected paired-end sequencing read.
[0036] This document further describes methods for detecting short genetic variants in a test sample, including: generating paired sequencing reads according to any of the above methods; comparing sequencing data related to the sequence of a third region of a polynucleotide with expected sequencing data of the expected sequence of the third region of the polynucleotide; and determining the presence or absence of a short genetic variant in a second region of the polynucleotide. In some embodiments, comparing the sequencing data related to the sequence of the third region of the polynucleotide with the expected sequencing data of the third region of the polynucleotide includes determining a match score, which indicates the likelihood that the sequencing data generated for the third region of the polynucleotide matches the expected sequencing data of the third region of the polynucleotide; and determining the presence or absence of a target short genetic variant in the second region of the polynucleotide includes using the determined match score. In some embodiments, the expected sequencing data of the third region of the polynucleotide is obtained by computer simulation of the sequencing and expected sequence of the third region of the polynucleotide. In some embodiments, the sequencing data related to the sequence of the first region or the sequencing data related to the sequence of the third region includes a flow signal representing a base count, which indicates the number of bases integrated at each flow position within a plurality of flow positions. In some embodiments, the flow signal includes a statistical parameter indicating the likelihood of a base count at at least one base count at each flow position. In some embodiments, the flow signal includes a statistical parameter indicating the likelihood of a plurality of base counts at each flow position. In some embodiments, the sequencing data related to the sequence of the third region includes a flow signal representing a base count, which indicates the number of bases integrated at each flow position within a plurality of flow positions, where the flow signal includes a statistical parameter indicating the likelihood of a plurality of base counts; and the method further includes selecting a statistical parameter corresponding to the base count of the expected sequence at each flow position in the sequencing data and determining a match score indicating the likelihood of the sequencing data set matching the expected sequence. In some embodiments, the match score is a combined value of the selected statistical parameters across flow positions in the sequencing data.
[0037] In some embodiments of the above method, the flow cycle order includes four separately repeated flows in the same order.
[0038] In some embodiments of the above method, the flow cycle order includes five or more separately repeated flows.
[0039] In some embodiments of the above method, generating paired sequencing reads further comprises: further extending the primer through the fourth region using nucleotides provided in a fourth region flow order, wherein (i) the primer is extended through the fourth region without detecting the presence or absence of a label of the nucleotide incorporated into the extended primer, (ii) a mixture of at least two different types of nucleotide bases is used in at least one step of the fourth region flow order, or (iii) the extension of the primer through the fourth region proceeds faster than the extension of the primer through the first region or the third region; and generating sequencing data related to the sequence of a fifth region of the polynucleotide by further extending the primer extended through the fourth region using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides. In some embodiments, the method further comprises correlating the sequencing data of the fifth region with the sequencing data of the first region or the sequencing data of the third region.
[0040] In some embodiments of the above method, the polynucleotide is amplified using rolling circle amplification.
[0041] The present disclosure also describes a method for detecting short genetic variants in a test sample, comprising: (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide that includes at least a first copy of the polynucleotide and a second copy of the polynucleotide; (b) hybridizing the RCA-amplified polynucleotide with a primer to form a hybridization template; (c) generating sequencing data related to the sequence of a first region of the polynucleotide within the first copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (d) further extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order, or (iii) the extension of the primer through the second region of the polynucleotide within the first copy of the polynucleotide proceeds faster than the extension of the primer through the first region; (e) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (f) comparing the sequencing data generated for the third region of the polynucleotide with the sequencing data expected for the expected sequence of the third region of the polynucleotide; (g) determining the presence of a short genetic variant in the second region of the polynucleotide; (h) generating sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; and (i) determining the identity of the short genetic variant in the second region of the polynucleotide. In some embodiments, the extension of the primer through the second region of the polynucleotide within the first copy of the polynucleotide proceeds faster than the extension of the primer through the first region of the polynucleotide within the first copy. In some embodiments, based on the determination of the presence of a short genetic variant in the second region of the polynucleotide, sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide is dynamically generated. In some embodiments, the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of labels of nucleotides incorporated into the extended primer. In some embodiments, at least a portion of the nucleotides used to extend the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides. In some embodiments, the nucleotides used to extend the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides. In some embodiments, a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order.In some embodiments, a mixture of three different types of nucleobases is used in at least one step of the second region flow order.
[0042] The present disclosure further describes a method for detecting short genetic variants in a test sample, comprising: (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide, the RCA-amplified polynucleotide comprising at least a first copy of the polynucleotide and a second copy of the polynucleotide; (b) hybridizing a primer to a first region of the polynucleotide within the first copy of the polynucleotide to form a hybridization template; (c) extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides; (e) comparing the sequencing data generated for the third region of the polynucleotide with expected sequencing data for an expected sequence of the third region of the polynucleotide; (f) determining the presence of a short genetic variant in the second region of the polynucleotide; (g) generating sequencing data related to the sequence of a second region of the polynucleotide within the second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides; and (h) determining the identity of the short genetic variant in the second region of the polynucleotide. In some embodiments, the first region comprises a naturally occurring sequence targeted by the primer. In some embodiments, based on the determination of the presence of a short genetic variant in the second region of the polynucleotide, sequencing data related to the sequence of a second region of the polynucleotide within the second copy of the polynucleotide is generated dynamically. In some embodiments, the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer. In some embodiments, at least a portion of the nucleotides used to extend the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides. In some embodiments, the nucleotides used to extend the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides. In some embodiments, a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order. In some embodiments, a mixture of three different types of nucleobases is used in at least one step of the second region flow order.
[0043] The present disclosure also describes a method for synchronizing sequencing primers within a sequencing cluster, comprising: (a) hybridizing the primers to copies of polynucleotides within the sequencing cluster; (b) extending the primers through a first region of the polynucleotide copies using labeled nucleotides flowing in a first region cycle; (c) extending the primers through a second region of the polynucleotide copies using one or more rephasing flows, wherein a mixture of at least two different types of nucleobases is used in at least one of the one or more rephasing flows; and (d) extending the primers through a third region of the polynucleotide copies using labeled nucleotides flowing in a third region cycle. In some embodiments, a mixture of three different types of nucleobases is used in at least one of the one or more rephasing flows. In some embodiments, the one or more rephasing flows include four or more flow steps. In some embodiments, the one or more rephasing flows include, in any order: (i) a first flow comprising a mixture containing A, C, and G nucleotides and omitting T nucleotides; (ii) a second flow comprising a mixture containing T, C, and G nucleotides and omitting A nucleotides; (iii) a third flow comprising a mixture containing T, A, and G nucleotides and omitting C nucleotides; and (iv) a fourth flow comprising a mixture containing T, A, and C nucleotides and omitting G nucleotides. In some embodiments, the method includes generating sequencing data related to the sequence of the first region by detecting the presence or absence of incorporated labeled nucleotides while extending the primers through the first region. In some embodiments, the method includes generating sequencing data related to the sequence of the third region by detecting the presence or absence of incorporated labeled nucleotides while extending the primers through the third region.
[0044] The present disclosure also describes a system, comprising one or more processors; and a non-transitory storage medium containing one or more programs that are executable by the one or more processors to receive information related to one or more coupled sequencing reads; and perform any one or more of the methods described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 A schematic diagram illustrating an exemplary method for generating coupled sequencing read pairs from a polynucleotide.
[0047] Figure 2 A schematic diagram illustrating an exemplary method for generating expected sequencing data using a reference sequence.
[0048] Figure 3 A schematic diagram illustrating how a coupled sequencing read pair is mapped to a reference sequence using distance information indicating the length of the second region of the coupled sequencing read pair when the third region of the coupled sequencing read pair maps to two different loci.
[0049] Figure 4 Illustrates how to map a paired - end sequencing read pair to a reference sequence when the third region of the paired - end sequencing read pair maps to a repetitive region, using distance information indicating the length of the second region of the paired - end sequencing read pair.
[0050] Figure 5 Schematic diagram showing how to use paired - end sequencing read pairs to detect insertions in a subject's genome.
[0051] Figure 6 Illustrates an exemplary method for detecting structural variants using paired - end sequencing read pairs.
[0052] Figure 7 Schematic diagram showing how to use paired - end sequencing read pairs to detect a structural variant in a subject's genome, where the structural variant is an insertion.
[0053] Figure 8 Schematic diagram showing how to use paired - end sequencing read pairs to detect a structural variant in a subject's genome, where the structural variant is a deletion.
[0054] Figure 9 Schematic diagram showing how to use paired - end sequencing read pairs to detect a structural variant in a subject's genome, where the structural variant is an inversion.
[0055] Figure 10 Schematic diagram showing how to use paired - end sequencing read pairs to detect a structural variant in a subject's genome, where the structural variant is a chromosome fusion.
[0056] Figure 11 Illustrates an exemplary method for detecting structural variants using paired - end sequencing read pairs.
[0057] Figure 12 Illustrates a schematic diagram that demonstrates an example of how to use distance information indicating the length of the second region of a paired - end sequencing read pair to detect structural variants using the paired - end sequencing read pair.
[0058] Figure 13 Illustrates an exemplary method for detecting variants between two sequencing regions of a paired - end sequencing read pair.
[0059] Figure 14A Shows sequencing data obtained by extending a primer with the sequence 5’ - TATGGTCGTCGA - 3’ (SEQ ID NO:15) using a repetitive flow - cycle order of T - A - C - G. The sequencing data represents the extended primer strand and is equally valid for sequencing information of the complementary template strand that can be easily determined. Figure 14B Shows Figure 14AThe sequencing data shown, given the sequencing data, selects the most likely sequence based on the highest likelihood at each flow position (as indicated by the asterisk). Figure 14C shows Figure 14A the sequencing data shown, where the traces represent two different candidate sequences (each represented by its complementary sequence): TATGGTCATCGA (SEQ ID NO:16) (solid circles) and TATGGTCGTCGA (SEQ ID NO:15) (open circles). The likelihood that the sequencing data matches a given sequence can be determined by the product of the likelihoods that each flow position matches a candidate sequence.
[0060] Figure 15A shows the alignment results of sequencing reads R1 (SEQ ID NO:15), R2 (SEQ ID NO:17), and R3 (SEQ ID NO:18) (each represented by the sequence of the extension primer) with two candidate sequences H1 (SEQ ID NO:19) and H2 (SEQ ID NO:20) (each represented by their complementary sequences). Figure 15B shows the sequencing data corresponding to R1, where the traces represent H1 (solid circles) and H2 (open circles). Figure 15C shows the sequencing data corresponding to R2, where the traces represent H1 (solid circles) and H2 (open circles). Figure 15D shows the sequencing data corresponding to R3, where the traces represent H1 (solid circles) and H2 (open circles).
[0061] Figure 16 shows the sequencing data of a putative nucleic acid molecule sequenced using an A-T-G-C flow cycle order. Traces can be generated using the potential haplotype sequences (each represented by its complementary sequence) TATGGTCG-TCGA (SEQ ID NO:21) (H1) and TATGGTCGATCG (SEQ ID NO:22) (H2), where H1 has a 1-base deletion relative to H2. The sequencing data has a better match with the H2 candidate sequence, and no indels are determined in this sequence.
[0062] Figure 17 shows an exemplary schematic for comparing paired coupled sequencing reads to determine the test variant status.
[0063] Figure 18 illustrates an example of a computing device according to one embodiment, which can be used to implement the methods described herein.
[0064] Figure 19AShows the signals from bases incorporated in the first and third regions after each flow sequencing cycle when extending a sequencing primer through a polynucleotide. No data is collected within the second region because primer extension through this region accelerates without base incorporation detected.
[0065] Figure 19B Shows the signals from bases incorporated in the first and third regions after each flow sequencing cycle when extending a sequencing primer through a polynucleotide. Data is collected through the second region but not shown to compress the size of the figure.
[0066] Figure 20A - 20E Shows, in an exemplary simulated sequencing scheme, the number of primers extended against the same polynucleotide template after a 100 - nucleotide flow ( Figure 20A ) and a re - phasing flow designed to synchronize primers within the sequencing clusters. The shown re - phasing flow order is a four - step order that includes nucleotide flow 101 ( Figure 20B ), flow 102 ( Figure 20C ), flow 103 ( Figure 20D ) and flow 104 ( Figure 20E ).
[0067] Figure 21A - 21E Shows, in another exemplary simulated sequencing scheme, the number of primers extended against the same polynucleotide template after a 100 - nucleotide flow ( Figure 21A ) and a re - phasing flow designed to synchronize primers within the sequencing clusters. The shown re - phasing flow order is a four - step order that includes nucleotide flow 101 ( Figure 21B ), flow 102 ( Figure 21C ), flow 103 ( Figure 21D ) and flow 104 ( Figure 21E ).
[0068] Figure 22A - 22E Shows, in another exemplary simulated sequencing scheme, the number of primers extended against the same polynucleotide template after a 100 - nucleotide flow ( Figure 22A ) and a re - phasing flow designed to synchronize primers within the sequencing clusters. The shown re - phasing flow order is a four - step order that includes nucleotide flow 101 ( Figure 22B ), flow 102 ( Figure 22C ), flow 103 ( Figure 22D ) and flow 104 ( Figure 22E ).
[0069] Figure 23 Shows the sensitivity of SNP alignment detection for four exemplary flow cycle orders, including three of which are extended flow cycle orders, given a random sequencing start position. At Figure 23In this case, the x-axis indicates the fraction of the mobile phase (or fragmentation start position), while the y-axis indicates the fraction of SNP arrangements with induced signal changes (i.e., new zeros or new non-zero signals) at more than two mobile positions.
[0070] Figure 24 A matrix is shown that shows the base detection sensitivity of various SNP variants detected using a simulated fast-forward sequencing protocol, where a second region of a synthetic polynucleotide is sequenced using a repeated four-step flow cycle, with each flow having a single nucleotide base.
[0071] Figure 25A Shows the average base integration of the flows in the first, second, and third regions for a simulated fast-forward sequencing protocol using a repeated four-step flow cycle, where each flow includes a mixture of three different nucleotide bases. Figure 25B A matrix of the detection sensitivity of variant base pairs to reference bases in terms of electric potential. Figure 25C Shows the distribution of base coverage in synthetic reads.
[0072] Figure 26A Shows the distribution of the sum of the total phase error (lag phase error plus lead phase error) accumulated on 10,000 simulated flow maps for a control protocol (105 rounds of T-G-C-A flow cycles) or a rephasing protocol (105 rounds of T-G-C-A flow cycles, where a rephasing flow containing a mixture of C and G is used after every 24th flow). The mean and standard deviation are shown in the legend. Also shown is the distribution integral of the control and rephasing protocols.
[0073] Figure 26B Shows the distribution of the sum of the total phase error (lag phase error plus lead phase error) accumulated on 10,000 simulated flow maps for a control protocol (105 rounds of T-G-C-A flow cycles) or a rephasing protocol (105 rounds of T-G-C-A flow cycles, where a rephasing flow containing a mixture of C and G is used after every 48th flow). The mean and standard deviation are shown in the legend. Also shown is the distribution integral of the control and rephasing protocols.
[0074] Figure 26C Shows the distribution of the sum of the total phase error (lag phase error plus lead phase error) accumulated on 10,000 simulated flow maps for a control protocol (105 rounds of T-G-C-A flow cycles) or a rephasing protocol (105 rounds of T-G-C-A flow cycles, where a rephasing flow containing a mixture of C and G is used after every 96th flow). The mean and standard deviation are shown in the legend. Also shown is the distribution integral of the control and rephasing protocols.
[0075] Figure 26DShows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow patterns for a control scenario (105 - cycle T - G - C - A flow cycle) or a re - phasing scenario (105 - cycle T - G - C - A flow cycle with a re - phasing flow containing a mixture of C and G after every 192nd flow). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0076] Figure 26E Shows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow patterns for a control scenario (105 - cycle T - G - C - A flow cycle) or a re - phasing scenario (105 - cycle T - G - C - A flow cycle with a re - phasing flow containing a mixture of C, G, and T after every 48th flow). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0077] Figure 26F Shows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow patterns for a control scenario (105 - cycle T - G - C - A flow cycle) or a re - phasing scenario (105 - cycle T - G - C - A flow cycle with a re - phasing flow containing a mixture of C, G, and T after every 96th flow). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0078] Figure 26G Shows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow patterns for a control scenario (105 - cycle T - G - C - A flow cycle) or a re - phasing scenario (105 - cycle T - G - C - A flow cycle with a first re - phasing flow containing a mixture of C, G, and T and a second re - phasing flow containing a mixture of A, C, and G after every 96th flow). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0079] Figure 26HShows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow diagrams for a control scenario (105 rounds of T - G - C - A flow cycles) or a re - phasing scenario (105 rounds of T - G - C - A flow cycles, where after every 192nd flow, a first re - phasing flow containing a mixture of C, G, and T and a second re - phasing flow containing a mixture of A, C, and G are used). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0080] Figure 26I Shows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow diagrams for a control scenario (105 rounds of T - G - C - A flow cycles) or a re - phasing scenario (105 rounds of T - G - C - A flow cycles, where after every 96th flow, a first re - phasing flow containing a mixture of C, G, and T, a second re - phasing flow containing a mixture of A, C, and T, a third re - phasing flow containing a mixture of A, G, and T, and a fourth re - phasing flow containing a mixture of A, C, and G are used). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios.
[0081] Figure 26J Shows the distribution of the sum of the total phase - setting errors (lag phase - setting error plus lead phase - setting error) accumulated over 10,000 simulated flow diagrams for a control scenario (105 rounds of T - G - C - A flow cycles) or a re - phasing scenario (105 rounds of T - G - C - A flow cycles, where after every 192nd flow, a first re - phasing flow containing a mixture of C, G, and T, a second re - phasing flow containing a mixture of A, C, and T, a third re - phasing flow containing a mixture of A, G, and T, and a fourth re - phasing flow containing a mixture of A, C, and G are used). The mean and standard deviation are shown in the legend. Also shown are the distribution integrals for the control and re - phasing scenarios. DETAILED DESCRIPTION OF THE INVENTION
[0083] Described herein are methods for generating paired - end sequencing reads from a polynucleotide and methods for analyzing such paired - end sequencing reads. The paired - end sequencing reads can be analyzed, for example, to map the paired - end sequencing reads to a reference sequence, to detect structural variants, to detect variants (e.g., SNPs) in the region between the paired ends of a polynucleotide, to detect transversions, or to determine or verify a consensus sequence.
[0084] A polynucleotide can hybridize with a sequencing primer, which sequences the first region (i.e., the 3' end) of the polynucleotide. The primer then extends through the second region of the polynucleotide, which can occur at a faster rate than the primer extension through the first region. The accelerated primer extension through the second region can be referred to as "fast-forward sequencing". As further discussed herein, since the primer extends through the second region (rather than the second region being completely skipped by the primer as occurs in more traditional paired-end sequencing), some information (possibly including some sequencing data) about the second region can be obtained even if the second region is not sequenced in the same manner as the first region. For example, the primer can be extended through the second region using only unlabeled nucleotides. Once the sequencing primer extends through the second region, the primer extends into the third region (i.e., the 5' end) of the polynucleotide to sequence the third region. The sequencing data for this region and the third region can be coupled to produce a coupled sequencing read pair for the polynucleotide, and additional sequencing data can be derived from the second region as further described herein.
[0085] In one example, a coupled sequencing read pair from a polynucleotide can be generated by the following steps: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of the first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through the second region using nucleotides provided in the second region flow order, wherein the extension of the primer through the second region proceeds faster than the extension of the primer in step (b); and (d) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. The sequencing data for the first region can be associated with the sequencing data for the third region, which represents a coupled sequencing read pair. The nucleotides used to extend the primer through the second region can be unlabeled.
[0086] In some embodiments, paired sequencing reads from a polynucleotide can be generated by the steps of: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer with labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using nucleotides provided in the second region flow order, wherein the primer extension through the second region is without detecting the presence or absence of labels of nucleotides incorporated into the extended primer; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) with labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. The sequencing data of the first region can be associated with the sequencing data of the third region, which indicates paired sequencing reads. The nucleotides used to extend the primer through the second region can be unlabeled.
[0087] In some embodiments, paired sequencing reads from a polynucleotide can be generated by the steps of: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer with labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using nucleotides provided in the second region flow order, wherein a mixture of at least two different types of nucleobases is used in at least one step of the flow order; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) with labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. The sequencing data of the first region can be associated with the sequencing data of the third region, which indicates paired sequencing reads. The nucleotides used to extend the primer through the second region can be unlabeled.
[0088] In some embodiments, primer extension proceeds through a second region to re-phase (i.e., synchronize) multiple sequencing reactions within a sequencing cluster. The chemical process of incorporating nucleotides into an extended primer is often imperfect, resulting in desynchronization between the strands within the sequencing cluster. As read lengths increase, desynchronization can lead to signal degradation and thus reduced accuracy when detecting the presence or absence of nucleotide incorporation into the extended primer. The result of re-synchronization is to counteract signal loss, which allows for longer effective read lengths. To re-phase the sequencing reactions, a re-phasing cycle is used to cause primer extension through the second region, where a mixture of at least two (e.g., two or three) different types of nucleotide bases is used in multiple steps of the flow order in the second region. In some embodiments, nucleotides incorporated during the re-phasing cycle may not be detected, which will result in gaps in the resulting reads. However, this read gap can be addressed when aligning the sequence to a reference sequence or other sequences.
[0089] A reference sequence can be used to extract sequencing data for the second region, even if the second region may not have been directly or fully sequenced. For example, by detecting the presence or absence of labeled nucleotides incorporated into the extended primer, sequencing data can be obtained from the first region and / or the third region of the polynucleotide. However, the primer can be extended through the second region using unlabeled nucleotides, or the presence or absence of incorporated nucleotides is not detected. Using unlabeled nucleotides (or by not allowing time for the sequencing system to detect incorporated labels) allows for faster primer extension through the second region, but the sequencing data cannot be directly determined. However, because the primer is extended through the second region using nucleotides provided in a predetermined flow order, variants in the second region can affect the sequencing data determined within the third region. A reference sequence can be used to determine the expected sequencing data (e.g., the expected flowgram), which is compared to the generated sequencing data (such as the detected flowgram) to detect variants, including variants within the second region. The comparison between the expected sequencing information (e.g., the expected flowgram) and the generated sequencing data (e.g., the generated flowgram) can be performed in the third region (to detect variants in the second region). This methodology provides a significant advantage over traditional paired-end sequencing methods, for which sequencing data at the 3' or 5' end of a polynucleotide is not affected by variants in the polynucleotide between the 3' and 5' ends of the polynucleotide.
[0090] Definitions
[0091] As used herein, the singular forms "a," "an," and "the" include plural referents unless the context clearly dictates otherwise.
[0092] As used herein, the term "about" in reference to a value or parameter includes (and describes) variations that are directed to the value or parameter itself. For example, a description of "about X" includes a description of "X".
[0093] "Expected sequencing data" refers to the sequencing data that would be expected if the sequence of the polynucleotide used to generate the paired sequencing reads or the sequence of a region of the polynucleotide matches a reference sequence.
[0094] "Flow order" refers to the order of individual, separate nucleotide flows used to sequence a nucleic acid molecule using non-terminating nucleotides. The flow order can be divided into cycles of repeating units, and the flow order of the repeating units is referred to as the "flow cycle order". "Flow position" refers to the sequential position of an individual nucleotide flow during the sequencing process.
[0095] The terms "individual", "patient", and "subject" are used interchangeably and refer to an animal, including a human.
[0096] As used herein, the term "label" refers to a detectable moiety that is or can be conjugated to another moiety (such as a nucleotide or nucleotide analogue). The label can emit a signal or modify a signal transmitted to the label such that the presence or absence of the label can be detected. In some cases, the conjugation can be through a linker, which can be cleavable, such as photocleavable (e.g., cleavable under ultraviolet light), chemically cleavable (e.g., by a reducing agent such as dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP)), or enzymatically cleavable (e.g., by an esterase, lipase, peptidase, or protease). In some embodiments, the label is a fluorophore.
[0097] "Non-terminating nucleotide" is a nucleic acid moiety that can be ligated to the 3' end of a polynucleotide using a polymerase or transcriptase, and to which another non-terminating nucleic acid can be ligated using a polymerase or transcriptase without the need to remove a protecting group or reversible terminator from the nucleotide. Naturally occurring nucleic acids are one class of non-terminating nucleic acids. Non-terminating nucleic acids can be labeled or unlabeled.
[0098] "Short genetic variant" is used herein to describe a genetic polymorphism (i.e., a mutation) that is 10 consecutive bases or fewer in length (i.e., 10, 9, 8, 7, 6, 5, 4, 3, 2, or 1 base in length). The term includes single nucleotide polymorphisms (SNPs), multinucleotide polymorphisms (MNPs), and indels that are 10 consecutive bases or fewer in length.
[0099] It should be understood that aspects and variations of the invention described herein include "consisting of" and / or "consisting essentially of" the aspects and variations.
[0100] When a range of values is provided, each intermediate value between the upper and lower limits of that range, as well as any other stated value or intermediate value in that range of states, is to be understood as being included within the scope of the present disclosure. Where the range includes an upper or lower limit, ranges excluding either of those included limits are also included in the present disclosure.
[0101] Some of the analysis methods described herein include mapping a sequence to a reference sequence, determining sequence information, and / or analyzing sequence information. It is well understood in the art that complementary sequences can be readily determined and / or analyzed, and the description provided herein encompasses analysis methods performed with reference to complementary sequences.
[0102] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. The description is presented to enable a person having ordinary skill in the art to make and use the invention, and the description is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles herein can be applied to other embodiments. Thus, the invention is not intended to be limited to the embodiments shown, but is to be accorded the widest scope consistent with the principles and features described herein.
[0103] The figures illustrate processes in accordance with various embodiments. In an exemplary process, some blocks are optionally combined, the order of some blocks is optionally changed, and some blocks are optionally omitted. In some instances, additional steps may be combined with the exemplary process. Thus, the operations as shown (and described in more detail below) are exemplary in nature and should not be considered limiting.
[0104] The disclosures of all publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. In the event of any conflict between any of the incorporated references and the present disclosure, the present disclosure shall control.
[0105] Flow sequencing method
[0106] Sequencing data can be generated using a flow sequencing method, the method comprising extending a primer that is bound to a template polynucleotide molecule according to a predetermined flow cycle, wherein in a predetermined flow cycle, a single type of nucleotide is accessible to the extending primer at any given flow position. In some embodiments, at least some specific types of nucleotides include a label, and when the labeled nucleotide is incorporated into the extending primer, the label generates a detectable signal. The resulting sequence incorporated into the extending primer by such nucleotides should be the reverse complement of the template polynucleotide molecule sequence. In some embodiments, for example, sequencing data is generated using a flow sequencing method, the method comprising extending a primer using labeled nucleotides and detecting the presence or absence of the labeled nucleotides incorporated into the extending primer. The flow sequencing method may also be referred to as a "native sequencing-by-synthesis" or "non-terminating sequencing-by-synthesis" method. Exemplary methods are described in U.S. Patent No. 8,772,473, which is incorporated herein by reference in its entirety. Although the following description is provided with reference to the flow sequencing method, it should be understood that other sequencing methods may be used to sequence all or a portion of the sequencing region.
[0107] Flow sequencing involves using nucleotides to extend a primer that is hybridized to a polynucleotide. If a complementary base is present in the template strand, nucleotides of a given base type (e.g., A, C, G, T, U, etc.) can be mixed with the hybridized template to extend the primer. The nucleotides can be, for example, non-terminating nucleotides. When the nucleotides are non-terminating, if there are more than one consecutive complementary bases in the template strand, more than one consecutive base can be incorporated into the extending primer strand. In contrast to non-terminating nucleotides are nucleotides having a 3' reversible terminator, where the blocking group is typically removed before successive nucleotides are ligated. If no complementary base is present in the template strand, primer extension stops until a nucleotide complementary to the next base in the template strand is introduced. At least a portion of the nucleotides can be labeled such that their incorporation can be detected. Most commonly, only a single nucleotide type is introduced at a time (i.e., discrete addition), although in certain embodiments two or three different types of nucleotides can be introduced simultaneously. This methodology can be contrasted with sequencing methods using reversible terminators, where primer extension stops after each single base extension, after which the terminator is reversed to allow incorporation of the next successive base.
[0108] Nucleotides can be introduced in a defined order during primer extension, which can be further divided into cycles. Nucleotides are added stepwise, which allows the incorporated nucleotides to be integrated at the end of the sequencing primer that is complementary to the bases present in the template strand. The cycles can have the same nucleotide order and number of different base types, or different nucleotide orders and / or different numbers of different base types. However, the set of bases corresponding to a given flow step (i.e., one or more different bases used simultaneously in a single flow step) is not repeated in the same cycle (as the term is used herein), which can provide a marker for differentiating different cycles. By way of example only, the order of the first cycle can be A-T-G-C, and the order of the second cycle can be A-T-C-G. Additionally, one or more cycles can omit one or more nucleotides. By way of example only, the order of the first cycle can be A-T-G-C, and the order of the second cycle can be A-T-C. Those skilled in the art can readily envision alternative orders. Between the introduction of different nucleotides, unincorporated nucleotides can be removed, for example by washing the sequencing platform with a wash solution.
[0109] A polymerase can be used to extend the sequencing primer by integrating one or more nucleotides at the primer terminus in a template-dependent manner. In some embodiments, the polymerase is a DNA polymerase. The polymerase can be a naturally occurring polymerase or a synthetic (e.g., mutant) polymerase. The polymerase can be added at the initial step of primer extension, but can optionally be added during sequencing as a supplemental polymerase, for example concomitant with the stepwise addition of nucleotides or after multiple flow cycles. Exemplary polymerases include DNA polymerase, RNA polymerase, thermostable polymerase, wild-type polymerase, modified polymerase, Bst DNA polymerase, Bst 2.0 DNA polymerase, Bst 3.0 DNA polymerase, Bsu DNA polymerase, Escherichia coli DNA polymerase I, T7 DNA polymerase, bacteriophage T4 DNA polymerase, Φ29 (phi29) DNA polymerase, Taq polymerase, Tth polymerase, Tli polymerase, Pfu polymerase, and SeqAmp DNA polymerase.
[0110] When determining the sequence of the template strand, the introduced nucleotides can include labeled nucleotides, and the presence or absence of the incorporated labeled nucleic acid can be detected to determine the sequence. The label can be, for example, an optically active label (e.g., a fluorescent label) or a radioactive label, and a detector can be used to detect the signal emitted or altered by the label. The presence or absence of the labeled nucleotides incorporated into the primer hybridized to the template polynucleotide can be detected, which allows for the determination of the sequence (e.g., by generating a flow map). In some embodiments, the labeled nucleotides are labeled with a fluorescent, luminescent, or other light-emitting moiety. In some embodiments, the label is linked to the nucleotide via a linker. In some embodiments, the linker is cleavable, for example, by a photochemical or chemical cleavage reaction. For example, the label can be cleaved after detection and before incorporation of successive nucleotides. In some embodiments, the label (or linker) links to the nucleobase or to another site on the nucleotide without interfering with the extension of the nascent DNA strand. In some embodiments, the linker includes a disulfide or a PEG-containing moiety.
[0111] In some embodiments, the introduced nucleotides include only unlabeled nucleotides, and in some embodiments, the nucleotides include a mixture of labeled and unlabeled nucleotides. For example, in some embodiments, the fraction of labeled nucleotides is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less compared to the total nucleotides. In some embodiments, the fraction of labeled nucleotides is about 100%, about 95% or more, about 90% or more, about 80% or more, about 70% or more, about 60% or more, about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more, or about 0.01% or more compared to the total nucleotides. In some embodiments, the fraction of labeled nucleotides is about 0.01% to about 100%, such as about 0.01% to about 0.025%, about 0.025% to about 0.05%, about 0.05% to about 0.1%, about 0.1% to about 0.25%, about 0.25% to about 0.5%, about 0.5% to about 1%, about 1% to about 1.5%, about 1.5% to about 2%, about 2% to about 2.5%, about 2.5% to about 3%, about 3% to about 4%, about 4% to about 5%, about 5% to about 10%, about 10% to about 20%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to less than 100% or about 90% to about 100%.
[0112] Sequencing data, such as a flowgram, can be generated based on the detection of incorporated nucleotides and the order of nucleotide incorporation. For example, flow template sequences: CTG and CAG, and a repeating flow cycle of T - A - C - G (i.e., sequential addition of T, A, C, and G nucleotides, which are incorporated into the primer only when the complementary base is present in the template polynucleotide). The resulting flowgram is shown in Table 1, where 1 indicates incorporation of the introduced nucleotide and 0 indicates non - incorporation of the introduced nucleotide. The flowgram can be used to determine the sequence of the template strand.
[0113] Table 1
[0114]
[0115] The flow diagram can be binary or non - binary. A binary flow diagram detects the presence (1) or absence (0) of incorporated nucleotides. A non - binary flow diagram can more quantitatively determine the number of incorporated nucleotides from each stepwise introduction. For example, the sequence CCG will incorporate two G bases, and any signal emitted by the labeled bases will have a greater intensity with the incorporation of a single base. This is shown in Table 1. A non - binary flow diagram also indicates the presence or absence of bases, but can provide additional information, including the number of bases incorporated at a given step.
[0116] Prior to generating sequencing data, the polynucleotide is hybridized to a sequencing primer to produce a hybridization template. The polynucleotide can be ligated to an adapter during sequencing library preparation. The adapter can include a hybridization sequence that hybridizes to the sequencing primer. For example, the hybridization sequence of the adapter can be a uniform sequence across multiple different polynucleotides, and the sequencing primer can be a uniform sequencing primer. This allows for multiplex sequencing of different polynucleotides in the sequencing library.
[0117] The polynucleotide can be attached to a surface (e.g., a solid support) for sequencing. The polynucleotide can be amplified (e.g., by bridge amplification or other amplification techniques) to generate a polynucleotide sequencing colony. The polynucleotides amplified within a cluster are substantially identical or complementary (some errors may be introduced during the amplification process such that a portion of the polynucleotide may not necessarily be identical to the original polynucleotide). Colony formation enables signal amplification so that the detector can accurately detect the incorporation of labeled nucleotides for each colony. In some cases, emulsion PCR is used to form colonies on beads, and the beads are distributed on the sequencing surface. Examples of systems and methods for sequencing can be found in U.S. Patent Serial Number 10,344,328, which is incorporated herein by reference in its entirety.
[0118] Primer extension hybridized to a polynucleotide proceeds through a first region, a second region, and a third region of the polynucleotide. Sequencing data associated with sequences within the first region and / or the third region can be generated as described above. However, an accelerated “fast forward” process is used to extend the primer through the second region (which is between the first region and the third region). That is, primer extension through the second region between the first region and the third region of the polynucleotide can occur more rapidly than primer extension through the first region and / or the third region. For example, primer extension through the second region can occur by extending the primer without detecting the presence or absence of labeled nucleotides incorporated into the extended primer. During flow sequencing, as described above, labeled nucleotides are incorporated into the extended primer, the hybridized template is washed, and a detector is used to detect signals from the nucleotide labels that indicate whether the nucleotide has been incorporated into the extended primer. However, the detection process takes time, and primer extension through the second region can be accelerated by skipping the detection process. In some embodiments, unlabeled nucleotides (or only unlabeled nucleotides) are used to extend the primer through the second region, which can further accelerate the rate of primer extension.
[0119] Primer extension through the second region can be optionally or additionally accelerated by using a mixture of at least two different types of nucleotides in at least one step of the flow order used during primer extension through the second region. For example, two different bases, such as G and C, can be used simultaneously in the same step, and if complementary C or G bases are present, they will extend the primer. This accelerates primer extension by incorporating consecutive bases into the primer, even if those bases have different base types. In some embodiments, at least one step of the flow order includes 2 different bases. In some embodiments, at least one step of the flow order includes 3 different bases. By way of example, consider the sequence of SEQ ID NO:1 and the corresponding flow order and flow diagram shown in Table 2. The flow order process for extending a sequencing primer hybridized to a polynucleotide containing SEQ ID NO:1 includes 5 cycles, where cycles 1, 4, and 5 are identical to each other, and cycles 2 and 3 are identical to each other (where cycles 1, 4, and 5 are different from cycles 2 and 3). In this example, each cycle has 4 steps, where cycles 1, 4, and 5 include the sequential and independent addition of A-C-T-G nucleotides, where a single base type is added at each cycle step. Cycles 2 and 3 include four cycle steps, where step 1 omits the A nucleotide (i.e., includes C, T, and G), step 2 omits the C nucleotide (i.e., includes A, T, and G), step 3 omits the T nucleotide (i.e., includes A, C, and G), and step 4 omits the G nucleotide (i.e., includes A, C, and T). Because cycles 2 and 3 include multiple different nucleobase types simultaneously during primer extension, the primer extends faster compared to using only a single base type at any given step. The flow diagram for extending the primer against the SEQ ID NO:1 template using this flow order shown in Table 2 results in the addition of up to 6 bases (cycle 3, step 3) during the fast-forward portion of primer extension. In contrast, Table 3 shows the flow diagram for the same SEQ ID NO:1 using A-C-T-G cycles, where a single nucleotide is used at each step (similar to cycles 1, 4, and 5 in Table 2). The flow order for extending the primer shown in Table 3 requires 10 four-step cycles to extend the primer through the polynucleotide, which is substantially slower than the 5 four-step cycles required to extend the primer through the polynucleotide using the flow order provided in Table 2.
[0120]
[0121] The fast-forward method is particularly suitable for accelerating primer extension through regions that are not directly sequenced. For example, referring to Table 2, cycles 1, 4, and 5 use labeled nucleotides in a stepwise manner to generate sequencing data related to the first region (cycle 1) and the third region (cycles 4 and 5), while the primer rapidly extends through the second region (cycles 2 and 3) between the first and third regions.
[0122] Primer extension using flow sequencing allows for long-range sequencing on the order of hundreds or even thousands of bases in length. The number of flow steps or cycles can be increased or decreased to obtain the desired sequencing length. Extension of the primer in the first region or the third region can include one or more flow steps of stepwise extension of the primer using nucleotides having one or more different base types. In some embodiments, extension of the primer in the first region or extension of the primer in the third region includes from 1 to about 1000 flow steps, such as from 1 to about 10 flow steps, about 10 to about 20 flow steps, about 20 to about 50 flow steps, about 50 to about 100 flow steps, about 100 to about 250 flow steps, about 250 to about 500 flow steps, or about 500 to about 1000 flow steps. The flow steps can be divided into identical or different flow cycles. The number of bases incorporated into the primer in the first region or the third region depends respectively on the sequence of the first region or the third region, and the flow order used to extend the primer in the first region or the third region. In some embodiments, the first region or the third region is from about 1 base to about 4000 bases in length, such as from about 1 base to about 10 bases in length, about 10 bases to about 20 bases in length, about 20 bases to about 50 bases in length, about 50 bases to about 100 bases in length, about 100 bases to about 250 bases in length, about 250 bases to about 500 bases in length, about 500 bases to about 1000 bases in length, about 1000 bases to about 2000 bases in length, or about 2000 bases to about 4000 bases in length.
[0123] Primer extension through the second region can be carried out by any number of flow steps. In some embodiments, primer extension through the second region omits labeled nucleotides, which further increases the feasible extension distance of the primer without polymerase stalling. In some embodiments, primer extension through the second region includes from 1 to about 10,000 flow steps, such as from 1 to about 10 flow steps, about 10 to about 20 flow steps, about 20 to about 50 flow steps, about 50 to about 100 flow steps, about 100 to about 250 flow steps, about 250 to about 500 flow steps, about 500 to about 1000 flow steps, about 1000 flow steps to about 2500 flow steps, about 2500 flow steps to about 5000 flow steps, or about 5000 flow steps to about 10,000 flow steps. In some embodiments, primer extension through the second region includes more than about 10,000 flow steps. The number of bases incorporated in the primer in the second region depends on the sequence of the second region and the flow order used to extend the primer in the second region. In some embodiments, the second region is from about 1 base to about 50,000 bases long, such as from about 1 base to about 10 bases long, about 10 bases to about 20 bases long, about 20 bases to about 50 bases long, about 50 bases to about 100 bases long, about 100 bases to about 250 bases long, about 250 bases to about 500 bases long, about 500 bases to about 1000 bases long, about 1000 bases to about 2000 bases long, about 2000 bases to about 2500 bases long, about 2500 to about 5000 bases long, about 5000 to about 10,000 bases long, about 10,000 to about 25,000 bases long, or about 25,000 to about 50,000 bases long. In some embodiments, the length of the second region is greater than about 50,000 bases.
[0124] The extension of the primer can occur through a first region, a second region, and a third region, wherein the primer is extended using labeled nucleotides through the first region and the third region. Detection of the nucleotides incorporated into the extended primer can be performed to generate sequencing data. The extension of the primer through the second region can occur at a faster rate than the extension of the primer through the first and / or third regions, for example, by not detecting the presence or absence of the label of the nucleotides incorporated into the extended primer, or by extending the primer with a mixture comprising at least two different types of nucleobases (wherein the extension of the primer through the first and / or third depends on fewer different types of nucleobases). The extension of the primer can be further extended in an alternating pattern. For example, after the primer is extended through the third region, it can be further extended into a fourth region. The extension of the primer through the fourth region can occur at a faster rate than the extension of the primer through the first and / or third regions, for example, by not detecting the presence or absence of the label of the nucleotides incorporated into the extended primer, or by extending the primer with a mixture comprising at least two different types of nucleobases. The primer can then be extended into a fifth region using labeled nucleotides, and sequencing data for the fifth region can be generated by detecting the nucleotides incorporated into the extended primer. This process can be repeated multiple times with variations in cycles as needed. The sequencing data from any two regions can be correlated to generate paired sequencing reads, and the paired sequencing reads can be analyzed as described herein (e.g., by treating the region between the selected regions as the "second region" as described for the analysis methods provided herein).
[0125] Figure 1Schematic diagram illustrating an exemplary method for generating paired sequencing reads from a polynucleotide, such as DNA. At 102, polynucleotide 104 is hybridized with primer 106 to form a hybridization template. In some embodiments, the polynucleotide includes an adapter region 108, which can be ligated to the 3' of the target polynucleotide during sequencing library preparation. The adapter region 108 can include a hybridization region, and the primer 106 can hybridize to the hybridization region of the adapter region 108. At step 110, sequencing data for the first region 112 of the polynucleotide 104 is generated by extending the primer 106 using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. The nucleotides used for extending the primer can also include unlabeled nucleotides, although labeled nucleotides are used to detect nucleotide incorporation for generating the sequencing data. In some embodiments, nucleotides are stepwise added in one or more cycles according to a first region flow order to extend the primer 106 through the first region 112, and the hybridized template can be washed after the cycle steps to remove unincorporated nucleotides before detecting the presence or absence of incorporated labeled nucleotides. At step 114, according to a second region flow order, the primer 106 is extended through the second region 116 of the polynucleotide 104. The primer 106 can be extended through the second region 116 at a faster rate than the primer extension in step 110. This accelerated primer extension can be referred to as the "fast forward" portion of the method. According to the second region flow order, nucleotides (which are unlabeled in some embodiments) are stepwise added to the hybridized template in one or more cycles. In some embodiments, more than one (e.g., two or three) different base types are used simultaneously in a given cycle step, which accelerates the primer extension. In some embodiments, the nucleotides are unlabeled, which allows for faster primer extension than labeled nucleotides. In some embodiments, the primer is extended without detecting the presence or absence of nucleotide labels. At step 118, sequencing data for the third region 118 of the polynucleotide 104 is generated by extending the primer 106 using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. The generation of the sequencing data for the third region 118 can be performed in a similar manner as described for generating the sequencing data for the first region 112. At step 122, the sequencing data generated for the first region 112 is associated with the sequencing data generated for the third region 120, which results in paired sequencing reads 124 for the polynucleotide 104. The associated sequencing data between the first region and the third region can include the sequences of the first region and the third region. The paired sequencing reads 124 include the sequencing data for the first region 112 and the third region 120, where the first region 112 and the third region 120 are separated by the second region 116, and the sequencing data for the second region 116 is not necessarily known.
[0126] It is not necessary to generate sequencing data for the first region of the polynucleotide according to some embodiments described herein. For example, sequencing primers can be used for targeted sequencing by hybridizing to a targeted region. In targeted sequencing, the first region of the polynucleotide is known, and the primers are designed to specifically bind to the first region. The primers can then be extended through the second and third regions as described, generating sequencing data for the third region. In some embodiments, a method for generating paired sequencing reads from a polynucleotide comprises: (a) hybridizing a primer to the first region of the polynucleotide to form a hybridization template; (b) extending the primer through the second region using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleotide bases is used in at least one step of the second region flow order; and (c) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (b) using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides.
[0127] A reference sequence can be used to determine the expected sequencing data (such as a flowgram) for the first region, the second region, and / or the third region. The sequences of the first and third regions can be determined from the sequencing data generated for those regions. For example, referring to Table 2, Cycle 1 is related to the first region, and its sequence can be easily determined as the complementary sequence of the bases (i.e., the base flow A-C-T-G corresponds to the sequence TGAC), and Cycles 4 and 5 are related to the third region, and its sequence is determined as CTGAC (i.e., the complementary sequence of G-A-C-T-G). Thus, using the sequencing data generated from the first region and / or the third region, the first region and / or the third region (or at least a portion of the first region and / or the third region) can be mapped to the reference sequence. Once mapped to the reference sequence, the expected sequencing data for the second region can be generated using the flow order for extending the primer through the second region and the reference sequence.
[0128] The expected sequencing data for the third region can also be determined using the reference sequence of the second region, the flow order of the second region, the flow order of the third region, and information about the sequence of the third region. Similarly, the expected sequencing data for the first region can be determined using the reference sequence of the second region, the flow order of the second region, the flow order of the first region, and information about the sequence of the first region. Information about the sequence of the third region (or the first region) can be obtained, for example, from a reference sequence (or a different reference sequence), or from the generated sequencing data, such as the sequencing data generated by extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides or the sequencing data obtained by other methods (e.g., independently sequencing the third region of the polynucleotide).
[0129] For example, the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region can be used to determine the expected sequencing data of the third region. The first region (or a part thereof) can be mapped to the reference sequence, and the reference sequence corresponding to the second region and the second region flow order can be used to determine the expected reference sequencing data of the second region. Similarly, the reference sequence of the third region and the third region flow order can be used to determine the expected reference sequencing data of the third region. A similar method can be used to determine the expected sequencing data of the first region. For example, the reference sequence of the second region, the second region flow order, the first region flow order, and the reference sequence of the first region can be used to determine the expected sequencing data of the first region. The third region (or a part thereof) can be mapped to the reference sequence, and the reference sequence corresponding to the second region and the second region flow order can be used to determine the expected reference sequencing data of the second region. Similarly, the reference sequence of the first region and the first region flow order can be used to determine the expected reference sequencing data of the first region.
[0130] In another example, the reference sequence of the second region, the second region flow order, the third region flow order, and the sequencing data related to the sequencing of the third region can be used to determine the expected sequencing data of the third region, and the sequencing data can be the same as or different from the sequencing data generated as described above. The first region (or a part thereof) can be mapped to the reference sequence, and the reference sequence corresponding to the second region and the second region flow order can be used to determine the expected reference sequencing data of the second region. The sequencing data of the third region can be used to determine the sequence of the third region. In addition, the sequence of the third region and the third region flow order can be used to determine the expected sequencing data of the third region.
[0131] Figure 2Schematic diagram illustrating an exemplary method for generating expected sequencing data. At step 202, paired sequencing reads are mapped to a reference sequence. Mapping paired sequencing reads can include mapping a first region (or a portion thereof) of the paired sequencing reads (or a portion thereof) to the reference sequence, mapping a third region (or a portion thereof) of the paired sequencing reads to the reference sequence, or mapping both the first region (or a portion thereof) and the third region (or a portion thereof) to the reference sequence. At step 204, the expected sequencing data (such as an expected flowgram) for the second region is determined using the second region flow order and the reference sequence. Given that the flow order and the reference sequencing are known, it is easy to obtain the determination of the expected sequencing data (i.e., if the second region of the polynucleotide matches the reference sequence, the sequencing data will be as expected). Additionally, the expected sequencing data for the second region can be used to determine the expected 5' end of the second region. The 5' end of the second region can vary depending on the flow order of the region and the sequence of the second region. Thus, the 3' end of the third region can also vary based on the second region flow order and the sequence of the second region, since the 3' end of the third region is adjacent to the 5' end of the second region. As shown in step 206, once the 3' end of the third region is established (e.g., as determined using the expected sequencing data of the second region), the expected sequencing data for the third region can be determined. As further described herein, the expected sequencing data for the third region can be used to determine variants, such as variants within the second region of the polynucleotide.
[0132] If the polynucleotide includes a variant within the second region, the generated sequencing data (e.g., flowgram) associated with the third region can be different (depending on the sequence context and variant size) from the expected sequencing data associated with the third region. Thus, in some embodiments, variants are detected based on the difference between the expected sequencing data and the generated sequencing data.
[0133] The reference sequence can be any suitable sequence of the same species as the polynucleotide, and there may be some differences between the reference sequence and the polynucleotide sequence. In some embodiments of the methods described herein, these differences or variants can be detected. In some embodiments, the test variant (i.e., the variant of interest) is included in the reference sequence, and in other embodiments, the test variant is omitted from the reference sequence. In some embodiments, the analysis can be performed with two different reference sequences, where one reference sequence includes the test variant and the other reference sequence omits the test variant. In some embodiments, the only difference between the two reference sequences is the presence or absence of the test variant.
[0134] The sensitivity of the variant detection method described herein can depend on the context of the variant and / or the flow order used to extend primers in the first, second, and / or third regions. In the first, second, and / or third regions, missed variants for a given flow order can be detected using a different flow order. Thus, in some embodiments of the methods described herein, more than one paired-end sequencing read is generated using different flow orders for extending primers through one or more of the first, second, and / or third regions of a polynucleotide.
[0135] The polynucleotides used in the methods described herein can be obtained from any suitable biological source, such as a tissue sample, blood sample, plasma sample, saliva sample, fecal sample, or urine sample. The polynucleotides can be DNA or RNA polynucleotides. In some embodiments, RNA polynucleotides are reverse transcribed into DNA polynucleotides prior to hybridization with sequencing primers. In some embodiments, the polynucleotides are cell-free DNA (cfDNA), such as circulating tumor DNA (ctDNA) or fetal cell-free DNA.
[0136] A polynucleotide library can be prepared by known methods. In some embodiments, the polynucleotides can be ligated to adapter sequences. The adapter sequences can include hybridization sequences that hybridize to primers extended during paired-end sequencing read generation.
[0137] In some embodiments, sequencing data is obtained without amplification of nucleic acid molecules prior to establishing sequencing colonies (also referred to as sequencing clusters). Methods for generating sequencing colonies include bridge amplification or emulsion PCR. Methods that rely on shotgun sequencing and consensus sequence determination typically use unique molecular identifiers (UMIs) to label nucleic acid molecules and amplify the nucleic acid molecules to generate many copies of the same independently sequenced nucleic acid molecules. The amplified nucleic acid molecules can then be attached to a surface and bridge amplified to generate independently sequenced sequencing clusters. The UMIs can then be used to associate the independently sequenced nucleic acid molecules. However, the amplification process can introduce errors into the nucleic acid molecules, for example due to the limited fidelity of DNA polymerase. In some embodiments, nucleic acid molecules are not amplified prior to amplification to generate colonies for obtaining sequencing data. In some embodiments, nucleic acid sequencing data is obtained without the use of unique molecular identifiers (UMIs).
[0138] In some embodiments, the flow-through sequencing method is used in conjunction with rolling circle amplification (RCA) sequencing. RCA allows for the formation of multiple copies of nucleic acid molecules covalently linked in a linear sequence. See, e.g., Dean et al., Rapid Amplification of Plasmid and Phage DNA Using Phi29 DNA Polymerase and Multiply-Primed Rolling Circle Amplification, Genome Research, vol. 11, pp. 1095-1099 (2001); and U.S. Patent No. 5,714,320, the respective contents of which are incorporated herein by reference. Because multiple copies of the nucleic acid molecule can be sequenced linearly, as sequencing proceeds, a given region can be sequenced alternately in a "dark" or patterned or "bright" mode. In some embodiments, the sequencing mode switch can be determined dynamically (and optionally, automatically). For example, a variant can be detected within a "dark" region, but the limited information generated prevents the specific variant from being determined. Thus, the sequencing flow can be dynamically adjusted to sequence the region of the nucleic acid molecule containing the variant in the bright mode.For example, a method of detecting short genetic variants in a test sample can include (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide that includes at least a first copy of the polynucleotide and a second copy of the polynucleotide; (b) hybridizing the RCA-amplified polynucleotide with a primer to form a hybridization template; (c) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (d) further extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order, or (iii) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide more rapidly than the primer is extended through the first region; (e) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (f) comparing the sequencing data generated for the third region of the polynucleotide with expected sequencing data of an expected sequence of the third region of the polynucleotide; (g) determining the presence of a short genetic variant in the second region of the polynucleotide; (h) generating sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; and (i) determining the identity of the short genetic variant in the second region of the polynucleotide. In some embodiments, based on determining the presence of a short genetic variant in the second region of the polynucleotide, sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide is dynamically generated.
[0139] Extended flow cycle
[0140] The flow cycle order need not be limited to four basic flow cycles (e.g., each of A, G, C, and T, in any repeating order), and can be an extended flow cycle having more than four basic types in the cycle. The extended cycle order can be repeated to obtain the desired number of cycles to extend the sequencing primer. For example, in some embodiments, the extended flow order includes 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or more individual nucleotide flows in the flow cycle order. The cycle can include at least one of each of A, G, C, and T, but one or more base types are repeated within the cycle before the cycle is repeated. The extended flow cycle can be used, for example, to extend the primer through a second region according to the methods described herein.
[0141] Compared to a flow cycle order with four repeating bases, the extended flow cycle order can be used to detect a greater proportion of small genomic variants (e.g., SNPs). For example, there are 192 substitution SNPs in the valid configuration of the form XYZ → XQZ, where Q ≠ Y (and Q, X, Y, and Z are each any one of A, C, G, and T). Of these, 168 can generate new signals (i.e., new non-zero signals or new zero signals) in a sequencing dataset (e.g., a flowgram). Given the same tail sequence relative to the reference in the variant, the new zero or non-zero signal in combination with a sensitive flow order can generate a propagated signal (e.g., a flow shift, which can extend more than the length of the cycle) for multiple flow positions. It should be noted that the insertion or deletion of homopolymers, rather than the change in homopolymer length, can cause the propagation of signal differences. The remaining 24 variants cause a change in homopolymer length at the affected flow position, but this change does not cause a propagated signal change. Thus, a theoretical maximum of 87.5% of SNPs can generate new signals different from the reference (or candidate) sequence for more than two flow positions. As described above, the propagated signal differences increase the likelihood of differences between the test sequencing dataset and an incorrectly matched candidate sequence. Additionally, the propagated signal changes depend on the flow order across the variant.
[0142] When using a flow order extended sequencing primer, sequencing of the randomly fragmented nucleic acid molecules in the test sample results in a random shift in the variant flow order content. That is, the flow position of the variant can change depending on the starting position of the nucleic acid molecule being sequenced. For all 87.5% of SNPs, not all flow cycle combinations are able to detect signal changes at more than two flow positions, even using all the sequencing starting positions in the nucleic acid molecule sequence. For example, for 41.7% of SNPs, the four-base flow cycle order T-A-C-G can result in the test sequencing data set differing from the reference sequencing data set at more than two flow positions. As further discussed herein, given a sufficiently high sequencing depth (i.e., sampling a sufficiently large number of starting positions), an extended flow cycle order has been designed such that all theoretical maxima of SNPs (i.e., 87.5% of possible SNPs, or all SNPs except those that result in changes in homopolymer length) can differ between the test sequencing data set and the reference sequencing data set at more than two flow positions.
[0143] The extended sequencing flow order can have different efficiencies (i.e., average number of integrations per flow when sequencing the human reference genome). In some embodiments, the flow order has an efficiency of about 0.6 or greater (such as about 0.62 or greater, about 0.64 or greater, about 0.65 or greater, about 0.66 or greater, or about 0.67 or greater). In some embodiments, the flow order has an efficiency of about 0.6 to about 0.7. Examples of flow cycle orders and corresponding estimated efficiencies are shown in Table 4.
[0144] In some embodiments, an extended sequencing flow order is selected to generate signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 50% to 87.5% of SNP alignments at at least 5% of random sequencing start positions. In some embodiments, an extended sequencing flow order is selected to generate signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 60% to 87.5% of SNP alignments at at least 5% of random sequencing start positions (i.e., "flow phase"). In some embodiments, an extended sequencing flow order is selected to generate signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 70% to 87.5% of SNP alignments at at least 5% of random sequencing start positions. In some embodiments, an extended sequencing flow order is selected to generate signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 80% to 87.5% of SNP alignments at at least 5% of random sequencing start positions.
[0145] In some embodiments, an extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 50% to 87.5% of SNP alignments at at least 10% of random sequencing start positions. In some embodiments, an extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 60% to 87.5% of SNP alignments at at least 10% of random sequencing start positions. An extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 70% to 87.5% of SNP alignments at at least 10% of random sequencing start positions. An extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 80% to 87.5% of SNP alignments at at least 10% of random sequencing start positions.
[0146] In some embodiments, an extended sequencing flow order is selected to create signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 50% to 87.5% of SNP alignments at at least 20% of random sequencing start positions. In some embodiments, an extended sequencing flow order is selected to create signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 60% to 87.5% of SNP alignments at at least 20% of random sequencing start positions. An extended sequencing flow order is selected to create signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 70% to 87.5% of SNP alignments at at least 20% of random sequencing start positions. An extended sequencing flow order is selected to create signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 80% to 87.5% of SNP alignments at at least 20% of random sequencing start positions.
[0147] In some embodiments, an extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 50% to 87.5% of SNP alignments at at least 30% of random sequencing start positions. In some embodiments, an extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 60% to 87.5% of SNP alignments at at least 30% of random sequencing start positions. An extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 70% to 87.5% of SNP alignments at at least 30% of random sequencing start positions. An extended sequencing flow order is selected to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set), the two sequencing data sets being associated with nucleic acid molecules that differ in SNPs in about 80% to 87.5% of SNP alignments at at least 30% of random sequencing start positions.
[0148] In some embodiments, the extended sequencing flow order is any of the extended sequencing flow orders in Table 4. "Shift sensitivity" refers to the maximum sensitivity to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set) over all possible SNP alignments. "Maximum shift sensitivity" refers to the maximum sensitivity to produce signal differences at more than two flow positions between two sequencing data sets (e.g., a test or target sequencing data set and a candidate or reference sequencing data set) over SNP alignments at the highest fraction of mobile phases that maintain that sensitivity.
[0149]
[0150]
[0151]
[0152] In some embodiments, for 50% or more of the possible SNP arrangements at the 5% random sequencing start position, the flow cycle order induces signal changes at more than two flow positions. In some embodiments, the induced signal change is a change in signal intensity, or a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the induced signal change is a new substantially zero (or new zero) or new substantially non-zero (or new non-zero) signal. In some embodiments, the flow cycle order has an efficiency of integrating 0.6 or more bases per flow. In some embodiments, the flow cycle is any of the flow cycle orders listed in Table 4.
[0153] Re-phasing flow
[0154] One or more re-phasing flows can be used as a second region or within a second region for re-phasing (i.e., synchronizing) parallel sequencing reactions within sequencing clusters. Sequencing clusters include multiple copies of polynucleotides tightly attached to a common surface (e.g., beads or flow cells). Clusters can be formed, for example, by attaching polynucleotides to the surface and amplifying the attached polynucleotides (e.g., by bridge amplification). Sequencing data can be collected from the sequencing clusters as a whole, since primers hybridized to each polynucleotide are extended simultaneously by integrating nucleotides based on the same template. However, the chemical process of integrating nucleotides into the extended primers is generally imperfect, resulting in desynchronization between the strands within the sequencing clusters. That is, some primers may lag behind other extended primers within the cluster. When the presence or absence of nucleotides integrated into the extended primers is detected as the read length increases, desynchronization can lead to signal decay and thus reduced accuracy. The result of re-synchronization can counteract the signal loss, which allows for longer effective read lengths. To re-phase the sequencing reaction, one or more re-phasing flows are used to extend the primers through a second region, where a mixture of at least two (e.g., two or three) different types of nucleotide bases is used in multiple steps of the flow order in the second region. In some embodiments, nucleotides integrated during the re-phasing flow may not be detected, which will result in gaps in the resulting reads. However, when the sequence is aligned with a reference or other sequence, this read gap can be disposed of. By including such a "catch-up flow", the lagging primers can catch up with other extended primers within the cluster.
[0155] A method for resynchronizing sequencing clusters (e.g., within a sequencing cluster) that contain multiple copies of a polynucleotide can include using a rephasing flow order to extend a primer hybridized to a polynucleotide copy, wherein a mixture of at least two different types of nucleobases is used in at least one step of the rephasing flow order. In some embodiments, a method for synchronizing sequencing primers within a sequencing cluster includes: (a) hybridizing a primer to a polynucleotide copy within the sequencing cluster; (b) extending the primer through a first region of the polynucleotide copy using labeled nucleotides according to a first region flow cycle; (c) extending the primer through a second region of the polynucleotide copy using one or more rephasing flows, wherein a mixture of at least two different types of nucleobases is used in each of the one or more rephasing flows; and (d) extending the primer through a third region of the polynucleotide copy using labeled nucleotides according to a third region flow cycle.
[0156] A method for generating sequencing reads from multiple copies of a polynucleotide (e.g., within a sequencing cluster) can include the resynchronization method. For example, a method for generating sequencing reads from multiple copies of a polynucleotide can include: (a) hybridizing the polynucleotide copy to a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide copy by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using nucleotides provided in one or more rephasing flows, wherein a mixture of at least two different types of nucleobases is used in each of the one or more rephasing flows; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides.
[0157] The rephasing flow order (or rephasing flow cycle) includes one or more steps that allow lagging primers to catch up with leading primers in the sequencing cluster. At least one step in the rephasing flow order (e.g., 1, 2, 3, 4, or more) includes a mixture of two or more (e.g., three) different types of nucleobases. In some embodiments, the rephasing flow order includes 1, 2, 3, 4, 5, or more flows, each including a mixture of two or three different types of nucleobases.
[0158] The reprogrammed flow order is configured to increase the portion of the synchronous extension primer after the reprogrammed flow order. In some embodiments, the reprogrammed flow order includes, in any order, (i) a flow step that includes a mixture of A, C, and G nucleotides and omits T (and / or U) nucleotides (also referred to as a "non-T" (and / or "non-U") step); (ii) a flow step that includes a mixture of T (and / or U), C, and G nucleotides and omits A nucleotides (also referred to as a "non-A" step); (iii) a flow step that includes a mixture of T (and / or U), A, and G nucleotides and omits C nucleotides (also referred to as a "non-C" step); and (iv) a flow step that includes a mixture of T (and / or U), A, and C nucleotides and omits G nucleotides (also referred to as a "non-G" step).
[0159] Other reprogrammed flows can be determined. For example, in some embodiments, the reprogrammed flow (in the reprogrammed flow order) includes, in any order, one or more of the following: (i) a flow step that includes a mixture of A and C nucleotides and omits G and T (and / or U) nucleotides; (ii) a flow step that includes a mixture of T (and / or U) and G nucleotides and omits A and C nucleotides; (iii) a flow step that includes a mixture of A and G nucleotides and omits T (and / or U) and C nucleotides; (iv) a flow step that includes a mixture of T (and / or U) and C nucleotides and omits A and G nucleotides; (v) a flow step that includes a mixture of A and T (and / or U) nucleotides and omits G and C nucleotides; (vi) a flow step that includes a mixture of C and G nucleotides and omits A and T (and / or U) nucleotides; (vii) a flow step that includes a mixture of A, G, and C nucleotides and omits T nucleotides; (viii) a flow step that includes a mixture of T (and / or U), A, and G nucleotides and omits C nucleotides; (ix) a flow step that includes a mixture of C, T (and / or U), and A nucleotides and omits G nucleotides; and / or (x) a flow step that includes a mixture of G, C, and T (and / or U) nucleotides and omits A nucleotides.
[0160] A mixture comprising all four types of non-terminating nucleotides (i.e., a mixture containing A, C, G, and T (and / or U)) can result in uncontrolled primer extension. However, a mixture of all four types of nucleotides can be used in a rephasing flow order where three base types are non-terminating nucleotides and one base type includes a reversible terminator. For example, in some embodiments, the rephasing flow order includes (i) a flow step comprising a mixture containing non-terminating A nucleotides, non-terminating C nucleotides, non-terminating G nucleotides, and T (and / or U) nucleotides containing a reversible terminator; or (ii) a flow step comprising a mixture containing (or consisting of) non-terminating T (and / or U) nucleotides, non-terminating A nucleotides, non-terminating C nucleotides, and G nucleotides containing a reversible terminator; or (iii) a flow step comprising a mixture containing (or consisting of) non-terminating G nucleotides, non-terminating T (and / or U) nucleotides, non-terminating A nucleotides, and C nucleotides containing a reversible terminator; or (iv) a flow step comprising a mixture containing (or consisting of) non-terminating C nucleotides, non-terminating G nucleotides, non-terminating T (and / or) nucleotides, and A nucleotides containing a reversible terminator. By integrating nucleotides based on the template strand to extend the primer until a nucleotide containing a reversible terminator is integrated, which synchronizes the extended primer at the base in the sequencing cluster with the reversible terminator. The reversible terminator can then be removed, and subsequent sequencing procedures can be performed with the synchronized primer.
[0161] In some embodiments, the rephasing flow order includes (i) a first rephasing flow of a mixture of bases including C, G, and T (and / or U) in any order (omitting the A base), and a second rephasing flow order of a mixture of A, C, and G bases (omitting the T and / or U bases).
[0162] The method described herein for synchronizing the extended primer within a sequencing cluster can be used in a sequencing-by-synthesis method that uses non-terminating nucleotides to extend the primer. In some embodiments, the method is used in combination with other methods described herein, such as the fast-forward sequencing method described herein (e.g., a sequencing method that produces a "dark" region).
[0163] Mapping paired sequencing reads to a reference sequence
[0164] Paired sequencing reads can be mapped to a reference sequence, which may or may not include a target test variant. Sequencing data of the first region or the third region can be used to derive the sequence of the first region or the third region, respectively. The first region or a part of the first region, or the third region or a part of the third region can be mapped to the reference sequence. The distance between the first region and the third region (i.e., the length of the second region) can be determined or estimated to provide an approximate locus for the unmapped third or first region. Using the approximate locus, the unmapped first or third region can then be easily mapped to the reference sequence.
[0165] A mapped sequence refers to the alignment of one sequence (such as the sequence of a region or a part thereof) with another sequence (such as a reference sequence). A mappable sequence is a sequence (such as the sequence of a region or a part thereof) that can map to another sequence (such as a reference sequence) according to a selected mapping threshold (i.e., mapping score). Thus, an unmappable sequence is a sequence that cannot be mapped to another sequence according to the selected mapping threshold (mapping score). The score can be predetermined (i.e., selected before mapping) based on the error risk tolerance. For example, when mapping one sequence to another sequence, the Smith-Waterman algorithm can be used, and the mapping threshold can be selected to distinguish between "mappable" sequences and "unmappable" sequences. For example, the mapping score threshold can be +5 or higher, +6 or higher, +8 or higher, +10 or higher, +12 or higher, +14 or higher, +16 or higher, +18 or higher, or +20 or higher, where the match score is +1, the mismatch score is -1, the gap opening score is -2, and the gap extension score is -2. Those skilled in the art can select other scores or penalty scores.
[0166] Sequences, such as one or more regions of paired sequencing reads, can be mapped using any suitable mapping software, such as GATK, Bowtie, Bowtie2, BWA, BWA-MEM, NovoAlign, SOAP2, SOAP3, and others, including other aligners based on the Burrows-Wheeler transform (BWT). See, e.g., Miller et al., Assembly algorithms for next-generation sequencing data, Genomics, vol. 95, pp. 315-327 (2010); Chaisson et al., De novo fragment assembly with short mate-paired reads: does the read length matter? Genome Research, vol. 19, pp. 336-346 (2009); Mielczarek et al., Review of alignment and SNP calling algorithms for next-generation sequencing data, J. Appl. Genetics, vol. 57, pp. 71-79 (2016); Nielsen et al., Genotype and SNP calling from next-generation sequencing data, Nature Reviews Genetics, vol. 2, pp. 443-451 (2011); and Hwang et al., Systematic comparison of variant calling pipelines using gold standard personal exome variants, Sci Rep., vol. 5, 17875 (2015); each of which is incorporated herein by reference for all purposes.
[0167] The use of distance information to approximate the locus of a polynucleotide region to a reference sequence is useful for detecting structural variants (such as insertions or deletions) within a second region of a polynucleotide or resolving multiple mappable loci within a genome (e.g., where the first region or the third region includes repetitive regions and other non-unique sequences). The distance information discussed herein relates to the amount of space between two points (e.g., the start and end of a region) and can be considered in different frames of reference. For example, distance information in physical space can refer to the number of bases or the physical distance (e.g., the number of micrometers in one-dimensional space if the polynucleotide is linearly positioned). Distance information in sequencing data space (e.g., flowgram space) can refer to the number of flow steps used to extend a primer within the space in a given flow order. If the sequence (or reference sequence) and the flow order are known, the distance information in physical space and the distance information in sequencing data space are analytically interchangeable.
[0168] The distance information indicates the length of the second region, although it need not be the exact length of the second region, since the unmapped region is ultimately mapped within the location approximated by the distance information. In one example, the second region flow order (or information related to the second region flow order) and the probability distribution of bases in the second region are used to determine the distance information. The probability distribution of bases in the second region can be, for example, the assumed distribution of bases across the entire genome, or it can be a more localized probability based on the mapped loci of the first or third region. Information related to the second region flow order can be, for example, the number of different types of nucleotide bases used simultaneously to extend a primer through the second region. By way of example, a three-base flow step is used in a repetitive cycle to extend a primer within the second region (e.g., using a cycle step of (non-A)-(non-C)-(non-T)-(non-G), where each cycle step includes three other bases) and assuming that the distribution of bases in the second region is approximately the same as that of the genome as a whole, it is expected that the primer extends approximately 4.7 bases at each step in the cycle. Thus, the length of the second region can be approximated as 4.7 times the number of steps in the second region flow order.
[0169] In some embodiments, the distance information is derived from the expected reference sequencing data of the second region. As discussed herein, the expected reference sequencing data of the second region can be determined using the reference sequence and the second region flow order. Once the first or third region of the polynucleotide is mapped to the reference sequence, the expected sequence information, including the expected sequence length, which provides the length between the first and third regions of the polynucleotide, is determined.
[0170] When more than one mappable position is available within the reference sequence, distance information can be used to map paired sequencing reads to the reference sequence. For example, in some embodiments, the first region may map to the reference sequence with high confidence, but the third region may map to multiple different positions within the reference sequence. In some embodiments, the third region may map to the reference sequence with high confidence, but the first region may map to multiple different positions within the reference sequence. In some embodiments, both the first region and the third region may map to multiple different positions within the reference sequence. The distance information of the second region can be used to select the correct pair of positions for the first region and the second region that map to the reference sequence. For example, a method of mapping paired sequencing reads to a reference sequence may include mapping a first region (or a portion thereof) and a third region (or a portion thereof) of the paired sequencing reads to the reference sequence at two or more different pairs of positions including a first position and a second position. Then, the distance information indicating the length of the second region of the polynucleotide can be compared with the distance information indicating the length between the first position and the second position. If the compared distance information is close to or matches each other, the correct pair of positions can be selected. However, if the length of the second region is significantly different from the distance between the first position and the second position, the pair of positions can be rejected.
[0171] Figure 3 Illustrates how distance information indicating the length of the second region of a paired sequencing read is used to map the paired sequencing read to the reference sequence. The paired sequencing read 304 includes a first region 306, a second region 308, and a third region 310. The first region 306 may map to a reference first region 312 of the reference sequence 302, but the third region 310 may map to a reference third region, option A, 314 and a reference third region, option B, 316. The length of the distance between the end of the reference first region 312 and the start of the reference third region, option A, 314 is n bases (based on the reference sequence), and the length of the distance between the end of the reference first region 312 and the start of the reference third region, option B, 316 is m bases (based on the reference sequence). The distance information of the second region indicates that the length of the second region is approximately n bases. Thus, it can be concluded that the third region 310 is properly mapped to the reference third region, option A, 314. A similar analysis can be performed even if there are multiple mappable loci for the first region and / or multiple mappable loci for the third region.
[0172] In addition, when the first region or the third region cannot be unambiguously mapped to an exact position due to a repetitive region at the locus of the first region or the third region, distance information can be used to map the paired sequencing reads to the reference sequence. Figure 4Describes how to map a paired - end sequencing read pair to a reference sequence using distance information indicating the length of the second region of the paired - end sequencing read pair when the third region of the paired - end sequencing read pair maps to a repetitive region. Figure 4 Shows reference sequence 402 and paired - end sequencing read pair 404. The paired - end sequencing read pair includes a first region 406, a second region 408, and a third region 410. The first region 406 can map to a specific locus within the reference first region 412, but the third region 410 can map anywhere within the repetitive region 414. By knowing the length of the second region 408, the third region 410 can be more accurately mapped within the repetitive region 414. For example, if the length of the second region 408 is approximately n bases, once the first region 406 is mapped, this distance information can be used to locate the third region 410. Similarly, this method can be used when the third region can be precisely mapped but the first region maps within a repetitive region.
[0173] Detection of structural variants
[0174] Paired - end sequencing read pairs generated from polynucleotides derived from a genome can be used to detect variants, such as structural variants within the genome. Structural variants can include insertions, deletions, inversions, and chromosomal fusion variants, which can be located within the first, second, or third regions of the polynucleotide, or can be located at positions that bridge the first, second, or third regions of the polynucleotide.
[0175] An insertion in the genome can be of any size, such as from 1 base in length to hundreds or thousands of kilobases or longer. Additionally, an insertion can be an endogenous insertion (i.e., the inserted sequence is derived from a locus elsewhere in the subject's genome), or it can be an exogenous insertion (such as an inserted sequence derived from a source outside the subject's genome, like a viral genome inserted into the subject's genome). Exogenous insertions result in nucleic acid sequences not present within the reference sequence, posing additional challenges for detecting or localizing exogenous insertion variants within the subject's genome. The methods described herein can be used to detect and / or localize exogenous insertions as well as other structural variants.
[0176] In one example, a method of detecting a structural variant (such as an exogenous insertion) within a genome using paired - end sequencing reads includes mapping a first region (or a portion thereof) of the paired - end sequencing reads to a reference sequence and attempting to map a third region (or a portion thereof) to the reference sequence. If the third region (or a portion thereof) is unmappable, the presence of an exogenous insertion can be identified. This is because the reference sequence does not include a sequence corresponding to the third region. Similarly, a method of detecting an exogenous insertion within a genome using paired - end sequencing reads can include mapping a third region (or a portion thereof) of the paired - end sequencing reads to a reference sequence and attempting to map a first region (or a portion thereof) to the reference sequence. If the first region (or a portion thereof) is unmappable, the presence of an exogenous insertion can be identified. This is because the reference sequence does not include a sequence corresponding to the first region. Additionally (and in either example), the locus of the exogenous insertion within the reference sequence can be determined based on expected distance information indicative of the length of a second region. Figure 5 A schematic diagram illustrating an exemplary method of detecting an exogenous insertion. The paired - end sequencing reads 502 include a first region 504, a second region 506, and a third region 508, where the second region 506 is between the first region 504 and the third region 508. The third region 508 includes an exogenous insertion element 510 present in the genome 512 of a subject, although not present in the reference sequence 514. A reference element 516 is present in both the genome 512 of the subject and the reference sequence 514, although spaced differently from a reference first region 518. The first region 504 maps to the reference first region 518 within the reference sequence. However, the third region 508 does not have a corresponding region to map within the reference sequence 514 (i.e., it is unmappable). This indicates that the sequence of the third region 508 is the result of an exogenous insertion in the subject's genome. The distance information of the second region 506 can also be used to determine the locus of the exogenous genome relative to the reference first region 518. That is, if the length of the second region 506 is approximately n bases, the exogenous insertion is located approximately n bases from the end of the first region 504.
[0177] In another example, paired sequencing reads can be used to detect structural variants (such as insertions, deletions, inversions, or chromosomal fusions) by using expected sequencing data and comparing the resulting sequencing data with the expected sequencing data. For example, one of the first region (or a portion thereof) or the third region (or a portion thereof) of the paired sequencing reads can be mapped to a reference sequence. Distance information indicating the length of the second region can be used to determine the locus within the reference sequence of the unmapped first region (or a portion thereof) or the unmapped third region (or a portion thereof). For example, the distance information can be determined as described herein. Once the locus of the unmapped first region (or a portion thereof) or the unmapped third region (or a portion thereof) is determined, the expected sequencing data of the reference sequence at that locus can be determined. For example, the expected sequence data can be determined based on the sequence of the second region, the flow order of the second region, information related to the sequence of the unmapped region, and the flow order of the unmapped region. Then the expected sequencing data can be compared with the sequencing data of the generated unmapped region. A difference between the sequencing data of the unmapped region and the expected sequencing data indicates a structural variant at that locus.
[0178] Figure 6 An exemplary method for detecting structural variants using paired sequencing reads is illustrated. In step 602, one of the first region or a portion thereof (or the third region or a portion thereof) is mapped to a reference sequence. In step 604, the expected locus of the third region or a portion thereof (or the first region or a portion thereof) within the reference sequencing is determined. That is, if the first region or a portion thereof is mapped during step 602, the expected locus of the third region or a portion thereof is determined at step 604, and if the third region or a portion thereof is mapped during step 602, the expected locus of the first region or a portion thereof is determined at step 604. In step 606, the expected sequencing data of the third region or a portion thereof (or the first region or a portion thereof) at the determined expected locus is determined. In step 608, the expected sequencing data of the third region or a portion thereof (or the first region or a portion thereof) is compared with the determined sequencing data of the third region or a portion thereof (or the first region or a portion thereof), wherein a difference between the determined sequencing data and the expected sequencing data indicates a structural variant.
[0179] Figure 7Illustrated is a schematic diagram of using paired sequencing reads for detecting structural variants in a subject's genome, where the structural variant is an insertion. The subject's genome 702 includes a first region 704 and an insertion 706 between a first reference region 708 and a second reference region 710. The reference sequence 712 includes the first region 704, the first reference region 708, and the second reference region 710, but does not include the insertion 706 between the first reference region 708 and the second reference region 710 (the insertion may correspond to a region found in another part of the reference region or may be a completely exogenous sequence). The paired sequencing read pair 714 includes a first region 716 (corresponding to the first region 704) and a third region 718 (corresponding to the insertion 706), separated by a second region 720. The first region 716 of the paired sequencing read pair 714 maps to the first region 704 of the reference sequence 712. The distance information indicates that the length of the second region 720 of the paired sequencing read pair 714 is approximately n bases in length. Thus, the starting point of the expected locus 722 of the third region 718 is determined to start approximately n bases from the end of the first region 704. Then the expected sequencing data for the expected locus can be determined as described herein. For example, the reference sequence 712 (e.g., the first region 704 to the expected locus and / or the reference sequence including the expected locus), the flow order of the second region, and the flow order of the third region can be used to determine the expected sequencing data for the expected locus. In Figure 7 the illustrated example, if the third region 718 is the second reference region 710, the expected sequencing data corresponds to the obtained sequencing data because the second reference region 710 is at the expected locus. If the expected sequencing data for the expected locus is different from the sequencing data of the third region 718 of the generated paired sequencing read pair 714 (which is the case in Figure 7 the example illustrated), then a structural variant is detected.
[0180] Figure 8A schematic diagram illustrating the use of paired sequencing reads to detect structural variants in a subject's genome, where the structural variant is a deletion. The subject's genome 802 includes a first region 804, a first reference region 806, and a second reference region 808. The reference sequence 810 includes the first region 804, the first reference region 806, and the second reference region 808, as well as an additional region 812 located between the first reference region 806 and the second reference region 808. Although the additional region 812 is present in the reference sequence 810, the additional region 812 has been deleted from the subject's genome 802. The paired sequencing reads 814 include a first region 816 (corresponding to the first region 804) and a third region 818 (corresponding to the second reference region 808), which separates a second region 820. The first region 816 of the paired sequencing reads 814 maps to the first region 804 of the reference sequence 810. The distance information indicates that the length of the second region 820 of the paired sequencing reads 814 is approximately n bases long. Thus, the starting point of the expected locus 822 of the third region 818 is determined to be approximately n bases from the end of the first region 804. Then the expected sequencing data for the expected locus can be determined as described herein. For example, the reference sequence 812 (e.g., the first region 804 to the expected locus and / or the reference sequence including the expected locus), the flow order of the second region, and the flow order of the third region can be used to determine the expected sequencing data for the expected locus. In Figure 8 In the illustrated example, if the third region 818 is the additional region 812 (deleted in the subject's genome), then the expected sequencing data corresponds to the obtained sequencing data because the additional region 812 is at the expected locus. If the expected sequencing data for the expected locus is different from the sequencing data of the third region 818 of the generated paired sequencing reads 814 (which is the case in the example illustrated in Figure 8 ), then a structural variant is detected.
[0181] Figure 9A schematic diagram showing the use of paired sequencing reads to detect structural variants in a subject's genome, where the structural variant is an inversion. The subject's genome 902 includes a first segment 904, a second segment 906, and a third segment 908. The reference sequence 910 also includes a first segment 904, a second segment 906, and a third segment 908. However, in the reference sequence 910, the second segment 906 is closer to the 5' end relative to the third segment 908, while in the subject's genome 902, the second segment 906 is closer to the 3' end relative to the third segment 908. Thus, the second segment 906 and the third segment 908 in the subject's genome 902 are inverted relative to the reference sequence 910. The paired sequencing reads 912 include a first region 914 (corresponding to the first segment 904) and a third region 916 (corresponding to the third segment 908), which separate a second region 918. The first region 914 of the paired sequencing reads 912 maps to the first segment 904 of the reference sequence 910. The distance information indicates that the length of the second region 918 of the paired sequencing reads 912 is approximately n bases long. Thus, it is determined that the start of the expected locus 920 of the third segment 908 begins approximately n bases from the end of the first segment 904. Then the expected sequencing data for the expected locus can be determined as described herein. For example, the reference sequence 910 (e.g., the first segment 904 to the expected locus and / or the reference sequence including the expected locus), the flow order of the second region, and the flow order of the third region can be used to determine the expected sequencing data for the expected locus. In Figure 9 the example illustrated, if the third region 916 corresponds to the second segment 906, the expected sequencing data corresponds to the obtained sequencing data because the second segment 906 (rather than the third segment 908) is at the expected locus in the reference sequence 910. If the expected sequencing data for the expected locus is different from the sequencing data of the third region 916 of the generated paired sequencing reads 912 (which is the case in Figure 9 the example illustrated), then a structural variant is detected.
[0182] Figure 10A schematic diagram showing the use of paired - end sequencing reads to detect structural variants in a subject's genome, where the structural variant is a chromosome fusion. The chromosome fusion is caused by a chromosomal rearrangement event in which a portion of a chromosome fuses with another portion of a chromosome (the same chromosome or a different chromosome). The reference sequence 1002 includes chromosome A, which includes a first segment 1004 and a second segment 1006, and chromosome B, which includes a third segment 1008. The subject's genome 1010 includes a chromosome fusion of chromosome A and chromosome B at points 1012 and 1014 of the reference genome 1002. This forms chromosome A / B, which includes the 3' end of chromosome A and the 5' end of chromosome B, and forms chromosome B / A, which includes the 3' end of chromosome B and the 5' end of chromosome A. Thus, chromosome A / B includes the first segment 1004 and the third segment 1008, and chromosome B / A includes the second segment 1006. The paired - end sequencing reads 1016 are derived from chromosome A / B of the subject's genome 1010 and include a first region 1018 (corresponding to the first segment 1004) and a third region 1020 (corresponding to the third segment 1008), separated by a second region 1022. The first region 1018 of the paired - end sequencing reads 1016 maps to the first fragment 1004 of the reference sequence 1002. The distance information indicates that the length of the second region 1022 of the paired - end sequencing reads 1016 is approximately n bases long. Thus, the start of the expected locus 1024 of the third fragment 1020 is determined to be approximately n bases from the end of the first fragment 1004. Then the expected sequencing data for the expected locus can be determined as described herein. For example, the chromosome A of the reference sequence 1002 (e.g., the first segment 1004 and / or the reference sequence between the first segment 1004 and the expected locus, the second segment 1006), the flow order of the second region 1022, and the flow order of the third region 1020 can be used to determine the expected sequencing data for the expected locus. In Figure 10 the example shown, the expected sequencing data corresponds to the sequencing data that would be obtained if the third region 1020 corresponded to the second segment 1006 because the second segment 1006 (rather than the third segment 1008) is at the expected locus in the reference sequence 1002. If the expected sequencing data for the expected locus is different from the sequencing data of the third region 1020 of the generated paired - end sequencing reads 1016 (which is the case in the example shown in Figure 10 ), then a structural variant is detected.
[0183] The connection of a structural variant (e.g., an insertion, deletion, chromosome fusion, or inversion) relative to the reference sequence does not need to span the entire first region or third region of the paired - end sequencing reads. In some embodiments, at least a portion of the structural variant ends within the first region or third region of the paired - end sequencing reads. The expected sequencing data will still be different from the sequencing data determined for the first or third region.
[0184] Detection of Variants within the Second Region
[0185] In some embodiments, paired sequencing reads are used to detect variants within the second region, even if nucleotide incorporation into a primer extending through the second region is not required. Detectable variants include structural variants (such as insertions, deletions, inversions, or chromosomal fusions) or single nucleotide polymorphisms (SNPs).
[0186] Methods for detecting structural variants (such as chromosomal fusions, inversions, insertions, or deletions) can include mapping a first region (or a portion thereof) and a third region (or a portion thereof) of the paired sequencing reads to a reference sequence. Distance information for an inversion that occurs entirely within the second region is generally considered with reference to the flow order of the second region (e.g., in the flow graph space), while distance information for chromosomal fusions, insertions, or deletions that do not occur entirely within the second region (e.g., at least partially within the first region or the third region) can be considered with reference to physical space or the flow order of the second region. Distance information (i.e., mapped distance information) between the first region mapped to the reference sequence and the third region mapped to the reference sequence can be determined. The mapped distance information indicates the distance between the mapped position of the first region mapped to the reference sequence and the mapped position of the third region mapped to the reference sequence, such as the number of base pairs between the first and third mapped regions. Expected distance information indicating the length of the second region of the paired sequencing reads can also be determined (e.g., using the flow order of the second region and the reference sequence, or as otherwise described herein). A comparison between the expected distance information and the mapped distance information can be used to detect structural variants. For example, if the expected distance is shorter than the mapped distance, it indicates a structural variant within the subject's genome, such as an insertion or chromosomal fusion variant. If the expected distance is longer than the mapped distance, it indicates a deletion variant within the subject's genome.
[0187] Figure 11 An exemplary method for detecting a structural variant is illustrated, which includes, in step 1102, mapping a first region (or a portion thereof) and a third region (or a portion thereof) of the paired sequencing reads to a reference sequence. In step 1104, mapped sequence distance information is determined, which indicates the distance between the first region mapped to the reference sequence and the third region mapped to the reference sequence. In step 1106, expected distance information for the second region is determined based on the sequence region flow order and information about the sequence of the second region (e.g., the sequence of the second region from the reference sequence). In step 1108, a structural variant is identified by comparing the expected distance information with the mapped distance information, wherein a difference between the mapped distance information and the expected distance information indicates a structural variant.
[0188] Figure 12Illustrates a schematic diagram demonstrating an example of how to detect structural variants using paired sequencing reads. The example shown depicts an insertion in the subject's genome, but the method is similarly applicable to other structural variants (e.g., deletions or chromosomal fusions). The reference sequence 1202 includes a first segment 1204 and a second segment 1206. The subject's genome 1208 also includes the first segment 1204 and the second segment 1206, but further includes an insert 1210 between the first segment 1204 and the second segment 1206. The paired sequencing reads 1212 generated from the subject's genome 1208 include a first region 1214 corresponding to the first segment 1204 and a third region 1216 corresponding to the second segment 1206. A second region 1218 separates the first region 1214 and the third region 1216. The sequences of the first region 1214 and the third region 1216 can be mapped to the reference sequence 1202 at the first segment 1204 and the second segment 1206, respectively. Once mapped, mapping distance information indicating the distance between the first region 1214 and the third region 1216 mapped to the reference sequence 1202 (i.e., the distance between the first segment 1204 and the second segment 1206 of the reference sequence 1202) is determined as distance n. Expected distance information for the length of the second region 1218 can also be determined as m. Then, the structural variant can be determined by comparing the mapped distance information n with the expected distance information m.
[0189] In another method for detecting variants (e.g., structural variants or SNPs) within a second region, the expected sequencing data is compared to the determined sequencing data. For example, in some embodiments, a method for detecting variants between two sequencing regions of a paired sequencing read (wherein a primer is extended through a first region using nucleotides provided in a first region flow order and / or a primer is extended through a third region using nucleotides provided in a third region flow order) includes mapping the first region (or a portion thereof) and / or the third region (or a portion thereof) to a reference sequence. The expected reference sequencing data for the other region or a portion thereof is then determined (i.e., if the first region or portion is mapped, the other region refers to the third region or a portion thereof; and if the third region or a portion thereof is mapped, the other region refers to the first region or a portion thereof). For example, the expected sequencing data can be determined using the reference sequence of the second region, the second region flow order, the reference sequence of the other region or a portion thereof (i.e., the third region or a portion thereof if the first region or a portion thereof is the mapped region, the first region or a portion thereof if the third region or a portion thereof is the mapped region), and the flow order of the other region or a portion thereof. In another example, the expected sequencing data is determined using the reference sequence of the second region, the second region flow order, the flow order of the other region, and the sequencing data associated with the sequence of the other region (which can be the same sequencing data generated when generating the paired sequencing read or sequencing data generated by other means). The expected sequencing data for the determined other region can be compared to the sequencing data for the generated other region. A difference between the expected and generated sequencing data indicates the presence of a variant.
[0190] In some embodiments, a method of detecting a variant (such as a structural variant (e.g., chromosomal fusion, inversion, insertion, or deletion) or SNP) between two sequencing regions of a paired sequencing read includes mapping a first region or a portion thereof to a reference sequence; determining expected sequencing data for a third region or a portion thereof using (1) the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region, or (2) the reference sequence of the second region, the second region flow order, the third region flow order, and the generated sequencing data related to the sequence of the third region; and detecting the presence of the variant by comparing the expected sequencing data for the third region with the generated sequencing data related to the sequence of the third region. In some embodiments, a method of detecting a variant (such as a structural variant (e.g., chromosomal fusion, inversion, insertion, or deletion) or SNP) between two sequencing regions of a paired sequencing read (wherein a nucleotide extension primer is provided in a first region flow order) includes mapping a third region or a portion thereof to a reference sequence; determining expected sequencing data for a first region or a portion thereof using (1) the reference sequence of the second region, the second region flow order, the first region flow order, and the reference sequence of the first region, or (2) the reference sequence of the second region, the second region flow order, the first region flow order, and the generated sequencing data related to the sequence of the first region; and detecting the presence of the variant by comparing the expected sequencing data for the first region with the generated sequencing data related to the sequence of the first region.
[0191] Figure 13 An exemplary method of detecting a variant between two sequencing regions of a paired sequencing read is illustrated. At step 1302, a first region or a portion thereof or a third region or a portion thereof of the paired sequencing read is mapped to a reference sequence. At step 1304, expected sequencing data for a third region or a portion thereof or a first region or a portion thereof is determined. At step 1306, the presence of the variant is detected by comparing the expected sequencing data for the first region or the third region with the generated sequencing data related to the sequence of the first region or the third region. Exemplary variant detection methods are provided in the examples.
[0192] Methods for detecting variants can use a reference sequence, which may or may not include the test variant. The test variant can be selected, for example, to identify a test variant within a second polynucleotide or from a biomarker panel. For example, the test variant can be used to determine the haplotype of a polynucleotide. Alleles or variants can be identified in a polynucleotide, and the methods described herein can be used to determine whether the polynucleotide that gave rise to a paired sequencing read pair is of the same haplotype or a different haplotype as the polynucleotide having the identified allele or variant. The test variant detected in the paired sequencing read pair can be related to the allele sequenced in the first or third region of the polynucleotide.
[0193] When detecting the presence of a test variant, the reference sequence can include the test variant, and the presence of the test variant within the subject's genome can be detected by comparing the expected test variant sequencing data for the third region or a portion thereof with the determined sequencing data for the third region or a portion thereof. If the expected test variant sequencing data matches the determined sequencing data, the test variant is detected within the reference sequence. For example, in some embodiments, a method for detecting a test variant between two sequencing regions of a paired sequencing read pair (wherein a primer has been extended through the first region using nucleotides provided in a first region flow order and / or a primer has been extended through the third region using nucleotides provided in a third region flow order) includes mapping the first region or a portion thereof to a reference sequence that includes the test variant. Then, the reference sequencing data expected for the test variant of the other region or a portion thereof is determined (i.e., if the first region or portion is mapped, the other region refers to the third region or a portion thereof). The reference sequencing data expected for the test variant can be determined, for example, using a reference sequence that includes the test variant of the second region, the second region flow order, the reference sequence of the other region or a portion thereof, and the flow order of the other region or a portion thereof. In another example, a reference sequence having the test variant of the second region, the second region flow order, the flow order of the other region, and sequencing data related to the sequence of the other region (which can be the same sequencing data generated when generating the paired sequencing read pair or sequencing data generated by other means) is used to determine the expected sequencing data. The determined reference sequencing data expected for the test variant of the other region can be compared with the generated sequencing data for the other region. A match between the expected and generated sequencing data indicates the presence of the test variant.
[0194] Detection of short genetic variants
[0195] The methods described herein can be used to detect short genetic variants (e.g., SNPs or short indels (less than 10 consecutive bases in length)) within a second region (e.g., when primer extension through the second region does not detect the presence or absence of a nucleotide label in the integrated extension primer, or when the primer is extended with a mixture comprising at least two different types of nucleobases). The short genetic variants within the second region can be detected by analyzing signals obtained upon detection of nucleotide incorporation in a downstream (e.g., third) region. The short genetic variants can be, for example, variants or mutations found within a subset of individuals, or variants or mutations unique to a single or specific individual. The short genetic variants can be germline variants or somatic variants.
[0196] Sequencing data can be generated based on the detection of incorporated nucleotides and the order of nucleotide incorporation. For example, take the flowing extension sequences (i.e., each reverse complementary sequence of the corresponding template sequence): CTG, CAG, CCG, CGT, and CAT (assuming no previous or subsequent sequences for the sequencing method), and a repeating flowing cycle of T-A-C-G (i.e., adding T, A, C, and G nucleotides in order in the repeating cycle). A particular type of nucleotide at a given flowing position will be incorporated into the primer only when the complementary base is present in the template polynucleotide. An exemplary resulting flowing map is shown in Table 5, where 1 indicates incorporation of the introduced nucleotide and 0 indicates non-incorporation of the introduced nucleotide. The flowing map can be used to derive the sequence of the template strand. For example, the sequencing data (e.g., flowing map) discussed herein represents the sequence of the extended primer strand, and its reverse complementary sequence can be readily determined to represent the sequence of the template strand. The asterisk (*) in Table 5 indicates that a signal may be present in the sequencing data if additional nucleotides are incorporated in the extended sequencing strand (e.g., the longer template strand).
[0197] Table 5
[0198]
[0199] The flowing map can be binary or non-binary. A binary flowing map detects the presence (1) or absence (0) of incorporated nucleotides. A non-binary flowing map can more quantitatively determine the number of nucleotides incorporated from each stepwise introduction. For example, the extension sequence of CCG will include incorporation of two C bases in the extended primer within the same C flow (e.g., at flowing position 3), and the signal emitted by the labeled base will have an intensity level greater than that corresponding to single base incorporation. This is shown in Table 5. The non-binary flowing map also indicates the presence or absence of bases and can provide additional information, including the number of bases that may be incorporated into each extended primer at a given flowing position. These values need not be integers. In some cases, these values can reflect the uncertainty and / or probability of the number of bases incorporated at a given flowing position.
[0200] In some embodiments, the sequencing data set includes flow signals representing base counts, where the base counts indicate the number of bases in the sequenced nucleic acid molecules integrated at each flow position. For example, as shown in Table 5, a primer extended with a CTG sequence using a T-A-C-G flow cycle order has a value of 1 at position 3, indicating a base count of 1 at that position (1 base is C, which is complementary to the G in the template strand being sequenced). Also in Table 5, a primer extended with a CCG sequence using a T-A-C-G flow cycle order has a value of 2 at position 3, indicating a base count of 2 for the extended primer at that position during that flow position. Here, the 2 bases refer to the C-C sequence at the start of the CCG sequence in the extended primer sequence, and it is complementary to the G-G sequence in the template strand.
[0201] The flow signals in the sequencing data set can include one or more statistical parameters that indicate the likelihood or confidence interval of one or more base counts at each flow position. In some embodiments, the flow signals are determined from analog signals detected during the sequencing process, such as fluorescence signals of one or more bases integrated into the sequencing primer during sequencing. In some cases, the analog signals can be processed to generate the statistical parameters. For example, machine learning algorithms can be used to correct context effects of the analog sequencing signals, as described in the published international patent application WO2019084158A1, which is hereby incorporated by reference in its entirety. Although zero or more bases are integrated at any given flow position, a given analog signal may not exactly match the analog signal. Thus, for a given detected signal, statistical parameters can be determined that indicate the likelihood of the number of bases integrated at the flow position. By way of example only, for the CCG sequence in Table 5, the likelihood that the flow signal indicates 2 bases are integrated at flow position 3 can be 0.999, and the likelihood that the flow signal indicates 1 base is integrated at flow position 3 can be 0.001. The sequencing data set can be formatted as a sparse matrix, where the flow signals include statistical parameters indicating the likelihood of multiple base counts at each flow position. By way of example only, primers extended with the following sequences using a repeating flow cycle order of T-A-C-G: TATGGTCGTCGA (SEQ ID NO:15) can produce Figure 14A the sequencing data set shown. The statistical parameters or likelihood values can vary, for example, based on noise or other artifacts present during detection of the analog signals during sequencing. In some embodiments, if the statistical parameter or likelihood is below a predetermined threshold, the parameter can be set to a predetermined non-zero value that is essentially zero (i.e., some very small value or negligible value) to assist with the statistical analysis further discussed herein, where a true zero value could cause computational errors or insufficiently distinguish levels of improbability, e.g., highly improbable (0.0001) and inconceivable (0).
[0202] Values indicative of the likelihood of a sequencing data set for a given sequence can be determined from a sequencing data set without sequence alignment. For example, in the case of given data, the most likely sequence can be determined by selecting the base counts with the highest likelihood at each flow cell position, as shown by the stars in Figure 14B (using the same data shown in Figure 14A ). Thus, the sequence of primer extension can be determined based on the most likely base counts at each flow cell position: TATGGTCGTCGA (SEQ ID NO: 15). From this, the reverse complementary sequence (i.e., the template strand) can be readily determined. In addition, given the sequence TATGGTCGTCGA (SEQ ID NO: 15) (or the reverse complementary sequence), the likelihood of this sequencing data set can be determined as the product of the likelihoods selected at each flow cell position.
[0203] A sequencing data set associated with a nucleic acid molecule can be compared to one or more (e.g., 2, 3, 4, 5, 6 or more) possible candidate sequences. A close match between the sequencing data set and a candidate sequence (based on a match score, discussed below) indicates that the sequencing data set may be from a nucleic acid molecule having the same sequence as the closely matched candidate sequence. In some embodiments, the sequence of the sequenced nucleic acid molecule can be mapped to a reference sequence (e.g., using the Burrows-Wheeler alignment (BWA) algorithm or other suitable alignment algorithms) to determine the locus (or loci) of the sequence. As described above, a sequencing data set in flow space can be readily converted to base space (or vice versa if the flow order is known), and mapping can be performed in flow space or base space. The locus (or loci) corresponding to the mapped sequence can be associated with one or more variant sequences, which can operate as candidate sequences (or haplotype sequences) for the analysis methods described herein. An advantage of the methods described herein is that in some cases, the sequence of the sequenced nucleic acid molecule does not need to be aligned with each candidate sequence using an alignment algorithm, which is typically computationally expensive. Instead, the match score for each candidate sequence can be determined using the sequencing data in flow space, which is a more computationally efficient operation.
[0204] The match score indicates the degree to which the sequencing data set supports a candidate sequence. For example, in the case of expected sequencing data for a given candidate sequence, a match score indicative of the likelihood of the sequencing data set matching the candidate sequence can be determined by selecting a statistical parameter (e.g., likelihood) at each flow cell position corresponding to the base counts at that flow cell position. The product of the selected statistical parameters can provide the match score. For example, assume the sequencing data set for the extended primer shown in Figure 14A and TATGGTC ACandidate primer extension sequences of TCGA (SEQ ID NO:16). Figure 14C (Shown Figure 14A in the same sequencing dataset as) shows the traces (solid circles) of the candidate sequences. As a comparison, the traces of the TATGGTC G TCGA (SEQ ID NO:15) sequence (see Figure 14B ) are shown as open circles in Figure 14C . The match scores indicating the likelihood that the sequencing data matches the first candidate sequence TATGGTCATCGA (SEQ ID NO:16) are substantially different from the match scores indicating the likelihood that the sequencing data matches the second candidate sequence TATGGTCGTCGA (SEQ ID NO:15), even though the sequences differ by only a single base change. As shown in Figure 14C , a difference between the traces is observed at flow position 12 and propagates for at least 9 flow positions (and potentially longer if the sequencing data extends across additional flow positions). This persistent propagation across one or more flow cycles can be referred to as a "flow shift" or "cycle shift", and is generally a very unlikely event if the sequencing dataset matches the candidate sequence.
[0205] A match score can then be determined between each sequencing dataset and the candidate sequence (or each candidate sequence). For example, the likelihood L(Rj|Hi) that the sequencing dataset matches a given candidate sequence can be determined using the likelihood (e.g., the product) of the selected base counts at each flow position of the given candidate sequence.
[0206] The match score can be used to classify the test sequencing data and / or the nucleic acid molecule associated with the test sequencing data. The classifier can indicate that the nucleic acid molecule includes a variant (e.g., the variant included in the candidate sequence), that the nucleic acid molecule does not include a variant, or can indicate a null determination. A null determination neither indicates the presence or absence of a variant in the nucleic acid molecule associated with the test sequencing data, but rather indicates that the match score cannot be used to make a determination with the desired statistical confidence. For example, if the match score is above the desired confidence threshold, the test sequencing data or nucleic acid molecule can be classified as having a variant. Conversely, for example, if the match score is below the desired confidence threshold, the test sequencing data or nucleic acid molecule can be classified as not having a variant.
[0207] The above analysis can be applied to select a candidate sequence from two or more different candidate sequences. A match score indicating the likelihood of the sequencing data set matching each candidate sequence can be determined. For example, for each candidate sequence, a statistical parameter corresponding to the base count of the candidate sequence at each flow cell position in the sequencing data set can be selected. In some embodiments, such analysis includes generating expected sequencing data for candidate sequencing, assuming that the candidate sequence is sequenced using the same flow order as the sequencing data set of the test nucleic acid molecule used to generate the sequencing. This can be generated by sequencing the nucleic acid molecule with the candidate sequence, or by generating a candidate sequencing data set on a computer based on the candidate sequence and the flow order. An exemplary candidate sequencing data set is shown below the test data sequencing data set in Figure 14C , where the first candidate sequence (TATGGTCATCGA (SEQ ID NO:16)) corresponds to the solid circle trace, and the second candidate sequence (TATGGTCGTCGA (SEQ ID NO:15)) corresponds to the open circle trace. In some embodiments, for example, if the match scores of two or more different candidate sequences are determined, the test sequencing data or nucleic acid molecule can be classified as a variant having one of the two or more candidate sequences, not having one of the two or more candidate sequences, or an indeterminate decision can be made between the two or more candidate sequences (e.g., if no decision can be made for any candidate sequence, or if the match score indicates two or more different variants at the same locus).
[0208] Once the matching scores of the sequencing datasets of the candidate sequences are determined, candidate sequences with short genetic variants can be selected based on the matching scores (e.g., generating a candidate sequence with the highest likelihood of match from two or more candidate sequences). The sequencing data generated from the nucleic acid molecules of the sequences with short genetic variants will match the candidate sequences with short genetic variants, and that candidate sequence can be selected, while the rejected (or unselected) candidate sequences do not include short genetic variants, as indicated by a lower likelihood of match (based on the determined matching scores of those candidate sequences). The unselected candidate sequences can differ from the selected candidate sequence (which best matches the sequencing dataset of the sequenced nucleic acid molecule) at two or more flow positions, and the two or more flow positions can be two or more consecutive flow positions or two or more non - consecutive flow positions. In some embodiments, the unselected candidate sequences differ from the selected candidate sequence at 3 or more, 4 or more, 5 or more, 6 or more, 7 or more, 8 or more, 9 or more, or 10 or more flow positions. In some embodiments, the unselected candidate sequences differ from the selected candidate sequence in 1 or more, 2 or more, 3 or more, 4 or more, or 5 or more flow cycles. In some embodiments, the unselected candidate sequences differ from the selected candidate sequence at X base positions, where the sequencing dataset associated with the nucleic acid molecule differs from the unselected candidate sequence at (X + 2) or more flow positions. An increase in the number of different flow positions between the selected candidate sequence and the unselected candidate sequence (where the sequencing dataset of the sequenced nucleic acid molecule best matches the selected candidate sequence) reduces the likelihood of the sequencing dataset of the sequenced nucleic acid molecule obtained from sequencing the nucleic acid molecule using the unselected candidate sequence.
[0209] The likelihood that the sequencing dataset of the sequenced nucleic acid molecule matches the unselected candidate sequence is preferably low, such as less than 0.05, less than 0.04, less than 0.03, less than 0.02, less than 0.01, less than 0.005, less than 0.001, less than 0.0005, or less than 0.0001. The likelihood that the sequencing dataset of the sequenced nucleic acid molecule matches the selected candidate sequence is preferably high, such as greater than 0.95, greater than 0.96, greater than 0.97, greater than 0.98, greater than 0.99, greater than 0.995, or greater than 0.999.
[0210] In some embodiments, a method for detecting short genetic variants in a test sample can include analyzing a plurality of test sequencing data sets, where each test sequencing data set is associated with separate test nucleic acid molecules in the test sample. For example, if the sequence of a nucleic acid molecule aligns with a reference sequence, the nucleic acid molecule at least partially overlaps at a locus. At least a portion of the nucleic acid molecule can have different sequencing start positions (relative to the locus), which results in different flow positions and / or different flow order contexts for a given base within the sequence. In this way, the same candidate sequence can be used to analyze a plurality of test sequencing data sets. For each candidate sequence, a match score can be determined that indicates the likelihood of the plurality of test sequencing data sets matching the candidate sequence, and the candidate sequence with the highest likelihood of a match (and thus, including the short genetic variant) can be selected. An exemplary analysis of detecting short genetic variants using a plurality of test sequencing data sets is shown in Figure 15A - 15D In Figure 15A the sequences corresponding to three sequenced test nucleic acid molecules (R1, R2, and R3, each represented by the sequence of an extended primer) are aligned with a reference sequence at overlapping loci associated with two candidate sequences (HI and H2). Figure 15B , Figure 15C and Figure 15D show exemplary sequencing data sets for R1, R2, and R3, respectively, and the selected statistical parameters at each flow position in the sequencing data set corresponding to the bases of H1 (solid circles) or H2 (hollow circles).
[0211] One or more determined match scores can be used to determine the presence (or identity) or absence of a short genetic variant in the test sample. In some embodiments, for example, a single nucleic acid molecule (or associated test sequencing data set) classified as having a variant may be sufficient to determine the presence, identity, or absence of the variant, e.g., if the match score indicates a match to the candidate sequence with a desired or preset confidence. In some embodiments, a predetermined number (e.g., 1 or more, 2 or more, 3 or more, 4 or more, 5 or more, etc.) of nucleic acid molecules (or test sequencing data sets associated with the nucleic acid molecules) are classified as having a variant before determining the variant for the test sample. In some embodiments, the number of nucleic acid molecules (or test sequencing data sets associated with the nucleic acid molecules) is dynamically selected based on the match score; for example, a single nucleic acid molecule classified as having a variant with a high-confidence match score can be used to determine the variant, or two or more nucleic acid molecules classified as having a variant with a lower-confidence match score can be used to determine the variant.
[0212] Optionally, individual match scores of the sequencing datasets are jointly analyzed to determine the match scores of multiple test sequencing datasets. For example, once the match scores of each test sequencing dataset for each candidate sequence have been determined using the methods described herein, known Bayesian methods can be used, such as the HaplotypeCaller algorithm included in the Genome Analysis Toolkit (GATK), to determine a match score indicative of the likelihood that the multiple test sequencing datasets match the candidate sequence, and the candidate sequence with the highest likelihood of a match can be selected. See, e.g., Depristo et al., A framework for variation discovery and genotyping using next-generation DNA sequencing data, Nature Genetics 43, 491-498 (2011); and Poplin et al., Scaling Accurate Genetic Variant Discovery to Tens of Kestomen of Samples, BioRxiv, www.bioRxiv.org / content / 10.1101 / 201178v3 (July 24, 2018); Hwang et al., Systematic Comparison of Variant Calling Pipelines Using Gold Standard Personal Exome Variants, Scientific Reports, Volume 5, Article No. 17875 (2015); the respective contents of which are incorporated herein by reference.
[0213] Assume Example 1 - SNP detection. Sequencing of the hypothetical nucleic acid molecule is performed using non-terminating nucleotides provided in separate nucleotide streams according to the flow cycle order A - T - G - C, resulting in Figure 14AThe test sequencing dataset shown in . Each value in the sequencing dataset indicates the likelihood that the base count indicated at each flow cell position is correct. Based on the sequencing dataset, a preliminary sequence was determined to be TATGGTCGTCGA (SEQ ID NO:15), which was mapped to a locus of a reference genome. The locus of the reference genome is associated with potential haplotype sequences TATGGTCGTCGA (SEQ ID NO:15) (H1) and TATGGTCATCGA (SEQ ID NO:16) (H2). For each haplotype, the likelihood values associated with the base counts of the haplotype sequence at each flow cell position were selected. The likelihood of the sequencing dataset for each given haplotype was determined by multiplying the likelihood values associated with the base counts of the haplotype sequence at each flow cell position. If H1 is the correct sequence, the log-likelihood of the sequencing dataset is -0.015, and if H2 is the correct sequence, the log-likelihood of the sequencing dataset is -27.008. Thus, the sequence of H1 was selected for the nucleic acid molecule.
[0214] Hypothetical Example 2 - Indel detection. Sequencing of a hypothetical nucleic acid molecule was performed using non-terminating nucleotide pairs provided in separate nucleotide streams according to the flow cycle order A - T - G - C, resulting in Figure 16 the test sequencing dataset shown in . Each value in the sequencing dataset indicates the likelihood that the base count indicated at each flow cell position is correct. Based on the sequencing dataset (i.e., by selecting the most likely base count at each flow cell position), a preliminary sequence was determined to be TATGGTCGATCG (SEQ ID NO:22), which was mapped to a locus of a reference genome. The locus of the reference genome is associated with potential haplotype sequences TATGGTCGTCGA (SEQ ID NO:21) (H1) and TATGGTCGATCG (SEQ ID NO:22) (H2). For each haplotype, the likelihood values associated with the base counts of the haplotype sequence at each flow cell position were selected. The likelihood of the sequencing dataset for each given haplotype was determined by multiplying the likelihood values associated with the base counts of the haplotype sequence at each flow cell position. If H1 is the correct sequence, the log-likelihood of the sequencing dataset is -24.009, and if H2 is the correct sequence, the log-likelihood of the sequencing dataset is -0.015. Thus, the sequence of H2 was selected for the nucleic acid molecule.
[0215] When the signal difference caused by a variant in the second (i.e., “dark”) region propagates to the third region (i.e., the region where nucleotide incorporation is detected), a flow shift caused by the variant in the second region can be detected in the third region. In the hypothetical example discussed above, for instance, cycle 3 can be considered the “dark” or second region (which can be any number of cycles), and cycles 4 and 5 can be the third region (which can also be any number of cycles).
[0216] Detection of transversions
[0217] A transversion is an SNP that exchanges a purine for a pyrimidine or vice versa. The methods described herein can be implemented to be particularly sensitive for detecting transversions within the second region of a paired sequencing read. For example, a second region flow order that includes alternating nucleotide pairs of pyrimidines (C+T) and purines (A+G) will be highly sensitive to transversions by primer extension through the second region.
[0218] For example, paired sequencing reads for detecting the presence of a base transversion in a polynucleotide can be generated by: (a) hybridizing the polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using a flow order that includes (1) cytosine and thymine and (2) adenine and guanine alternating nucleotide pairs; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides. Transversions can be detected in the second region even without detecting the presence or absence of labels of nucleotides incorporated into the primer extended through the second region.
[0219] Paired sequencing reads generated for transversion detection can be used to detect a transversion by mapping the first region or a portion thereof (or the third region or a portion thereof) of the paired sequencing reads; determining the expected sequencing data for the third region or a portion thereof (or the first region or a portion thereof) using the second region flow order, the third region flow order, and a reference sequence; and detecting the presence of a base transversion based on the difference between the expected reference sequencing data for the third region and the generated sequencing data for the third region.
[0220] The expected reference sequencing data for the third region or a portion thereof (or the first region or a portion thereof) can be determined, for example, by using the second region flow order, the third region flow order, the reference sequence of the second region, and the reference sequence of the third region. In some embodiments, the expected reference sequencing data for the third region is determined using the second region flow order, the third region flow order, the reference sequence of the second region, and the sequence data generated that is related to the sequence of the third region, where the sequence data generated that is related to the sequence of the third region is the same or different sequence data generated when generating the paired sequencing reads.
[0221] Variant verification
[0222] Multiple at least partially overlapping paired sequencing reads can be used to verify the variant status. Since sequencing errors may occasionally occur during the normal process of nucleotide incorporation into the extension primer (e.g., due to polymerase errors or read errors), variant verification can help minimize the reporting of false positives or false negatives. Additionally, the sensitivity of the methods described herein can vary depending on the content and flow order of the variant used when extending the primer through the second region. Thus, to minimize false positive or false negative errors, overlapping or at least partially overlapping paired sequencing read pairs can be compared to verify the variant. The multiple paired sequencing read pairs used to verify the variant can include different starting points (e.g., different first region starting points, different second region starting points, and / or different third region starting points), or can be generated using different second region flow orders.
[0223] A target test variant can be selected, and multiple overlapping paired sequencing read pairs are analyzed to determine the status of the test variant within the paired sequencing read pair (e.g., whether the variant is present or absent). The overlapping paired sequencing read pairs include loci corresponding to the test variant locus. In some embodiments, the test variant is within the first region of at least a portion of the paired sequencing read pair. In some embodiments, the test variant is within the second region of at least a portion of the paired sequencing read pair. In some embodiments, the test variant is within the third region of at least a portion of the paired sequencing read pair.
[0224] A tolerance threshold can be selected to determine whether the test variant is present or absent at the locus. If the test variant is positively identified in more than a predetermined threshold of the multiple paired sequencing reads, then, for example, the test variant is positively determined. The threshold can be set according to the needs of the risk tolerance. For example, the tolerance threshold can be 60% or higher, 70% or higher, 80% or higher, 90% or higher, or 95% or higher of the paired sequencing read pairs that identify the test variant.
[0225] Figure 17Shows an exemplary schematic for comparing paired sequencing reads to determine the status of a test variant. Multiple overlapping paired sequencing reads 1402 are aligned with a reference sequence 1404. At locus 1406, four of the five overlapping paired sequencing reads allow identification of a variant not identified in one of the paired sequencing reads. Specifically, paired sequencing reads 1408, 1410, 1414, and 1416 each include a variant identified at loci 1418, 1420, 1424, and 1426, respectively. The locus of the variant at each paired sequencing read is aligned with locus 1406 of the reference sequence 1404. Paired sequencing read 1412 does not identify the variant at locus 1422 (e.g., due to a sequencing read error or due to the content of the variant and the second region as well as the flow order used to generate paired sequencing read 1412).
[0226] Construction or validation of a consensus sequence
[0227] Paired sequencing reads generated according to the methods described herein can be used to generate one or more consensus sequences by assembling the paired sequencing reads. Paired-end sequencing has previously been used to assemble consensus sequences, but the limited information available for the region between the sequenced ends of the polynucleotide results in lower quality consensus sequences with frequently misaligned sequences. See, e.g., Zerbino et al., Velvet: Algorithms for de novo short read assembly using deBruinn graphs, Genome Research, Vol. 18, pp. 821-820 (2008), which is incorporated herein by reference for all purposes. The methods described herein allow extraction of substantially more information from the unsequenced second region between the first and third regions of the sequencing. This additional information allows for a more robust and accurate consensus sequence.
[0228] In one instance, distance information indicative of the length of a second region of a paired sequencing read is used to assemble one or more consensus sequences. The distance information can be determined as described herein. In one instance, the distance information is determined using the second region flow order (or information related to the second region flow order) and the probability distribution of bases in the second region. The probability distribution of bases in the second region can be, for example, the assumed distribution of bases across the genome or can be a more localized probability based on the mapped loci of the first or third regions. Information related to the second region flow order can be, for example, the number of different types of nucleotide bases used simultaneously to extend a primer through the second region. For example, a three-base flow step is used in iterative cycles to extend a primer within the second region (e.g., using a cycle step of (not A)-(not C)-(not T)-(not G), where each cycle step includes three additional bases) and assuming that the distribution of bases in the second region is approximately the same as across the genome, it is expected that the primer extends approximately 4.7 bases at each step in the cycle. Thus, the length of the second region can be approximated as 4.7 times the number of steps in the second region flow order.
[0229] In some embodiments, the distance information is derived from the expected reference sequencing data for the second region. As discussed herein, the expected reference sequencing data for the second region can be determined using a reference sequence and the second region flow order. Once the first or third region of a polynucleotide is mapped to a reference sequence, the expected sequence information, including the expected sequence length, which provides the length between the first and third regions of the polynucleotide, is determined.
[0230] Mate - paired sequencing reads can be used to verify one or more consensus sequences or a portion of one or more consensus sequences. Given the available data, consensus sequence assembly can result in multiple possible sequence assemblies, and traditional paired - end sequencing data can be used to challenge which of these possible sequences is the correct consensus sequence. Because additional information can be extracted from the second region of mate - paired sequencing reads, consensus sequence verification is more robust using the methods described herein. To verify a consensus sequence, the first region or a portion thereof (or the third region or a portion thereof) can be mapped to the selected consensus sequence. The expected sequencing data for other regions or portions thereof (i.e., if the first region or a portion thereof is mapped, the third region or a portion thereof, or if the third region or a portion thereof is mapped, the first region or a portion thereof). For example, the expected sequencing data can be determined as described herein. In one instance, the second region flow order, the selected consensus sequence, and the first region flow order (if the expected sequencing data is for the first region or a portion thereof) or the third region flow order (if the expected sequencing data is for the third region or a portion thereof) are used to determine the expected sequencing data. The expected sequencing data can then be compared to the sequencing data of the mate - paired sequencing reads generated at the corresponding region to verify the consensus sequence portion. Expected sequencing data that matches the generated sequencing data indicates that the consensus sequence portion is correctly assembled. Expected sequencing data that does not match the generated sequencing data indicates that the consensus sequence portion is incorrectly assembled.
[0231] In some embodiments, more than one consensus sequence is constructed or verified. For example, certain organisms are polyploid (e.g., a healthy human is a diploid organism and each chromosome has two copies (except for the sex chromosomes in males). Consensus sequences can be assembled corresponding to one or more chromosome copies (e.g., consensus sequences can be assembled for each chromosome pair in a human sequence). The process of assigning mate - paired sequencing reads to the corresponding chromosomes of a polyploid organism can be referred to as haplotype analysis. The methods described herein can be used to improve the accuracy or efficiency of haplotype analysis. For example, information from the second region of mate - paired sequencing reads described herein can be used to associate a test variant with the first chromosome or the second chromosome (or other additional chromosomes from a polyploid organism).
[0232] Systems, devices, and reports
[0233] The operations described above (including those referenced Figures 1 - 17 are optionally implemented by Figure 18 the components depicted in Figure 18The components depicted therein can be used to implement other processes, such as combinations or sub - combinations of all or part of the above - mentioned operations. Those of ordinary skill in the art will also be clear on how the methods, techniques, systems, and devices described herein can be combined with each other, in whole or in part, and whether those methods, techniques, systems, and / or devices are implemented and / or provided by Figure 18 the depicted components.
[0234] Figure 18 FIG. illustrates an example of a computing device according to one embodiment. Device 1800 can be a network - connected host computer. Device 1800 can be a client computer or a server. As Figure 18 shown, device 1800 can be any suitable type of microprocessor - based device, such as a personal computer, a workstation, a server, or a handheld computing device (portable electronic device), such as a phone or a tablet. The device can include, for example, one or more of a processor 1810, an input device 1820, an output device 1830, a storage device 1840, and a communication device 1860. Input device 1820 and output device 1830 can generally correspond to those described above and can be connected to or integrated with the computer.
[0235] Input device 1820 can be any suitable device that provides input, such as a touch screen, a keyboard or keypad, a mouse, or a voice - recognition device. Output device 1830 can be any suitable device that provides output, such as a touch screen, a tactile device, or a speaker.
[0236] Memory 1840 can be any suitable device that provides storage, such as an electrical, magnetic, or optical memory, including RAM, cache, a hard - drive, or a removable storage disk. Communication device 1860 can include any suitable device capable of sending and receiving signals over a network, such as a network interface chip or device. The components of the computer can be connected in any suitable manner, such as via a physical bus or a wireless connection.
[0237] Software 1850 that can be stored in memory 1840 and executed by processor 1810 can include, for example, programming that embodies the functions of the present disclosure (e.g., as embodied in the devices described above).
[0238] Software 1850 can also be stored and / or transmitted within any non - transitory computer - readable storage medium for use by or in integration with an instruction - execution system, apparatus, or device, such as those described above, which can obtain instructions related to the software from the instruction - execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a computer - readable storage medium can be any medium, such as storage device 1840, that can contain or store a program for use by or in integration with an instruction - execution system, apparatus, or device.
[0239] The software 1850 can also be propagated within any transmission medium for use in or in integration with an instruction execution system, apparatus, or device, such as those described above, which can obtain instructions related to the software from the instruction execution system, apparatus, or device and execute the instructions. In the context of the present disclosure, a transmission medium can be any medium that can transmit, propagate, or convey a program for use in or in integration with an instruction execution system, apparatus, or device. A transmission readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared wired or wireless propagation media.
[0240] The device 1800 can be connected to a network, which can be any suitable type of interconnected communication system. The network can implement any suitable communication protocol and can be protected by any suitable security protocol. The network can include any suitable arrangement of network links that can effect the transmission and reception of network signals, such as wireless network connections, T1 or T3 lines, cable networks, DSL, or telephone lines.
[0241] The device 1800 can implement any operating system suitable for operation on a network. The software 1850 can be written in any suitable programming language, such as C, C++, Java, or Python. In various embodiments, the application software embodying the functions of the present disclosure can be deployed in different configurations, such as in a client / server arrangement or, for example, as a web-based application or web service via a web browser.
[0242] The methods described herein optionally further include reporting information determined using an analysis method and / or generating a report that includes information determined using an analysis method. For example, in some embodiments, the method further includes reporting or generating a report that includes information related to the identification of variants in a polynucleotide derived from a subject (e.g., within the genome of a subject). The information reported or within the report can be related to, for example, the locus of paired sequencing reads mapped to a reference sequence, detected variants (such as detected structural variants or detected SNPs), one or more assembled consensus sequences, and / or validation statistics for one or more assembled consensus sequences. The report can be distributed to a recipient, or the information can be reported to a recipient, such as a clinician, subject, or researcher.
[0243] Exemplary Embodiments
[0244] The following embodiments are exemplary and are not intended to limit the scope of the claimed invention.
[0245] Embodiment 1. A method for generating paired sequencing reads from a polynucleotide, comprising:
[0246] (a) hybridizing the polynucleotide with a primer to form a hybridization template;
[0247] (b) Generating sequencing data related to the sequence of the first region of a polynucleotide by extending a primer with labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides;
[0248] (c) Further extending the primer extended in step (b) through a second region using the nucleotides provided in the second region flow order, wherein (i) the primer is extended through the second region without detecting the presence or absence of the label of the nucleotides incorporated into the extended primer; (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order, or (iii) the extension of the primer through the second region proceeds faster than the extension of the primer in step (b); and
[0249] (d) Generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (c) with labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides.
[0250] Embodiment 2. The method according to Embodiment 1, wherein the extension of the primer through the second region proceeds faster than the extension of the primer through the first region.
[0251] Embodiment 3. The method according to Embodiment 1 or 2, further comprising correlating the sequencing data of the first region with the sequencing data of the third region.
[0252] Embodiment 4. A method for generating paired sequencing reads from a polynucleotide, comprising:
[0253] (a) Hybridizing a primer to a portion of the first region of a polynucleotide to form a hybridization template;
[0254] (b) Extending the primer through a second region using the nucleotides provided in the second region flow order, wherein (i) the primer is extended through the second region without detecting the presence or absence of the label of the nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; and
[0255] (c) Generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (b) with labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides.
[0256] Embodiment 5. The method according to Embodiment 4, wherein the first region contains a naturally occurring sequence targeted by the primer.
[0257] Embodiment 6. The method according to any one of Embodiments 1 - 5, wherein the primer is extended through the second region without detecting the presence or absence of the label of the nucleotides incorporated into the extended primer.
[0258] Embodiment 7. The method of any one of Embodiments 1-6, wherein at least a portion of the nucleotides used to extend the primer through the second region are unlabeled nucleotides.
[0259] Embodiment 8. The method of any one of Embodiments 1-6, wherein the nucleotides used to extend the primer through the second region are unlabeled nucleotides.
[0260] Embodiment 9. The method of any one of Embodiments 1-8, wherein a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order.
[0261] Embodiment 10. The method of any one of Embodiments 1-9, wherein the second region flow order comprises five or more nucleotide flows.
[0262] Embodiment 11. The method of any one of Embodiments 10, wherein each of the nucleotide flows comprises a single nucleobase.
[0263] Embodiment 12. The method of Embodiment 10 or 11, wherein for 50% or more of the possible SNP arrangements at 5% or more of the random sequencing start positions, the second region flow order induces signal changes at more than two flow positions.
[0264] Embodiment 13. The method of any one of Embodiments 10-12, wherein the second region flow order has an efficiency of integrating 0.6 or more bases per flow.
[0265] Embodiment 14. The method of any one of Embodiments 1-13, further comprising determining the expected sequencing data of the second region using a reference sequence and the second region flow order.
[0266] Embodiment 15. The method of any one of Embodiments 1-14, wherein the primer is extended through the third region using nucleotides provided in a third region flow order, and the method further comprises using the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region to determine the expected sequencing data of the third region.
[0267] Embodiment 16. The method of Embodiment 15, wherein the third region flow order comprises five or more nucleotide flows.
[0268] Embodiment 17. The method of Embodiment 16, wherein each of the nucleotide flows comprises a single nucleobase.
[0269] Embodiment 18. The method of embodiment 16 or 17, wherein for 50% or more of the possible SNP arrangements at 5% or more of the random sequencing start positions, the third region flow order induces signal changes at more than two flow positions.
[0270] Embodiment 19. The method of any one of embodiments 16 - 18, wherein the second region flow order has an efficiency of integrating 0.6 or more bases per flow.
[0271] Embodiment 20. The method of any one of embodiments 1 - 19, wherein the primer is extended through the third region using nucleotides provided in the third region flow order, and the method further includes using the reference sequence of the second region, the second region flow order, the third region flow order, and the sequencing data related to the sequence of the third region to determine the expected sequencing data of the third region, wherein the sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region.
[0272] Embodiment 21. The method of any one of embodiments 14 - 20, wherein the expected reference data of the second region or the third region includes a binary or non - binary flow map.
[0273] Embodiment 22. The method of any one of embodiments 14 - 21, further including using the second region flow order and the second reference sequence of the second region to determine the expected test variant sequencing data of the second region, wherein the second reference sequence contains the test variant.
[0274] Embodiment 23. The method of embodiment 22, wherein the primer is extended through the third region using nucleotides provided in the third region flow order, and the method further includes using the second reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region to determine the expected test variant sequencing data of the third region.
[0275] Embodiment 24. The method of embodiment 22, wherein the primer is extended through the third region using nucleotides provided in the third region flow order, and the method further includes using the second reference sequence of the second region, the second region flow order, the third region flow order, and the sequencing data related to the sequence of the third region to determine the expected test variant sequencing data of the third region, wherein the sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region.
[0276] Embodiment 25. The method of any one of embodiments 22 - 24, wherein the expected reference sequencing data of the second region or the third region includes a binary or non - binary flow map.
[0277] Embodiment 26. A method of mapping paired - end sequencing reads to a reference sequence, comprising:
[0278] Mapping the first region or a portion thereof, or the third region or a portion thereof, of the paired sequencing reads generated by the method according to any one of Embodiments 1-25 to a reference sequence; and
[0279] Mapping the unmapped first region or a portion thereof, or the unmapped third region or a portion thereof, to the reference sequence using distance information indicating the length of the second region.
[0280] Embodiment 27. A method for detecting a structural variant, comprising:
[0281] Mapping the first region or a portion thereof or the third region or a portion thereof of the paired sequencing reads generated by the method according to any one of Embodiments 1-25 to a reference sequence;
[0282] Using distance information indicating the length of the second region, determining an expected locus within the reference sequence for the unmapped first region or a portion thereof or the unmapped third region or a portion thereof;
[0283] Determining expected sequencing data for the sequence at the expected locus based on the reference sequence; and
[0284] Detecting a structural variant by comparing the sequencing data of the unmapped first region or a portion thereof or the unmapped third region or a portion thereof with the expected sequencing data, wherein a difference between the sequencing data of the unmapped first region or a portion thereof or the unmapped third region or a portion thereof and the expected sequencing data indicates a structural variant.
[0285] Embodiment 28. A method for detecting a structural variant, comprising:
[0286] Mapping the first region or a portion thereof or the third region or a portion thereof of the paired sequencing reads generated by the method according to any one of Embodiments 1-25 to a reference sequence, wherein the unmapped first region or the unmapped third region is unmappable within the reference sequence.
[0287] Embodiment 29. The method of Embodiment 28, further comprising determining the locus of the structural variant within the reference sequence based on expected distance information indicating the length of the second region.
[0288] Embodiment 30. The method according to any one of Embodiments 27-29, wherein the unmapped first region or a portion thereof or the unmapped third region or a portion thereof is within an insertion relative to the reference sequence.
[0289] Embodiment 31. The method according to any one of Embodiments 27-29, wherein the unmapped first region or a portion thereof or the unmapped third region or a portion thereof bridges to the start or end of an insertion relative to the reference sequence.
[0290] Embodiment 32. A method for detecting a structural variant, comprising:
[0291] Mapping a first region or a portion thereof and a third region or a portion thereof of a paired sequencing read generated by the method according to any one of Embodiments 1-25 to a reference sequence;
[0292] Determining mapping distance information between the mapped first region and the mapped third region; and
[0293] Detecting a structural variant by comparing the mapped distance information with the expected distance information of a second region, wherein a difference between the mapped distance information and the expected distance information indicates a structural variant.
[0294] Embodiment 33. The method according to any one of Embodiments 27-32, wherein the structural variant is a chromosomal fusion, inversion, insertion or deletion.
[0295] Embodiment 34. The method according to any one of Embodiments 27-32, wherein the variant is an insertion or deletion within the second region.
[0296] Embodiment 35. The method according to any one of Embodiments 26-32, wherein information related to the flow order of the second region and the probability distribution of bases in the second region are used to determine the distance information.
[0297] Embodiment 36. The method of Embodiment 35, wherein the information related to the flow order of the second region is the number of different types of nucleotide bases that are simultaneously used to extend a primer through the second region.
[0298] Embodiment 37. The method according to Embodiment 35 or 36, wherein the probability distribution of bases in the second region is determined by the base distribution within the genome.
[0299] Embodiment 38. The method according to any one of Embodiments 26-35, wherein the distance information is derived from expected sequencing data of the second region determined using the reference sequence and the flow order of the second region.
[0300] Embodiment 39. The method of Embodiment 38, wherein the expected sequencing data includes a binary or non-binary flow map.
[0301] Embodiment 40. A method for mapping a paired sequencing read to a reference sequence, comprising:
[0302] Mapping a first region or a portion thereof and a third region or a portion thereof of a paired sequencing read generated by the method according to any one of Embodiments 1-25 to the reference sequence at two or more different position pairs including a first position and a second position; and
[0303] For the two or more position pairs, the correct position pair is selected using first distance information indicating the length of the second region and second distance information indicating the distance between the first position and the second position.
[0304] Embodiment 41. The method of embodiment 40, wherein the first distance information is determined using information related to the flow order of the second region and the probability distribution of bases in the second region.
[0305] Embodiment 42. The method of embodiment 41, wherein the information related to the flow order of the second region is the number of different types of nucleotide bases that are simultaneously used to extend the primer through the second region.
[0306] Embodiment 43. The method of embodiment 41 or 42, wherein the base probability distribution in the second region is determined by the base distribution within the genome.
[0307] Embodiment 44. The method of embodiment 40, wherein the first distance information is derived from the expected sequencing data of the second region determined using a reference sequence and the flow order of the second region.
[0308] Embodiment 45. The method of embodiment 44, wherein the expected reference sequencing data includes a binary or non-binary flow map.
[0309] Embodiment 46. A method for detecting a variant between two sequencing regions of a coupled sequencing read pair generated according to any one of embodiments 1-25, wherein the primer to be extended is extended through a third region using nucleotides provided in the third region flow order, the method comprising:
[0310] Mapping the first region or a portion thereof to a reference sequence;
[0311] Determining the expected sequencing data of the third region or a portion thereof using (1) the reference sequence of the second region, the flow order of the second region, the flow order of the third region, and the reference sequence of the third region, or (2) the reference sequence of the second region, the flow order of the second region, the flow order of the third region, and the generated sequencing data related to the sequence of the third region, wherein the generated sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region; and
[0312] Detecting the presence of a variant by comparing the expected sequencing data of the third region with the generated sequencing data related to the sequence of the third region.
[0313] Embodiment 47. The method of embodiment 46, wherein the variant is a structural variant.
[0314] Embodiment 48. The method of embodiment 47, wherein the structural variant is a chromosomal fusion, inversion, insertion, or deletion.
[0315] Embodiment 49. The method of Embodiment 46, wherein the variant is a single nucleotide polymorphism (SNP).
[0316] Embodiment 50. The method of any one of Embodiments 46-49, wherein the method is used to detect a test variant and the reference sequence contains the test variant.
[0317] Embodiment 51. The method of Embodiment 50, wherein the test variant is selected by identifying the test variant within a second polynucleotide.
[0318] Embodiment 52. The method of Embodiment 50 or 51, including correlating the detected test variant with an allele sequenced in the first or third region of the polynucleotide.
[0319] Embodiment 53. A method of generating a paired sequencing read for detecting the presence of a base transversion in an unsequenced region of a polynucleotide, comprising:
[0320] (a) hybridizing a polynucleotide with a primer to form a hybridization template;
[0321] (b) generating sequencing data related to the sequence of the first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides;
[0322] (c) further extending the primer extended in step (b) through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and
[0323] (d) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides.
[0324] Embodiment 54. A method of generating paired sequencing reads from a polynucleotide, comprising:
[0325] (a) hybridizing a primer with a first region of a polynucleotide to form a hybridization template;
[0326] (b) extending the primer through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and
[0327] (c) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (b) using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides.
[0328] Embodiment 55. The method of embodiment 54, wherein the first region comprises a naturally occurring sequence targeted by a primer.
[0329] Embodiment 56. The method of embodiment 54 or 55, wherein primer extension is through the second region without detecting the presence or absence of a label of a nucleotide incorporated into the extended primer.
[0330] Embodiment 57. A method for detecting the presence of a base transversion in an unsequenced region of a polynucleotide, comprising:
[0331] mapping a first region or a portion thereof and a third region or a portion thereof of a pair of coupled sequencing reads generated according to any one of embodiments 54-56 to a reference sequence, wherein primer extension is through the third region using nucleotides provided in the third region flow order;
[0332] determining the expected sequencing data for the third region using the second region flow order, the third region flow order, and the reference sequence; and
[0333] detecting the presence of a base transversion based on a difference between the expected sequencing data for the third region and the generated sequencing data for the third region.
[0334] Embodiment 58. The method of embodiment 57, wherein the expected sequencing data for the third region is determined using the second region flow order, the third region flow order, the reference sequence for the second region, and the reference sequence for the third region.
[0335] Embodiment 59. The method of embodiment 57, wherein the expected sequencing data for the third region is determined using the second region flow order, the third region flow order, the reference sequence for the second region, and the generated sequence data related to the sequence of the third region, wherein the generated sequence data related to the sequence of the third region is the same or different sequence data generated for the third region.
[0336] Embodiment 60. The method of any one of embodiments 57-59, wherein the expected sequencing data for the third region comprises a binary or non-binary flowgram.
[0337] Embodiment 61. A method for generating one or more consensus sequences, comprising assembling a plurality of pairs of coupled sequencing reads generated according to any one of embodiments 1-25.
[0338] Embodiment 62. The method of embodiment 61, wherein distance information indicating the length of the second region of the plurality of pairs of coupled sequencing reads is used to assemble one or more consensus sequences.
[0339] Embodiment 63. The method of embodiment 61, wherein the distance information is determined using information related to the second region flow order and the base probability distribution in the second region.
[0340] Embodiment 64. The method of embodiment 63, wherein the information related to the second region flow order is the number of different types of nucleotide bases that are simultaneously used to extend the primer through the third region.
[0341] Embodiment 65. The method of embodiment 63 or 64, wherein the probability distribution of bases in the second region is determined by the base distribution within the genome.
[0342] Embodiment 66. The method of embodiment 62, wherein the distance information is derived from the expected reference sequencing data of the second region determined using the reference sequence and the second region flow order.
[0343] Embodiment 67. The method of embodiment 66, wherein the expected reference sequencing data includes a binary or non-binary flow map.
[0344] Embodiment 68. The method of any one of embodiments 61-67, further comprising validating a portion of a consensus sequence selected from one or more consensus sequences using a selected coupled sequencing read related to the portion of the selected consensus sequence, wherein the primer that extends through the third region when generating the selected coupled sequencing read is extended using nucleotides provided in the third region flow order, and the validation includes:
[0345] Determining the expected sequencing data of the third region of the selected coupled sequencing read using the second region flow order, the third region flow order, and the portion of the selected consensus sequence; and
[0346] Validating the portion of the selected consensus sequence by comparing the expected sequencing data of the third region of the selected coupled sequencing read with the generated sequencing data of the third region.
[0347] Embodiment 69. A method for validating a test variant status, comprising:
[0348] Comparing the status of a variant on a plurality of overlapping coupled sequencing read pairs generated according to any one of embodiments 1-25, the plurality of overlapping coupled sequencing read pairs including loci corresponding to the locus of the test variant;
[0349] Validating the status of the variant based on the comparison.
[0350] Embodiment 70. The method of embodiment 69, wherein the first region or the third region of the selected coupled sequencing read overlaps with at least a portion of the second region of other coupled sequencing reads among the plurality of overlapping coupled sequencing reads.
[0351] Embodiment 71. The method of embodiment 69 or 70, wherein the variant status of the selected coupled sequencing read indicates a variant in the first region or the third region of the selected coupled sequencing read.
[0352] Embodiment 72. The method of embodiment 71, wherein the second region of the selected coupled sequencing read overlaps with at least a portion of the second regions of other coupled sequencing reads among the plurality of overlapping coupled sequencing reads.
[0353] Embodiment 73. The method according to embodiment 71 or 72, wherein the variant status of the selected coupled sequencing read indicates a variant in the second region of the selected coupled sequencing read.
[0354] Embodiment 74. A method for detecting short genetic variants in a test sample, comprising:
[0355] generating a pair of coupled sequencing reads according to any one of embodiments 1-25;
[0356] comparing sequencing data related to the sequence of the third region of the polynucleotide with expected sequencing data of the expected sequence of the third region of the polynucleotide; and
[0357] determining the presence or absence of a short genetic variant in the second region of the polynucleotide.
[0358] Embodiment 75. The method of embodiment 74, wherein:
[0359] comparing sequencing data related to the sequence of the third region of the polynucleotide with expected sequencing data of the expected sequence of the third region of the polynucleotide, which includes determining a match score indicating the likelihood that the sequencing data generated for the third region of the polynucleotide matches the expected sequencing data of the third region of the polynucleotide; and
[0360] determining the presence or absence of a short genetic variant in the second region of the polynucleotide, including using the determined match score.
[0361] Embodiment 76. The method of embodiment 74 or 75, wherein the expected sequencing data of the third region of the polynucleotide is obtained by sequencing and the expected sequence of the third region of the polynucleotide on a computer.
[0362] Embodiment 77. The method according to any one of embodiments 1-76, wherein the sequencing data related to the sequence of the first region or the sequencing data related to the sequence of the third region contains a flow signal representing a base count, and the base count indicates the number of bases integrated at each flow position within a plurality of flow positions.
[0363] Embodiment 78. The method of embodiment 77, wherein the flow signal contains a statistical parameter indicating the likelihood of a base count at each flow position.
[0364] Embodiment 79. The method of embodiment 78, wherein the flow signal comprises a statistical parameter of base count probabilities indicative of a plurality of base counts at each flow position.
[0365] Embodiment 80. The method of embodiment 75 or 76, wherein:
[0366] Sequencing data related to the sequence of the third region includes a flow signal representing a base count indicative of the number of bases incorporated at each flow position within a plurality of flow positions, wherein the flow signal comprises a statistical parameter of base count probabilities indicative of the plurality of base counts; and
[0367] The method further comprises selecting, at each flow position in the sequencing data, the statistical parameter corresponding to the base count of the expected sequence at that flow position and determining a match score indicative of the likelihood that the sequencing data set matches the expected sequence.
[0368] Embodiment 81. The method of embodiment 80, wherein the match score is a combined value of the selected statistical parameters across flow positions in the sequencing data.
[0369] Embodiment 82. The method of any one of embodiments 1-81, wherein the flow cycle order comprises 4 separate flows repeated in the same order.
[0370] Embodiment 83. The method of any one of embodiments 1-81, wherein the flow cycle order comprises 5 or more separate flows.
[0371] Embodiment 84. The method of any one of embodiments 1-83, wherein generating the paired sequencing reads further comprises:
[0372] using nucleotides provided in a fourth region flow order to further extend the primer through the fourth region, wherein (i) the primer is extended through the fourth region without detecting the presence or absence of a label of the nucleotide incorporated into the extended primer, (ii) a mixture of at least two different types of nucleotide bases is used in at least one step of the fourth region flow order, or (iii) the primer is extended through the fourth region more rapidly than the primer is extended through the first region or the third region; and
[0373] generating sequencing data related to the sequence of the fifth region of the polynucleotide by further extending the primer extended through the fourth region with a labeled nucleotide and detecting the presence or absence of the incorporated labeled nucleotide.
[0374] Embodiment 85. The method of embodiment 84, further comprising correlating the sequencing data of the fifth region with the sequencing data of the first region or the sequencing data of the third region.
[0375] Embodiment 86. The method of any one of Embodiments 1-85, wherein rolling circle amplification is used to amplify the polynucleotide.
[0376] Embodiment 87. A method for detecting short genetic variants in a test sample, comprising:
[0377] (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide comprising at least a first copy of the polynucleotide and a second copy of the polynucleotide;
[0378] (b) hybridizing the RCA-amplified polynucleotide with a primer to form a hybridization template;
[0379] (c) generating sequencing data related to the sequence of a first region of the polynucleotide within the first copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides;
[0380] (d) further extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order, or (iii) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide faster than the primer is extended through the first region;
[0381] (e) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides;
[0382] (f) comparing the sequencing data generated for the third region of the polynucleotide with the sequencing data expected for the expected sequence of the third region of the polynucleotide;
[0383] (g) determining the presence of a short genetic variant in the second region of the polynucleotide;
[0384] (h) generating sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; and
[0385] (i) determining the identity of the short genetic variant in the second region of the polynucleotide.
[0386] Embodiment 88. The method of embodiment 87, wherein primer extension proceeds faster through a second region of a polynucleotide within a first copy of the polynucleotide than through a first region of the polynucleotide within the first copy of the polynucleotide.
[0387] Embodiment 89. A method for detecting a short genetic variant in a test sample, comprising:
[0388] (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide comprising at least a first copy of the polynucleotide and a second copy of the polynucleotide;
[0389] (b) hybridizing a primer to a first region of a polynucleotide within a first copy of the polynucleotide to form a hybridization template;
[0390] (c) extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of a label of the nucleotide incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order;
[0391] (d) further extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides to generate sequencing data related to the sequence of a third region of the polynucleotide;
[0392] (e) comparing the sequencing data generated for the third region of the polynucleotide with the sequencing data expected for the expected sequence of the third region of the polynucleotide;
[0393] (f) determining the presence of a short genetic variant in the second region of the polynucleotide;
[0394] (g) extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides to generate sequencing data related to the sequence of a second region of the polynucleotide within a second copy of the polynucleotide; and
[0395] (h) determining the identity of the short genetic variant in the second region of the polynucleotide.
[0396] Embodiment 90. The method of embodiment 89, wherein the first region comprises a naturally occurring sequence targeted by the primer.
[0397] Embodiment 91. The method of any one of embodiments 87-90, wherein based on the determination of the presence of a short genetic variant in the second region of the polynucleotide, sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide is dynamically generated.
[0398] Embodiment 92. The method of any one of embodiments 87-91, wherein the primer is extended through a second region of a polynucleotide within a first copy of the polynucleotide without detecting the presence or absence of a label on the nucleotide incorporated into the extended primer.
[0399] Embodiment 93. The method of any one of embodiments 87-92, wherein at least a portion of the nucleotides used to extend the primer through a second region of a polynucleotide within a first copy of the polynucleotide are unlabeled nucleotides.
[0400] Embodiment 94. The method of any one of embodiments 87-92, wherein the nucleotides used to extend the primer through a second region of a polynucleotide within a first copy of the polynucleotide are unlabeled nucleotides.
[0401] Embodiment 95. The method of any one of embodiments 87-94, wherein a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order.
[0402] Embodiment 96. The method of any one of embodiments 87-95, wherein a mixture of three different types of nucleobases is used in at least one step of the second region flow order.
[0403] Embodiment 97. A method for synchronizing sequencing primers within a sequencing cluster, comprising:
[0404] (a) hybridizing the primer to a polynucleotide copy within the sequencing cluster;
[0405] (b) extending the primer through a first region of the polynucleotide copy using labeled nucleotides according to a first region flow cycle;
[0406] (c) extending the primer through a second region of the polynucleotide copy using one or more rephasing flows, wherein a mixture of at least two different types of nucleobases is used in at least one of the one or more rephasing flows; and
[0407] (d) extending the primer through a third region of the polynucleotide copy using labeled nucleotides according to a third region flow cycle.
[0408] Embodiment 98. The method of embodiment 97, wherein a mixture of three different types of nucleobases is used in at least one of the one or more rephasing flows.
[0409] Embodiment 99. The method of embodiment 97 or 98, wherein the one or more rephasing flows comprise four or more flow steps.
[0410] Embodiment 100. The method of Embodiment 99, wherein the one or more rephasing flows include, in any order:
[0411] (i) a first flow comprising a mixture that includes nucleotides A, C, and G and omits nucleotide T;
[0412] (ii) a second flow comprising a mixture that includes nucleotides T, C, and G and omits nucleotide A;
[0413] (iii) a third flow comprising a mixture that includes nucleotides T, A, and G and omits nucleotide C; and
[0414] (iv) a fourth flow comprising a mixture that includes nucleotides T, A, and C and omits nucleotide G.
[0415] Embodiment 101. The method of any one of Embodiments 97 - 100, comprising generating sequencing data related to the sequence of a first region by detecting the presence or absence of an incorporated labeled nucleotide while extending a primer through the first region.
[0416] Embodiment 102. The method of any one of Embodiments 97 - 101, comprising generating sequencing data related to the sequence of a third region by detecting the presence or absence of an incorporated labeled nucleotide while extending a primer through the third region.
[0417] Embodiment 103. A system comprising:
[0418] one or more processors; and
[0419] a non - transitory storage medium that includes one or more programs, the one or more programs being executable by the one or more processors to:
[0420] receive information related to one or more coupled sequencing reads; and
[0421] perform the method of any one of Embodiments 26 - 52 and 57 - 86.
[0422] Embodiment 104. The system of Embodiment 103, wherein one or more coupled sequencing reads are generated according to the method of any one of Embodiments 1 - 25, 53 - 56, and 87 - 96. Examples
[0423] The present application can be better understood by referring to the following non-limiting examples provided as exemplary embodiments of the present application. The following examples are provided to more fully illustrate the embodiments, but should in no way be construed as limiting the broad scope of the present application. Although certain embodiments of the present application have been shown and described herein, it is obvious that these embodiments are provided by way of example only. Many variations, changes, and substitutions can be envisioned by those skilled in the art without departing from the spirit and scope of the invention. It should be understood that various alternatives to the embodiments described herein can be used to practice the methods described herein.
[0424] Example 1
[0425] A nucleic acid construct of 262 bases was sequenced using a flow sequencing method that included a fast-forward region and again using a standard flow sequencing method (i.e., one that did not include a fast-forward region). The polynucleotide was ligated with adapter sequences and tethered to beads, which were amplified and associated with the sequencing surface. A sequencing primer was hybridized to a hybridization region within the adapter sequence, which allowed initiation of the flow sequencing method. In the first method, 62 bases were sequenced by alternating flow of a single type of fluorescently labeled non-terminating nucleotide to extend the sequencing primer, and nucleotide incorporation after each step was determined using a fluorescence detector. The next 177 bases were exposed to alternating flow of unlabeled non-terminating nucleotides, where three of the four nucleotides were present in each flow (i.e., the "fast-forward" mode) to allow primer extension through the second region. After the primer extended through the "dark" (i.e., nucleotides with undetected incorporation) second region, an additional 23 bases were sequenced, with alternating flow of a single type of fluorescently labeled non-terminating nucleotide, and nucleotide incorporation after each step was determined using a fluorescence detector. The results are shown in Figure 19A which shows the number of flow steps on the horizontal channel and the measurement of the sequencing signal (i.e., normalized fluorescence signal) in the vertical channel. This method produced high-quality sequencing data that followed the fast-forward mechanism.
[0426] The same 262-base construct was fully sequenced using the standard flow sequencing method without an intermediate fast-forward protocol. That is, alternating flow of a single type of fluorescently labeled non-terminating nucleotide was used to sequence all 262 bases, and nucleotide incorporation after each step was determined using a fluorescence detector. The results are shown in Figure 19B which omits data from the corresponding 177-base region to compress the figure.
[0427] Using the fast-forward flow sequencing method advanced the sequencing construct faster than using the standard flow sequencing method. Sequencing data from both ends of the polynucleotide can be correlated to generate paired sequencing reads and analyzed.
[0428] Example 2
[0429] The detection of variants within SEQ ID NO:4 is described in this example (having a C→G single nucleotide polymorphism variant at base position 15 relative to the reference sequence SEQ ID NO:1). By hybridizing a primer to the hybridization sequence at the 5'-end of SEQ ID NO:4 and extending the primer using a flow sequencing method, paired sequencing reads of SEQ ID NO:4 can be generated. In this example, 5 cycles are used, where cycle 1 is used to extend the primer through the first region, cycles 2 and 3 are used to extend the primer through the second region, and cycles 4 and 5 are used to extend the primer through the third region. Cycles 1, 4, and 5 use labeled nucleotides to extend the primer, and the incorporation of nucleotides into the primer is detected after each cycle step. In contrast, during cycles 2 and 3, the incorporation of nucleotides into the primer can be skipped. Each cycle has 4 steps, where cycles 1, 4, and 5 include the sequential and independent addition of A-C-T-G labeled nucleotides, where a single base type is added at each cycle step, and the incorporation of the labeled nucleotides is detected after each step. Cycles 2 and 3 are implemented in a "fast forward" mode and include 4 cycle steps, where step 1 omits the A nucleotide (i.e., includes C, T, and G), step 2 omits the C nucleotide (i.e., includes A, T, and G), step 3 omits the T nucleotide (i.e., includes A, C, and G), and step 4 omits the G nucleotide (i.e., includes A, C, and T). Nucleotide incorporation is not detected during the fast forward mode of cycles 2 and 3. Since cycles 2 and 3 include multiple different nucleotide base types simultaneously during primer extension, the primer extends faster compared to using only a single base type at any given step. The flow diagrams of SEQ ID NO:1 (reference sequence) and SEQ ID NO:4 (SNP sequence) are shown in Table 6. The sequencing data indicates that the third region (cycles 4 and 5) of SEQ ID NO:1 is 3'-CTGAC-5' (SEQ ID NO:5), and the third region (cycles 4 and 5) of SEQ ID NO:4 is 3'-CCTGC-5' (SEQ ID NO:7). The difference between the sequencing data between SEQ ID NO:1 and SEQ ID NO:4 indicates the presence of a variant within the second region.
[0430]
[0431]
[0432]
[0433]
[0434] Example 3
[0435] The detection of the variant within SEQ ID NO:8 is described in this example (which includes an ATC insertion after base position 23 relative to the reference sequence SEQ ID NO:1). A coupled sequencing read pair can be generated for SEQ ID NO:1 and SEQ ID NO:8 using a flow sequencing method that includes a fast-forward portion through the second region. In this example, 5 cycles are used, where cycle 1 is for extending the primer through the first region, cycles 2 and 3 are for extending the primer through the second region, and cycles 4 and 5 are for extending the primer through the third region. Cycles 1, 4, and 5 use labeled nucleotides to extend the primer and detect the incorporation of the nucleotides into the primer after each cycle step. In contrast, during cycles 2 and 3, the incorporation of nucleotides into the primer can be skipped. Each cycle has 4 steps, where cycles 1, 4, and 5 include the sequential and independent addition of A-C-T-G labeled nucleotides, where a single base type is added at each cycle step and the incorporation of the labeled nucleotides is detected after each step. Cycles 2 and 3 are implemented in a "fast-forward" mode and include 4 cycle steps, where step 1 omits the A nucleotide (i.e., includes C, T, and G), step 2 omits the C nucleotide (i.e., includes A, T, and G), step 3 omits the T nucleotide (i.e., includes A, C, and G), and step 4 omits the G nucleotide (i.e., includes A, C, and T). Nucleotide incorporation is not detected during the fast-forward mode of cycles 2 and 3. Since cycles 2 and 3 include multiple different nucleotide base types simultaneously during primer extension, the primer extends faster compared to using only a single base type at any given step. The flowgrams of SEQ ID NO:1 (reference sequence) and SEQ ID NO:8 are shown in Table 7. The sequencing data indicates that the third region (cycles 4 and 5) of SEQ ID NO:1 is 3'-CTGAC-5' (SEQ ID NO:5), and the third region (cycles 4 and 5) of SEQ ID NO:8 is 3'-AC-5'. The difference between the sequencing data between SEQ ID NO:1 and SEQ ID NO:8 indicates the presence of a variant within the second region.
[0436] Example 4
[0437] The detection of a variant within SEQ ID NO:9 (which includes a deletion of the GCCTGCA (SEQ ID NO:13) bases after base position 17 relative to the reference sequence SEQ ID NO:1) is described in this example. Coupled sequencing read pairs can be generated for SEQ ID NO:1 and SEQ ID NO:9 using a flow sequencing method that includes a fast-forward portion through a second region. In this example, 5 cycles are used, where cycle 1 is for primer extension through the first region, cycles 2 and 3 are for primer extension through the second region, and cycles 4 and 5 are for primer extension through the third region. Cycles 1, 4, and 5 use labeled nucleotides to extend the primer, and the incorporation of nucleotides into the primer is detected after each cycle step. In contrast, during cycles 2 and 3, the incorporation of nucleotides into the primer can be skipped. Each cycle has 4 steps, where cycles 1, 4, and 5 include the sequential and independent addition of A-C-T-G labeled nucleotides, where a single base type is added at each cycle step, and the incorporation of the labeled nucleotides is detected after each step. Cycles 2 and 3 are implemented in a "fast-forward" mode and include 4 cycle steps, where step 1 omits the A nucleotide (i.e., includes C, T, and G), step 2 omits the C nucleotide (i.e., includes A, T, and G), step 3 omits the T nucleotide (i.e., includes A, C, and G), and step 4 omits the G nucleotide (i.e., includes A, C, and T). Nucleotide incorporation is not detected during the fast-forward mode of cycles 2 and 3. Because cycles 2 and 3 include multiple different nucleotide base types simultaneously during primer extension, the primer extends faster compared to using only a single base type at any given step. The flowgrams for SEQ ID NO:1 (reference sequence) and SEQ ID NO:9 are shown in Table 8. The sequencing data indicates that the third region (cycles 4 and 5) of SEQ ID NO:1 is 3'-CTGAC-5' (SEQ ID NO:5), and the third region (cycles 4 and 5) of SEQ ID NO:9 is 3'-AC-5'. The difference in the sequencing data between SEQ ID NO:1 and SEQ ID NO:8 indicates the presence of a variant within the second region.
[0438] Example 5
[0439] The detection of a variant within SEQ ID NO:12 is described in this example (which includes an inversion of the base GCCTGCA (SEQ ID NO:13) after base position 17 relative to the reference sequence SEQ ID NO:1). Coupled sequencing read pairs can be generated for SEQ ID NO:1 and SEQ ID NO:12 using a flow sequencing method that includes a fast-forward portion through the second region. In this example, 5 cycles are used, where cycle 1 is for primer extension through the first region, cycles 2 and 3 are for primer extension through the second region, and cycles 4 and 5 are for primer extension through the third region. Cycles 1, 4, and 5 use labeled nucleotides to extend the primer and detect the incorporation of nucleotides into the primer after each cycle step. In contrast, during cycles 2 and 3, the incorporation of nucleotides into the primer can be skipped. Each cycle has 4 steps, where cycles 1, 4, and 5 include the sequential and independent addition of A-C-T-G labeled nucleotides, where a single base type is added at each cycle step and the incorporation of the labeled nucleotides is detected after each step. Cycles 2 and 3 are implemented in a "fast-forward" mode and include 4 cycle steps, where step 1 omits the A nucleotide (i.e., includes C, T, and G), step 2 omits the C nucleotide (i.e., includes A, T, and G), step 3 omits the T nucleotide (i.e., includes A, C, and G), and step 4 omits the G nucleotide (i.e., includes A, C, and T). Nucleotide incorporation is not detected during the fast-forward mode of cycles 2 and 3. Because cycles 2 and 3 include multiple different nucleobase types simultaneously during primer extension, the primer extends faster compared to using only a single base type at any given step. The flowgrams for SEQ ID NO:1 (reference sequence) and SEQ ID NO:12 are shown in Table 9. The sequencing data indicates that the third region (cycles 4 and 5) of SEQ ID NO:1 is 3'-CTGAC-5' (SEQ ID NO:5), and the third region (cycles 4 and 5) of SEQ ID NO:12 is 3'-G-5'. The difference between the sequencing data for SEQ ID NO:1 and SEQ ID NO:12 indicates the presence of a variant within the second region.
[0440] Example 6
[0441] In the synthesis - while - sequencing method, nucleotides are usually not fully incorporated into the extended primer. Over time, within a sequencing cluster, the primers may become out of sync, causing signal attenuation and a decrease in the confidence for making base incorporation determinations. By assuming a sequencing cluster with 10,000 identical template strands and sequencing the template strands with non - terminating nucleotides assuming an A - C - T - G flow order, where each flow has a single nucleotide. The probability of incorporation failure (i.e., when a template - indicated nucleotide should have been incorporated but was not incorporated into the extended primer strand) is set to 0.5%. Figure 20A Shows the number of primers (strands) extended at each read base after 100 flow steps, where the 100th flow has a G non - terminating nucleotide. The sequencing cluster includes templates hybridized to a leading sequencing primer, where a G nucleotide is incorporated into the extended primer such that the next expected incorporated nucleotide is A; templates hybridized to a first lagging primer, where a G nucleotide is incorporated into the extended primer such that the next expected incorporated nucleotide is C; and templates hybridized to a second lagging primer, where no nucleotide from the 100th flow is incorporated into the extended primer. The first lagging primer and the second lagging primer represent primers that failed to incorporate the expected nucleotide into the extended primer at some point during the sequencing process.
[0442] Simulating the synchronization of extended primers using a re - phasing flow order with a synchronized flow order. At flow 101, primers are extended ( Figure 20B ) using a mixture of G, C, and A non - terminating nucleotides, which extends the first and second lagging primers until they are synchronized with the leading primer. Since flow 101 does not contain a T nucleotide, it does not extend further. The simulated synchronized flow order continues with flow 102, which has a mixture of G, C, and T non - terminating nucleotides ( Figure 20C ), flow 103, which has a mixture of G, T, and A non - terminal nucleotides ( Figure 20D ); flow 104, which has a mixture of T, A, and C non - terminating nucleotides ( Figure 20E ).
[0443] As shown in Figure 21A - 21E and Figure 22A - 22E , the simulated synchronized flow order is tested using additional sequences. Other successful simulations are performed using the synchronized flow order and different template sequences.
[0444] Example 7
[0445] More than one million extended sequencing flow orders were tested on a computer to determine their likelihood of inducing signal changes at more than two flow positions across the set of all possible SNPs (XYZ → XQZ, where Q ≠ Y (and Q, X, Y, and Z are each any one of A, C, G, and T)). The extended flow orders were designed to have a minimum 12-base sequence, have all valid 2-base flow arrangements, and remove flow orders with sequential base repeats. All possible starting positions for the flow orders were tested to evaluate the sensitivity of the extended flow orders to induce signal changes at more than two flow positions. Figure 23 And Table 4 shows exemplary results of this analysis. In Figure 23 , the x-axis indicates the fraction of the flow phase (or fragmentation start position), and the y-axis indicates the fraction of SNP arrangements that induce signal changes at more than two flow positions. For approximately 10% of the reads (or flow start positions), several flow orders induce two or more signal differences at all possible (87.5%) SNP arrangements. The four-base periodic flow induces a cyclic shift in only 42% of the possible SNPs, but it does so for all reads or flow phases. A final assessment of efficiency was performed on a subset of one million reads against the human reference genome to establish feasibility. This is a practical measurement of how effectively the flow order extends the sequence given the patterns and biases in actual tissue.
[0446] Example 8
[0447] To test the sensitivity of fast-forward sequencing to detect SNPs, a sequencing method was simulated by computer to sequence approximately 11,400 synthetic nucleic acid molecules within the hg 38 reference genome, where each synthetic nucleic acid molecule is a 2-kilobase fragment with a random starting point within the reference genome. A 502-bp segment was generated from each synthetic sequencing read, and all three possible single-base mutations queried at each base within the ∼502-bp segment (i.e., a total of 500 × ∼1.14M × 3 possible variants (i.e., ABC → ADC, where B ≠ D)) were queried for SNP detection. For each SNP variant ABC → ADC, the SNP was considered undetectable when (A = B and D = C) or (A = D and B = C) because neither SNP would generate a new zero or new non-zero signal in the flowgram. Figure 24 A matrix of variant base pair reference base detection sensitivity is shown in
[0448] Then computer sequencing of the synthetic nucleic acid molecule is performed using a four-step flow cycle, where each flow includes a mixture of three nucleotides in the middle (second) region. The first region of the synthetic nucleic acid molecule is sequenced using an 80-nucleotide flow according to the four-step flow cycle, where each step includes a single nucleotide base type. The sequencing primer extends 54 ± 7 bases (~0.675 bases per flow) in the 80 flows in the first region. The second region of the synthetic nucleic acid molecule is sequenced using 200 nucleotides according to the four-step flow cycle, where each step includes a mixture of three nucleotide base types and omits one nucleotide base type (i.e., (i) A, C, T, and no G; (ii) G, A, C, and no T; (iii) T, G, A, and no C; and (iv) C, T, G, and no A). The sequencing primer extends 915 ± 89 bases (~4.575 bases per flow) in the 200 flows in the second region. The third region of the synthetic nucleic acid molecule is sequenced using an 80-nucleotide flow according to the four-step flow cycle, where each step includes a single nucleotide base type. The sequencing primer extends 54 ± 7 bases (~0.675 bases per flow) in the 80 flows in the third region. The flow map of the third (downstream) region of each synthetic variant nucleic acid molecule is compared to the flow map of the third region of the corresponding synthetic wild-type nucleic acid molecule. New non-zero flow map entries and / or new zero flow map entries in the third region of the synthetic variant nucleic acid molecule, compared to the corresponding synthetic wild-type nucleic acid molecule, indicate detection of an SNP introduced into the second region. Figure 25A Shows the average base incorporation across flows in the first, second, and third regions. Figure 25B Shows a matrix of variant base pair reference base detection sensitivity. Figure 25C Shows the distribution of base coverage in synthetic reads.
[0449] Example 9
[0450] The rephasing effect of rephasing flow steps using mixtures of two or three different nucleotide bases was studied using a simulated sequencing methodology.
[0451] Approximately 10,000 synthetic sequencing reads, each 600 bp in length, were generated by selecting random starting sites from the human genome. In the control group, simulated flow maps were generated by computer sequencing the synthetic sequencing reads using 105 rounds of a T-G-C-A flow cycle (420 total flows). The probability of lagging rephasing (i.e., the fraction of nucleotides that should integrate when the template indicates a nucleotide but do not integrate into the extending primer strand for each correctly integrated nucleotide) was set to 0.2%, and the probability of leading rephasing (i.e., the fraction of sequencing reads that have an extra nucleotide integrated into the extending primer after each flow) was set to 0.5%. The average read length of the control group was 322 bp ± 18 bp.
[0452] In a series of test groups, simulated flow maps were generated by computer sequencing of synthetic sequencing reads using 105 rounds of T-G-C-A flow cycles (420 flows in total), with the difference being one of the following conditions: (1) After every 24th flow, a rephasing flow containing a mixture of C and G was inserted ( Figure 26A ); (2) After every 48th flow, a rephasing flow containing a mixture of C and G was inserted ( Figure 26B ); (3) After every 96th flow, a rephasing flow containing a mixture of C and G was inserted ( Figure 26C ); (4) After every 192nd flow, a rephasing flow containing a mixture of C and G was inserted ( Figure 26D ); (5) After every 48th flow, a rephasing flow containing a mixture of C, G, and T was inserted, followed by a single A flow (to avoid redundant flows), and then the T-G-C-A cycle was resumed according to the control protocol ( Figure 26E ); (6) After every 96th flow, a rephasing flow containing a mixture of C, G, and T was inserted, followed by a single A flow (to avoid redundant flows), and then the T-G-C-A cycle was resumed according to the control protocol ( Figure 26F ); (7) After every 96th flow, a rephasing flow containing a mixture of C, G, and T was inserted, and then a rephasing flow containing a mixture of A, C, and G was inserted ( Figure 26G ); (8) After every 192nd flow, a rephasing flow containing a mixture of C, G, and T was inserted, and then a rephasing flow containing a mixture of A, C, and G was inserted ( Figure 26H ); (9) After every 96th flow, a rephasing flow containing a mixture of C, G, and T was inserted, followed by a rephasing flow containing a mixture of A, C, and T, followed by a rephasing flow containing a mixture of A, G, and T, followed by a rephasing flow containing a mixture of A, C, and G ( Figure 26I ); or (10) After every 192nd flow, a rephasing flow containing a mixture of C, G, and T was inserted, followed by a rephasing flow containing a mixture of A, C, and T, followed by a rephasing flow containing a mixture of A, G, and T, followed by a rephasing flow containing a mixture of A, C, and G ( Figure 26J ).
[0453] Compared to the control, after full-round computer sequencing, the use of any of the tested rephasing flows resulted in a significant reduction in the total phasing error (i.e., the sum of the fraction of chains with lag phasing error and the fraction of chains with lead phasing error relative to the nominal sequencing strand with no introduced lag or lead error), with minimal loss of sequencing data. Figures 26A - 26JShows the distribution of the sum of the total phasing errors for the control scenario and each corresponding rephasing flow scenario. Using a rephasing flow with a mixture of C and G after every 24th flow reduced the average total cumulative phasing error to 31.2 ± 9.6% (compared to 51.5 ± 1.3% for the control)( Figure 26A ), after every 48th flow, the average total cumulative phasing error was reduced to 36.9 ± 9.7%( Figure 26B ), after every 96th flow, the average total cumulative phasing error was reduced to 40.2 ± 10.1%( Figure 26C ), and after every 192nd flow, the average total cumulative phasing error was reduced to 42.8 ± 10.4%( Figure 26D ), while each rephasing flow only produced an average primer extension of ~1 bp (i.e., a sequencing gap). Using a rephasing flow with a mixture of C, G, and T after every 48th flow reduced the average total cumulative phasing error to 28.5 ± 10.6%( Figure 26E ), and after every 96th flow, the average total cumulative phasing error was reduced to 31.1 ± 12.2%( Figure 26F ), while each rephasing flow only produced an average primer extension of ~5 bp. Using a first rephasing flow with a mixture of C, G, and T and a second rephasing flow with a mixture of A, C, and G after every 96th flow reduced the average total cumulative phasing error to 25.3 ± 10.6%( Figure 26G ), and after every 192nd flow, the average total cumulative phasing error was reduced to 26.6 ± 12.6%( Figure 26H ), while each rephasing doublet flow only produced an average primer extension of ~9 bp. Using a first rephasing flow with a mixture of C, G, and T, a second rephasing flow with a mixture of A, C, and T, a third rephasing flow with a mixture of A, G, and T, and a fourth rephasing flow with a mixture of A, C, and G after every 96th flow reduced the average total cumulative phasing error to 20.6 ± 9.4%( Figure 26I ), and after every 192nd flow, the average total cumulative phasing error was reduced to 20.9 ± 11.2%( Figure 26J ), while each rephasing quadruplet flow only produced an average primer extension of ~18 bp.
Claims
1. A method for generating paired sequencing reads from a polynucleotide, comprising: (a) hybridizing a polynucleotide with a primer to form a hybridization template; (b) generating sequencing data related to the sequence of a first region of the polynucleotide using nucleotides provided in one or more flow steps of a first region flow order, wherein a single type of nucleotide is used in any given flow of the one or more flow steps of the first region flow order, and the flow steps include extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using nucleotides provided in 10 or more flow steps of a second region flow order, wherein the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer; and wherein (i) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; and (d) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides.
2. The method of claim 1, wherein the extension of the primer through the second region proceeds faster than the extension of the primer through the first region.
3. The method of claim 1, further comprising correlating the sequencing data of the first region with the sequencing data of the third region.
4. The method of claim 1, wherein at least a portion of the nucleotides used to extend the primer through the second region are unlabeled nucleotides.
5. The method of claim 1, wherein the nucleotides used to extend the primer through the second region are unlabeled nucleotides.
6. The method of claim 1, wherein the second region flow order has an efficiency of incorporating 0.6 or more bases per flow.
7. The method of claim 1, further comprising determining the expected sequencing data of the second region using a reference sequence and the second region flow order.
8. The method of claim 1, wherein the primer is extended through a third region using nucleotides provided in a third region flow order, and the method further comprises using a reference sequence of the second region, the second region flow order, the third region flow order, and a reference sequence of the third region to determine the expected sequencing data of the third region.
9. The method of claim 8, wherein the third region flow order comprises five or more nucleotide flows.
10. The method of claim 9, wherein each of the nucleotide flows of the third region flow order comprises a single nucleobase.
11. The method of any one of claims 10, wherein the second region flow order has an efficiency of incorporating 0.6 or more bases per flow.
12. The method of claim 1, wherein the primer is extended through the third region using nucleotides provided in a third region flow order, the method further comprising determining expected sequencing data for the third region using a reference sequence of the second region, a second region flow order, a third region flow order, and sequencing data related to the sequence of the third region, wherein the sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region.
13. The method of claim 7, wherein the expected reference data for the second region or the third region comprises a binary or non-binary flow map.
14. The method of claim 7, further comprising determining expected test variant sequencing data for the second region using a second region flow order and a second reference sequence of the second region, wherein the second reference sequence comprises a test variant.
15. The method of claim 14, wherein the primer is extended through the third region using nucleotides provided in a third region flow order, the method further comprising determining expected test variant sequencing data for the third region using the second reference sequence of the second region, a second region flow order, a third region flow order, and a reference sequence of the third region.
16. The method of claim 14, wherein the primer is extended through the third region using nucleotides provided in a third region flow order, the method further comprising determining expected test variant sequencing data for the third region using the second reference sequence of the second region, a second region flow order, a third region flow order, and sequencing data related to the sequence of the third region, wherein the sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region.
17. The method of claim 14, wherein the expected reference sequencing data for the second region or the third region comprises a binary or non-binary flow map.
18. A method of mapping paired coupled sequencing reads to a reference sequence, comprising: mapping a first region or a portion thereof, or a third region or a portion thereof, of the paired coupled sequencing reads generated by the method of any one of claims 1-17 to the reference sequence; and mapping an unmapped first region or a portion thereof, or an unmapped third region or a portion thereof, to the reference sequence using distance information indicating the length of the second region.
19. A method of detecting a structural variant, comprising: mapping a first region or a portion thereof, or a third region or a portion thereof, of the paired coupled sequencing reads generated by the method of any one of claims 1-17 to the reference sequence; using distance information indicating the length of the second region to determine an expected locus within the reference sequence for an unmapped first region or a portion thereof, or an unmapped third region or a portion thereof; determining expected sequencing data for the sequence at the expected locus based on the reference sequence; and detecting a structural variant by comparing the sequencing data of an unmapped first region or a portion thereof, or an unmapped third region or a portion thereof, with the expected sequencing data, wherein a difference between the sequencing data of an unmapped first region or a portion thereof, or an unmapped third region or a portion thereof, and the expected sequencing data indicates a structural variant.
20. A method of detecting a structural variant, comprising: Mapping the first region or a portion thereof or the third region or a portion thereof of a paired sequencing read generated by the method according to any one of claims 1-17 to a reference sequence, wherein the unmapped first region or the unmapped third region is unmappable within the reference sequence.
21. A method for detecting a structural variant, comprising: mapping the first region or a portion thereof and the third region or a portion thereof of a paired sequencing read generated by the method according to any one of embodiments 1-17 to a reference sequence; determining mapping distance information between the mapped first region and the mapped third region; and detecting a structural variant by comparing the mapped distance information with expected distance information of a second region, wherein a difference between the mapped distance information and the expected distance information indicates a structural variant.
22. The method of claim 19, wherein the structural variant is a chromosome fusion, inversion, insertion or deletion.
23. The method of claim 19, wherein the variant is an insertion or deletion within the second region.
24. A method for mapping a paired sequencing read to a reference sequence, comprising: mapping the first region or a portion thereof and the third region or a portion thereof of a paired sequencing read generated by the method according to any one of claims 1-17 to the reference sequence at two or more different position pairs including a first position and a second position; and for the two or more position pairs, using first distance information indicating the length of the second region and second distance information indicating the distance between the first position and the second position to select the correct position pair.
25. A method for detecting a variant between two sequencing regions of a paired sequencing read generated by the method according to any one of claims 1-17, wherein an extended primer is extended through a third region using nucleotides provided in a third region flow order, comprising: mapping the first region or a portion thereof to a reference sequence; using (1) the reference sequence of the second region, the second region flow order, the third region flow order, and the reference sequence of the third region, or (2) the reference sequence of the second region, the second region flow order, the third region flow order, and the generated sequencing data related to the sequence of the third region, to determine the expected sequence data of the third region or a portion thereof, wherein the generated sequencing data related to the sequence of the third region is the same or different sequencing data generated for the third region; and detecting the presence of a variant by comparing the expected sequencing data of the third region with the generated sequencing data related to the sequence of the third region.
26. The method of claim 25, wherein the variant is a structural variant.
27. The method of claim 26, wherein the structural variant is a chromosome fusion, inversion, insertion or deletion.
28. The method of claim 25, wherein the variant is a single nucleotide polymorphism (SNP).
29. The method of claim 25, wherein the method is used to detect a test variant, and the reference sequence includes the test variant.
30. The method of claim 29, wherein the test variant is selected by identifying the test variant within a second polynucleotide.
31. The method of claim 29, comprising associating the detected integrated variant with an allele sequenced in the first or third region of the polynucleotide.
32. A method of generating paired sequencing reads for detecting the presence of base transversions in an unsequenced region of a polynucleotide, comprising: (a) hybridizing a primer to the polynucleotide to form a hybridization template; (b) generating sequencing data related to the sequence of the first region of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (c) further extending the primer extended in step (b) through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and (d) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (c) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides.
33. A method of generating paired sequencing reads from a polynucleotide, comprising: (a) hybridizing a primer to the first region of the polynucleotide to form a hybridization template; (b) extending the primer through a second region using a flow order of alternating nucleotide pairs comprising (1) cytosine and thymine and (2) adenine and guanine; and (c) generating sequencing data related to the sequence of the third region of the polynucleotide by further extending the primer extended in step (b) using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides.
34. The method of claim 33, wherein the first region comprises a naturally occurring sequence targeted by the primer.
35. The method of claim 33, wherein the primer is extended through the second region without detecting the presence or absence of labels of nucleotides incorporated into the extended primer.
36. A method of detecting the presence of base transversions in an unsequenced region of a polynucleotide, comprising: mapping the first region or a portion thereof and the third region or a portion thereof of the paired sequencing reads generated according to any one of claims 33-35 to a reference sequence, wherein the primer is extended through the third region using nucleotides provided in the third region flow order; using the second region flow order, the third region flow order, and the reference sequence to determine the expected sequencing data for the third region; and detecting the presence of base transversions based on the difference between the expected sequencing data for the third region and the generated sequencing data for the third region.
37. The method of claim 36, wherein the second region flow order, the third region flow order, and the reference sequence of the second region and the reference sequence of the third region are used to determine the expected sequencing data for the third region.
38. The method of claim 36, wherein the second region flow order, the third region flow order, and the reference sequence of the second region and the generated sequence data related to the sequence of the third region are used to determine the expected sequencing data for the third region, wherein the generated sequence data related to the sequence of the third region is the same or different sequence data generated for the third region.
39. The method of claim 36, wherein the expected sequencing data for the third region comprises a binary or non-binary flow chart.
40. A method of generating one or more consensus sequences, comprising assembling a plurality of paired sequencing reads generated according to any one of claims 1-17.
41. The method of claim 40, further comprising validating a portion of a consensus sequence selected from one or more consensus sequences using selected paired sequencing reads related to a portion of the selected consensus sequence, wherein when generating the selected paired sequencing reads, primers are extended through a third region using nucleotides provided in a third region flow order, the validation comprising: using a second region flow order, a third region flow order, and a portion of the selected consensus sequence to determine expected sequencing data for a third region of the selected paired sequencing reads; and validating the portion of the selected consensus sequence by comparing the expected sequencing data for the third region of the selected paired sequencing reads with the generated sequencing data for the third region.
42. A method of validating a test variant status, comprising: comparing variant status on a plurality of overlapping paired sequencing reads generated according to any one of claims 1-17, the plurality of overlapping paired sequencing reads including loci corresponding to loci of a test variant; validating the variant status based on the comparison.
43. A method for detecting short genetic variants in a test sample, comprising: generating paired sequencing reads according to any one of claims 1-17; comparing sequencing data related to a sequence of a third region of a polynucleotide with expected sequencing data for an expected sequence of the third region of the polynucleotide; and determining the presence or absence of a short genetic variant in a second region of the polynucleotide.
44. The method of any one of claims 1-17 and 32-35, wherein the sequencing data related to the sequence of the first region or the sequencing data related to the sequence of the third region comprises flow signals representing base counts, the base counts indicating the number of bases incorporated at each of a plurality of flow positions.
45. The method of any one of claims 1-17 and 32-35, wherein the flow cycle order comprises 4 individually separate flows repeated in the same order.
46. The method of any one of claims 1-17 and 32-35, wherein the flow cycle order comprises 5 or more individually separate flows.
47. The method of any one of claims 1-17 and 32-35, wherein generating paired sequencing reads further comprises: further extending primers through a fourth region using nucleotides provided in a fourth region flow order, wherein the primers are extended through the fourth region without detecting the presence or absence of labels of nucleotides incorporated into the extended primers, and wherein a mixture of at least two different types of nucleotide bases is used in at least one step of the fourth region flow order; and generating sequencing data related to a sequence of a fifth region of a polynucleotide by further extending primers extended through the fourth region using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides.
48. The method of claim 47, further comprising correlating the sequencing data of the fifth region with the sequencing data of the first region or the third region. The method of any one of claims 1-17 and 32-35, wherein rolling circle amplification is used to amplify the polynucleotide.
50. A method for detecting short genetic variants in a test sample, comprising: (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide comprising at least a first copy of the polynucleotide and a second copy of the polynucleotide; (b) hybridizing the RCA-amplified polynucleotide with a primer to form a hybridization template; (c) generating sequencing data related to the sequence of a first region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in one or more flow steps of a first region flow order, wherein a single type of nucleotide is used in any given flow of the one or more flow steps of the first region flow order, the flow steps comprising extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (d) further extending the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of labels of nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; (e) generating sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; (f) comparing the sequencing data generated for the third region of the polynucleotide with the sequencing data expected for the expected sequence of the third region of the polynucleotide; (g) determining the presence of a short genetic variant in the second region of the polynucleotide; (h) generating sequencing data related to the sequence of the second region of the polynucleotide within the second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of incorporated labeled nucleotides; and (i) determining the identity of the short genetic variant in the second region of the polynucleotide.
51. The method of claim 50, wherein the extension of the primer through the second region of the polynucleotide within the first copy of the polynucleotide proceeds faster than the extension of the primer through the first region of the polynucleotide within the first copy of the polynucleotide.
52. A method for detecting short genetic variants in a test sample, comprising: (a) amplifying a polynucleotide using rolling circle amplification (RCA) to generate an RCA-amplified polynucleotide comprising at least a first copy of the polynucleotide and a second copy of the polynucleotide; (b) hybridizing a primer with a first region of the polynucleotide within the first copy of the polynucleotide to form a hybridization template; (c) Extend the primer through a second region of the polynucleotide within the first copy of the polynucleotide using nucleotides provided in a second region flow order, wherein (i) the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer, or (ii) a mixture of at least two different types of nucleobases is used in at least one step of the second region flow order; (d) Generate sequencing data related to the sequence of a third region of the polynucleotide by further extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides; (e) Compare the sequencing data generated for the third region of the polynucleotide with the expected sequencing data of the expected sequence of the third region of the polynucleotide; (f) Determine the presence of a short genetic variant in the second region of the polynucleotide; (g) Generate sequencing data related to the sequence of a second region of the polynucleotide within a second copy of the polynucleotide by extending the primer using labeled nucleotides and detecting the presence or absence of the incorporated labeled nucleotides; and (h) Determine the identity of the short genetic variant in the second region of the polynucleotide.
53. The method of claim 52, wherein the first region comprises a naturally occurring sequence targeted by the primer.
54. The method of any one of claims 50 - 53, wherein based on determining the presence of a short genetic variant in the second region of the polynucleotide, sequencing data related to the sequence of a second region of the polynucleotide within a second copy of the polynucleotide is dynamically generated.
55. The method of any one of claims 50 - 53, wherein the primer is extended through the second region of the polynucleotide within the first copy of the polynucleotide without detecting the presence or absence of a label of the nucleotides incorporated into the extended primer.
56. The method of any one of claims 50 - 53, wherein at least a portion of the nucleotides for extending the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides.
57. The method of any one of claims 50 - 53, wherein the nucleotides for extending the primer through the second region of the polynucleotide within the first copy of the polynucleotide are unlabeled nucleotides.
58. A method for synchronizing sequencing primers within a sequencing cluster, comprising: (a) Hybridize the primer to a polynucleotide copy within the sequencing cluster; (b) Extend the primer through a first region of the polynucleotide copy using labeled nucleotides according to a first region flow order; (c) Extend the primer through a second region of the polynucleotide copy using one or more rephasing flows, wherein a mixture of at least two different types of nucleobases is used in at least one of the one or more rephasing flows; and (d) Extend the primer through a third region of the polynucleotide copy using labeled nucleotides according to a third region flow order.
59. The method of claim 58, wherein a mixture of three different types of nucleobases is used in at least one of the one or more rephasing flows.
60. The method of claim 58 or 59, wherein the one or more rephasing flows comprise four or more flow steps. The method of claim 60, wherein one or more of the rephasing flows occur in any order comprising: (i) a first flow comprising a mixture that includes nucleotides A, C, and G and omits nucleotide T; (ii) a second flow comprising a mixture that includes nucleotides T, C, and G and omits nucleotide A; (iii) a third flow comprising a mixture that includes nucleotides T, A, and G and omits nucleotide C; and (iv) a fourth flow comprising a mixture that includes nucleotides T, A, and C and omits nucleotide G.
62. The method of claim 58 or 59, comprising generating sequencing data related to the sequence of the first region by detecting the presence or absence of an incorporated labeled nucleotide while extending a primer through the first region.
63. The method of claim 58 or 59, comprising generating sequencing data related to the sequence of the third region by detecting the presence or absence of an incorporated labeled nucleotide while extending a primer through the third region.
Citation Information
Patent Citations
Methods for biological sample processing and analysis
US10344328B2
Rolling circle synthesis of oligonucleotides and amplification of select randomized circular oligonucleotides
US5714320A
Mostly natural DNA sequencing by synthesis
US8772473B2
Methods and systems for sequence calling
WO2019084158A1
Methods and systems for sequencing long nucleic acids
CN103917654A