RNA sequencing method
By hybridizing the polynucleotides of mRNA with primers, sequencing the sequences of barcode regions and target regions using a mixture of labeled nucleotides and unlabeled nucleotides, the efficiency and accuracy problems caused by poly(A) tail and poly(T) regions in eukaryotic mRNA sequencing are solved, and efficient and economical mRNA sequencing is achieved.
Patent Information
- Application Number
- CN202080063117.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-07-10
- Filing Date
- 2020-07-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-07-10
AI Technical Summary
The prior art presents challenges in direct mRNA sequencing, especially when targeting mRNA molecules to reverse transcription into complementary DNA (cDNA) molecules and sequenced, the poly(A) tail of eukaryotic mRNA and its complementary poly(T) regions lead to limited efficiency and accuracy of sequencing methods.
By hybridizing the mRNA-derived polynucleotide with the primer, multiple hybrid templates are formed, including the barcode region, the homopolymer region and the target region. Use a mixture of labeled nucleotides and unlabeled nucleotides to determine the sequence of the barcode region and the target region, avoiding the extension of primers within the homopolymer region, and reducing sequencing costs and time.
Efficient sequencing of targeted region sequences in mRNA molecules is achieved, reducing the sequencing requirement for homopolymer regions, and improving sequencing efficiency and economicality.
Smart Images

Figure CN114423873B_ABST
Abstract
Description
[0001] Cross - reference to related applications
[0002] This application claims priority to U.S. Provisional Application No. 62 / 872,558, filed Jul. 10, 2019, which is incorporated herein by reference for all purposes. Field of the Invention
[0003] Described herein are methods for sequencing a target region within an mRNA molecule.
[0004] Background
[0005] Messenger RNA (mRNA) sequencing can provide information about real - time gene expression, gene expression profiles in different tissues, and gene expression levels. Changes in the expression level of a particular gene can vary in response to environmental stimuli at different developmental stages, or in the context of a causal relationship with a disease.
[0006] Direct mRNA sequencing is challenging. Targeted mRNA molecules are typically reverse - transcribed into complementary DNA (cDNA) molecules using reverse transcriptase. The cDNA molecules are then sequenced to provide sequencing information about the original mRNA molecule. Eukaryotic mRNAs typically contain a poly(A) tail at the 3' end of the mRNA molecule, which is reflected as a poly(T) region at the 5' end of the cDNA molecule. Optionally, a DNA strand complementary to the cDNA molecule is synthesized, generating a poly(A) tail at the 3' end of the DNA molecule.
[0007] Certain high - throughput sequencing methods utilize non - terminating nucleotides to sequence nucleic acid molecules. These sequencing methods may be referred to as "flow - through sequencing", "natural synthesis sequencing", or "non - terminating synthesis sequencing" methods (see, e.g., U.S. Patent No. 8,772,473, which is incorporated herein by reference in its entirety). Summary of the Invention
[0008] Described herein are methods for determining the sequence of a target region from an mRNA molecule.
[0009] In some embodiments, a method of determining the sequence of a target region from an mRNA molecule includes (a) hybridizing a plurality of polynucleotides derived from the mRNA with primers to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, wherein the barcode region omits the identical bases present in the homopolymer region; (b) determining the sequence of the barcode region using labeled nucleotides lacking the complementary bases of the identical bases present in the homopolymer region; (c) extending the primer within the homopolymer region using nucleotides complementary to the bases present in the homopolymer region; and (d) determining the sequence of the target region using labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) comprise non-terminating labeled nucleotides. In some embodiments, the nucleotides used in step (c) comprise non-terminating nucleotides. In some embodiments, the nucleotides used in step (d) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) are mixed with unlabeled nucleotides of the identical bases. In some embodiments, the nucleotides complementary to the bases present in the homopolymer region comprise unlabeled nucleotides. In some embodiments, a mixture of unlabeled nucleotides and labeled nucleotides is used to determine the sequence of the target region. In some embodiments, different bases of the labeled nucleotides used to determine the target region are used discretely. In some embodiments, different bases of the labeled nucleotides used to determine the barcode region are used discretely. In some embodiments, the target region of each polynucleotide is related to a unique barcode region. In some embodiments, no more than 50% of the total nucleotides used in step (b), step (c), or step (d) are labeled. In some embodiments, no more than 0.1% of the total nucleotides used in step (c) are labeled. In some embodiments, the method further comprises repeating step (c) one or more times until the primer extends to the end of the homopolymer region, wherein unincorporated nucleotides are removed between the repeating steps.
[0010] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides and unlabeled nucleotides at a first ratio of labeled nucleotides to total nucleotides; (c) extending the primer using a labeled nucleotide complementary to the base present in the homopolymer region at a second ratio of labeled nucleotides to total nucleotides, wherein the second ratio is greater than the first ratio, and wherein primer extension stops within the homopolymer region; (d) extending the primer to the end of the homopolymer region using an unlabeled nucleotide complementary to the base present in the homopolymer region; (e) determining the sequence of the target region using labeled nucleotides. In some embodiments, the labeled nucleotides and unlabeled nucleotides used in step (b) comprise non-terminating labeled nucleotides and non-terminating unlabeled nucleotides. In some embodiments, the labeled nucleotides used in step (c) comprise non-terminating labeled nucleotides. In some embodiments, the unlabeled nucleotides used in step (d) comprise non-terminating unlabeled nucleotides. In some embodiments, the labeled nucleotides used in step (e) comprise non-terminating labeled nucleotides. In some embodiments, the sequence of the target region is determined using labeled nucleotides and unlabeled nucleotides. In some embodiments, 50% or less of the nucleotides used in step (b) or (e) are labeled. In some embodiments, more than 50% of the nucleotides used in step (c) are labeled. In some embodiments, 0.1% or less of the nucleotides used in step (d) are labeled. In some embodiments, all of the nucleotides used in step (d) are unlabeled. In some embodiments, step (d) includes repeating the removal of unincorporated nucleotides and addition of new nucleotides one or more times until the primer is extended to the end of the homopolymer region. In some embodiments, the target region of each polynucleotide is related to a unique barcode region. In some embodiments, different bases of the labeled nucleotides used for determining the target region are used discretely. In some embodiments, different bases of the labeled nucleotides used for determining the barcode region are used discretely.
[0011] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA with primers to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, and wherein the target region is associated with a unique barcode region; (b) determining the sequence of the barcode region using labeled nucleotides in a plurality of predetermined cycles, wherein the predetermined cycles and the barcode region of the polynucleotide are configured such that the primer extends to the end of the barcode region across the plurality of polynucleotides before extending into the homopolymer region; (c) extending the primer within the homopolymer region using unlabeled nucleotides; and (d) determining the sequence of the target region using labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) comprise non-terminating labeled nucleotides. In some embodiments, the unlabeled nucleotides used in step (c) comprise non-terminating unlabeled nucleotides. In some embodiments, the labeled nucleotides used in step (d) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) or step (d) are mixed with unlabeled nucleotides of the same base. In some embodiments, 50% or less of the nucleotides used in step (b) or step (d) are labeled. In some embodiments, 0.1% or less of the nucleotides used in step (c) are labeled. In some embodiments, all of the nucleotides used in step (c) are unlabeled. In some embodiments, the method further comprises repeating step (c) one or more times until the primer extends to the end of the homopolymer region, wherein unincorporated nucleotides are removed prior to repeating step (c). In some embodiments, different bases of the labeled nucleotides used for determining the target region are used discretely. In some embodiments, different bases of the labeled nucleotides used for determining the target region are used discretely. In some embodiments, the target region of each polynucleotide is associated with a unique barcode region.
[0012] In some embodiments, a method of determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a first primer to form a first plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides; (c) hybridizing the plurality of polynucleotides to a second primer to form a second plurality of hybridization templates, wherein the second primer comprises a homopolymer region comprising a plurality of consecutive and identical bases complementary to the bases in the homopolymer region of the polynucleotide; (d) determining the sequence of the target region using labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used in step (d) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) or step (d) are mixed with unlabeled nucleotides of the same base. In some embodiments, steps (c) and (d) are performed before steps (a) and (b). In some embodiments, steps (a) and (b) are performed before steps (c) and (d). In some embodiments, the method includes removing the first primer after step (b) or removing the second primer after step (d). In some embodiments, the second primer comprises a 3' anchor at the 3' end of the primer, the 3' anchor comprising bases other than the bases present in the homopolymer region of the second primer. In some embodiments, the second primer comprises a 5' anchor at the 5' end of the primer, wherein the anchor is covalently bound to the homopolymer region of the second primer via a linker. In some embodiments, the 5' anchor comprises a nucleic acid segment comprising a sequence identical to at least a portion of the first primer. In some embodiments, the linker comprises one or more nucleic acids. In some embodiments, the linker is a PEG phosphoramidite linker. In some embodiments, the method further includes extending the second primer within the homopolymer region using nucleotides. In some embodiments, the nucleotides used to extend the second primer within the homopolymer region comprise non-terminating nucleotides. In some embodiments, the nucleotides used to extend the second primer within the homopolymer region comprise unlabeled nucleotides. In some embodiments, a mixture of unlabeled nucleotides and labeled nucleotides is used to determine the sequence of the target region or the barcode region. In some embodiments, different bases of the labeled nucleotides used to determine the target region are used discretely. In some embodiments, different bases of the labeled nucleotides used to determine the target region are used discretely. In some embodiments, the target region of each polynucleotide is associated with a unique barcode region.
[0013] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates; the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; wherein the primer comprises a first primer segment, a second primer segment, and a cleavable linker between the first primer segment and the second primer segment, the second primer segment comprising a homopolymer region that comprises a plurality of bases complementary to the bases in the homopolymer region of the polynucleotide; (b) determining the sequence of the target region using labeled nucleotides; (c) cleaving the primer at the cleavable linker; and (d) determining the sequence of the barcode region using labeled nucleotides. In some embodiments, the labeled nucleotides used to determine the sequence of the target region in step (b) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used to determine the sequence of the barcode region in step (d) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used in step (b) or step (d) are mixed with unlabeled nucleotides of the same base. In some embodiments, the cleavable linker comprises one or more nucleic acids. In some embodiments, the cleavable linker comprises a uracil base, and wherein the primer is cleaved by contacting the primer with a uracil-specific nuclease. In some embodiments, the method further includes extending the second primer segment within the homopolymer region using nucleotides. In some embodiments, the nucleotides used to extend the second primer segment within the homopolymer region comprise non-terminating nucleotides. In some embodiments, the nucleotides used to extend the second primer segment within the homopolymer region comprise unlabeled nucleotides. In some embodiments, a mixture of unlabeled nucleotides and labeled nucleotides is used to determine the sequence of the target region or the barcode region. In some embodiments, different bases of the labeled nucleotides used to determine the target region are used discretely. In some embodiments, different bases of the labeled nucleotides used to determine the target region are used discretely. In some embodiments, the target region of each polynucleotide is associated with a unique barcode region.
[0014] In some embodiments, methods for determining the sequence of a target region from an mRNA molecule include: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, wherein the polynucleotides comprise a homopolymer region and a target region, the homopolymer region comprising a plurality of consecutive and identical bases, and the target region comprising a sequence related to the target region from the mRNA molecule; (b) extending the primer into the homopolymer region using nucleotides of the same base, wherein the primer stops within the homopolymer region and unincorporated nucleotides are removed; (c) repeating step (b) one or more times to extend the primer through the homopolymer region; (d) determining the sequence of the target region using labeled nucleotides. In some embodiments, the nucleotides used to extend the primer in step (b) comprise non-terminating labeled nucleotides. In some embodiments, the labeled nucleotides used to determine the sequence of the target region in step (d) comprise non-terminating labeled nucleotides. In some embodiments, the nucleotides used in step (b) comprise labeled nucleotides. In some embodiments, the nucleotides used in step (b) comprise unlabeled nucleotides. In some embodiments, the sequence of the target region is determined using a mixture of unlabeled and labeled nucleotides. In some embodiments, the polynucleotide further comprises a barcode region, and wherein the target region is associated with a unique barcode region, and the method further comprises determining the sequence of the barcode region using labeled nucleotides.
[0015] In some embodiments of any of the above methods, the bases in the homopolymer region of the polynucleotide are adenine or thymine bases.
[0016] In some embodiments of any of the above methods, the homopolymer region of the polynucleotide comprises at least 8 consecutive and identical bases. In some embodiments, the homopolymer region of the polynucleotide comprises at least 50 consecutive and identical bases.
[0017] In some embodiments of any of the above methods, the barcode region comprises a sample barcode.
[0018] In some embodiments of any of the above methods, the method further comprises associating the determined sequence of the barcode region with the determined sequence related to the mRNA coding region from the same target nucleic acid molecule.
[0019] In some embodiments of any of the above methods, the polynucleotide is a cDNA molecule.
[0020] In some embodiments of any of the above methods, the target region of the mRNA molecule comprises the coding region of the mRNA molecule.
[0021] In some embodiments of any of the above methods, the target region of the mRNA molecule comprises the 3'-untranslated region or the 5'-untranslated region of the mRNA molecule.
[0022] BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 Exemplary methods for obtaining polynucleotides that can be sequenced using the methods described herein are illustrated.
[0024] Figure 2A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule is shown, where the barcode region omits identical bases present in the homopolymer region.
[0025] Figure 2B An exemplary method for determining the sequence of a target region from an mRNA molecule is shown in graphical form, where the barcode region omits identical bases present in the homopolymer region.
[0026] Figure 3A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule is shown, where primer extension stops within the homopolymer region.
[0027] Figure 3B An exemplary method for determining the sequence of a target region from an mRNA molecule is shown in graphical form, where primer extension stops within the homopolymer region.
[0028] Figure 4A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule is shown, where the flow cycling of configured polynucleotides and different barcode regions causes the primer to extend to the end of the barcode region spanning multiple polynucleotides before extending into the homopolymer region.
[0029] Figure 4B An exemplary method for determining the sequence of a target region from an mRNA molecule is shown in graphical form, where multiple polynucleotides have sequences of different barcode regions but the same flow length.
[0030] Figure 5A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule using two primers is shown.
[0031] Figure 5B An exemplary method for determining the sequence of a target region from an mRNA molecule using two primers is shown in graphical form.
[0032] Figure 5C A flow chart of another exemplary method for determining the sequence of a target region from an mRNA molecule using two primers is shown, where compared to Figure 5A the exemplary method shown, the sequences of the two primers hybridized to the polynucleotide are reversed.
[0033] Figure 5D An exemplary method for determining the sequence of a target region from an mRNA molecule using two primers is shown in graphical form, where the sequences of the two primers hybridized to the polynucleotide are opposite compared to the exemplary method shown in Figure 5A The sequences of the two primers hybridized to the polynucleotide are opposite compared to the exemplary method shown in
[0034] Figure 6A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule using a cleavable primer is shown.
[0035] Figure 6B An exemplary method for determining the sequence of a target region from an mRNA molecule using a cleavable primer is shown in graphical form.
[0036] Figure 7A A flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule is shown, where primer extension stops and restarts within a homopolymer region of the polynucleotide.
[0037] Figure 7B An exemplary method for determining the sequence of a target region from an mRNA molecule is shown in graphical form, where primer extension stops and restarts within a homopolymer region of the polynucleotide.
[0038] Figure 8 The sequencing flow matrix and signal traces for three exemplary sequences are shown, the exemplary sequences comprising a barcode region, a homopolymer region, and a target region. DETAILED DESCRIPTION OF THE INVENTION
[0040] Described herein are methods for determining the sequence of a target region (e.g., coding region, 3' untranslated region (3'-UTR), and / or 5' untranslated region (5'-UTR)) within an mRNA molecule. The methods allow for sequencing of the target region within the mRNA molecule while avoiding sequencing of all or most of the homopolymer region (e.g., poly-A region or its complement, poly-T region). The poly-A region of an mRNA molecule is generally considered uninteresting and uninformative as it does not provide important information regarding the mRNA sequence or gene expression levels. Thus, there are significant benefits in reducing the cost or time required to extend a sequencing primer through the homopolymer region to reach the more informative regions of the polynucleotide. For example, cost can be reduced by extending the sequencing primer while reducing or eliminating the use of labeled nucleotides, which are significantly more expensive than their non-labeled nucleotide counterparts. Additionally, the sequencing primer extension time can be reduced by skipping the detection step for each incorporated base and allowing the primer to continuously extend through the homopolymer region.
[0041] The sequenced polynucleotides can include a barcode region, which can include, for example, a sample barcode and / or a unique molecular identifier (UMI). A homopolymer region can be located in the polynucleotide between the barcode region and the region of interest of the mRNA molecule. This results in a situation where the two regions whose sequences need to be known (i.e., the barcode region and the region of interest of the mRNA molecule) are separated by a non-region of interest (the homopolymer region). Some of the methods described herein allow for sequencing of both the barcode region and the target mRNA region while minimizing disruption caused by the presence of the homopolymer region.
[0042] In some embodiments, there is a method for determining the sequence of a region of interest from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the region of interest from the mRNA molecule, wherein the barcode region omits the identical bases present in the homopolymer region; (b) determining the sequence of the barcode region using labeled nucleotides lacking the complementary bases to the identical bases present in the homopolymer region; (c) extending the primer within the homopolymer region using nucleotides complementary to the bases present in the homopolymer region; and (d) determining the sequence of the target region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0043] In some embodiments, there is a method for determining the sequence of a region of interest from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the region of interest from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides and unlabeled nucleotides at a first ratio of labeled nucleotides to total nucleotides; (c) extending the primer within the homopolymer region using labeled nucleotides complementary to the bases present in the homopolymer region at a second ratio of labeled nucleotides to total nucleotides, wherein the second ratio is greater than the first ratio, and wherein primer extension stops within the homopolymer region; (d) extending the primer to the end of the homopolymer region using unlabeled nucleotides complementary to the bases present in the homopolymer region; and (e) determining the sequence of the target region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0044] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, and wherein the target region is associated with a unique barcode region; (b) determining the sequence of the barcode region using labeled nucleotides in a plurality of predetermined cycles, wherein the predetermined cycles and the barcode region of the polynucleotide are configured such that the primer extends to the end of the barcode region across the plurality of polynucleotides before extending into the homopolymer region; (c) extending the primer within the homopolymer region using unlabeled nucleotides; and (d) determining the sequence of the target region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0045] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a first primer to form a first plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides; (c) hybridizing the plurality of polynucleotides to a second primer to form a second plurality of hybridization templates, wherein the second primer comprises a homopolymer region comprising a plurality of consecutive and identical bases complementary to the bases in the homopolymer region of the polynucleotide; and (d) determining the sequence of the target region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0046] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridization templates; the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; wherein the primer comprises a first primer segment, a second primer segment that comprises a homopolymer region that comprises a plurality of bases complementary to the bases in the homopolymer region of the polynucleotide, and a cleavable linker between the first primer segment and the second primer segment; (b) determining the sequence of the target region using labeled nucleotides; (c) cleaving the primer at the cleavable linker; and (d) determining the sequence of the barcode region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0047] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of mRNA-derived polynucleotides with primers to form a plurality of hybridization templates, wherein the polynucleotides comprise a homopolymer region containing a plurality of consecutive and identical bases and a target region comprising a sequence related to the target region from the mRNA molecule; (b) extending the primers with nucleotides of the same base into the homopolymer region, wherein the primers stop within the homopolymer region and unincorporated nucleotides are removed; (c) repeating step (b) one or more times to extend the primers through the homopolymer region; and (d) determining the sequence of the target region using labeled nucleotides. Optionally, the nucleotides are non-terminating nucleotides.
[0048] Definitions
[0049] As used herein, unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" include plural referents.
[0050] References herein to "about" a value or parameter include (and describe) variations that are directed to the value or parameter itself. For example, a description of "about X" includes a description of "X".
[0051] The terms "individual", "patient", and "subject" are used synonymously and refer to an animal including a human.
[0052] As used herein, the term "label" refers to a detectable moiety that is conjugated or can be conjugated to another moiety, such as a nucleotide or nucleotide analogue. The label can emit a signal or modify a signal delivered to the label, such that the presence or absence of the label can be detected. In some cases, the conjugation can be through a linker, which can be cleavable, such as photocleavable (e.g., cleavable under ultraviolet light), chemically cleavable (e.g., by a reducing agent such as dithiothreitol (DTT), tris(2-carboxyethyl)phosphine (TCEP)), or enzymatically cleavable (e.g., by an esterase, lipase, peptidase, or protease).
[0053] A "non-terminating nucleotide" is a nucleic acid moiety that can be attached to the 3' end of a polynucleotide using a polymerase or transcriptase, and to which another non-terminating nucleic acid can be attached using a polymerase or transcriptase without the need to remove a protecting group or reversible terminator from the nucleotide. Naturally occurring nucleic acids are non-terminating nucleic acids. Non-terminating nucleic acids can be labeled or unlabeled.
[0054] It should be understood that aspects and variations of the invention described herein include aspects and variations of "consisting of" and / or "consisting essentially of".
[0055] When numerical ranges are provided, it is to be understood that each intermediate value between the upper and lower limits of such range, and any other stated value or intermediate value in such stated range, is included within the scope of the present disclosure. Where the stated range includes the upper and lower limits, ranges excluding either of those included ranges are also included in the present disclosure.
[0056] The section headings used herein are for organizational purposes only and should not be construed as limiting the subject matter described. This description is provided to enable a person of ordinary skill in the art to make and use the invention and is provided in the context of a patent application and its requirements. Various modifications to the described embodiments will be apparent to those skilled in the art, and the general principles herein can be applied to other embodiments. Thus, the invention is not intended to be limited to the embodiments shown but is to be accorded the widest scope consistent with the principles and features described herein.
[0057] Figures 1 - 8 Processes in accordance with various embodiments are shown. In an exemplary process, some of the blocks are optionally combined, the order of some of the blocks is optionally changed, and some of the blocks are optionally omitted. In some examples, additional steps may be performed in conjunction with the exemplary process. Thus, the operations illustrated (and described in more detail below) are exemplary in nature and should not be considered limiting.
[0058] The disclosures of all publications, patents, and patent applications mentioned herein are hereby incorporated by reference in their entirety. If any reference incorporated by reference conflicts with the present disclosure, the present disclosure shall control.
[0059] Polynucleotide
[0060] The polynucleotides used in the methods described herein may comprise homopolymer regions and target regions. In some embodiments, the polynucleotide further comprises a barcode region. Generally, the polynucleotide is derived from an mRNA molecule, for example by generating a complementary DNA (cDNA) molecule from a reverse transcriptase. In some embodiments, the polynucleotide is the complement of a cDNA molecule. Multiple polynucleotides may be derived from multiple mRNA molecules, and different polynucleotides may comprise sequences reflecting different mRNA molecules. Since different mRNA molecules may comprise different sequences of the region of interest and / or different lengths of homopolymers (e.g., poly-A regions), the polynucleotides may also comprise different target regions and / or different lengths of homopolymers.
[0061] The target region of the polynucleotide contains a sequence related to the region of interest from an mRNA molecule. For example, the sequence can be identical or complementary to the region of interest from the mRNA molecule. The region of interest can include one or more 3'-UTRs, 5'-UTRs, and / or coding regions from the mRNA molecule, or a portion of any of these regions.
[0062] The polynucleotide contains a homopolymer region (e.g., a poly-A or poly-T region) that represents the poly-A region in the mRNA molecule or its complement. In some embodiments, the length of the homopolymer region of the polynucleotide is about 5 bases or more, about 8 bases or more, about 10 bases or more, about 20 bases or more, about 30 bases or more, about 40 bases or more, about 50 or more bases, about 100 or more bases, or about 200 or more bases. In some embodiments, the length of the homopolymer region of the polynucleotide is from about 5 bases to about 200 bases, such as from about 5 bases to about 8 bases, from about 8 bases to about 10 bases, from about 10 bases to about 20 bases, from about 20 bases to about 30 bases, from about 30 bases to about 40 bases, from about 40 bases to about 50 bases, from about 50 bases to about 100 bases, or from about 100 bases to about 200 bases.
[0063] In some embodiments, the polynucleotide contains a barcode region (e.g., a sample barcode and / or UMI). In some embodiments, the polynucleotides are derived from different mRNA molecules and thus can contain different target regions of interest from different mRNA molecules and, if present, can contain different barcode regions to uniquely identify the original mRNA molecule. In some embodiments, the barcode region contains a sample barcode. The sample barcode can be unique for a given sample and thus can be common among capture primers used for the same sample (e.g., from the same patient, the same tissue of a patient, or from the same cell). Thus, polynucleotides from different samples can be combined for multiplex sequencing, and the origin of any given sequenced polynucleotide can be determined based on the sample barcode.
[0064] In some embodiments, the homopolymer region of the polynucleotide is located between the target region and the barcode region. For example, in some embodiments, the 5' end of the homopolymer region is adjacent to the 3' end of the target region, and the 3' end of the homopolymer region is adjacent to the 5' end of the barcode region. In some embodiments, the 3' end of the homopolymer region is adjacent to the 5' end of the target region, and the 5' end of the homopolymer region is adjacent to the 3' end of the barcode region. In some embodiments, there are no intervening bases between the barcode region and the homopolymer region. In some embodiments, there are no intervening bases between the homopolymer region and the target region.
[0065] mRNA can be isolated from a biological sample (e.g., a sample derived from a blood sample, a cell sample, a tissue sample, or other biological sample) using a capture primer (e.g., a DNA capture primer). For example, a blood sample can be drawn or a tissue sample can be biopsied to obtain a biological sample. In some embodiments, the biological sample is a single cell, and the mRNA from the single cell can be labeled with a common sample barcode (e.g., single cell RNA sequencing or “scRNA-seq”). In some embodiments, the capture primer comprises a poly-T sequence that hybridizes to the poly-A tail at the 3' end of the mRNA. The capture primer can be bound to a surface, such as a bead or other attachment surface, and unbound nucleic acid molecules or other biological debris can be removed. In some embodiments, the capture primer can comprise a barcode region. The barcode region can comprise a unique molecular identifier (UMI) that is associated with the mRNA molecule that hybridizes to that particular capture primer. Thus, multiple different mRNA molecules can be captured using the unique barcode region.
[0066] A reverse transcriptase can be used to extend the capture primer hybridized to the mRNA molecule to generate a cDNA molecule. If the barcode is to be present in the final polynucleotide, the cDNA molecule comprises a poly-T homopolymer region as well as the barcode region of the capture primer. In some embodiments, the complement of the cDNA molecule is generated, and the cDNA molecule and / or the complement of the cDNA molecule is used as a template for sequencing. In some embodiments, the length of the poly-T homopolymer region in the capture probe is from about 5 bases to about 50 bases, such as from about 10 to about 30 bases, or about 20 bases.
[0067] Figure 1Exemplary methods of obtaining polynucleotides that can be sequenced using the methods described herein are illustrated. The mRNA molecule 102 includes a target region 104 and a poly-A homopolymer region 106 near the 3'-end of the target region 102. The target region can include a 3'-UTR 108, a coding region 110, and a 5'-UTR. The mRNA molecule 102 obtained from a sample can hybridize with a capture probe 116, which is typically fused to a surface 114, such as a bead or other suitable surface. The capture probe 116 includes a linker region 118 that attaches the capture probe 116 to the surface 114, a barcode region 120 (optional), and a poly-T homopolymer region 122. The poly-T homopolymer region 122 of the capture probe hybridizes to the poly-A homopolymer region 106 of the mRNA molecule. The surface 114 containing the attached capture probe 116 hybridized to the mRNA molecule 102 can be washed to remove uncaptured polynucleotides or other materials from the sample. Using the mRNA molecule 102 as a template, the capture probe 116 is extended with a reverse transcriptase 124. This results in a cDNA molecule 126 that includes a target region 128 complementary to the target region of the mRNA molecule 102, a polyT homopolymer region 122, and a barcode region 120. The linker region 118 can be cleaved from the surface 114, thereby releasing the cDNA molecule 126. Adapters 130 and 132 can be ligated to the 3' and / or 5'-ends of the cDNA molecule 126. A complement 134 of the cDNA molecule 126 can be generated, which can hybridize with a sequencing primer 136 to generate a hybridization template. In any of the embodiments, the hybridization template can be subjected to any of the sequencing methods described herein.
[0068] Flow sequencing method
[0069] Using the methods described herein, the sequence of a target region from an mRNA molecule can be determined. The polynucleotide derived from the mRNA hybridizes to a sequencing primer to generate a hybridization template for sequencing. The methods described herein can include sequencing the polynucleotide using a flow sequencing method, which can also be referred to as a "natural synthesis sequencing" or "non-terminating synthesis sequencing" method. Exemplary methods are described in U.S. Patent No. 8,772,473, which is incorporated herein by reference in its entirety. Although the following description is provided with reference to the flow sequencing method, it should be understood that other sequencing methods can be used to sequence all or part (e.g., one or more regions, such as barcode regions and / or target regions). Flow sequencing involves using nucleotides to extend a primer hybridized to a polynucleotide. If a complementary base is present in the template strand, nucleotides of a given base type (e.g., A, C, G, T, U, etc.) can be mixed with the hybridization template to extend the primer. For example, the nucleotides can be non-terminating nucleotides. When the nucleotides are non-terminating, if there are more than one consecutive complementary bases in the template strand, more than one consecutive base can be incorporated into the extended primer strand. Non-terminating nucleotides contrast with nucleotides having 3' reversible terminators, where the blocking group is typically removed before ligating consecutive nucleotides. If no complementary base is present in the template strand, primer extension stops until a nucleotide complementary to the next base in the template strand is introduced. At least some of the nucleotides can be labeled so that incorporation can be detected. Although two or three different types of nucleotides can be introduced simultaneously in some embodiments, most commonly only one nucleotide type is introduced at a time (i.e., discrete addition). This method can be contrasted with sequencing methods using reversible terminators, where primer extension stops after the extension of each single base, and then the terminator is reversed to allow incorporation of the next subsequent base.
[0070] The nucleotides can be introduced in a determined order, which can be further divided into cycles. The nucleotides are added stepwise, which allows the added nucleotides to be incorporated at the end of the sequencing primer at the complementary bases present in the template strand. The cycles can have the same nucleotide order and number of different base types (i.e., symmetric cycles) or different nucleotide orders and / or different numbers of different base types (i.e., asymmetric cycles). However, there are no repeated bases within the same cycle, which provides a marker for distinguishing different cycles. By way of example only, the order of the first cycle can be A-T-G-C and the order of the second cycle can be A-T-C-G. Additionally, one or more cycles can omit one or more nucleotides. By way of example only, the order of the first cycle can be A-T-G-C and the order of the second cycle can be A-T-C. Alternative orders can be readily envisioned by those skilled in the art. Between the introduction of different nucleotides, unincorporated nucleotides can be removed, for example, by washing the sequencing platform with a wash solution.
[0071] Polymerases can be used to extend sequencing primers by incorporating one or more nucleotides at the primer terminus in a template-dependent manner. In some embodiments, the polymerase is a DNA polymerase. In some embodiments, the polymerase is an RNA polymerase. The polymerase can be a naturally occurring polymerase or a synthetic (e.g., mutant) polymerase. The polymerase can be added at the initial step of primer extension, although a supplemental polymerase can optionally be added during sequencing, e.g., by stepwise addition of nucleotides or after multiple flow cycles.
[0072] The presence or absence of incorporated labeled nucleic acids can be detected to determine the sequence. The label can be, for example, an optically active label (e.g., a fluorescent label) or a radioactive label, and a detector can be used to detect the signal emitted or altered by the label.
[0073] A flow chart can be generated based on the detection of incorporated nucleotides and the order of nucleotide introduction. As an example, consider repeated flow cycles of the template sequences: CTG, CAG, and TACG. The resulting flow chart is shown in Table 1, where 1 represents the incorporation of an introduced nucleotide and 0 represents the non-incorporation of an introduced nucleotide. The flow chart can be used to determine the sequence of the template strand.
[0074] Table 1
[0075]
[0076] When determining the sequence of the template strand, the introduced nucleotides can include labeled nucleotides. For example, the label can be a fluorescent label. The presence or absence of the labeled nucleotides incorporated into the primer hybridized to the template polynucleotide can be detected, which allows for the determination of the sequence (e.g., by generating a flow chart). In some embodiments, the labeled nucleotides are labeled with a fluorescent, luminescent, or other light-emitting moiety. In some embodiments, the label is attached to the nucleotide via a linker. In some embodiments, the linker is cleavable, e.g., by a photochemical or chemical cleavage reaction. For example, the label can be cleaved after detection and prior to the incorporation of successive nucleotide(s). In some embodiments, the label (or linker) is attached to the nucleobase or to another site on the nucleotide that does not interfere with the extension of the nascent DNA strand. In some embodiments, the linker contains a disulfide or a PEG-containing moiety.
[0077] In some embodiments, the introduced nucleotides contain only unlabeled nucleotides, and in some embodiments, the nucleotides contain a mixture of labeled and unlabeled nucleotides. For example, in some embodiments, the labeled nucleotide portion is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more or about 0.01% or more compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is about 0.01% to about 100%, such as about 0.01% to about 0.025%, about 0.025% to about 0.05%, about 0.05% to about 0.1%, about 0.1% to about 0.25%, about 0.25% to about 0.5%, about 0.5% to about 1%, about 1% to about 1.5%, about 1.5% to about 2%, about 2% to about 2.5%, about 2.5% to about 3%, about 3% to about 4%, about 4% to about 5%, about 5% to about 10%, about 10% to about 20%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to less than 100% or about 90% to about 100%.
[0078] Prior to sequencing, the polynucleotide can be hybridized to a primer immobilized on a solid support. The primer can hybridize to an adapter region on the 3' and / or 5' end of the polynucleotide. The polynucleotide can be amplified (e.g., by bridge amplification or other amplification techniques) after hybridization to generate polynucleotide sequencing colonies. Colony formation allows signal amplification, so that the detector can accurately detect the incorporation of labeled nucleotides in each colony.
[0079] Other methods for sequencing and / or analyzing sequencing data that can be used in accordance with the methods described herein are described in U.S. Patent Application No. 16 / 864,971; U.S. Patent Application No. 16 / 864,981; International PCT Application No. PCT / US2020 / 031196; the content of each of which is incorporated herein by reference.
[0080] Control of primer extension through barcode region design
[0081] The barcode region is an artificial construct used to label individual mRNA molecules from a biological sample. Thus, in some embodiments, the barcode region is designed to omit the same bases present in the homopolymer region. For example, the barcode region is sequenced using a flow sequencing method that allows the sequencing primer to extend through the barcode region. However, since the barcode region does not contain the bases present in the homopolymer region (e.g., adenine or thymine bases), the bases complementary to the bases in the homopolymer region can be omitted from the flow sequencing cycle order, while the barcode region is sequenced and the primer extends within the barcode region. Once the primer extends to the start of the homopolymer region, nucleotides complementary to the bases present in the homopolymer region (e.g., non-terminating nucleotides) can be used to extend the primer to extend the primer within the homopolymer region. The nucleotides used to extend the primer within the homopolymer region can be unlabeled, labeled, or a mixture of labeled and unlabeled nucleotides (e.g., non-terminating unlabeled nucleotides). In some embodiments, the nucleotides used to extend the primer within the homopolymer region comprise unlabeled nucleotides, such as non-terminating unlabeled nucleotides. In some embodiments, the nucleotides used to extend the primer within the homopolymer region comprise unlabeled nucleotides and labeled nucleotides (e.g., unlabeled non-terminating nucleotides and labeled non-terminating nucleotides), wherein the proportion of labeled nucleotides used to extend the primer within the homopolymer region of the polynucleotide is less than the proportion of labeled nucleotides used to extend the primer within the barcode region and / or the target region of the polynucleotide relative to the total nucleotides.
[0082] In some embodiments, the homopolymer region is a poly-A region, and the barcode region omits adenine bases. In some embodiments, the homopolymer region is a poly-T region, and the barcode region omits thymine bases.
[0083] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, wherein the barcode region omits the identical bases present in the homopolymer region; (b) determining the sequence of the barcode region using labeled nucleotides lacking the complementary bases of the identical bases present in the homopolymer region; (c) extending the primer within the homopolymer region using nucleotides complementary to the bases present in the homopolymer region; and (d) determining the sequence of the target region using labeled nucleotides.
[0084] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, wherein the barcode region omits the identical bases present in the homopolymer region; (b) determining the sequence of the barcode region using labeled non-terminating nucleotides lacking the complementary bases of the identical bases present in the homopolymer region; (c) extending the primer within the homopolymer region using non-terminating nucleotides complementary to the bases present in the homopolymer region; and (d) determining the sequence of the target region using labeled non-terminating nucleotides.
[0085] A sequencing primer hybridizes with a plurality of polynucleotides derived from the mRNA to form a plurality of hybridization templates. The polynucleotide may comprise an adaptor region at the 3' and / or 5' end of the polynucleotide, and the adaptor region may hybridize with the sequencing primer.
[0086] Once a hybridization template is formed, primer extension through the barcode region can be carried out, for example, using a sequencing-by-synthesis method. During sequencing of the barcode region, the primer is extended by contacting the hybridization template with nucleotides (including labeled nucleotides and optionally unlabeled nucleotides) in a defined order. However, the defined order omits the complementary bases of the same base present in the homopolymer region while the primer extends within the barcode region. Nucleotide types (e.g., A, C, G, or T, except for the base complementary to the base in the homopolymer region) can be contacted with the hybridization template discretely. That is, one nucleotide is added at a time in a predetermined order. In some embodiments, two or three different nucleotide types are used simultaneously. A DNA polymerase can be used in a template-dependent manner to incorporate nucleotides at the 3'-end of the primer, thereby extending the primer. Template-dependent extension of the primer occurs if the complementary base (i.e., complementary to the added nucleotide) is present in the template polynucleotide chain at the position opposite the position of the newly incorporated base. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. The process is repeated using the order of nucleotides until the sequence of the barcode region is determined, where the complementary bases of the same base present in the homopolymer region are omitted from the order. In this way, the primer extends through the barcode region but does not extend into the homopolymer region. Thus, primer extension stops before starting in the homopolymer region until a base complementary to the homopolymer region is introduced.
[0087] Since the sequence of the homopolymer region is usually meaningless, primer extension can be carried out within the homopolymer region without detecting the presence or absence of nucleotide incorporation, for example, by using unlabeled nucleotides or by not detecting the label (e.g., fluorescence). Labeled nucleotides are generally more expensive and have an increased risk of primer extension stopping, so labeled nucleotides can be included as a smaller proportion of the total nucleotides (compared to the portion used to determine the sequence of the barcode region) or eliminated during primer extension through the homopolymer region. If primer extension stops within the homopolymer region, unincorporated nucleotides can be removed, for example, by washing the hybridization template, and fresh nucleotides complementary to the bases present in the homopolymer region can be added.
[0088] In some embodiments, unlabeled nucleotides are used to extend the primer within the homopolymer region. In some embodiments, labeled nucleotides are used to extend the primer within the homopolymer region. In some embodiments, labeled and unlabeled nucleotides are used to extend the primer within the homopolymer region. For example, in some embodiments, the labeled nucleotide portion is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more, or about 0.01% or more compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is from about 0.01% to about 100% compared to the total nucleotides, such as from about 0.01% to about 0.025%, from about 0.025% to about 0.05%, from about 0.05% to about 0.1%, from about 0.1% to about 0.25%, from about 0.25% to about 0.5%, from about 0.5% to about 1%, from about 1% to about 1.5%, from about 1.5% to about 2%, from about 2% to about 2.5%, from about 2.5% to about 3%, from about 3% to about 4%, from about 4% to about 5%, from about 5% to about 10%, from about 10% to about 20%, from about 20% to about 30%, from about 30% to about 40%, from about 40% to about 50%, from about 50% to about 60%, from about 60% to about 70%, from about 70% to about 80%, from about 80% to about 90%, from about 90% to less than 100%, or from about 90% to about 100%.
[0089] Instead of sequencing the barcode region, sequencing the target region can involve the ordered use of all four bases (A, T, G, C) without omitting the identical bases present in the homopolymer region, as all four bases may be present in the target region. A polymerase can be used to incorporate nucleotides into the 3' end of a primer in a template-dependent manner, thereby extending the primer from the end of the homopolymer region to the target region. Unincorporated nucleotides can be removed, for example, by washing the hybridized template. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. The process is repeated with the nucleotides in order until the sequence of the target region is determined, thereby allowing the sequence of the region of interest to be inferred from the mRNA molecule.
[0090] Figure 2A A flowchart showing an exemplary method for determining the sequence of a region of interest from an mRNA molecule, where the barcode region omits the identical bases present in the homopolymer region. Figure 2B An exemplary method is shown in graphical form. In step 202, a polynucleotide having a homopolymer region, a barcode region, and a target region is hybridized to a primer to form a hybridized template, where the target region contains a sequence related (e.g., complementary or identical) to the region of interest of the mRNA. Different polynucleotides contain different barcode regions and different target regions such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3' end of the polynucleotide can contain an adapter region, and the primer can hybridize to the adapter region of the polynucleotide. The adapter region is adjacent to the 3' end of the barcode region. As Figure 2B shown, the homopolymer region is located between the target region and the barcode region, where the barcode region is adjacent to the 3' end of the homopolymer region. However, the barcode region omits the identical bases present in the homopolymer region. For example, if the homopolymer region is a poly-A sequence, different barcodes contain different T, C, and / or G sequences but omit A. Similarly, if the homopolymer region is a poly-T sequence, different barcode polynucleotides contain different A, C, and / or G sequences. The polynucleotide can also contain an adapter region, and the primer hybridizes to the adapter region. Thus, during primer extension, the polynucleotide functions as the template strand for primer strand extension.
[0091] In Figure 2A and Figure 2B step 204, nucleotides lacking the complementary bases of the identical bases present in the homopolymer region are used to determine the sequence of the barcode region. For example, the barcode region can be sequenced using a flow sequencing method. At least some of the nucleotides are labeled. In some embodiments, the nucleotides contain labeled nucleotides and unlabeled nucleotides.
[0092] In Figure 2A and Figure 2BIn step 206, the primer is extended within the homopolymer region using nucleotides complementary to the bases present in the homopolymer region. The nucleotides can be labeled, unlabeled, or a mixture of labeled and unlabeled nucleotides. In some embodiments, the presence or absence of nucleotide incorporation during primer extension within the homopolymer region is not detected.
[0093] As Figure 2A and Figure 2B shown in step 208, once the primer has been extended through the homopolymer region, the sequence of the target region is determined using nucleotides, e.g., using a flow sequencing method. At least a portion of the nucleotides are labeled during sequencing. In some embodiments, the nucleotides comprise labeled nucleotides and unlabeled nucleotides. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. The process is repeated using the order of the nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0094] Primer extension within the homopolymer region stops
[0095] In some embodiments of the method for determining the sequence of the target region from an mRNA molecule, primer extension within the homopolymer region is stopped by using a higher proportion of labeled nucleotides compared to the proportion of labeled nucleotides used when determining the sequence of other regions of the polynucleotide (e.g., the barcode region and / or the target region). The nucleotides used to extend the primer can comprise labeled nucleotides or a mixture of labeled and unlabeled nucleotides. As the proportion of labeled nucleotides relative to the total amount of nucleotides (i.e., the sum of labeled and unlabeled nucleotides) increases, the likelihood that the polymerase used to extend the primer stops increases. Thus, primer extension within the homopolymer region can be intentionally stopped using a higher proportion of labeled nucleotides, unincorporated nucleotides can be removed, and primer extension can be continued using unlabeled nucleotides.
[0096] For example, the method for determining the sequence of the target region from an mRNA molecule can include: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides and unlabeled nucleotides at a first proportion of labeled nucleotides relative to the total nucleotides; (c) extending the primer using labeled nucleotides complementary to the bases present in the homopolymer region at a second proportion of labeled nucleotides relative to the total nucleotides, wherein the second proportion is greater than the first proportion, and wherein primer extension stops within the homopolymer region; (d) extending the primer to the end of the homopolymer region using unlabeled nucleotides complementary to the bases present in the homopolymer region; and (e) determining the sequence of the target region using labeled nucleotides.
[0097] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA with primers to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled non-terminating nucleotides and unlabeled non-terminating nucleotides at a first ratio of labeled non-terminating nucleotides to total non-terminating nucleotides; (c) extending the primer using labeled non-terminating nucleotides complementary to the bases present in the homopolymer region at a second ratio of labeled non-terminating nucleotides to total non-terminating nucleotides, wherein the second ratio is greater than the first ratio, and wherein primer extension stops within the homopolymer region; (d) extending the primer to the end of the homopolymer region using unlabeled non-terminating nucleotides complementary to the bases present in the homopolymer region; and (e) determining the sequence of the target region using labeled non-terminating nucleotides.
[0098] In some embodiments, the barcode region is adjacent to the 3' end of the homopolymer region. In some embodiments, there are no intervening bases between the barcode region and the homopolymer region. The primer can then extend through the barcode region and optionally into the homopolymer region before the ratio of labeled nucleotides is increased.
[0099] Once the hybridization templates are formed, e.g., using a flow sequencing method, the primer can extend through the barcode region. During sequencing of the barcode region, the primer is extended by contacting the hybridization templates with nucleotides (comprising labeled nucleotides and, optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted with the hybridization templates discretely, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate the nucleotides into the 3' end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides can be removed, e.g., by washing the hybridization templates. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the barcode region is determined.
[0100] Once the primer extends into the homopolymer region, the proportion of labeled nucleotides used for primer extension can be increased to deliberately stop primer extension. Since an increase in the proportion of labeled nucleotides relative to the total nucleotides increases the likelihood of polymerase stoppage, when primer extension is stopped within the homopolymer region, the proportion of labeled nucleotides relative to the total nucleotides is higher than the proportion of labeled nucleotides used to determine the sequences of the barcode region and the target region. That is, the proportion of labeled nucleotides relative to the total nucleotides used to determine the sequence of the barcode region is lower than the proportion of labeled nucleotides relative to the total nucleotides used for primer extension within the homopolymer region, and thus primer extension stops within the homopolymer region.
[0101] The probability of polymerase stalling depends on the type of polymerase, the structure of the labeled nucleotide (e.g., the structure of the label or the structure of the linker between the label and the nucleotide), and the ratio of the labeled nucleotide used during primer extension to the total nucleotides. Polymerase stalling is generally undesirable during sequence determination (e.g., when determining the sequence of a barcode region or a target region). However, since the sequence of homopolymer regions is usually not important, in some embodiments, the polymerase is intentionally stalled. Thus, the ratio of the labeled nucleotide used to extend the primer within the homopolymer region and thereby stall primer extension within the homopolymer region is generally greater than the ratio of the labeled nucleotide used when determining the sequence (e.g., the sequence of the barcode region and / or the target region). Those skilled in the art can readily select the ratio of the labeled nucleotide for inducing or restricting primer extension stalling. For example, in some embodiments, the labeled nucleotide portion is about 90% or less, about 80% or less, about 70% or less, about 60% or less, about 50% or less, about 40% or less, about 30% or less, about 20% or less, about 10% or less, about 5% or less, about 4% or less, about 3% or less, about 2.5% or less, about 2% or less, about 1.5% or less, about 1% or less, about 0.5% or less, about 0.25% or less, about 0.1% or less, about 0.05% or less, about 0.025% or less, or about 0.01% or less compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is about 50% or more, about 40% or more, about 30% or more, about 20% or more, about 10% or more, about 5% or more, about 4% or more, about 3% or more, about 2.5% or more, about 2% or more, about 1.5% or more, about 1% or more, about 0.5% or more, about 0.25% or more, about 0.1% or more, about 0.05% or more, about 0.025% or more, or about 0.01% or more compared to the total nucleotides. In some embodiments, the labeled nucleotide portion is about 0.01% to about 100%, such as about 0.01% to about 0.025%, about 0.025% to about 0.05%, about 0.05% to about 0.1%, about 0.1% to about 0.25%, about 0.25% to about 0.5%, about 0.5% to about 1%, about 1% to about 1.5%, about 1.5% to about 2%, about 2% to about 2.5%, about 2.5% to about 3%, about 3% to about 4%, about 4% to about 5%, about 5% to about 10%, about 10% to about 20%, about 20% to about 30%, about 30% to about 40%, about 40% to about 50%, about 50% to about 60%, about 60% to about 70%, about 70% to about 80%, about 80% to about 90%, about 90% to less than 100%, or about 90% to about 100%.
[0102] After primer extension has stopped within the homopolymer region, primer extension can be restarted using unlabeled nucleotides that are complementary to the bases present in the homopolymer region. Optionally, labeled and unlabeled nucleotides can be used, although the proportion of labeled nucleotides used to continue primer extension within the homopolymer region after stopping primer extension should be lower than the proportion used when stopping primer extension, and can be further lower than the proportion used when determining the sequence of a defined region (e.g., a barcode region and / or a target region). In some embodiments, no labeled nucleotides are used when extending the primer within the homopolymer region after stopping the primer within the homopolymer region. In some embodiments, less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, less than about 1%, less than about 0.5%, less than about 0.1%, less than about 0.05%, or less than about 0.01% of the total nucleotides used to continue extending the primer within the homopolymer region after primer extension has stopped are labeled nucleotides.
[0103] Although the likelihood of primer extension stopping is reduced by decreasing or eliminating the proportion of labeled nucleotides, primer extension can still stop within the homopolymer region after reducing the proportion of labeled nucleotides relative to the proportion included when intentionally stopping primer extension. To restart primer extension after stopping, unincorporated nucleotides and polymerase can be removed (e.g., by washing the hybridization template), and then the hybridization template can be contacted again with fresh, unlabeled nucleotides and polymerase. This process can be repeated 1, 2, 3, 4, or more times until primer extension proceeds through the homopolymer region.
[0104] Unlabeled nucleotides (or a small fraction of labeled nucleotides) can be used to extend the primer to the end of the homopolymer region to the target region. Since the target region contains sequences related to the region of interest of the mRNA, once the primer has been extended to the target region, sequencing can be restarted, for example, using a flow sequencing method. A polymerase can be used to incorporate labeled nucleotides in a template-dependent manner at the 3' end of the primer, thereby extending the primer from the end of the homopolymer region and into the target region. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the region of interest to be inferred from the mRNA molecule.
[0105] Figure 3A A flowchart of an exemplary method for determining the sequence of a region of interest from an mRNA molecule is shown, where primer extension stops within a homopolymer region. Figure 3BAn exemplary method is shown in graphical form. At step 302, a polynucleotide having a homopolymer region, a barcode region, and a target region hybridizes with a primer to form a hybridization template, where the target region contains a sequence related (e.g., complementary or identical) to a region of the mRNA of interest. Different polynucleotides contain different barcode regions and different target regions such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3' end of the polynucleotide may contain an adapter region, and the primer may hybridize with the adapter region of the polynucleotide. The adapter region may be near the 3' end of the barcode region. As Figure 3B shown, the homopolymer region is located between the target region and the barcode region, where the barcode region is near the 3' end of the homopolymer region.
[0106] At Figure 3A and Figure 3B step 304, the sequence of the barcode region is determined using labeled nucleotides, e.g., using a flow sequencing method. In some embodiments, both labeled and unlabeled nucleotides are used. In some embodiments, the presence or absence of nucleotide incorporation during primer extension within the homopolymer region is not detected.
[0107] At Figure 3A and Figure 3B step 306, the primer is extended within the homopolymer region using labeled nucleotides complementary to the bases present in the homopolymer region, thereby stopping primer extension within the homopolymer region. In some embodiments, both nucleotides contain labeled and unlabeled nucleotides. The proportion of labeled nucleotides used to extend the primer relative to the total nucleotides is greater than the proportion used to extend the primer within the barcode region, thereby stopping primer extension within the homopolymer region.
[0108] At Figure 3A and Figure 3B step 308, after primer extension within the homopolymer region has stopped, primer extension within the homopolymer region is continued using unlabeled nucleotides complementary to the bases present in the homopolymer region. Optionally, labeled nucleotides are also used to extend the primer within the homopolymer region, although the proportion of labeled nucleotides used to continue primer extension within the homopolymer region after stopping primer extension should be lower than the proportion used when stopping primer extension, and may be further lower than the proportion used when determining the sequence of a region (e.g., the barcode region and / or the target region).
[0109] As Figure 3A and Figure 3BAs shown in step 310 therein, once primer extension has passed through the homopolymer region, the sequence of the target region is determined using nucleotides, for example, using a flow sequencing method. At least some of the nucleotides are labeled during sequencing. In some embodiments, the nucleotides include labeled nucleotides and unlabeled nucleotides. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. The process is repeated using the order of the nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0110] Barcode region with a common flow length
[0111] Configure the flow cycles and different barcode regions of the polynucleotide such that the primer extends to the end of the barcode region spanning multiple polynucleotides before extending into the homopolymer region. This does not require the barcodes to have the same physical length (i.e., the same number of bases), only that the flow lengths determined by the barcode sequence and cycle order of the different barcode regions are the same. The flow length of a region is the number of cycles required to extend the primer from the start of the region to the end of the region. In this way, the primer extends to the start of the homopolymer region within the same flow cycle, such that primer extension within the homopolymer region starts simultaneously across the polynucleotides.
[0112] In some embodiments, the barcode region is near the 3' end of the homopolymer region. In some embodiments, there are no intervening bases between the barcode region and the homopolymer region. The last base in the barcode region is typically different from the bases within the homopolymer region.
[0113] The cycles and the order of the cycles can be predetermined before starting the sequencing of the barcode region. However, the cycles used during the sequencing of the barcode region can be the same (i.e., symmetric cycles) or different (i.e., asymmetric cycles). Additionally, a given cycle does not need to include all four different bases (i.e., A, T, C, G), but can include two, three, or four different bases. That is, the cycles can be symmetric cycles or asymmetric cycles. However, as described above, there are no base repeats within the same cycle.
[0114] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, and wherein the target region is associated with a unique barcode region; (b) determining the sequence of the barcode region using labeled nucleotides in a plurality of predetermined cycles, wherein the predetermined cycles and the barcode region of the polynucleotide are configured such that the primer extends to the end of the barcode region spanning a plurality of polynucleotides before extending into the homopolymer region; (c) extending the primer within the homopolymer region using unlabeled nucleotides; and (d) determining the sequence of the target region using labeled nucleotides.
[0115] In some embodiments, there is a method for determining the sequence of a target region from an mRNA molecule, comprising: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule, and wherein the target region is associated with a unique barcode region; (b) determining the sequence of the barcode region using labeled non-terminating nucleotides in a plurality of predetermined cycles, wherein the predetermined cycles and the barcode region of the polynucleotide are configured such that the primer extends to the end of the barcode region spanning a plurality of polynucleotides before extending into the homopolymer region; (c) extending the primer within the homopolymer region using unlabeled non-terminating nucleotides; and (d) determining the sequence of the target region using labeled non-terminating nucleotides.
[0116] Once a hybridization template is formed, primer extension can be carried out through the barcode region, for example, using a flow sequencing method. During sequencing of the barcode region, the primer is extended by contacting the hybridization template with nucleotides (including labeled nucleotides and, in some embodiments, unlabeled nucleotides). Different nucleotide types (e.g., A, C, G, or T) can be discretely contacted with the hybridization template within a given cycle, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate nucleotides into the 3'-end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. The barcode region is sequenced using a plurality of predetermined cycles. The cycles used for sequencing the barcode region can have the same (i.e., symmetric) or different (i.e., asymmetric) orders of nucleotide types, which are optionally used discretely. This process is repeated for the predetermined cycles until the sequence of the barcode region is determined. The predetermined cycles and the sequence of the barcode region are configured such that the primer extends to the end across multiple polynucleotide barcode regions before extending into a homopolymer region. This controls primer extension such that primer extension into the homopolymer region of any given polynucleotide does not occur without primer extension into the homopolymer regions of all polynucleotides in the plurality of polynucleotides.
[0117] Since primer extension is controlled at the end of the barcode region and before extending the primer into a homopolymer region, the proportion of labeled nucleotides used for extending the primer within the homopolymer region can be controlled. For example, primer extension within the homopolymer region can use a different proportion of labeled nucleotides than that used for extending the primer within the barcode region and / or the target region. In some embodiments, unlabeled nucleotides are used for extending the primer within the homopolymer region. In some embodiments, labeled nucleotides are used for extending the primer within the homopolymer region. In some embodiments, both labeled and unlabeled nucleotides are used for extending the primer within the homopolymer region. In some embodiments, no labeled nucleotides are used when extending the primer within the homopolymer region after stopping the primer within the homopolymer region. In some embodiments, less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, less than about 1%, less than about 0.5%, less than about 0.1%, less than about 0.05%, or less than about 0.01% of the total nucleotides are used for extending within the homopolymer region.
[0118] Although the likelihood of primer extension termination is reduced by decreasing or eliminating the proportion of labeled nucleotides, primer extension can still terminate within the homopolymer region after reducing the proportion of labeled nucleotides relative to the proportion included when intentionally terminating primer extension. To restart primer extension after termination, unincorporated nucleotides and polymerase can be removed (e.g., by washing the hybridization template), and then the hybridization template can be contacted again with fresh, unlabeled nucleotides and polymerase. This process can be repeated 1, 2, 3, 4, or more times until primer extension proceeds through the homopolymer region.
[0119] The primer can be extended to the end of the homopolymer region and the start of the target region. Since the target region contains a sequence related to the region of the mRNA of interest, the sequence of the target region is determined, for example, using a flow sequencing method. A polymerase can be used to incorporate labeled nucleotides into the 3' end of the primer in a template-dependent manner, thereby extending the primer from the end of the homopolymer region to the target region. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. This process is repeated using the order of the nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the region of interest to be inferred from the mRNA molecule.
[0120] Figure 4A A flowchart showing an exemplary method for determining the sequence of a region of interest from an mRNA molecule is shown, where the flow cycle and different barcode regions of the polynucleotide are configured such that the primer extends to the end of the barcode region spanning multiple polynucleotides before extending into the homopolymer region. Figure 4B An exemplary method is shown in graphical form, where multiple polynucleotides have sequences of different barcode regions but the same flow length. In step 402, a polynucleotide having a homopolymer region, a barcode region, and a target region is hybridized with a primer to form a hybridization template, the target region containing a sequence related (e.g., complementary or identical) to the region of the mRNA of interest. Different polynucleotides contain different barcode regions and different target regions such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3' end of the polynucleotide can contain an adapter region, and the primer can hybridize with the adapter region of the polynucleotide. The adapter region is adjacent to the 3' end of the barcode region. As Figure 4B shown in 402, the homopolymer region is located between the target region and the barcode region, where the barcode region is adjacent to the 3' end of the homopolymer region.
[0121] In Figure 4AIn step 404, for example using a flow sequencing method, the sequence of the barcode region is determined using labeled nucleotides. In some embodiments, labeled and unlabeled nucleotides are used. The polynucleotide is configured with a predetermined cycle and different barcode regions such that the primer extends to the end of the barcode region spanning multiple polynucleotides before extending into the homopolymer region.
[0122] Returning to Figure 4B , in 404a, the primer hybridized to the polynucleotide is extended within the barcode region by contacting the hybridized template with different types of labeled nucleotides, optionally using these nucleotide types discretely. When a complementary base is present in the template strand, the polymerase extends the primer in a template-dependent manner. Unincorporated nucleotides are removed, and the presence or absence of incorporated nucleotides is detected to obtain sequencing information. The process is repeated using the predetermined cycle to continue primer extension until the primer extends to the starting point of the homopolymer region. As Figure 4B shown in 404b, before the primer extends into the homopolymer region, the primer extends through the barcode region spanning multiple polynucleotides.
[0123] The primer does not extend within the homopolymer region spanning the polynucleotides until the primer has extended through the different barcode regions. As Figure 4A and Figure 4B shown in 406, however, once the primer has extended through the barcode region, nucleotides can be used to extend the primer within the homopolymer region. The nucleotides can be labeled, unlabeled, or a mixture of labeled and unlabeled nucleotides. Since the sequence of the homopolymer region is usually meaningless, the primer can be extended within the homopolymer region without detecting the presence or absence of nucleotide incorporation, for example by using unlabeled nucleotides or by not detecting the label (e.g., fluorescence). Labeled nucleotides are generally more expensive and have an increased risk of primer extension termination, so labeled nucleotides can be included as a smaller proportion of the total nucleotides (compared to the portion used to determine the sequence of the barcode region) or eliminated during primer extension through the homopolymer region. If primer extension stops within the homopolymer region, unincorporated nucleotides can be removed, for example by washing the hybridized template, and fresh nucleotides complementary to the bases present in the homopolymer region can be added.
[0124] As Figure 4A and Figure 4B shown in step 410, once the primer has extended through the homopolymer region, the sequence of the target region is determined using nucleotides, for example using a flow sequencing method. At least a portion of the nucleotides are labeled during sequencing. In some embodiments, the nucleotides include labeled nucleotides and unlabeled nucleotides. Then the presence or absence of incorporated labeled nucleotides is detected to determine the sequence at that position. The process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0125] Dual primer sequencing
[0126] In some embodiments, two primers are used to determine the sequence of a target region from an mRNA molecule, where the first primer is extended to sequence a barcode region and the second primer is extended to sequence a target region. The second primer hybridizes within a homopolymer region of the polynucleotide to avoid or limit primer extension within the homopolymer region. The order of primer hybridization and extension is reversible. That is, extension of the first primer to determine the sequence of the barcode region can occur before or after extension of the second primer to determine the sequence of the target region. Thus, the descriptions of "first primer" and "second primer" provided below are for clarity only, and the sequences of the primers can be reversed.
[0127] The second primer contains a homopolymer region that contains bases complementary to the bases in the homopolymer of the polynucleotide containing the target region. For example, if the homopolymer region of the polynucleotide containing the target region is a poly-T region, the second primer can contain a poly-A region, or if the homopolymer region of the polynucleotide containing the target region is a poly-A region, the second primer can contain a poly-T region. This allows the second primer to hybridize efficiently to the homopolymer region of the polynucleotide. The homopolymer region of the primer can be longer, shorter, or approximately the same length as the homopolymer region of the polynucleotide.
[0128] In some embodiments, the 3' end of the second primer does not hybridize to the 5' end of the homopolymer region of the polynucleotide containing the target region spanning multiple polynucleotides. Thus, the primer hybridizing to the homopolymer region can be extended within the homopolymer region, aligning the 3' end of the primer with the 5' end of the homopolymer region of the polynucleotide. Primer extension through the homopolymer region can be carried out using labeled, unlabeled, or a mixture of labeled nucleotides. If primer extension stops within the homopolymer region, unincorporated nucleotides can be removed, for example, by washing the hybridization template, and fresh nucleotides complementary to the bases present in the homopolymer region can be added. Once the primer has been extended through the homopolymer region of the polynucleotide, the target region of the polynucleotide can be sequenced by further extending the primer within the target region.
[0129] In some embodiments, the second primer comprises a homopolymer region and a 3'-anchor, the homopolymer region comprising bases complementary to the bases in the homopolymer of the polynucleotide comprising the target region. The 3'-anchor increases the likelihood of hybridization of the 3'-end of the homopolymer region of the second primer to the 5'-end of the homopolymer region of the polynucleotide. For example, the 3'-anchor can comprise bases other than the bases present in the homopolymer region of the second primer. Then the variable base can hybridize to the first base after the 5'-end of the homopolymer region of the polynucleotide (i.e., the first base of the target region). For example, if the homopolymer region of the polynucleotide is a poly-A sequence, the second primer can comprise a 5'-(poly-T)-V-3' sequence, where V is any base other than T (i.e., G, C, or A).
[0130] In some embodiments, the second primer comprises a homopolymer region and a 5'-anchor, the homopolymer region comprising bases complementary to the bases in the homopolymer of the polynucleotide comprising the target region. Optionally, the second primer can comprise a 3'-anchor and a 5'-anchor. For example, the 5'-anchor can comprise a polynucleotide sequence complementary to an adaptor sequence on the polynucleotide (e.g., it can be located near the 3'-end of the barcode region). By the 5'-anchor, the homopolymer region of the second primer is more likely to hybridize closer to the 3'-end of the homopolymer region. The homopolymer region and the 5'-anchor can be linked by a linker such as a PEG phosphoramidite linker or an abasic deoxyribose or abasic ribose linker. For example, the length of the linker can be approximately the same length as the barcode region, longer than the barcode region, or shorter than the barcode region.
[0131] The first primer can hybridize to the polynucleotide comprising the barcode region and the target region at a position near the 3'-end of the barcode region. For example, the polynucleotide can comprise a 3'-adaptor region, and the first primer can hybridize to the adaptor region. The first primer extends within the barcode region to determine the sequence of the barcode region.
[0132] In some embodiments, the first primer is removed from the hybridization template before hybridizing the second primer to the polynucleotide. For example, a high pH solution such as a sodium hydroxide solution can be used to remove the primer. In some embodiments, the first primer is not removed from the hybridization template before hybridizing the second primer to the polynucleotide.
[0133] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a first primer to form a first plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled nucleotides; (c) hybridizing the plurality of polynucleotides to a second primer to form a second plurality of hybridization templates, wherein the second primer comprises a homopolymer region comprising a plurality of consecutive and identical bases complementary to the bases in the homopolymer region of the polynucleotide; and (d) determining the sequence of the target region using labeled nucleotides. In some embodiments, the second primer is extended within the homopolymer region using nucleotides (which may be labeled, unlabeled, or a mixture thereof). In some embodiments, the first primer is hybridized to the polynucleotide before the second primer is hybridized to the polynucleotide. In some embodiments, the second primer is hybridized to the polynucleotide before the first primer is hybridized to the polynucleotide.
[0134] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a first primer to form a first plurality of hybridization templates, the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; (b) determining the sequence of the barcode region using labeled non-terminating nucleotides; (c) hybridizing a plurality of polynucleotides to a second primer to form a second plurality of hybridization templates, wherein the second primer comprises a homopolymer region comprising a plurality of consecutive and identical bases complementary to the bases in the homopolymer region of the polynucleotide; and (d) determining the sequence of the target region using labeled non-terminating nucleotides. In some embodiments, the second primer is extended within the homopolymer region using non-terminating nucleotides (which may be labeled, unlabeled, or a mixture thereof). In some embodiments, the first primer is hybridized to the polynucleotide before the second primer is hybridized to the polynucleotide. In some embodiments, the second primer is hybridized to the polynucleotide before the first primer is hybridized to the polynucleotide.
[0135] Figure 5A A flowchart showing an exemplary method for using two primers to determine the sequence of a target region from an mRNA molecule. Figure 5BAn exemplary method is shown in graphical form. In step 502, a polynucleotide having a homopolymer region, a barcode region, and a target region containing a sequence related to (e.g., complementary or identical to) the target mRNA region is hybridized with a first primer to form a first hybridization template. Different polynucleotides contain different barcode regions and different target regions such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3' end of the polynucleotide may contain an adaptor region, and the primer may hybridize to the adaptor region of the polynucleotide. The adaptor region may be near the 3' end of the barcode region. As Figure 5B shown, the homopolymer region is located between the target region and the barcode region, where the barcode region is near the 3' end of the homopolymer region.
[0136] In Figure 5A and Figure 5B step 506, for example using a flow sequencing method, the sequence of the barcode region is determined by extending the first primer using labeled nucleotides. The labeled nucleotides are optionally combined with unlabeled nucleotides. During sequencing of the barcode region, the primer is extended by contacting the hybridization template with nucleotides (including labeled nucleotides and optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) may contact the hybridization template discretely, or two or three different nucleotide types may be used simultaneously. A polymerase may be used to incorporate nucleotides into the 3' end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides may be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the barcode region is determined.
[0137] In Figure 5A and Figure 5B step 506, the polynucleotide is hybridized with a second primer containing a homopolymer region having bases complementary to the bases present in the homopolymer region of the polynucleotide, thereby forming a second hybridization template. Optionally, the first primer is removed before hybridizing the second primer to the polynucleotide. In some embodiments, as shown in step 508, the second primer is extended within the homopolymer region such that the 3' end of the primer is located at the 5' end of the homopolymer region of the polynucleotide. For example, this can be accomplished by template-dependent extension of the primer. For example, using a polymerase, unlabeled and / or labeled nucleotides complementary to the bases present in the homopolymer region may be contacted with the second hybridization template and incorporated at the end of the second primer, thereby extending the second primer. Step 508 is not always required because, for example, by using a 3' anchor, the second primer can be configured such that the 3' end of the primer is located at the 5' end of the homopolymer region of the polynucleotide upon hybridization.
[0138] As Figure 5A and 5BAs shown in step 510 therein, for example, using the flow sequencing method described herein, the second primer is extended to the target region using labeled (and, optionally, unlabeled) nucleotides to determine the sequence of the target region. During sequencing of the barcode region, the primer is extended by contacting the hybridization template with nucleotides (including labeled nucleotides and optionally unlabeled nucleotides) in a defined order. The nucleotide types (e.g., A, C, G, or T) are optionally contacted discretely with the second hybridization template, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate the nucleotides into the 3' end of the second primer in a template-dependent manner to extend the second primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0139] Figure 5C FIG. shows a flow chart of another exemplary method for determining the sequence of a target region from an mRNA molecule using two primers, wherein compared to the Figure 5A exemplary method shown, the sequences of the two primers hybridized to the polynucleotide are opposite. Figure 5D is shown graphically in an exemplary method. In Figure 5C and Figure 5D In step 512, a polynucleotide having a homopolymer region, a barcode region, and a target region (the polynucleotide) containing a sequence related to (e.g., complementary or identical to) the target mRNA region hybridizes to the first primer to form a hybridization template. The first primer contains a homopolymer region having bases complementary to the bases present in the homopolymer region of the polynucleotide. In some embodiments, as shown in step 514, the second primer is extended within the homopolymer region such that the 3' end of the primer is located at the 5' end of the homopolymer region of the polynucleotide. For example, this can be accomplished by template-dependent extension of the primer. For example, using a polymerase, unlabeled and / or labeled nucleotides complementary to the bases present in the homopolymer region can be contacted with the second hybridization template and incorporated into the end of the second primer to extend the first primer.
[0140] As Figure 5C and Figure 5D As shown in step 516 therein, for example, using a sequencing-by-synthesis method, labeled (and optionally, unlabeled) nucleotides are used to extend the first primer to the target region, thereby determining the sequence of the target region. During the sequencing of the target region, the first primer is extended by contacting the first hybridization template with nucleotides (including labeled nucleotides and optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted with the first hybridization template discretely, or two or more different nucleotide types can be used simultaneously. A polymerase can be used to incorporate nucleotides into the 3'-end of the first primer in a template-dependent manner, thereby extending the first primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0141] In Figure 5C and Figure 5D step 518 therein, the polynucleotide hybridizes with the second primer to form a second hybridization template. Optionally, the extended first primer is removed before the second primer hybridizes with the polynucleotide. Different polynucleotides contain different barcode regions and different target regions, such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3'-end of the polynucleotide can contain an adapter region, and the primer can hybridize with the adapter region of the polynucleotide. The adapter region can be adjacent to the 3'-end of the barcode region. As Figure 5D shown, a homopolymer region is located between the target region and the barcode region, where the barcode region is adjacent to the 3'-end of the homopolymer region.
[0142] In Figure 5C and Figure 5D step 520, for example, using a sequencing-by-synthesis method, the sequence of the barcode region is determined by extending the second primer using labeled nucleotides. The labeled nucleotides are optionally combined with unlabeled nucleotides. During the sequencing of the barcode region, the second primer is extended by contacting the second hybridization template with nucleotides (including labeled nucleotides and optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted with the hybridization template discretely, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate nucleotides into the 3'-end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the barcode region is determined.
[0143] Primer with a cleavable linker
[0144] In some embodiments, primers with cleavable linkers are used to determine the sequence of a target region from an mRNA molecule. The primer comprises a first primer segment, a second primer segment, the second primer segment comprising a homopolymer region that comprises a plurality of bases complementary to bases in a homopolymer region of a polynucleotide, and a cleavable linker located between the first primer segment and the second primer segment. The first primer segment is near the 3' end of the primer, and the second primer segment is near the 5' end of the primer.
[0145] For example, using a flow sequencing method, the second primer segment can be extended to sequence a target region of a polynucleotide. As further discussed herein, in some embodiments, the second primer segment is extended to the end of the homopolymer region of the polynucleotide before the primer is extended to the target region and the sequence of the primer region is determined. The cleavable linker can then be cleaved, and the first segment of the primer can be extended, for example using a flow sequencing method, to sequence a barcode region of the polynucleotide.
[0146] The second primer segment comprises a homopolymer region that comprises bases complementary to bases in the homopolymer of the polynucleotide that comprises the target region. For example, if the homopolymer region of the polynucleotide that comprises the target region is a poly-T region, the second primer can comprise a poly-A region, or if the homopolymer region of the polynucleotide that comprises the target region is a poly-A region, the second primer can comprise a poly-T region. This allows the second primer to hybridize efficiently to the homopolymer region of the polynucleotide. The homopolymer region of the primer can be longer, shorter, or approximately the same length as the homopolymer region of the polynucleotide.
[0147] In some embodiments, the 3' end of the second primer segment does not hybridize to the 5' end of the homopolymer region of the polynucleotide that comprises a target region spanning multiple polynucleotides. The second primer segment can be extended within the homopolymer region such that the 3' end of the primer is aligned with the 5' end of the homopolymer region of the polynucleotide. The second primer segment can be extended through the homopolymer region using labeled, unlabeled, or a mixture of labeled nucleotides. If primer extension stops within the homopolymer region, unincorporated nucleotides can be removed, for example by washing the hybridization template, and fresh nucleotides complementary to the bases present in the homopolymer region can be added. Once the second primer segment has been extended through the homopolymer region of the polynucleotide, the target region of the polynucleotide can be sequenced by further extending the primer within the target region.
[0148] In some embodiments, the second primer segment comprises a homopolymer region and a 3' anchor, where the homopolymer region comprises bases complementary to the bases in the homopolymer of the polynucleotide comprising the target region. The 3' anchor increases the likelihood of hybridization of the 3' end of the homopolymer region of the second primer segment to the 5' end of the homopolymer region of the polynucleotide. For example, the 3' anchor can comprise bases other than those present in the homopolymer region of the second primer segment. The variable base can then hybridize to the first base after the 5' end of the homopolymer region of the polynucleotide (i.e., the first base of the target region). For example, if the homopolymer region of the polynucleotide is a poly-A sequence, the second primer can comprise a 5'-(poly-T)-V-3' sequence, where V is any base other than T (i.e., G, C, or A).
[0149] The cleavable linker can be any suitable linker that covalently links the first primer segment and the second primer segment, and its cleavage can be controlled. In some embodiments, the cleavable linker comprises a polynucleotide sequence containing uracil (U) bases. A uracil-specific cleavage enzyme (e.g., the uracil-specific excision reagent available from New England BioLabs
[0150] enzyme) can be used to controllably cleave the uracil bases. Since uracil bases are unique to RNA, there is no risk of cleaving DNA polynucleotides or primers.
[0151] The first primer segment is configured to hybridize to a sequence near the 5' end of the barcode region (e.g., an adapter sequence). Once the primer is cleaved, e.g., by using a flow sequencing method, the 3' end of the first primer segment can be extended to determine the sequence of the barcode region.
[0152] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule comprises: (a) hybridizing a plurality of polynucleotides derived from mRNA to primers to form a plurality of hybridization templates; the polynucleotides comprise a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; wherein the primer comprises a first primer segment, a second primer segment, the second primer segment comprising a homopolymer region that comprises a plurality of bases complementary to the bases in the homopolymer region of the polynucleotide, and a cleavable linker between the first primer segment and the second primer segment; (b) determining the sequence of the target region using labeled nucleotides; (c) cleaving the primer at the cleavable linker; and (d) determining the sequence of the barcode region using labeled nucleotides.
[0153] In some embodiments, a method of determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates; the polynucleotides comprising a barcode region, a homopolymer region comprising a plurality of consecutive and identical bases, and a target region comprising a sequence related to the target region from the mRNA molecule; wherein the primer comprises a first primer segment, a second primer segment, the second primer segment comprising a homopolymer region comprising a plurality of bases complementary to the bases in the homopolymer region of the polynucleotide, and a cleavable linker between the first primer segment and the second primer segment; (b) determining the sequence of the target region using labeled non-terminating nucleotides; (c) cleaving the primer at the cleavable linker; and (d) determining the sequence of the barcode region using labeled non-terminating nucleotides.
[0154] Figure 6A A flow chart showing an exemplary method for determining the sequence of a target region from an mRNA molecule using a cleavable primer is shown. Figure 6B An exemplary method is shown in graphical form. In step 602, a polynucleotide having a homopolymer region, a barcode region, and a target region is hybridized with a primer to form a hybridization template, the target region comprising a sequence related (e.g., complementary or identical) to the target mRNA region. The primer comprises a first primer segment, a second primer segment, the second primer segment comprising a homopolymer region comprising a plurality of bases complementary to the bases in the homopolymer region of the polynucleotide, and a cleavable linker located between the first primer segment and the second primer segment. Different polynucleotides comprise different barcode regions and different target regions such that different target regions across the polynucleotides are associated with unique molecular identifiers. The 3' end of the polynucleotide may comprise an adapter region, and the first primer segment may hybridize to the adapter region of the polynucleotide. The adapter region is adjacent to the 3' end of the barcode region. As Figure 6B shown, the homopolymer region is located between the target region and the barcode region, where the barcode region is adjacent to the 3' end of the homopolymer region.
[0155] In some embodiments, as shown in step 604, the second primer segment extends within the homopolymer region such that the 3' end of the primer is located at the 5' end of the homopolymer region of the polynucleotide. For example, this can be accomplished by template-dependent extension of the primer. For example, using a polymerase, unlabeled and / or labeled nucleotides complementary to the bases present in the homopolymer region can be contacted with the second hybridization template and incorporated at the end of the second primer to extend the second primer. Step 504 is not always required because, for example, the second primer can be configured using a 3' anchor such that the 3' end of the primer is located at the 5' end of the homopolymer region of the polynucleotide upon hybridization.
[0156] AsFigure 6A and Figure 6A As shown in step 606 of Figure 6A and Figure 6A , for example, using a flow sequencing method, primers are extended to the target region using labeled (and optionally, unlabeled) nucleotides to determine the sequence of the target region. During sequencing of the target region, the first primer is extended by contacting the hybridization template with nucleotides (including labeled nucleotides and optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted discretely with the first hybridization template, or two or three different nucleotide types can be used simultaneously. A polymerase can be used in a template-dependent manner to incorporate nucleotides at the 3' end of the primer, thereby extending the first primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0157] In Figure 6A and Figure 5B In step 608 of Figure 6A and Figure 5B , the primer is cleaved at the cleavable linker. The cleavable linker can be contacted with a suitable cleavage reagent, which releases the 3' end of the first primer segment. Optionally, the second primer segment is removed, for example, using a high pH solution (e.g., a solution containing sodium hydroxide).
[0158] In Figure 6A and Figure 6B In step 610 of Figure 6A and Figure 6B , for example, using a flow sequencing method, the sequence of the barcode region is determined by extending the first primer segment using labeled nucleotides. The labeled nucleotides are optionally combined with unlabeled nucleotides. During sequencing of the barcode region, the second primer is extended by contacting the second hybridization template with nucleotides (including labeled nucleotides and optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted discretely with the hybridization template, or two or three different nucleotide types can be used simultaneously. A polymerase can be used in a template-dependent manner to incorporate nucleotides at the 3' end of the primer, thereby extending the primer. Unincorporated nucleotides can be removed, for example, by washing the hybridization template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the barcode region is determined.
[0159] Restarting a stopped primer extension
[0160] Primer extension within a long homopolymer region results in the cessation of primer extension. As the ratio of labeled nucleotides to total nucleotides increases, primer extension is more likely to stop because polymerases are generally more suited to incorporate natural or near-natural nucleotides. However, even when given a sufficiently long homopolymer region (such as the poly-A region of some mRNA molecules), primer extension is prone to stopping when using natural nucleotides for primer extension.
[0161] Removing unincorporated nucleotides and contacting the hybridized template with fresh nucleotides can restart primer extension. In some cases, the hybridized template is contacted with fresh nucleotides. This process can be repeated 1, 2, 3, 4, 5 or more times to ensure that the primer is fully extended through the homopolymer region. This strategy can be used in combination with other strategies described herein, where primer extension may stop during primer extension within a homopolymer region. For example, when primer extension is intentionally stopped within a homopolymer region (e.g., by increasing the ratio of labeled nucleotides used to extend the primer), or when using a dual-primer or cleavable primer system, when the polynucleotide has a barcode region, primer extension can stop and restart within the homopolymer region (e.g., omitting the barcode region of bases present in the homopolymer region, a barcode region having a common flow length with other barcode regions in multiple polynucleotides).
[0162] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridized templates, wherein the polynucleotides comprise a homopolymer region and a target region, the homopolymer region comprises a plurality of consecutive and identical bases, and the target region comprises a sequence related to the target region from the mRNA molecule; (b) extending the primer to the homopolymer region using nucleotides of the same base, wherein the primer stops within the homopolymer region, and removing unincorporated nucleotides; (c) repeating step (b) one or more times to extend the primer through the homopolymer region; and (d) determining the sequence of the target region using labeled nucleotides.
[0163] In some embodiments, a method for determining the sequence of a target region from an mRNA molecule includes: (a) hybridizing a plurality of polynucleotides derived from the mRNA to a primer to form a plurality of hybridized templates, wherein the polynucleotides comprise a homopolymer region and a target region, the homopolymer region comprises a plurality of consecutive and identical bases, and the target region comprises a sequence related to the target region from the mRNA molecule; (b) extending the primer to the homopolymer region using non-terminating nucleotides of the same base, wherein the primer stops within the homopolymer region, and removing unincorporated non-terminating nucleotides; (c) repeating step (b) one or more times to extend the primer through the homopolymer region; and (d) determining the sequence of the target region using labeled non-terminating nucleotides.
[0164] Figure 7A Shows a flow chart of an exemplary method for determining the sequence of a target region from an mRNA molecule, where primer extension stops and restarts within a homopolymer region of a polynucleotide. Figure 7B An exemplary method is shown in graphical form. Although the exemplary method described below refers to a polynucleotide having a homopolymer region and a target region, the polynucleotide can also contain additional regions, such as a barcode region near the 3' end of the homopolymer region.
[0165] In step 702, a polynucleotide having a homopolymer region and a target region is hybridized with a primer to form a hybridization template, the target region containing a sequence related (e.g., complementary or identical) to the target mRNA region (and optionally, a barcode region, Figure 7B not shown). The 3' end of the polynucleotide can contain an adaptor region, and the primer can hybridize to the adaptor region of the polynucleotide. If present, the adaptor region can be near the 3' end of the barcode region or near the 3' end of the homopolymer region.
[0166] If a barcode region is present, the sequence of the barcode region can be determined using nucleotides, e.g., using a flow sequencing method. During sequencing of the barcode region, the primer is extended by contacting the hybridization template with nucleotides (comprising labeled nucleotides and, optionally, unlabeled nucleotides) in a defined order. Nucleotide types (e.g., A, C, G, or T) can be contacted discretely with the hybridization template, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate the nucleotides into the 3' end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides can be removed, e.g., by washing the hybridization template. The presence or absence of the incorporated labeled nucleotides is then detected to determine the sequence at that position. This process is repeated using the order of nucleotides until the sequence of the barcode region is determined.
[0167] In Figure 7A and Figure 7B step 704, the primer is extended within the homopolymer region using nucleotides complementary to the bases present in the homopolymer region. In some embodiments, the nucleotides comprise labeled nucleotides, unlabeled nucleotides, or labeled and unlabeled nucleotides. As described above, as the proportion of labeled nucleotides increases, the likelihood of primer extension stopping increases. However, even when the nucleotides are completely unlabeled, primer extension can still stop when extended within a longer homopolymer region.
[0168] In Figure 7A step 706, unincorporated nucleotides are removed. For example, the hybridized template can be washed using a wash buffer. This step effectively "resets" the hybridization template such that primer extension can be restarted after stopping.
[0169] In Figure 7A and Figure 7B step 708, primer extension within the homopolymer region is restarted by repeating steps 704 and 706. Nucleotides complementary to the bases present in the homopolymer region are used to further extend the primer within the homopolymer region. The nucleotides can be unlabeled, labeled, or a mixture of labeled and unlabeled nucleotides. In some embodiments, when extending the primer within the homopolymer region after stopping the primer within the homopolymer region, labeled nucleotides are not used. In some embodiments, less than about 5%, less than about 4%, less than about 3%, less than about 2%, less than about 1%, less than about 1%, less than about 0.5%, less than about 0.1%, less than about 0.05%, or less than about 0.01% of the total nucleotides used to continue extending the primer within the homopolymer region after primer extension has stopped are labeled nucleotides.
[0170] Steps 704 and 706 can be repeated 1, 2, 3, 4, 5, or more times to extend the primer through the homopolymer region. For the method to be effective, primer extension does not need to actually stop or pause on all of the polynucleotides in a plurality of polynucleotides. Unincorporated nucleotides can be removed, and fresh nucleotides can be used to restart the polymerase or pre-empt stall after a pause. For example, the removal of unincorporated nucleotides and further primer extension using fresh nucleotides can be repeated periodically at regular or irregular intervals, regardless of the actual number of pauses, such as every 15 seconds, every 30 seconds, every 60 seconds, every 2 minutes, every 5 minutes, etc.
[0171] As Figure 7A and Figure 7B shown in step 710, once the primer has been extended through the homopolymer region, the sequence of the target region is determined using nucleotides, for example using a flow sequencing method. At least a portion of the nucleotides are labeled during sequencing. In some embodiments, the nucleotides comprise labeled nucleotides and unlabeled nucleotides. During sequencing of the barcode region, the primer is extended by contacting the hybridized template with nucleotides (comprising labeled nucleotides and, optionally, unlabeled nucleotides) in a defined order. Different nucleotide types (e.g., A, C, G, or T) can be contacted with the hybridized template discretely, or two or three different nucleotide types can be used simultaneously. A polymerase can be used to incorporate the nucleotides into the 3' end of the primer in a template-dependent manner, thereby extending the primer. Unincorporated nucleotides can be removed, for example, by washing the hybridized template. Then the presence or absence of the incorporated labeled nucleotides is detected to determine the sequence at that position. The process is repeated using the order of nucleotides until the sequence of the target region is determined, thereby allowing the sequence of the target region to be inferred from the mRNA molecule.
[0172] Methods for resuming stopped sequencing primer extension, including additional methods within homopolymer regions and non-homopolymer regions, are described in International PCT Application No. PCT / US2020 / 031196, the content of which is incorporated herein by reference for all purposes. Examples
[0173] The present application can be better understood by reference to the following non-limiting examples provided as exemplary embodiments of the present application. The following examples are provided to more fully illustrate the embodiments, however, and should in no way be construed as limiting the broad scope of the present application. Although certain embodiments of the present application have been shown and described herein, it is apparent that these embodiments are provided by way of example only. Many variations, changes, and substitutions may occur to those skilled in the art without departing from the spirit and scope of the invention. It should be understood that various alternatives to the embodiments described herein may be employed when practicing the methods described herein.
[0174] Example 1
[0175] Lyse sample cells using lysis buffer and separate the liquid lysate by centrifugation. Incubate the separated liquid lysate with capture beads, each capture bead containing a DNA oligomer that contains a barcode region (containing a common sample barcode and UMI) that is fused to the 5' of a poly-T region containing 10 - 30 consecutive thymine bases. Use the capture beads to separate mRNA molecules in the liquid lysate by separating and washing the capture beads. Reverse transcribe the mRNA molecules using a capture probe DNA oligomer and reverse transcriptase to generate a cDNA library. Each cDNA molecule in the cDNA library contains a barcode region with a sample barcode and UMI, a homopolymer poly-T region, and a target region complementary to the mRNA region of interest (which contains at least a portion of the coding region). The barcode region is specifically designed to omit adenine bases.
[0176] Prepare the cDNA library for sequencing by ligating adapter sequences to the 5' and 3' ends of the cDNA polynucleotide to generate a sequencing library. Apply the sequencing library to a sequencing array surface that contains DNA oligonucleotides attached to the surface. The DNA oligonucleotides contain sequences that hybridize to the adapter regions of the cDNA molecules. Sequencing colonies are formed by bridge amplification, which generates copies of the cDNA molecules and the complements of the cDNA molecules.
[0177] A sequencing primer is applied to a sequencing surface and hybridizes to an adapter region at the 3' end of the cDNA molecule complement. A DNA polymerase is also applied to the sequencing surface and the DNA polymerase binds to the hybridized template. A first solution containing a first nucleotide (e.g., deoxy-A, deoxy-G, or deoxy-C), such as a non-terminating nucleotide, is used and the sequencing surface is washed with a wash buffer to remove unincorporated nucleotides. The nucleotides contain approximately 2.5% fluorescently labeled nucleotides and approximately 97.5% unlabeled nucleotides. The presence or absence of base incorporation across the colonies is detected using a fluorescence detector. The process is repeated using a second solution and a third solution, each containing a different (i.e., second and third) nucleotide to complete a flow cycle, and the flow cycle is repeated to sequence the barcode region. The nucleotides in the second and third solutions contain approximately 2.5% fluorescent label and approximately 97.5% unlabeled nucleotides. During this process, thymine bases are not included in the cycle. After each imaging step, the fluorescent label is cleaved from the growing polynucleotide.
[0178] Once the sequencing primer extends through the barcode region, a fourth solution containing 100% unlabeled thymidine triphosphate is applied to the sequencing surface, which initiates primer extension through the homopolymer region. The surface is allowed to react for a period of time before washing the surface with a wash buffer and the process is repeated to eliminate stopped primer extensions. This allows the primer to extend through the poly-A homopolymer region and extend to the start of the target region of the cDNA molecule complement.
[0179] The first, second, and third solutions are used to sequence the target region of the target region, except for a fifth solution containing thymine nucleotides (approximately 2.5% fluorescently labeled and approximately 97.5% unlabeled). The solutions are applied to the sequencing surface separately, the surface is washed, and the presence or absence of base incorporation is detected before applying the next solution in the cycle, for a series of cycles. After each imaging step, the fluorescent label is cleaved from the growing polynucleotide.
[0180] Example 2
[0181] Lyse the sample cells using a lysis buffer and separate the liquid lysate by centrifugation. Incubate the separated liquid lysate with capture beads, each of which contains a DNA oligomer that contains a barcode region (containing a common sample barcode and UMI) that is fused to the 5' of a poly-T region containing 10 - 30 consecutive thymine bases. Use the capture beads to isolate mRNA molecules in the liquid lysate by separating and washing the capture beads. Reverse transcribe the mRNA molecules using a capture probe DNA oligomer and reverse transcriptase to generate a cDNA library. Each cDNA molecule in the cDNA library contains a barcode region with a sample barcode and UMI, a homopolymer poly-T region, and a target region complementary to the region of the mRNA of interest (which contains at least a portion of the coding region).
[0182] Prepare the cDNA library for sequencing by ligating adapter sequences to the 5' and 3' ends of the cDNA polynucleotides to generate a sequencing library. Apply the sequencing library to the surface of a sequencing array that contains DNA oligonucleotides attached to the surface. The DNA oligonucleotides contain sequences that hybridize to the adapter regions of the cDNA molecules. Sequencing colonies are formed by bridge amplification, which generates copies of the cDNA molecules and the complements of the cDNA molecules.
[0183] Apply a sequencing primer to the sequencing surface, which hybridizes to the adapter region at the 3' end of the cDNA molecule complement. DNA polymerase is also applied to the sequencing surface. Use a first solution containing a first nucleotide (e.g., deoxy-A, deoxy-G, deoxy-C, or deoxy-T), e.g., a non-terminating nucleotide, and wash the sequencing surface with a wash buffer to remove unincorporated nucleotides. The nucleotides contain approximately 2.5% fluorescently labeled nucleotides and approximately 97.5% unlabeled nucleotides. Detect the presence or absence of base incorporation across the colonies using a fluorescence detector. Repeat the process using second, third, and fourth solutions, each containing a different (i.e., second, third, and fourth) nucleotide to complete a flow cycle, and repeat the flow cycle to sequence the barcode region. The nucleotides in the second and third solutions contain approximately 2.5% fluorescent label and approximately 97.5% unlabeled nucleotides. After each imaging step, cleave the fluorescent label from the growing polynucleotide.
[0184] Once the sequencing primer extends through the barcode region, apply a fifth solution containing 80% labeled thymidine triphosphate to the sequencing surface, which initiates primer extension through the homopolymer region and then quickly stops. Then wash the sequencing surface to remove unincorporated nucleotides.
[0185] After primer extension stops, a sixth solution containing 100% unlabeled thymidine triphosphate is applied to the sequencing surface, which initiates primer extension through the homopolymer region. Before washing the surface with the wash buffer, the surface is allowed to react for a period of time, and the process is repeated to eliminate the stopped primer extension. This allows primer extension to pass through the poly-A homopolymer region and extend to the start of the target region of the cDNA molecule complement.
[0186] The first, second, third, and fourth solutions are used to sequence the target region of the target area. The solutions are applied to the sequencing surface separately, the surface is washed, and the presence or absence of base incorporation is detected before applying the next solution in a cycle, for a series of cycles. After each imaging step, the fluorescent label is cleaved from the growing polynucleotide.
[0187] Example 3
[0188] The sample cells are lysed using lysis buffer, and the liquid lysate is separated using centrifugation. The separated liquid lysate is incubated with capture beads, each capture bead containing a DNA oligomer that contains a barcode region (containing a common sample barcode and UMI) that is fused to the 5' of a poly-T region containing 10 - 30 consecutive thymine bases. The mRNA molecules in the liquid lysate are separated using the capture beads by separating and washing the capture beads. The mRNA molecules are reverse transcribed using a capture probe DNA oligomer and reverse transcriptase to generate a cDNA library. Each cDNA molecule in the cDNA library contains a barcode region with a sample barcode and UMI, a homopolymer poly-T region, and a target region complementary to the mRNA region of interest (which contains at least a portion of the coding region).
[0189] The barcode regions are specifically designed in conjunction with the flow cell design such that the barcode regions have the same flow length. Exemplary barcodes with corresponding flowcharts are shown in Table 2. The sequencing primers for each of the four barcodes reach the end of the barcode within the fourth cycle.
[0190] Table 2
[0191]
[0192] The cDNA library is prepared for sequencing by ligating adapter sequences to the 5' and 3' ends of the cDNA polynucleotide to generate a sequencing library. The sequencing library is applied to a sequencing array surface that contains DNA oligonucleotides attached to the surface. The DNA oligonucleotides contain sequences that hybridize to the adapter regions of the cDNA molecules. Sequencing colonies are formed by bridge amplification, which produces copies of the cDNA molecules and the complements of the cDNA molecules.
[0193] A sequencing primer is applied to a sequencing surface and hybridizes to an adapter region at the 3' end of the complement of the cDNA molecule. A flow sequencing method is used to query the barcode region, and the order of application of nucleotides is configured such that all barcodes in the sequencing library are sequenced before the primer extends into a homopolymer region. Briefly, DNA polymerase is applied to the sequencing surface, a first solution containing a first nucleotide (e.g., deoxy-A, deoxy-G, or deoxy-C), such as a non-terminating nucleotide, is applied, and then the surface is washed with a wash buffer to remove unincorporated nucleotides. The nucleotides contain approximately 2.5% fluorescently labeled nucleotides and approximately 97.5% unlabeled nucleotides. The presence or absence of base incorporation across the colonies is detected using a fluorescence detector. The process is repeated using second, third, and fourth solutions in addition to the first solution, each solution containing a different (i.e., second, third, and fourth) nucleotide to complete a flow cycle. The barcode region is sequenced using multiple flow cycles (which can be asymmetric). The nucleotides in the second and third solutions contain approximately 2.5% fluorescent labeling and approximately 97.5% unlabeled nucleotides.
[0194] The primer is extended into the homopolymer region during barcode region sequencing, but the order of the barcode region and discretely applied nucleotides is designed such that the time at which the primer reaches the homopolymer region is controlled. Once the sequencing primer extends through the barcode region and reaches the start of the homopolymer region, a fifth solution containing 100% unlabeled thymidine triphosphate is applied to the sequencing surface, which initiates primer extension through the homopolymer region. The surface is allowed to react for a period of time before washing the surface with a wash buffer, and the process is repeated to eliminate stopped primer extensions. This allows the primer to extend through the poly-A homopolymer region and into the start of the target region of the complement of the cDNA molecule.
[0195] Then the first, second, third, and fourth solutions are used to sequence the target region of the target region. The solutions are applied to the sequencing surface separately, the surface is washed, and the presence or absence of base incorporation is detected before the next solution is applied in a cycle, for a series of cycles.
[0196] Example 4
[0197] Lyse the sample cells using lysis buffer and separate the liquid lysate by centrifugation. Incubate the separated liquid lysate with capture beads, each of which contains a DNA oligomer that contains a barcode region (containing a common sample barcode and UMI) that is fused to the 5' end of a poly-T region containing 10 - 30 consecutive thymine bases. Use the capture beads to isolate mRNA molecules in the liquid lysate by separating and washing the capture beads. Reverse transcribe the mRNA molecules using a capture probe DNA oligomer and reverse transcriptase to generate a cDNA library. Each cDNA molecule in the cDNA library contains a barcode region with a sample barcode and UMI, a homopolymer poly-T region, and a target region complementary to the mRNA region of interest (which contains at least a portion of the coding region).
[0198] Prepare the cDNA library for sequencing by ligating adapter sequences to the 5' and 3' ends of the cDNA polynucleotides to generate a sequencing library. Apply the sequencing library to a sequencing array surface that contains DNA oligonucleotides attached to the surface. The DNA oligonucleotides contain sequences that hybridize to the adapter regions of the cDNA molecules. Sequencing colonies are formed by bridge amplification, which generates copies of the cDNA molecules and the complements of the cDNA molecules.
[0199] Apply a first sequencing primer to the sequencing surface that hybridizes to the adapter region at the 3' end of the cDNA molecule complement. Apply DNA polymerase to the sequencing surface. Apply a first solution containing a first nucleotide (e.g., deoxy-A, deoxy-G, deoxy-C, or deoxy-T), i.e., a non-terminating nucleotide, and wash the sequencing surface with a wash buffer to remove unincorporated nucleotides. The nucleotides contain approximately 2.5% fluorescently labeled nucleotides and approximately 97.5% unlabeled nucleotides. Detect the presence or absence of base incorporation across the colonies using a fluorescence detector. Repeat the process with second, third, and fourth solutions, each containing a different (i.e., second, third, and fourth) nucleotide to complete a flow cycle, and repeat the flow cycle to sequence the barcode region.
[0200] The nucleotides in the second and third solutions contain approximately 2.5% fluorescent label and approximately 97.5% unlabeled nucleotides.
[0201] A second sequencing primer is applied to the sequencing surface and comprises a poly-T homopolymer region and a 5' anchor containing variable bases other than thymine (i.e., A, C, or G). That is, the second sequencing primer is actually a mixture of three different primers with different sequences, which differ only by the variable base at the 5' end of the primer. The second sequencing primer hybridizes to the homopolymer region of the cDNA complement, where the 5' anchor hybridizes to the first base of the target region. First, second, third, and fourth solutions are used to extend the second sequencing primer and sequence the target region of the cDNA molecule complement. The solutions are applied to the sequencing surface separately, the surface is washed, and the presence or absence of base incorporation is detected before the next solution is applied in a cycle, for a series of cycles.
[0202] Even when separate primers are used to sequence the barcode region and the target region, sequences can be associated based on the spatial coordinates on the sequencing surface.
[0203] Example 5
[0204] Sample cells are lysed using a lysis buffer and the liquid lysate is separated by centrifugation. The separated liquid lysate is incubated with capture beads, each of which contains a DNA oligomer that contains a barcode region (containing a common sample barcode and UMI) that is fused to the 5' end of a poly-T region containing 10 - 30 consecutive thymine bases. The capture beads are used to isolate mRNA molecules in the liquid lysate by separating and washing the capture beads. The mRNA molecules are reverse transcribed using a capture probe DNA oligomer and reverse transcriptase to generate a cDNA library. Each cDNA molecule in the cDNA library contains a barcode region with a sample barcode and UMI, a homopolymer poly-T region, and a target region complementary to the mRNA region of interest (which contains at least a portion of the coding region).
[0205] The cDNA library is prepared for sequencing by ligating adapter sequences to the 5' and 3' ends of the cDNA polynucleotide to generate a sequencing library. The sequencing library is applied to a sequencing array surface that contains DNA oligonucleotides attached to the surface. The DNA oligonucleotides contain sequences that hybridize to the adapter regions of the cDNA molecules. Sequencing colonies are formed by bridge amplification, which produces copies of the cDNA molecules and the complements of the cDNA molecules.
[0206] Apply a sequencing primer to a sequencing surface, which comprises a first primer segment and a second primer segment, where the first primer segment hybridizes to an adaptor region at the 3' end of the cDNA molecule complement and the second primer segment hybridizes to a homopolymer region of the cDNA molecule complement. Separating the first primer segment and the second primer segment is a uracil RNA moiety. Apply a first solution containing 100% unlabeled thymidine triphosphate to the sequencing surface, which extends the second primer segment to the end of the homopolymer region of the cDNA molecule complement. Then remove unincorporated nucleotides by washing the sequencing surface with a wash buffer.
[0207] Apply a second solution containing nucleotides (e.g., deoxy-A, deoxy-G, or deoxy-C) to the sequencing surface and wash the sequencing surface with a wash buffer to remove unincorporated nucleotides. The nucleotides contain approximately 2.5% fluorescently labeled nucleotides and approximately 97.5% unlabeled nucleotides. Detect the presence or absence of base incorporation across the colony using a fluorescence detector. Repeat the process with third, fourth, and fifth solutions, each containing a different nucleotide to complete a flow cycle, and repeat the flow cycle to sequence the target region. Although the second solution does not contain thymine, any one of the third, fourth, or fifth solutions can contain a thymine base. The nucleotides in the second and third solutions contain approximately 2.5% fluorescent label and approximately 97.5% unlabeled nucleotides.
[0208] By applying a solution containing a uracil-specific excision reagent enzyme (available from New England BioLabs) to cleave the primer, which specifically cleaves the uracil moiety between the first and second primer segments. Then remove the cleavage solution using a wash buffer.
[0209] Then use the second, third, fourth, and fifth solutions to extend the first primer segment and sequence the barcode region of the cDNA molecule complement. Apply the solutions to the sequencing surface separately, wash the surface, and detect the presence or absence of base incorporation before applying the next solution in a cycle, for a series of cycles.
[0210] Even when separate primers are used to sequence the barcode region and the target region, sequences can be associated based on spatial coordinates on the sequencing surface.
[0211] Example 6
[0212] Sequence three synthetic test sequences using a non-terminating synthesis sequencing method. The test cDNA sequences contain a start region, a barcode region, a homopolymer region (CC), and a target region. Participate Figure 8。The barcode region of Sequence 3 contains homopolymers (e.g., the T-T-T homopolymer in Flow 14, the G-G-G homopolymer in Flow 16, the A-A-A homopolymer in Flow 18, and the A-A-A-A-A homopolymer in Flow 30), but these homopolymers are different from the C-C homopolymer region. Different flow orders of nucleotides are used to sequence (1) the start region (using the T-A-C-G flow order, 2 cycles), (2) the barcode region (using the T-C-A flow order, 9 cycles), and (3) the homopolymer region and the target region (using the G-A-T-C flow order, 2 cycles). The barcodes are designed to omit C nucleotides, which are included in the homopolymer region. Thus, the flow order used to sequence the barcodes omits the complementary G nucleotides. In addition, the flow order designed to sequence the homopolymer region and the target region sequences through the homopolymer region starting from G nucleotides before sequencing the target region.
[0213] Figure 8 A flow matrix showing the resulting sequencing data for each of the three sequences at each flow position is shown. The data are provided based on the nucleotides flowing into the sequencing reaction, which are complementary to the nucleotides of the sequencing molecule. Below the flow matrix are the signal traces for each sequencing molecule. The greater the signal intensity at any given position, the more nucleotides incorporated at the specified flow position.
[0214] The barcodes for all three sequences were successfully sequenced in Flows 9 - 30. The barcode flow cycles continue through Flows 31 - 35, but the barcodes are already fully sequenced and no nucleotides are incorporated in these flows. This even occurs at Flow 30 of Sequence 3, which contains the incorporation of five consecutive T nucleotides (no additional T nucleotides are incorporated at position 33). At Flow 36, G nucleotides are introduced, which allows the sequencing strand to extend through the homopolymer region (C-C). An additional benefit of this protocol is that the sequencing of all three sequences is synchronized at Flow 36.
Claims
1. A method for determining the sequence of a target region from an mRNA molecule, comprising: hybridizing a plurality of polynucleotides derived from the mRNA with a primer to form a plurality of hybridization templates, wherein the polynucleotides comprise a homopolymer region containing a plurality of consecutive and identical bases containing the poly-A region derived from the mRNA molecule and a target region containing a sequence related to the target region from the mRNA molecule, wherein the primer hybridizes with the homopolymer region of the polynucleotide sequence or an adaptor sequence; extending the primer within the homopolymer region using nucleotides of the same base complementary to the bases present in the homopolymer region, wherein the primer stops within the homopolymer region and unincorporated nucleotides are removed; repeating the extension one or more times in the absence of detected nucleotide incorporation into the primer to extend the primer through the homopolymer region; and after the primer has extended through the homopolymer region, determining the sequence of the target region using labeled nucleotides.
2. The method according to claim 1, wherein the nucleotides used to extend the primer within the homopolymer region comprise non-terminating labeled nucleotides.
3. The method according to claim 1, wherein the labeled nucleotides used to determine the sequence of the target region comprise non-terminating labeled nucleotides.
4. The method according to claim 1, wherein the nucleotides used to extend the primer within the homopolymer region comprise labeled nucleotides.
5. The method according to claim 1, wherein the nucleotides used to extend the primer within the homopolymer region comprise unlabeled nucleotides.
6. The method according to claim 1, wherein the sequence of the target region is determined using a mixture of unlabeled nucleotides and the labeled nucleotides.
7. The method according to claim 1, wherein the polynucleotide further comprises a barcode region, and wherein the target region is associated with a unique barcode region, and the method further comprises determining the sequence of the barcode region using labeled nucleotides.
8. The method according to any one of claims 1-7, wherein the bases in the homopolymer region of the polynucleotide are adenine or thymine bases.
9. The method according to any one of claims 1-7, wherein the homopolymer region of the polynucleotide comprises at least 8 consecutive and identical bases.
10. The method according to any one of claims 1-7, wherein the homopolymer region of the polynucleotide comprises at least 50 consecutive and identical bases.
11. The method according to claim 7, wherein the barcode region comprises a sample barcode.
12. The method according to claim 7, further comprising associating the determined sequence of the barcode region with the determined sequence of the target region related to the same polynucleotide.
13. The method according to any one of claims 1-7, wherein the polynucleotide is a cDNA molecule.
14. The method according to any one of claims 1-7, wherein the target region of the mRNA molecule comprises the coding region of the mRNA molecule.
15. The method according to any one of claims 1-7, wherein the target region of the mRNA molecule comprises the 3'-untranslated region or the 5'-untranslated region of the mRNA molecule.
16. The method according to any one of claims 1-7, wherein the nucleotides for extending the primer within the homopolymer region comprise a mixture of labeled nucleotides and unlabeled nucleotides.
17. The method according to claim 16, wherein the sequence of the target region is determined using a mixture of labeled nucleotides and unlabeled nucleotides.
18. The method according to claim 17, wherein the proportion of labeled nucleotides for extending the primer within the homopolymer region is higher than the proportion of labeled nucleotides relative to the total nucleotides for determining the sequence of the target region.
19. The method according to any one of claims 1-7, wherein the primer is initially extended within the homopolymer region using a first proportion of labeled nucleotides relative to the total nucleotides, and repeating the extension to extend the primer through the homopolymer region comprises using a second proportion of labeled nucleotides relative to the total nucleotides, wherein the first proportion is higher than the second proportion.
Citation Information
Patent Citations
Methods for detecting nucleic acid variants
US20200372971A1
Fast-forward sequencing by synthesis methods
US20200377937A1
Mostly natural DNA sequencing by synthesis
US8772473B2
Methods and systems for analyte detection and analysis
US10273528B1
Accurate molecular barcoding
US20170314067A1