Chemical compositions and methods for utilizing them
Sequencing probes with a target-binding and barcode domain enable rapid, enzyme-free nucleic acid sequencing, addressing the inefficiencies of current methods by providing low error rates and long readout lengths suitable for clinical applications.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BRUKER SPATIAL BIOLOGY INC
- Filing Date
- 2024-07-18
- Publication Date
- 2026-07-29
AI Technical Summary
Current nucleic acid sequencing methods require enzymatic amplification and polymerization, which are costly and time-consuming, necessitating a need for a rapid, enzyme-free, and amplification-free sequencing approach.
The use of sequencing probes with a target-binding domain and a barcode domain, where each position in the barcode corresponds to at least two nucleotides, allowing for long readout lengths and low error rates, and enabling rapid sequencing without amplification or enzymes.
This method provides rapid, enzyme-free, and amplification-free nucleic acid sequencing with low error rates and long readout lengths, particularly suitable for clinical applications.
Smart Images

Figure 0007897284000027 
Figure 0007897284000028 
Figure 0007897284000029
Abstract
Description
[Technical Field]
[0001] Cross-references of related applications This application claims priority and benefit from U.S. Provisional Application No. 62 / 671,091, filed on 14 May 2018, and U.S. Provisional Application No. 62 / 836,327, filed on 19 April 2019. The contents of each of these patent applications are incorporated herein by reference in their entirety.
[0002] Array List This application contains a sequence list submitted in ASCII format via EFS-Web, the entire list of which is incorporated herein by reference. A copy of this ASCII file was created on 13 May 2019 and named "NATE-039_001WO_SeqList.txt", with a size of 25,129 bytes. [Background technology]
[0003] Currently, a variety of methods exist for nucleic acid sequencing (i.e., the process of determining the precise order of nucleotides within a single nucleic acid molecule). Current methods require enzymatic amplification of nucleic acids (e.g., PCR) and / or amplification by cloning. Further enzymatic polymerization is required to generate a signal detectable by photodetectors. Such amplification and polymerization steps are costly and / or time-consuming. Therefore, there is a need in this field for a rapid nucleic acid sequencing method that does not require amplification or enzymes. This disclosure addresses this need. [Overview of the project]
[0004] This disclosure provides sequencing probes, methods, kits, and apparatus that offer long readout lengths, low error rates, rapid sequencing, and enzyme-, amplification, and library-free nucleic acid sequencing. The sequencing probes described herein include a barcode domain, where each position within the barcode domain corresponds to at least two nucleotides in a target-binding domain. Furthermore, the methods, kits, and apparatus have the ability to rapidly obtain results from a sample. These features are particularly useful for sequencing in clinical settings. This disclosure is an improvement on the disclosures disclosed in U.S. Patent Application Publication No. 2016 / 0194701, which are incorporated herein by reference in their entirety.
[0005] This disclosure provides a probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least eight nucleotides that hybridize to a target nucleic acid, with at least six nucleotides within the target-binding domain identifying the corresponding nucleotides in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain not identifying the corresponding nucleotides in the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least three attachment sites, each attachment site comprising at least one attachment site comprising at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule, the synthetic skeleton comprising L-DNA, and each of the at least three attachment sites comprising the target-binding domain Corresponding to two of the six nucleotides mentioned above, each of the three attachment sites has a different nucleic acid sequence, and the nucleic acid sequence of each of the three attachment sites determines the position and identity of the two corresponding nucleotides of the six nucleotides in the target nucleic acid to which the target binding domain binds; a first complementary primary nucleic acid molecule hybridizes to a first attachment site among the three attachment sites, the first complementary primary nucleic acid molecule comprises at least two domains and a cleavable linker, the first domain hybridizes to a first attachment site of the barcode domain, and the second domain can hybridize to at least one complementary secondary nucleic acid molecule, the linker modification is, [ka] It is one of the above, and the linker modification is located between the first domain and the second domain.
[0006] The probe can contain approximately 60 nucleotides. The probe can contain a single-stranded DNA synthesis skeleton and a double-stranded DNA spacer between the targeting binding domain and the barcoding domain. The single-stranded DNA synthesis skeleton can contain L-DNA. The single-stranded DNA synthesis skeleton can contain approximately 27 nucleotides. The double-stranded DNA spacer can contain L-DNA. The double-stranded DNA spacer can contain approximately 25 nucleotides in length.
[0007] The number of nucleotides in the target-binding domain of the probe can be greater than the number of attachment sites in the barcode domain of the probe. The target-binding domain can contain eight nucleotides, and the barcode domain can contain three attachment sites. At least one nucleotide in the target-binding domain that does not identify a corresponding nucleotide in the target nucleic acid molecule can be located before the above six nucleotides in the target-binding domain, and at least one nucleotide in the target-binding domain that does not identify a corresponding nucleotide in the target nucleic acid molecule can be located after the above six nucleotides in the target-binding domain.
[0008] A single attachment site within the barcode domain may contain a single attachment region. At least one nucleic acid sequence at each attachment site within the barcode domain may contain approximately nine nucleotides. At least one nucleic acid sequence at an attachment site may contain a 3'-terminal guanosine nucleotide. At least one nucleic acid sequence at each attachment site may contain at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, along with a 3'-terminal guanosine nucleotide. Each nucleotide in at least one nucleic acid sequence at an attachment site can be L-DNA. Each nucleotide in at least eight nucleotides of the target-binding domain can be D-DNA.
[0009] A primary nucleic acid molecule can serve as a complementary nucleic acid molecule, and this primary nucleic acid molecule can directly bind to at least one attachment region within at least one attachment site of the barcode domain. The primary nucleic acid molecule may contain at least two domains, the first domain which can bind to at least one attachment region within at least one attachment site of the barcode domain, and the second domain which can bind to at least one complementary secondary nucleic acid molecule. The first domain of the primary nucleic acid molecule may contain L-DNA. The second domain of the primary nucleic acid molecule may contain D-DNA. The first domain of the primary nucleic acid molecule may contain a 5' cytosine nucleotide. The first domain of the primary nucleic acid molecule may contain at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 5' cytosine nucleotide. A cleavable linker may be located between the first domain and the second domain of the primary nucleic acid molecule. The cleavable linker may contain at least one cleavable moiety. One possible cleavable portion is the portion that can be cut with light (photocleavable moiety).
[0010] The primary nucleic acid molecule can hybridize to at least one attachment region within at least one attachment site of the barcode domain, and can also hybridize to at least one secondary nucleic acid molecule. The primary nucleic acid molecule can hybridize to four secondary nucleic acid molecules.
[0011] A secondary nucleic acid molecule may contain at least two domains, the first domain which can bind to a complementary sequence in at least one primary nucleic acid molecule, and the second domain which can bind to (a) a first detectable label and at least a second detectable label, or (b) at least one complementary tertiary nucleic acid molecule, or (c) a combination thereof. The secondary nucleic acid molecule may contain a cleavable linker which can be located between the first and second domains. The cleavable portion can be cleaved with light. The secondary nucleic acid molecule may hybridize to at least one primary nucleic acid molecule and to at least one tertiary nucleic acid molecule. The secondary nucleic acid molecule may hybridize to (a) at least one primary nucleic acid molecule, (b) at least one tertiary nucleic acid molecule, and (c) a first detectable label and at least a second detectable label. Each secondary nucleic acid molecule may hybridize to one tertiary nucleic acid molecule. The first detectable label and at least the second detectable label may have the same emission spectrum or may have different emission spectra.
[0012] A tertiary nucleic acid molecule may contain at least two domains, the first of which can bind to a complementary sequence within a secondary nucleic acid molecule, and the second of which can bind to a first detectable label and at least a second detectable label. The tertiary nucleic acid molecule contains a cleavable linker, which may be located between the first and second domains. The cleavable linker can be cleaved by light. The tertiary nucleic acid molecule can hybridize to at least one secondary nucleic acid molecule, which may contain a first detectable label and at least a second detectable label. The first detectable label and at least the second detectable label may have the same emission spectrum or different emission spectra.
[0013] At least first and second detectable labels located on a secondary nucleic acid molecule may have the same emission spectrum, and at least first and second detectable labels located on a tertiary nucleic acid molecule may have the same emission spectrum, while the emission spectrum of the detectable label on the secondary nucleic acid molecule may differ from the emission spectrum of the detectable label on the tertiary nucleic acid molecule.
[0014] A primary nucleic acid molecule can be hybridized into four secondary nucleic acid molecules, each of which contains four primary detectable labels, and then hybridized into one tertiary nucleic acid molecule, which contains five detectable labels. The emission spectrum of the primary detectable labels on the secondary nucleic acid molecules may differ from the emission spectrum of the secondary detectable labels on the tertiary nucleic acid molecule.
[0015] The present disclosure provides a method for determining the nucleotide sequence of a nucleic acid, comprising: (1) hybridizing the target-binding domain of at least one first probe of claim 1 to a first region of a target nucleic acid, wherein the target nucleic acid is optionally immobilized on a substrate at one or more positions; (2) hybridizing a first complementary nucleic acid molecule containing at least one first detectable label and at least one second detectable label to a first attachment position among the at least three attachment positions of the barcode domain; (3) identifying the at least one first detectable label and at least one second detectable label of the first complementary nucleic acid molecule hybridized to the first attachment position; (4) removing the at least one first detectable label and at least one second detectable label hybridized to the first attachment position; and (5) identifying the at least one third detectable label and at least one second (6) Hybridize a second complementary nucleic acid molecule containing a detectable label 4 to the second attachment site of the barcode domain among the above at least three attachment sites; (7) Identify at least one third detectable label and at least one fourth detectable label of the second complementary nucleic acid molecule hybridized to the second attachment site; (8) Remove at least one third detectable label and at least one fourth detectable label hybridized to the second attachment site; (9) Hybridize a third complementary nucleic acid molecule containing at least one fifth detectable label and at least one sixth detectable label to the third attachment site of the barcode domain among the above at least three attachment sites; (10) Identify at least one fifth detectable label and at least one sixth detectable label of the third complementary nucleic acid molecule hybridized to the third attachment site;(10) A method is provided comprising determining the nucleotide sequences of at least six nucleotides of an optionally immobilized target nucleic acid hybridized to the at least six nucleotides of the target-binding domain of the at least one first probe, based on the attributes of the at least one first detectable label, the at least one second detectable label, the at least one third detectable label, the at least one fourth detectable label, the at least one fifth detectable label, and the at least one sixth detectable label.
[0016] The preceding method further comprises: (11) removing at least one first probe from a first region of an optionally immobilized target nucleic acid; (12) hybridizing the target-binding domain of at least one second probe according to claim 1 to a second region of an optionally immobilized target nucleic acid (provided that the target-binding domains of the first probe and at least two probes are different); (13) hybridizing a fourth complementary nucleic acid molecule containing at least one seventh detectable label and at least one eighth detectable label to a first attachment site among at least three attachment sites of the barcode domain of at least one second probe; (14) identifying at least one seventh detectable label and at least one eighth detectable label of the fourth complementary nucleic acid molecule hybridized to the first attachment site; (15) removing at least one seventh detectable label and at least one eighth detectable label of the fourth complementary nucleic acid molecule hybridized to the first attachment site; and (16) at least one (17) Hybridize a fifth complementary nucleic acid molecule containing a ninth detectable label and at least one tenth detectable label to a second attachment site among the at least three attachment sites on the barcode domain of at least one second probe; (18) Identify at least one ninth detectable label and at least one tenth detectable label of the fifth complementary nucleic acid molecule hybridized to the second attachment site; (19) Remove at least one ninth detectable label and at least one tenth detectable label hybridized to the second attachment site; (20) Hybridize a sixth complementary nucleic acid molecule containing at least one eleventh detectable label and at least one twelfth detectable label to a third attachment site among the at least three attachment sites on the barcode domain of at least one second probe; (20) Identify at least one eleventh detectable label and at least one twelfth detectable label of the sixth complementary nucleic acid molecule hybridized to the third attachment site;(21) This may include determining the nucleotide sequences of at least six nucleotides of the optionally immobilized target nucleic acid that has hybridized to the at least six nucleotides of the target binding domain of the at least one second probe, based on the attributes of the at least one seventh detectable label, the at least one eighth detectable label, the at least one ninth detectable label, the at least one tenth detectable label, the at least one eleventh detectable label, and the at least one twelfth detectable label.
[0017] This method may further include assembling nucleotides in an identified linear order from at least a first region and at least a second region of the optionally immobilized target nucleic acid, thereby identifying the sequence of the optionally immobilized target nucleic acid.
[0018] Steps (4) and (5) can occur sequentially or simultaneously. Steps (7) and (8) can occur sequentially or simultaneously.
[0019] The first detectable label and the second detectable label may have the same emission spectrum or different emission spectra. The third detectable label and the fourth detectable label may have the same emission spectrum or different emission spectra. The fifth detectable label and the sixth detectable label may have the same emission spectrum or different emission spectra.
[0020] The first, second, and third complementary nucleic acid molecules may contain a cleavable linker. The cleavable linker can be cleaved with light.
[0021] The first complementary nucleic acid molecule may include one primary nucleic acid molecule, four secondary nucleic acid molecules, and four tertiary nucleic acid molecules, where the primary nucleic acid molecule hybridizes into four secondary nucleic acid molecules, each of which contains four primary detectable labels, and also hybridizes into one tertiary nucleic acid molecule, each of which contains five secondary detectable labels.
[0022] The primary nucleic acid molecule may contain at least two domains, namely a first domain that hybridizes to the first attachment site of the barcode domain and a second domain that hybridizes to four secondary nucleic acid molecules. The primary nucleic acid molecule may also contain a cleavable linker located between the first and second domains.
[0023] A secondary nucleic acid molecule may contain at least two domains, namely a first domain that hybridizes to a second domain of the primary nucleic acid molecule, and a second domain containing four first detectable labels that hybridizes to a tertiary nucleic acid molecule. The secondary nucleic acid molecule may also contain a cleavable linker located between the first and second domains.
[0024] Removal of at least one first detectable label and at least one second detectable label hybridized at the first attachment site may include cleaving a cleavable linker between the first and second domains of the primary nucleic acid, or cleaving a cleavable linker between the first and second domains of each secondary nucleic acid, or any combination thereof.
[0025] This disclosure provides a composition comprising at least one molecular complex, wherein the at least one molecular complex comprises (A) a target nucleic acid molecule obtained from a biological sample, and (B) at least two nucleic acid molecular complexes, the first of these two nucleic acid molecular complexes comprising a partially double-stranded first nucleic acid molecule, one strand of which comprises a target-specific domain that hybridizes to a first portion of the target nucleic acid molecule, a double-stranded domain that anneals to the other strand of the partially double-stranded first nucleic acid molecule, and at least one first affinity moiety, the other strand of the partially double-stranded first nucleic acid molecule comprising a double-stranded domain that anneals to the other strand of the partially double-stranded first nucleic acid molecule, a substrate-specific domain that hybridizes to a complementary nucleic acid attached to a substrate, and at least one A composition is provided comprising a second affinity moiety, the second complex comprising a partially double-stranded second nucleic acid molecule, one strand of the partially double-stranded second nucleic acid molecule comprising a target-specific domain (the first and second parts do not overlap) that hybridizes to the second part of the target nucleic acid molecule, and a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule, the other strand of the partially double-stranded second nucleic acid molecule comprising a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule, a sample-specific domain for identifying the biological sample from which the target nucleic acid was supplied, a first single-stranded purified sequence, a first cleavable moiety located between the double-stranded domain and the sample-specific domain, and a second cleavable moiety located between the sample-specific domain and the first single-stranded purified sequence.
[0026] This disclosure provides a composition comprising at least one molecular complex, wherein the at least one molecular complex comprises (A) a target nucleic acid molecule obtained from a biological sample, and (B) at least two nucleic acid molecular complexes, the first of these two nucleic acid molecular complexes comprising a partially double-stranded first nucleic acid molecule, one strand of this partially double-stranded first nucleic acid molecule comprising a target-specific domain that hybridizes to a first portion of the target nucleic acid molecule, a double-stranded domain that anneals to the other strand of this partially double-stranded first nucleic acid molecule, and at least one first affinity moiety, the other strand of this partially double-stranded first nucleic acid molecule comprising a double-stranded domain that anneals to the other strand of this partially double-stranded first nucleic acid molecule and is functionally linked to the 3' end of the target nucleic acid molecule, and hybridizes to complementary nucleic acids attached to a substrate A composition is provided comprising a substrate-specific domain and at least one second affinity moiety, wherein the second complex comprises a partially double-stranded second nucleic acid molecule, one strand of this partially double-stranded second nucleic acid molecule comprising a target-specific domain (the first and second parts do not overlap) that hybridizes to the second part of the target nucleic acid molecule, and a double-stranded domain that anneals to the other strand of this partially double-stranded second nucleic acid molecule, the other strand of this partially double-stranded second nucleic acid molecule comprising a double-stranded domain that anneals to the other strand of this partially double-stranded second nucleic acid molecule and is functionally linked to the 5' end of the target nucleic acid molecule, a sample-specific domain for identifying the biological sample from which the target nucleic acid was supplied, a first single-stranded purified sequence, and a first cleavable moiety located between the double-stranded domain and the sample-specific domain.
[0027] This disclosure provides a composition comprising at least one molecular complex, wherein the at least one molecular complex comprises (A) a target nucleic acid molecule obtained from a biological sample and (B) at least two nucleic acid molecular complexes, the first of these two nucleic acid molecular complexes comprising a partially double-stranded first nucleic acid molecule, one strand of this partially double-stranded first nucleic acid molecule comprising a target-specific domain that hybridizes to a first portion of the target nucleic acid molecule, a double-stranded domain that anneals to the other strand of this partially double-stranded first nucleic acid molecule, and at least one first affinity moiety, the other strand of this partially double-stranded first nucleic acid molecule anneals to the other strand of this partially double-stranded first nucleic acid molecule and also anneals to the 3rd portion of the target nucleic acid molecule A composition is provided comprising a functionally ligated double-stranded domain at its terminus, a substrate-specific domain that hybridizes to a complementary nucleic acid attached to a substrate, and at least one second affinity moiety, wherein the second complex comprises a partially double-stranded second nucleic acid molecule, one strand of which comprises a target-specific domain that hybridizes to a second portion of a target nucleic acid molecule (the first and second portions do not overlap) and a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule, the other strand of which comprises a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule and is functionally ligated to the 5' terminus of the target nucleic acid molecule.
[0028] The present disclosure also provides a composition comprising: a flat solid support substrate; a first layer on the flat solid support substrate; and a second layer on the first layer, wherein the second layer comprises a plurality of nanowells, each nanowell providing access to an exposed portion of the first layer, and each nanowell comprises a plurality of first oligonucleotides covalently bonded to the exposed portion of the first layer.
[0029] This disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least eight nucleotides that hybridize to a target nucleic acid, with at least six nucleotides within the target-binding domain identifying corresponding nucleotides in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain not identifying corresponding nucleotides in the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least three attachment sites, each attachment site comprising at least one attachment region containing at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule, the nucleic acid sequences of the at least three attachment sites determining the positions and attributes of the at least six nucleotides in the target nucleic acid to which the target-binding domain binds, and each of the at least three attachment sites having a different nucleic acid sequence.
[0030] This disclosure also provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least eight nucleotides that hybridize to a target nucleic acid, with at least six nucleotides in the target-binding domain identifying corresponding nucleotides in the target nucleic acid molecule, and at least two nucleotides in the target-binding domain not identifying corresponding nucleotides in the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least three attachment sites, each of which comprises at least one attachment site containing at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule, each of the at least three attachment sites corresponding to two of the at least six nucleotides in the target-binding domain, and each of the at least three attachment sites having a different nucleic acid sequence, the nucleic acid sequence of each of the at least three attachment sites determining the position and attributes of the two corresponding nucleotides in the at least six nucleotides in the target nucleic acid to which the target-binding domain binds.
[0031] This disclosure provides a complex comprising a) a composition comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least eight nucleotides that hybridize to a target nucleic acid, with at least six nucleotides within the target-binding domain identifying corresponding nucleotides in the target nucleic acid molecule and at least two nucleotides within the target-binding domain not identifying corresponding nucleotides in the target nucleic acid molecule; and the barcode domain comprises a synthetic skeleton and at least three attachment sites, each attachment site comprising at least one attachment region containing at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule. The nucleic acid sequence of each of the at least three attachment sites determines the position and attributes of two corresponding nucleotides of the at least six nucleotides in the target nucleic acid to which the target binding domain binds, and each of the at least three attachment sites has a different nucleic acid sequence; a first complementary primary nucleic acid molecule hybridizes to a first attachment site among the at least three attachment sites, the first complementary primary nucleic acid molecule comprises at least two domains and a cleavable linker, the first domain hybridizes to a first attachment site of the barcode domain, the second domain can hybridize to at least one complementary secondary nucleic acid molecule, and the cleavable linker is [ka] It is one of the following and is located between the first domain and the second domain.
[0032] The present disclosure provides a method for determining the nucleotide sequence of a nucleic acid, comprising: (1) hybridizing the target-binding domain of a first sequencing probe of the present disclosure to a first region of a target nucleic acid, the target nucleic acid optionally immobilized on a substrate at one or more positions; (2) hybridizing a first complementary nucleic acid molecule containing at least one first detectable label and at least one second detectable label to a first attachment position among the at least three attachment positions of the barcode domain; (3) identifying at least one first detectable label and at least one second detectable label of the first complementary nucleic acid molecule hybridized to the first attachment position; (4) removing at least one first detectable label and at least one second detectable label hybridized to the first attachment position; and (5) identifying at least one third detectable label and at least one second (6) Hybridize a second complementary nucleic acid molecule containing a detectable label 4 to the second attachment site of the barcode domain among the above at least three attachment sites; (7) Identify at least one third detectable label and at least one fourth detectable label of the second complementary nucleic acid molecule hybridized to the second attachment site; (8) Remove at least one third detectable label and at least one fourth detectable label hybridized to the second attachment site; (9) Hybridize a third complementary nucleic acid molecule containing at least one fifth detectable label and at least one sixth detectable label to the third attachment site of the barcode domain among the above at least three attachment sites; (10) Identify at least one fifth detectable label and at least one sixth detectable label of the third complementary nucleic acid molecule hybridized to the third attachment site;(10) A method is provided comprising determining the nucleotide sequences of at least six nucleotides of an optionally immobilized target nucleic acid hybridized to the at least six nucleotides of the target binding domain of a first sequencing probe, based on the attributes of the at least one first detectable label, the at least one second detectable label, the at least one third detectable label, the at least one fourth detectable label, the at least one fifth detectable label, and the at least one sixth detectable label.
[0033] The present disclosure provides a method for determining the nucleotide sequence of a nucleic acid, comprising: (1) hybridizing the target-binding domain of a first sequencing probe according to claim 113 or 114 to a first region of a target nucleic acid, wherein the target nucleic acid is optionally immobilized on a substrate at one or more positions; (2) hybridizing a first complementary nucleic acid molecule comprising at least one first detectable label and at least one second detectable label to a first attachment position among the at least three attachment positions of the barcode domain; and (3) hybridizing the first attachment position to the first (1) Identify at least one first detectable label and at least one second detectable label of a complementary nucleic acid molecule; (4) Based on the attributes of the at least one first detectable label and the at least one second detectable label, identify the positions and attributes of the first and second nucleotides in the optionally immobilized target nucleic acid that hybridize to two of the at least six nucleotides of the target binding domain; (5) Identify at least one first detectable label and at least one second detectable label that hybridized to the first attachment site. (6) Remove any possible labels; hybridize a second complementary nucleic acid molecule containing at least one third detectable label and at least one fourth detectable label to the second attachment site of the barcode domain among the at least three attachment sites; (7) Identify the at least one third detectable label and at least one fourth detectable label of the second complementary nucleic acid molecule hybridized to the second attachment site; (8) Based on the attributes of the at least one third detectable label and the at least one fourth detectable label, identify the target binding domain (9) Identify the positions and attributes of a third and fourth nucleotide in an optionally immobilized target nucleic acid that hybridizes to two of the above six nucleotides; (10) Remove at least one third detectable label and at least one fourth detectable label that hybridized to the second attachment site; (11) Hybridize a third complementary nucleic acid molecule containing at least one fifth detectable label and at least one sixth detectable label to the third attachment site of the barcode domain among the above three attachment sites;A method is provided comprising: (11) identifying at least one fifth detectable label and at least one sixth detectable label of a third complementary nucleic acid molecule hybridized to a third attachment site; and (12) identifying the positions and attributes of the fifth and sixth nucleotides in an optionally immobilized target nucleic acid hybridized to two of the at least six nucleotides of the target-binding domain, based on the attributes of the at least one fifth detectable label and the at least one sixth detectable label; and determining the nucleotide sequences of at least six nucleotides of an optionally immobilized target nucleic acid hybridized to the at least six nucleotides of the target-binding domain of a first sequencing probe.
[0034] This disclosure provides a method for identifying the presence of a predetermined nucleotide sequence within a target nucleic acid, comprising: (1) hybridizing the target-binding domain of a first sequencing probe of this disclosure to a first region of the target nucleic acid, wherein the target nucleic acid is optionally immobilized on a substrate at one or more positions; (2) hybridizing a first complementary nucleic acid molecule containing at least one first detectable label and at least one second detectable label to a first attachment position among the at least three attachment positions of the barcode domain; and (3) hybridizing the first attachment position to (4) Identify at least one first detectable label and at least one second detectable label of the reduced first complementary nucleic acid molecule; (5) Remove at least one first detectable label and at least one second detectable label that have hybridized to the first attachment site; (6) Hybridize the second complementary nucleic acid molecule containing at least one third detectable label and at least one fourth detectable label to the second attachment site of the barcode domain among the at least three attachment sites; (7) Hybridize to the second attachment site A method is provided for revealing the presence of a predetermined nucleotide sequence based on the attributes of the at least one first detectable label, the at least one second detectable label, the at least one third detectable label, the at least one fourth detectable label, and the at least one sixth detectable label, by identifying at least one third detectable label and at least one fourth detectable label of a soyed second complementary nucleic acid molecule; (7) removing the at least one third detectable label and at least one fourth detectable label hybridized to the second attachment site; (8) hybridizing the third complementary nucleic acid molecule containing at least one fifth detectable label and at least one sixth detectable label to the third attachment site of the barcode domain; and (9) identifying at least one fifth detectable label and at least one sixth detectable label of the third complementary nucleic acid molecule hybridized to the third attachment site.
[0035] This disclosure provides a kit comprising: (A) a first nucleic acid molecule complex comprising a partially double-stranded first nucleic acid molecule (one strand of this partially double-stranded first nucleic acid molecule comprises a target-specific domain that hybridizes to a first portion of a target nucleic acid molecule, a double-stranded domain that anneals to the other strand of this partially double-stranded first nucleic acid molecule, and at least one first affinity moiety; the other strand of this partially double-stranded first nucleic acid molecule comprises a double-stranded domain that anneals to the other strand of this partially double-stranded first nucleic acid molecule, a substrate-specific domain that hybridizes to a complementary nucleic acid attached to a substrate, and at least one second affinity moiety); and (B) a second nucleic acid molecule complex comprising a partially double-stranded second nucleic acid molecule (one strand of this partially double-stranded second nucleic acid molecule is A kit is provided comprising: a target-specific domain that hybridizes to the second portion of the target nucleic acid molecule (the first and second portions do not overlap); a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule, the other strand of the partially double-stranded second nucleic acid molecule comprising: a double-stranded domain that anneals to the other strand of the partially double-stranded second nucleic acid molecule; a sample-specific domain that identifies the biological sample from which the target nucleic acid was obtained; a substrate-specific domain that hybridizes to a complementary nucleic acid attached to the substrate; a first single-stranded purified sequence; a first cleavable portion located between the double-stranded domain and the sample-specific domain; and a second cleavable portion located between the sample-specific domain and the first single-stranded purified sequence.
[0036] This disclosure provides a kit comprising: (A) a first single-stranded nucleic acid molecule comprising a target-specific domain that hybridizes to a first portion of a target nucleic acid molecule, a double-stranded domain that anneals to the double-stranded domain of a second nucleic acid molecule, and at least one first affinity moiety; (B) a second single-stranded nucleic acid molecule comprising a double-stranded domain that anneals to the double-stranded domain of the first single-stranded nucleic acid molecule, a substrate-specific domain that hybridizes to a complementary nucleic acid attached to a substrate, and at least one second affinity moiety; and (C) a target-specific domain that hybridizes to the second portion of the target nucleic acid (provided that the first portion... (D) The kit also provides a third single-stranded nucleic acid molecule containing a double-stranded domain that anneals to the double-stranded domain of a fourth single-stranded nucleic acid molecule (the second part of which does not overlap), and a fourth single-stranded nucleic acid molecule containing a double-stranded domain that anneals to the double-stranded domain of the third single-stranded nucleic acid molecule, a sample-specific domain that identifies the biological sample from which the target nucleic acid was obtained, a first single-stranded purified sequence, a first cleavable portion located between the double-stranded domain and the sample-specific domain, and a second cleavable portion located between the sample-specific domain and the first single-stranded purified sequence.
[0037] Any of the above embodiments can be combined with any other embodiment.
[0038] Unless otherwise specified, all scientific and technical terms have the same meaning as those generally understood by those skilled in the art to which this disclosure belongs. In this specification, the singular form includes the plural unless the context makes it clear that they are different; for example, “one” and “it” are understood to be singular or plural, and the term “or” is understood to be inclusive. For example, “one element” means one or more elements. Throughout this specification, “contains” or its variation “contains” is understood to include one element, one integer, one process, or a group of elements, a group of integers, or a group of processes described, but not any other elements, integers, processes, or groups of elements, a group of integers, or a group of processes. "Approximately" can be understood to mean within 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.5%, 0.1%, 0.05%, or 0.01% of the stated value. Unless otherwise made clear from the context, all figures presented herein are modified by the term "approximately."
[0039] When carrying out or testing this disclosure, similar or equivalent methods and materials to those described herein may be used, but suitable methods and materials are described below. All publications, patent applications, patents, and other references referenced herein are incorporated herein by reference in their entirety. References cited herein do not constitute prior art of the claimed invention. In the event of any conflict, this specification shall prevail, including definitions. In addition, materials, methods, and examples are for illustrative purposes only and are not intended to limit them. Other features and advantages of this disclosure will become apparent from the following detailed description and claims.
[0040] The patent or application file contains at least one color drawing. A copy of the published version of this patent or application, accompanied by the color drawing, will be provided by the Patent and Trademark Office upon request and payment of the required fees.
[0041] The above features and further details will become clearer when combined with the attached drawings, as will be explained in the following detailed description. [Brief explanation of the drawing]
[0042] [Figure 1] Figure 1 shows a typical sequencing probe of this disclosure. [Figure 2] Figure 2 shows the designs of the standard probe, the three-part sequencing probe, and the one-part linker probe of this disclosure. [Figure 3] Figure 3 shows a typical reporter complex of the present disclosure hybridized to a typical sequencing probe of the present disclosure. [Figure 4] Figure 4 shows a schematic diagram of a typical reporter probe of this disclosure. [Figure 5] Figure 5 is a schematic diagram of some representative reporter probes of this disclosure, including tertiary nucleic acids of various configurations. [Figure 6] Figure 6 is a schematic diagram of some representative reporter probes of this disclosure that include branched tertiary nucleic acids. [Figure 7] Figure 7 shows the location in a typical reporter probe of this disclosure where the linker can be modified to cut. [Figure 8] Figure 8 is a schematic diagram illustrating the capture of a target nucleic acid using the two capture probe systems of this disclosure. [Figure 9] Figure 9 shows the results from an experiment utilizing the method of the present invention to capture and detect a multi-cancer panel consisting of 100 targets using FFPE samples. [Figure 10] Figure 10 is a schematic diagram of one cycle of the sequencing method of this disclosure. [Figure 11-1] Figure 11 shows a schematic diagram of one cycle of the sequencing method of this disclosure and the corresponding imaging data recovered during this cycle. [Figure 11-2]Figure 11 shows a schematic diagram of one cycle of the sequencing method of this disclosure and the corresponding imaging data recovered during this cycle. [Figure 12] Figure 12 shows a typical sequencer probe pool configuration of this disclosure, in which eight different sequencer probe pools are designed using eight different color combinations. [Figure 13] Figure 13 compares the barcode domain design disclosed in U.S. Patent Application Publication No. 2016 / 0194701 with the barcode domain design of this disclosure. [Figure 14] Figure 14 is a schematic diagram of the sequencing cycle of the present disclosure, in which one location in the barcode is darkened using modification of a severable linker. [Figure 15] Figure 15 is a schematic diagram of a typical sequencing cycle of the present disclosure, in which one position within the barcode domain is darkened by the substitution of the primary nucleic acid. [Figure 16] Figure 16 is a schematic diagram illustrating how the sequencing method of this disclosure enables sequencing of the same base of a target nucleic acid using different sequencing probes. [Figure 17] Figure 17 illustrates how multiple base calls at a specific nucleotide position on a target nucleic acid are recorded from one or more sequencing probes, combined to form a consensus sequence, thereby increasing the accuracy of the final base call. [Figure 18-1] Figure 18 shows the results from a sequencing experiment, obtained using the sequencing method described herein and analyzed using the Assembly Algorithm. Regarding the graphs, the sequences shown correspond to sequence numbers 3, 4, 6, 8, 7, and 5, starting from the top left graph and moving clockwise. Regarding the table on the right, the sequences correspond to sequence numbers 3, 4, 7, 8, 6, and 5, from top to bottom. [Figure 18-2]Figure 18 shows the results from a sequencing experiment, obtained using the sequencing method described herein and analyzed using the Assembly Algorithm. Regarding the graphs, the sequences shown correspond to sequence numbers 3, 4, 6, 8, 7, and 5, starting from the top left graph and moving clockwise. Regarding the table on the right, the sequences correspond to sequence numbers 3, 4, 7, 8, 6, and 5, from top to bottom. [Figure 19] Figure 19 shows a schematic diagram of an experimental design for multiple capture and sequencing of oncogene targets from FFPE samples. [Figure 20] Figure 20 shows a schematic diagram of direct RNA sequencing and results from experiments investigating the compatibility of RNA molecules using the sequencing method disclosed herein. [Figure 21] Figure 21 shows the results of sequencing RNA and DNA molecules having the same nucleotide sequence using the sequencing method of this disclosure. [Figure 22] Figure 22 shows a comparison of the performance of the standard sequencing probe and the three-part sequencing probe of this disclosure. [Figure 23] Figure 23 shows the effect of LNA substitution within a representative target-binding domain of this disclosure using individual probes. [Figure 24] Figure 24 shows the effect of LNA substitution in a representative target-binding domain of this disclosure when using a pool of nine probes. [Figure 25] Figure 25 shows the effect of substitutions by modified nucleotides and nucleic acid analogs within representative target-binding domains of this disclosure. [Figure 26] Figure 26 shows experimental results quantifying the raw accuracy of the sequencing method of this disclosure. [Figure 27] Figure 27 shows experimental results for determining the accuracy of the sequencing method of this disclosure when sequencing nucleotides in a target nucleic acid using two or more sequencing probes. [Figure 28]Figure 28 is a schematic diagram of the sequencing probe of this disclosure, including a pocket oligo. [Figure 29] Figure 29 is a schematic diagram of the sequencing probe of this disclosure, which includes PEG linker regions between each attachment site. [Figure 30] Figure 30 is a schematic diagram of the sequencing probe of this disclosure, which includes non-basic regions between each attachment site. [Figure 31] Figure 31 shows a typical reporter complex of the present disclosure indirectly hybridized to a typical sequencing probe of the present disclosure via a connector oligo. [Figure 32] Figure 32 is an explanatory diagram of the parity scheme used in the method of this disclosure. [Figure 33] Figure 33 is a schematic diagram of the present invention, which consists of a capture probe, an adapter oligonucleotide, and a lawn oligonucleotide. [Figure 34] Figure 34 is a schematic diagram of the c5 probe complex and c3 probe complex of this disclosure that hybridize to a target nucleic acid. [Figure 35] Figure 35 is a schematic diagram of the target nucleic acid-c3 probe-c5 probe complex of this disclosure after digestion using FEN1. [Figure 36] Figure 36 is a schematic diagram of the target nucleic acid-c3 probe-c5 probe complex of this disclosure after ligation. [Figure 37] Figure 37 is a schematic diagram illustrating the cleavage of the target nucleic acid-C3 probe-C5 probe complex of this disclosure via a user. [Figure 38] Figure 38 is a schematic diagram of the target nucleic acid-c3 probe-c5 probe complex of this disclosure after cleavage by USER. [Figure 39] Figure 39 is a schematic diagram illustrating the cleavage of the target nucleic acid-c3 probe-c5 probe complex of this disclosure by UV light. [Figure 40] Figure 40 is a schematic diagram of the target nucleic acid-c3 probe-c5 probe complex of this disclosure, which has been attached to a substrate via complementary nucleic acids, after being cleaved by UV light. [Figure 41] Figure 41 is a schematic diagram of the c3.2 probe complex and c5.2 probe complex of this disclosure hybridized to a target nucleic acid. [Figure 42] Figure 42 is a schematic diagram of the target nucleic acid complex of this disclosure after the c3.2 probe complex and the c5.2 probe complex have been linked. [Figure 43] Figure 43 is a schematic diagram of the cleavage and release of a single-stranded purified sequence within the nucleic acid complex of this disclosure. [Figure 44] Figure 44 is a schematic diagram of the target nucleic acid complex of the present disclosure immobilized on the substrate of the present disclosure. [Figure 45] Figure 45 is a schematic diagram illustrating the cleavage and release of substrate-specific domains after the target nucleic acid complex of the present disclosure is immobilized on the substrate of the present disclosure. [Figure 46] Figure 46 is a schematic diagram of the target nucleic acid complex of the present disclosure after the release of the substrate-specific domain immobilized on the substrate of the present disclosure. [Figure 47] Figure 47 is a schematic cross-sectional view of a typical array of the present invention. [Figure 48] Figure 48 is a schematic cross-sectional view of a typical array of the present invention that includes pyramidal nanowells. [Figure 49] Figure 49 is a schematic diagram of a typical array of the present disclosure, which includes multiple cylindrical nanowells arranged in a random pattern. [Figure 50] Figure 50 is a schematic diagram of a typical array of the present disclosure, which includes cylindrical nanowells arranged in a regular grid with a constant pitch. [Figure 51] Figure 51 is a schematic diagram of a typical array of the present disclosure, in which a single target nucleic acid complex is immobilized in each nanowell. [Figure 52] Figure 52 is a schematic diagram of a typical array of the present disclosure, in which a single target nucleic acid complex is immobilized in each nanowell, thereby preventing the immobilization of other target nucleic acid complexes. [Figure 53] Figure 53 is a schematic diagram of the sequencing probe of this disclosure, which consists entirely of L-DNA and includes an attachment region having a 3' terminal L-dG nucleotide. [Figure 54] Figure 54 is a schematic diagram of the sequencing probe of this disclosure, which consists entirely of D-DNA and includes pocket oligos between attachment region 1 (spot 1) and attachment region 2 (spot 2), and between attachment region 2 (spot 2) and attachment region 3 (spot 3). [Figure 55] Figure 55 is a schematic diagram of a synthetic target nucleic acid immobilized on the surface of a solid substrate using a capture probe and a lone oligonucleotide in combination with a protein lock. [Figure 56] Figure 56 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probe and the D-pocket sequencing probe of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the theoretical sequencing diversity, observed sequencing diversity, and observed sequencing coverage for the LG-mediated sequencing probe and the D-pocket sequencing probe. The red boxes indicate the expected problem areas with respect to sequencing. [Figure 57] Figure 57 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probe and the D-pocket sequencing probe of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the theoretical sequencing diversity, observed sequencing diversity, and observed sequencing coverage for the LG-mediated sequencing probe and the D-pocket sequencing probe. The red boxes indicate anticipated problem areas with respect to sequencing. [Figure 58]Figure 58 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probe and the D-pocket sequencing probe of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the theoretical sequencing diversity, observed sequencing diversity, and observed sequencing coverage for the LG-mediated sequencing probe and the D-pocket sequencing probe. The red boxes indicate anticipated problem areas with respect to sequencing. [Figure 59] Figure 59 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probes and D-pocket sequencing probes of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the observed sequencing diversity and observed sequencing coverage for the LG-mediated sequencing probes and D-pocket sequencing probes. [Figure 60] Figure 60 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probes and D-pocket sequencing probes of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the observed sequencing diversity and observed sequencing coverage for the LG-mediated sequencing probes and D-pocket sequencing probes. [Figure 61] Figure 61 is a series of charts showing the results of sequencing experiments using the LG-mediated sequencing probes and D-pocket sequencing probes of this disclosure. The x-axis represents the specific nucleotides of the target nucleic acid being sequenced. The top chart shows the observed sequencing diversity and observed sequencing coverage for the LG-mediated sequencing probes and D-pocket sequencing probes. [Figure 62]Figure 62 is a series of histograms showing the total number of barcode events and the number of valid 3-spot readouts in sequencing experiments using the LG-intervening sequencing probe and D-pocket sequencing probe of this disclosure. [Figure 63] Figure 63 is a series of graphs showing the total number of on-target events, invalid events, off-target events, events with 1 error in b1-b6, events with 2 errors in b1-b6, events with 3 errors in b1-b6, events with 4 errors in b1-b6, events with 5 errors in b1-b6, and events with 6 errors in b1-b6 in sequencing experiments using the LG-intervening sequencing probe and D-pocket sequencing probe of this disclosure. [Figure 64-1] Figure 64 is a series of graphs showing the total number of on-target events, invalid events, off-target events, events with 1 error in b1-b6, events with 2 errors in b1-b6, events with 3 errors in b1-b6, events with 4 errors in b1-b6, events with 5 errors in b1-b6, and events with 6 errors in b1-b6 in sequencing experiments using the LG-intervening sequencing probe and D-pocket sequencing probe of this disclosure. [Figure 64-2] Figure 64 is a series of graphs showing the total number of on-target events, invalid events, off-target events, events with 1 error in b1-b6, events with 2 errors in b1-b6, events with 3 errors in b1-b6, events with 4 errors in b1-b6, events with 5 errors in b1-b6, and events with 6 errors in b1-b6 in sequencing experiments using the LG-intervening sequencing probe and D-pocket sequencing probe of this disclosure. [Figure 64-3]Figure 64 is a series of graphs showing the total number of on-target events, invalid events, off-target events, events with 1 error in b1-b6, events with 2 errors in b1-b6, events with 3 errors in b1-b6, events with 4 errors in b1-b6, events with 5 errors in b1-b6, and events with 6 errors in b1-b6 in sequencing experiments using the LG-intervening sequencing probe and D-pocket sequencing probe of this disclosure. [Figure 65] Figure 65 is a chart showing the 1-spotter event (where only one of the three possible reporter probes successfully recorded), the 2-spotter event (where only two of the three possible reporter probes successfully recorded), and the 3-spotter event (where all three possible reporter probes successfully recorded) in each cycle of a sequencing experiment using the D-pocket sequencing probes (cycles 1-50) and LG-intervening sequencing probes (cycles 51-100) of this disclosure. [Figure 66-1] Figure 66 is a series of charts showing the results of sequencing experiments using the LG-intervened sequencing probe and the D-pocket sequencing probe of this disclosure. The leftmost figure shows the number of on-target events, novel hexamer events, redundant hexamer events, off-target events, and invalid events in each cycle of the sequencing experiment. Cycles 1-50 were performed using the D-pocket sequencing probe of this disclosure, and cycles 51-100 were performed using the LG-intervened sequencing probe of this disclosure. [Figure 66-2]Figure 66 is a series of charts showing the results of sequencing experiments using the LG-intervened sequencing probe and the D-pocket sequencing probe of this disclosure. The leftmost figure shows the number of on-target events, novel hexamer events, redundant hexamer events, off-target events, and invalid events in each cycle of the sequencing experiment. Cycles 1-50 were performed using the D-pocket sequencing probe of this disclosure, and cycles 51-100 were performed using the LG-intervened sequencing probe of this disclosure. [Figure 66-3] Figure 66 is a series of charts showing the results of sequencing experiments using the LG-intervened sequencing probe and the D-pocket sequencing probe of this disclosure. The leftmost figure shows the number of on-target events, novel hexamer events, redundant hexamer events, off-target events, and invalid events in each cycle of the sequencing experiment. Cycles 1-50 were performed using the D-pocket sequencing probe of this disclosure, and cycles 51-100 were performed using the LG-intervened sequencing probe of this disclosure. [Figure 66-4] Figure 66 is a series of charts showing the results of sequencing experiments using the LG-intervened sequencing probe and the D-pocket sequencing probe of this disclosure. The leftmost figure shows the number of on-target events, novel hexamer events, redundant hexamer events, off-target events, and invalid events in each cycle of the sequencing experiment. Cycles 1-50 were performed using the D-pocket sequencing probe of this disclosure, and cycles 51-100 were performed using the LG-intervened sequencing probe of this disclosure. [Figure 67] Figure 67 is a schematic diagram of a target nucleic acid immobilized on a solid substrate using the method and composition of the present disclosure. The target nucleic acid is immobilized using a protein lock between the biotin moiety located on the capture probe and the lone oligonucleotide and neutraavidin moiety. [Modes for carrying out the invention]
[0043] This disclosure provides sequencing probes, reporter probes, methods, kits, and apparatus that enable rapid nucleic acid sequencing without enzymes, amplification, or libraries, and that offer long readout lengths and low error rates.
[0044] Composition of the present disclosure
[0045] This disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises any construct listed in Table 1. A typical target-binding domain comprises at least eight nucleotides capable of binding to a target nucleic acid, at least six nucleotides within the target-binding domain capable of identifying a corresponding (complementary) nucleotide in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain not identifying a corresponding nucleotide in the target nucleic acid molecule; any of the above at least six nucleotides within the target-binding domain can be a modified nucleotide or nucleotide analog, and the at least two nucleotides within the target-binding domain that do not identify a corresponding nucleotide in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain. A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which a complementary nucleic acid molecule can bind, each of the at least three attachment sites corresponds to two of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the nucleic acid sequence at each of the at least three attachment sites determines the position and attributes of the two corresponding nucleotides among the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0046] In another embodiment, a representative target-binding domain may contain at least six nucleotides capable of hybridizing to a target nucleic acid, and these at least six nucleotides within the target-binding domain may identify the corresponding (complementary) nucleotides within the target nucleic acid molecule; any of these at least six nucleotides within the target-binding domain may be a modified nucleotide or a nucleotide analogue.
[0047] This disclosure also provides a sequencing probe comprising a target-binding domain and a barcode domain; the target-binding domain comprises at least 10 nucleotides capable of binding to a target nucleic acid, wherein at least 6 nucleotides within the target-binding domain can identify corresponding (complementary) nucleotides within the target nucleic acid molecule, and at least 4 nucleotides within the target-binding domain do not identify corresponding nucleotides within the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and includes at least 3 attachment sites, each attachment site comprising at least 1 attachment region containing at least 1 nucleic acid sequence to which a complementary nucleic acid molecule can bind, each of the at least 3 attachment sites corresponding to 2 nucleotides of the at least 6 nucleotides within the target-binding domain, each of the at least 3 attachment sites having a different nucleic acid sequence, the nucleic acid sequence of each of the at least 3 attachment sites determining the positions and attributes of the 2 corresponding nucleotides of the at least 6 nucleotides within the target nucleic acid to which the target-binding domain binds.
[0048] This disclosure also provides a group of sequencing probes, including any of the sequencing probes disclosed herein.
[0049] The target-binding domain, barcode domain, and backbone of the disclosed sequencing probe, as well as complementary nucleic acid molecules (e.g., reporter molecules or reporter complexes), are described in more detail below.
[0050] The sequencing probes of this disclosure include a target domain and a barcode domain. Figure 1 is a schematic diagram of a typical sequencing probe of this disclosure. Figure 1 shows that the target domain can bind to a target nucleic acid. The target nucleic acid can be any nucleic acid that the sequencing probe of this disclosure can hybridize. The target nucleic acid can be DNA or RNA. The target nucleic acid can be obtained from a biological sample derived from the subject. The terms “target-binding domain” and “sequencing domain” are used interchangeably herein.
[0051] The target-binding domain can contain a series of nucleotides (e.g., a polynucleotide). The target-binding domain can contain DNA, RNA, or a combination thereof. If the target-binding domain is a polynucleotide, it binds to the target nucleic acid by hybridizing to a complementary portion of the target nucleic acid in the target nucleic acid, as shown in Figure 1.
[0052] The target-binding domain of a sequencing probe can be designed to control the likelihood and rate of hybridization and / or dehybridization of the sequencing probe. Generally, the lower the probe's Tm, the faster and more likely it is to dehybridize from the target nucleic acid. Therefore, using probes with lower Tm will result in fewer probes binding to the target nucleic acid.
[0053] The length of the target-binding domain partially influences the likelihood of a probe hybridizing to a target nucleic acid and the likelihood of it remaining hybridized. Generally, the longer the target-binding domain (the more nucleotides it contains), the less likely the complementary sequence is to be present within the target nucleotide. Conversely, the shorter the target-binding domain, the greater the likelihood the complementary sequence is present within the target nucleotide. For example, the probability of a tetrameric sequence being located within a target nucleic acid is 1 / 256, while the probability of a hexamer sequence being located within a target nucleic acid is 1 / 4096. As a result, a set of shorter probes may have a greater chance of binding to more positions on a given nucleic acid of a given length compared to a set of longer probes.
[0054] In several situations, for example, when detecting mutations or SNP alleles, it is preferable to increase the coverage of the target nucleic acid, or a portion of the target nucleic acid (particularly the portion of interest), by increasing the number of reads in a given length of nucleic acid by having a probe with a shorter target-binding domain.
[0055] The target-binding domain can consist of any length or number of nucleotides. The target-binding domain can be any of the following: at least 12 nucleotides, at least 10 nucleotides, at least 8 nucleotides, at least 6 nucleotides, or at least 3 nucleotides.
[0056] Each nucleotide within the target-binding domain can identify (or encode) a complementary nucleotide of the target molecule. Alternatively, some nucleotides within the target-binding domain can identify (or encode) a complementary nucleotide of the target molecule, while some nucleotides within the target-binding domain do not identify (or encode) a complementary nucleotide of the target molecule.
[0057] The target-binding domain may contain at least one native base. The target-binding domain may not contain a native base. The target-binding domain may contain at least one modified nucleotide or nucleic acid analog. The target-binding domain may not contain a modified nucleotide or nucleic acid analog. The target-binding domain may contain at least one universal base. The target-binding domain may not contain a universal base. The target-binding domain may contain at least one degenerate base. The target-binding domain may not contain a degenerate base.
[0058] The target-binding domain may contain any combination of native bases (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more native bases), modified nucleotides or nucleic acid analogs (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more modified nucleotides or nucleic acid analogs), universal bases (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more universal bases), or degenerate bases (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more degenerate bases). The native bases, modified nucleotides or nucleic acid analogs, universal bases, and degenerate bases of individual target-binding domains can be arranged in any order when present in combination.
[0059] Non-limiting examples of the term “modified nucleotide” or “nucleic acid analog” include locked nucleic acids (LNA), bridged nucleic acids (BNA), propyne-modified nucleic acids, zip nucleic acids (ZNA®), isoguanine, isocytosine, 6-amino-1-(4-hydroxy-5-hydroxymethyl-tetrahydrofuran-2-yl)-1,5-dihydro-pyrazolo[3,4-d]pyrimidine-4-one (PPG), and 2'-modified nucleic acids (such as 2'-O-methyl nucleic acid). The target-binding domain may contain 0 to 6 (e.g., 0, 1, 2, 3, 4, 5, 6) modified nucleotides or nucleic acid analogs. The modified nucleotide or nucleic acid analog is preferably locked nucleic acid (LNA).
[0060] In this specification, a non-limiting example of the term “locked nucleic acids (LNA)” includes modified RNA nucleotides in which the ribose moiety contains a methylene crosslink connecting the 2' oxygen and 4' carbon atoms. This methylene crosslink locks the ribose into the 3' end conformation (also known as the north conformation) found in type A RNA double strands. The term inaccessible RNA can be used interchangeably with LNA. In this specification, a non-limiting example of the expression “crosslinked nucleic acids (BNA)” includes modified RNA molecules containing a five- or six-membered crosslink structure with a fixed 3' end conformation (also known as the north conformation). The crosslink structure connects the 2' oxygen of the ribose to the 4' carbon of the ribose. A variety of different crosslink structures are possible, including carbon atoms, nitrogen atoms, and hydrogen atoms. In this specification, non-limiting examples of the term “propyne-modified nucleic acid” include pyrimidines with propyne modification at the C5 position of the nucleic acid base, namely cytosine and thymine / uracil. In this specification, non-limiting examples of the term “Zip Nucleic Acid (ZNA®)” include oligonucleotides conjugated to a cationic spermine moiety.
[0061] In this specification, the non-limiting term “universal base” includes nucleotide bases that do not follow the Watson-Crick base pairing rules but can bind to any of the four normal bases (A, T / U, C, G) located on the target nucleic acid. In this specification, the non-limiting term “degenerate base” includes nucleotide bases that do not follow the Watson-Crick base pairing rules but can bind to at least two of the four normal bases (A, T / U, C, G) located on the target nucleic acid, rather than all four. Degenerate bases may also be called fluctuating bases, and these terms are used interchangeably in this specification.
[0062] The typical sequencing probe shown in Figure 1 exhibits a target-binding domain containing a 6-length nucleotide (hexameric) sequence (b1- b2- b3- b4- b5- b6) that specifically hybridizes to complementary nucleotides 1-6 of the target nucleic acid to be sequenced. This hexameric portion (b1- b2- b3- b4- b5- b6) of the target-binding domain identifies (or codes for) the complementary nucleotides (1- 2- 3- 4- 5- 6) in the target sequence. Each of these hexameric sequences is flanked by a base (N). The base represented by (N) can independently be a universal base or a degenerate base. Typically, the base represented by (N) is independently one of the canonical bases. The base indicated by (N) does not identify (or code for) a complementary nucleotide to which it binds in the target sequence and is independent of the nucleic acid sequence of the (hexamer) sequence (b1-b2-b3-b4-b5-b6).
[0063] The sequencing probes shown in Figure 1 can be used in combination with the sequencing method of this disclosure to sequence target nucleic acids, utilizing only hybridization reactions and requiring no covalent chemistry, enzymes, or amplification. A total of 4096 sequencing probes are required to sequence all possible hexameric sequences within a target nucleic acid molecule (46 (=4096).
[0064] Figure 1 shows an example of the configuration of the target-binding domain of a sequencing probe according to this disclosure. Table 1 shows several other configurations of the target-binding domain according to this disclosure. One preferred target-binding domain is called the “6 LNA” target-binding domain and contains six LNAs at positions b1–b6 of the target-binding domain. Each of these six LNAs is flanked by a base (N). As used herein, the base (N) can be a universal base / degenerate base independent of the nucleic acid sequence (b1–b2–b3–b4–b5–b6) or a canonical base. In other words, while the bases b1–b2–b3–b4–b5–b6 may be specific to any given target sequence, the (N) base can be a universal base / degenerate base or any of the four canonical bases that are not specific to the target specified by b1–b2–b3–b4–b5–b6. For example, if the target sequence being investigated is CAGGCATA, the bases b1-b2-b3-b4-b5-b6 of the target-binding domain are thought to become TCCGTA, while each of the (N) bases of the target-binding domain can independently become A, C, T, or G. Therefore, the resulting target-binding domain can be any of the sequences ATCCGTAG, TTCCGTAC, GTCCGTAG, or any other of the 16 possible sequences. Alternatively, two (N) bases can be located before the 6 LNA. Or, furthermore, two (N) bases can be located after the 6 LNA.
[0065] [Table 1]
[0066] Table 1 also lists "decamer" target-binding domains containing 10 target-specific native bases. Table 1 also lists "octamer" target-binding domains containing 8 target-specific native bases.
[0067] Table 1 further describes the “Natural I” target-binding domain containing six native bases at positions b1–b6. Two (N) bases are adjacent to each of these six native bases. Alternatively, all four (N) bases could be located in front of the six native bases. Alternatively, all four (N) bases could be located behind the six native bases. If any number of the four (N) bases (i.e., 1, 2, 3, or 4) could be located in front of the six native bases, the remaining (N) bases would be located behind them.
[0068] Table 1 further describes the “Natural II” target-binding domain, which contains six native bases at positions b1–b6. One (N) base is adjacent to each of these six native bases. Alternatively, both (N) bases could be positioned in front of the six native bases. Alternatively, both (N) bases could be positioned behind the six native bases. Typically, the (N) bases in the Natural II target-binding domain are degenerate bases.
[0069] Table 1 also describes the “2LNA” target-binding domain, which contains a combination of two LNAs and four native bases at positions b1–b6 of the target-binding domain. The two LNAs and four native bases can appear in any order. For example, positions b3 and b4 can be LNAs, while positions b1, b2, b5, and b6 are native bases. One (N) base is adjacent to each of the bases b1–b6. Alternatively, two (N) bases can be located before the bases b1–b6. Or, two (N) bases can follow the bases b1–b6.
[0070] Table 1 further describes the “4LNA” target-binding domain, which contains a combination of four LNAs and two native bases at positions b1–b6 of the target-binding domain. The four LNAs and two native bases can appear in any order. For example, positions b2–b5 can be LNAs, while positions b1 and b6 are native bases. One (N) base is adjacent to each of the bases b1–b6. Alternatively, two (N) bases can be located before the bases b1–b6. Or, two (N) bases can follow the bases b1–b6.
[0071] Table 1 further describes a “6LNA” target-binding domain containing six LNAs at positions b1-b6 of the target-binding domain. One (N) base can be adjacent to each of the bases b1-b6.
[0072] Table 1 further describes “octameric” target-binding domains that contain individual native bases or LNAs at any position b1-b6 of the target-binding domain. One (N) base can be adjacent to each side of bases b1-b6.
[0073] The target-binding domain can also include a minor-groove binder. The minor-groove binder is a chemical modification of the oligonucleotide, adding a chemical portion that allows the oligonucleotide to bind to the minor groove of the target nucleotide with which it hybridizes. While not strictly theoretical, including the minor-groove binder increases the affinity of the target-binding domain to the target nucleic acid, thereby raising the melting temperature of the target-binding domain-target nucleic acid double strand. A higher binding affinity allows for the use of a smaller target-binding domain.
[0074] The target-binding domain may also contain one or more twisted intercalating nucleic acids (TINAs). TINAs are nucleic acid molecules that stabilize the formation of Hoogsteen triple-stranded DNA from double-stranded oligonucleotides and triple-stranding oligonucleotides. By stabilizing double-stranded oligonucleotides using TINAs, the specificity and sensitivity of oligonucleotide probes to target nucleic acids can be improved.
[0075] The target-binding domain can also contain nucleic acid molecules containing 2'-O-methyl-modified bases. 2'-O-methyl-modified bases are nucleoside modifications of RNA, where a methyl group is added to the 2'-hydroxyl group of ribose to generate a 2'-methoxy group. 2'-O-methyl-modified bases provide excellent protection against base hydrolysis and digestion by nucleases. Without being constrained by theory, the addition of 2'-O-methyl-modified bases also increases the melting temperature of nucleic acid double strands.
[0076] The target-binding domain can also include covalently bound stilbene modifications. Stilbene modifications can increase the stability of nucleic acid double helix.
[0077] The sequencing probes of this disclosure include a synthetic skeleton. The target-binding domain (also referred to herein as the sequencing domain) and the barcode domain are functionally linked. The target-binding domain and the barcode domain can be covalently linked to form part of a single synthetic skeleton. The target-binding domain and the barcode domain can be linked via a linker (e.g., a nucleic acid linker, a chemical linker). The synthetic skeleton can include any material (e.g., polysaccharides, polynucleotides, polymers, plastics, fibers, peptides, peptide nucleic acids, polypeptides). The synthetic skeleton is preferably rigid. The synthetic skeleton can include a single-stranded DNA molecule. The skeleton can include a "DNA origami" consisting of six DNA double helices (see, e.g., Lin et al., "Submicron Geometrically Encoded Fluorescent Barcodes Self-Assembled from DNA," Nature Chemistry; October 2012; Vol. 4(10): pp. 832-839). Barcodes can be created from DNA origami tiles (Jungmann et al., "Multiplexed 3D Cell Super-Resolution Imaging Using DNA-PAINT and Exchange-PAINT," Nature Methods, Vol. 11, No. 3, 2014).
[0078] The sequencing probes of this disclosure may include a partially double-stranded synthetic skeleton. The sequencing probes may include a single-stranded DNA synthetic skeleton and a double-stranded DNA spacer between the target-binding domain and the barcode domain. The double-stranded DNA spacer may include at least one modified nucleotide or nucleic acid analog. Typical modified nucleotides or nucleic acid analogs useful in double-stranded DNA spacers are isoguanine and isocytosine. Alternatively, L-DNA may be independently included as each nucleic acid containing the double-stranded DNA spacer. In some embodiments, the double-stranded DNA spacer may include L-DNA. The double-stranded DNA spacer may consist mainly of L-DNA.
[0079] A double-stranded DNA spacer can contain approximately 1 to 100 nucleotides in length. Alternatively, a double-stranded DNA spacer can contain approximately 25 nucleotides in length.
[0080] The synthetic skeleton can contain L-DNA. The synthetic skeleton can consist of L-DNA. The synthetic skeleton can consist mainly of L-DNA. A single-stranded DNA synthetic skeleton can contain approximately 10 to 100 nucleotides in length. A single-stranded DNA synthetic skeleton can contain approximately 52 nucleotides in length. A single-stranded DNA synthetic skeleton can contain approximately 27 nucleotides in length.
[0081] The barcode domain can contain L-DNA. The barcode domain can consist of L-DNA. The barcode domain can consist mainly of L-DNA. The barcode domain can contain approximately 27 nucleotides, or approximately 52 nucleotides, or approximately 99 nucleotides, or approximately 74 nucleotides. The barcode domain can contain approximately 27 nucleotides, or approximately 52 nucleotides, or approximately 99 nucleotides, or approximately 74 nucleotides in length.
[0082] Sequencing probes may include a single-stranded DNA synthesis backbone and a polymer spacer that has mechanical properties similar to double-stranded DNA and is located between the target-binding domain and the barcode domain. Typical polymer spacers include polyethylene glycol (PEG) type polymers.
[0083] As a double-stranded DNA spacer, it is possible to have a length of approximately 1 to 100 nucleotides, approximately 2 to 50 nucleotides, or approximately 20 to 40 nucleotides. Preferably, the double-stranded DNA spacer has a length of approximately 36 nucleotides.
[0084] One sequencing probe in this disclosure is named the “standard probe” and is shown in the left panel of Figure 2. The standard probe in Figure 2 contains a barcode domain covalently bound to a target-binding domain, so that the target-binding domain and the barcode domain reside within the same single-stranded oligonucleotide. In the left panel of Figure 2, the single-stranded oligonucleotide binds to a stem oligonucleotide to create a 36-nucleotide double-stranded spacer region called the stem sequence. Each sequencing probe in the probe pool can utilize this architecture to hybridize to the same stem sequence.
[0085] In another embodiment, the barcode domain and the region that binds to the stem oligonucleotide of the standard probe can consist of canonical bases, modified nucleotides, or nucleic acid analogs. Typical modified nucleotides or nucleic acid analogs useful in the barcode domain and the region that binds to the stem oligonucleotide of the standard probe are isoguanine and isocytosine. Alternatively, L-DNA can also be used as the nucleic acid in the barcode domain and the region that binds to the stem oligonucleotide of the standard probe. For example, the barcode domain and the region that binds to the stem oligonucleotide of the standard probe can consist entirely of L-DNA. In another example, the barcode domain and the region that binds to the stem oligonucleotide of the standard probe can consist of multiple compartments of L-DNA separated by compartments of non-basic single-stranded nucleic acids or compartments of polymers with mechanical properties similar to double-stranded DNA (such as PEG, described in more detail below).
[0086] Another sequencing probe in this disclosure is named the “three-part probe” and is shown in the center figure of Figure 2. The three-part probe in Figure 2 includes a barcode domain attached to a target-binding domain via a linker. In this example, the linker is a single-stranded stem oligonucleotide that hybridizes a single-stranded oligonucleotide containing the target-binding domain to a single-stranded oligonucleotide containing the barcode domain, creating a 36-nucleotide double-stranded spacer region that bridges the barcode domain (18 nucleotides) and the target-binding domain (18 nucleotides). Using this typical probe configuration, it is possible to design each barcode to hybridize to only one stem sequence in order to prevent exchange of barcode domains. Furthermore, after hybridizing each barcode domain to its corresponding stem oligonucleotide, different sequencing probes can be pooled together.
[0087] In another embodiment, each nucleic acid contained in a single-stranded stem oligonucleotide can be a normal nucleic acid, a modified nucleotide, or a nucleic acid analog. Typical modified nucleotides or nucleic acid analogs useful in single-stranded stem oligonucleotides are isoguanine and isocytosine. Alternatively, each nucleic acid contained in a single-stranded stem oligonucleotide can independently be L-DNA.
[0088] In another embodiment, the nucleic acids included in the region on the barcode domain into which the single-stranded stem oligonucleotide hybridizes can be canonical bases, modified nucleotides, or nucleic acid analogs. Typical modified nucleotides or nucleic acid analogs useful in single-stranded stem oligonucleotides are isoguanines and isocytosines. Alternatively, L-DNA can independently be included as the nucleic acids included in the region on the barcode domain into which the single-stranded stem oligonucleotide hybridizes.
[0089] In another embodiment, the nucleic acids contained in the region on the single-stranded oligonucleotide containing the target-binding domain to which the single-stranded stem oligonucleotide hybridizes can be canonical bases, modified nucleotides, or nucleic acid analogs. Typical modified nucleotides or nucleic acid analogs useful for single-stranded stem oligonucleotides are isoguanines and isocytosines. Alternatively, L-DNA can independently be used as the nucleic acids contained in the region on the single-stranded oligonucleotide containing the target-binding domain to which the single-stranded stem oligonucleotide hybridizes.
[0090] Another sequencing probe in this disclosure is named a "one-part linker probe" and is shown in the right-hand panel of Figure 2. The one-part linker probe in Figure 2 contains a barcode domain attached to a target-binding domain via a linker. In this example, the linker is a PEG molecule. Alternatively, the linker could be trans-stilbene. Furthermore, any polymer with mechanical properties similar to double-stranded DNA is possible as the linker. A typical polymer spacer is a polyethylene glycol (PEG) type polymer.
[0091] The sequencing probes of this disclosure may contain approximately 60 nucleotides. The sequencing probes of this disclosure may contain approximately 107 nucleotides. The sequencing probes of this disclosure may have a length of approximately 60 nucleotides or a length of approximately 107 nucleotides. The nucleotides comprising the sequencing probe may individually be canonical bases, modified nucleotides, or nucleic acid analogs, including L-DNA and D-DNA.
[0092] The barcode domain contains multiple attachment sites (e.g., one, two, three, four, five, six, seven, eight, nine, ten, or more). The number of attachment sites can be less than, the same as, or more than the number of nucleotides contained in the target-binding domain. The target-binding domain can contain more nucleotides than the number of attachment sites contained in the skeletal domain (e.g., one, two, three, four, five, six, seven, eight, nine, ten, or more). The target-binding domain can contain eight nucleotides, and the barcode domain contains three attachment sites. The target-binding domain can contain ten nucleotides, and the barcode domain contains three attachment sites.
[0093] The barcode domain has no length limit, as long as there is sufficient space for at least three attachment locations, as described below. The terms “attachment location,” “location,” and “spot” are used interchangeably herein. The terms “barcode domain” and “reporting domain” are used interchangeably herein.
[0094] Each attachment site within the barcode domain corresponds to two nucleotides (dinucleotides) within the target-binding domain, and therefore corresponds to complementary dinucleotides in the target nucleic acid that hybridize to those dinucleotides within the target-binding domain. As a non-restrictive example, the first attachment site within the barcode domain corresponds to the first and second nucleotides in the target-binding domain (for example, in Figure 1, R1 is the first attachment site within the barcode domain, and this R1 corresponds to dinucleotides b1 and b2 in the target-binding domain, which in turn identify dinucleotides 1 and 2 of the target nucleic acid). The second attachment site within the barcode domain corresponds to the third and fourth nucleotides in the target-binding domain (for example, in Figure 1, R2 is the second attachment site within the barcode domain, and this R2 corresponds to dinucleotides b3 and b4 in the target-binding domain, which in turn identify dinucleotides 3 and 4 of the target nucleic acid). The third attachment site within the barcode domain corresponds to the fifth and sixth nucleotides within the target-binding domain (for example, in Figure 1, R3 is the third attachment site within the barcode domain, and this R3 corresponds to dinucleotides b5 and b6 within the target-binding domain, which in turn identify dinucleotides 5 and 6 of the target nucleic acid). In yet another non-limiting example, the first, second, and third attachment sites within the barcode domain collectively correspond to the first through sixth nucleotides within the target-binding domain (for example, in Figure 1, these are nucleotides b1-b6 within the target-binding domain, which in turn identify the six nucleotides of the target nucleic acid).
[0095] Each attachment location within a barcode domain contains at least one attachment area (e.g., 1 to 50, or more). Some locations within a barcode domain may have more attachment areas than others (e.g., a first attachment location may have 3 attachment areas while a second attachment location has 2). Alternatively, each attachment location within a barcode domain may have the same number of attachment areas. Each attachment location within a barcode domain may contain one attachment area. Each attachment location within a barcode domain may contain two or more attachment areas. At least one of the at least three attachment locations within a barcode domain may contain a different number of attachment areas than the other two attachment locations within the barcode domain. In some embodiments, each attachment location within a barcode domain may contain one attachment area.
[0096] Each attachment region contains at least one copy (i.e., 1 to 50, e.g., 10 to 30) of the nucleic acid sequence to which a complementary nucleic acid molecule (e.g., DNA or RNA) can reversibly bind. The nucleic acid sequences of multiple attachment regions at a single attachment site can be identical to each other. Therefore, the complementary nucleic acid molecules that bind to these attachment regions are identical to each other. Alternatively, the nucleic acid sequences of multiple attachment regions at a single site are not identical to each other. Therefore, the complementary nucleic acid molecules that bind to these attachment regions are not identical to each other.
[0097] The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 6 to 20 nucleotides in length. The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 12 nucleotides in length. The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 16 nucleotides in length. The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 14 nucleotides in length. The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 8 nucleotides in length. The nucleic acid sequence containing each attachment region within the barcode domain can be approximately 9 nucleotides in length.
[0098] The attachment site, or attachment region, or at least one nucleic acid sequence of the attachment region may contain at least one super T base (5-hydroxybutyl-2'-deoxyuridine). The attachment site, or attachment region, or at least one nucleic acid sequence of the attachment region may contain at least one 3'-terminal super T base (5-hydroxybutyl-2'-deoxyuridine). The attachment site, or attachment region, or at least one nucleic acid sequence of the attachment region may contain at least one 5'-terminal super T base (5-hydroxybutyl-2'-deoxyuridine).
[0099] Each nucleic acid contained within each attachment region of a barcode domain can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, or six nucleotides within the attachment region of a barcode domain can be modified nucleotides or nucleotide analogs. A typical ratio of modified nucleotides or nucleotide analogs to normal bases within a barcode domain is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in attachment regions within a barcode domain are isoguanine or isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine or isocytosine) can improve the efficiency and accuracy of reporter binding to appropriate attachment regions within a barcode domain, while minimizing binding to other locations (including targets).
[0100] One or more attachment sites within a barcode domain can contain L-DNA. L-DNA is left-handed and is a mirror image version of the naturally occurring right-handed D-DNA. L-DNA is more stable and resistant to enzymatic digestion. Because L-DNA cannot hybridize to D-LNA, it can improve the efficiency and accuracy of reporter binding to appropriate attachment sites within the barcode domain, while also preventing reporter binding to other locations on the sequencing probe. In some embodiments, L-DNA can be each nucleotide in at least one nucleic acid sequence at the attachment site.
[0101] Each nucleic acid contained within each attachment region in the barcode domain can independently contain one of the following: an adenine base, a cytosine base, a guanine base, or a thymine base. Alternatively, each nucleic acid contained within each attachment region in the barcode domain can independently contain one of the following: an adenine base, a guanine base, or a thymine base.
[0102] Each nucleic acid sequence contained within each attachment region in the barcode domain may contain at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 3'-terminal guanosine nucleotide. Each nucleic acid sequence contained within each attachment region in the barcode domain may consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 3'-terminal guanosine nucleotide. Each nucleic acid sequence contained within each attachment region in the barcode domain may mainly consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 3'-terminal guanosine nucleotide.
[0103] Each nucleic acid sequence contained within each attachment region in the barcode domain may contain at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 5'-terminal guanosine nucleotide. Each nucleic acid sequence contained within each attachment region in the barcode domain may consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 5'-terminal guanosine nucleotide. Each nucleic acid sequence contained within each attachment region in the barcode domain may mainly consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one cytosine nucleotide, or any combination thereof, and a 5'-terminal guanosine nucleotide.
[0104] In some embodiments, at least one attachment region among at least one attachment site of the barcode domain may contain a 3'-terminal guanosine nucleotide. In some embodiments, at least one attachment region among at least two attachment sites of the barcode domain may contain a 3'-terminal guanosine nucleotide. In some embodiments, at least one attachment region among at least three attachment sites of the barcode domain may contain a 3'-terminal guanosine nucleotide. L-DNA is possible as the 3'-terminal guanosine nucleotide.
[0105] In some embodiments, at least one attachment region among at least one attachment site of the barcode domain may contain a 5'-terminal guanosine nucleotide. In some embodiments, at least one attachment region among at least two attachment sites of the barcode domain may contain a 5'-terminal guanosine nucleotide. In some embodiments, at least one attachment region among at least three attachment sites of the barcode domain may contain a 5'-terminal guanosine nucleotide. As the 5'-terminal guanosine nucleotide, L-DNA is possible (e.g., L-deoxyguanosine (L-dG)). The terminal L-dG nucleotide maintains stability by reducing cross-junctional hybridization between the attachment region and / or attachment site, while also providing base stacking interactions.
[0106] One or more attachment regions can be integrated with a polynucleotide backbone. That is, the backbone is a single polynucleotide, and the attachment regions are part of the sequence of that single polynucleotide. One or more attachment regions can be linked to a modified monomer (e.g., a modified nucleotide) within the synthetic backbone, and the attachment regions can be branched from the synthetic backbone. One attachment site can contain two or more attachment regions, some of which branch from the synthetic backbone, while others are integrated with it. At least one attachment region within at least one attachment site can be integrated with the synthetic backbone. Each attachment region within each of the at least three attachment sites can be integrated with the synthetic backbone. At least one attachment region within at least one attachment site can be branched from the synthetic backbone. Each attachment region within each of the at least three attachment sites can be branched from the synthetic backbone.
[0107] Each attachment site within a barcode domain corresponds to one of 16 possible dinucleotides, namely adenine-adenine, adenine-thymine / uracil, adenine-cytosine, adenine-guanine, thymine / uracil-adenine, thymine / uracil-thymine / uracil, thymine / uracil-cytosine, thymine / uracil-guanine, cytosine-adenine, cytosine-thymine / uracil, cytosine-cytosine, cytosine-guanine, guanine-adenine, guanine-thymine / uracil, guanine-cytosine, and guanine-guanine. Therefore, one or more attachment regions located within a single attachment site of a barcode domain correspond to one of the 16 possible dinucleotides and contain nucleic acid sequences specific to the dinucleotide that the attachment region corresponds to. Attachment regions located within different attachment sites of a barcode domain contain their own unique nucleic acid sequences, even if these attachment sites within the barcode domain correspond to the same dinucleotide. For example, given a sequencing probe of this disclosure containing a target-binding domain having a hexamer encoding the sequence AGAGAC, the barcode domain of this sequencing probe would contain three attachment sites, with the first attachment site corresponding to an adenine-guanine dinucleotide, the second attachment site corresponding to an adenine-guanine dinucleotide, and the third attachment site corresponding to an adenine-cytosine dinucleotide. In this example, the attachment region at position 1 of the probe is thought to contain a unique nucleic acid sequence different from the nucleic acid sequence of the attachment region at position 2, even if both attachment sites 1 and 2 correspond to an adenine-guanine dinucleotide. The sequences of specific attachment sites are designed, and it is investigated whether the complementary nucleic acids of each attachment site interact with other attachment sites. In addition, there are no constraints on the nucleotide sequence of the complementary nucleic acid. It is preferable that this nucleotide sequence has no substantial homology (e.g., 50% to 99.9%) to known nucleotide sequences. Doing so limits undesirable hybridization of the complementary nucleic acid and the target nucleic acid.
[0108] Figure 1 shows a diagram of a representative sequencing probe of this disclosure, including a representative barcode domain. The barcode domain illustrated in Figure 1 includes three attachment sites R1, R2, and R3. Each attachment site corresponds to a specific dinucleotide present in the hexameric sequence (b1-b6) of the target-binding domain. In this example, R1 corresponds to positions b1 and b2, R2 corresponds to positions b3 and b4, and R3 corresponds to positions b5 and b6. Thus, each site decodes a specific dinucleotide present in the hexameric sequence of the target-binding domain, thereby enabling the identification of two specific bases (A, C, G, or T) present in each dinucleotide.
[0109] In the typical barcode domain shown in Figure 1, each attachment site contains a single attachment region integrated into the synthetic skeleton. Each attachment region of the three attachment sites described above contains a specific nucleotide sequence corresponding to the individual dinucleotide encoded by that attachment site. For example, attachment site R1 contains an attachment region having a specific sequence corresponding to the attributes of dinucleotides b1-b2.
[0110] A barcode domain may further contain one or more binding regions. A barcode domain may contain at least one single-stranded nucleic acid sequence adjacent to or near at least one attachment site. A barcode domain may contain at least two single-stranded nucleic acid sequences adjacent to or near at least two attachment sites. A barcode domain may contain at least three single-stranded nucleic acid sequences adjacent to or near at least three attachment sites. These adjacent regions are known as "toe-holds" and can be used to increase the exchange rate of oligonucleotides hybridized adjacent to toe-holds by providing additional binding sites for single-stranded oligonucleotides (e.g., "toe-hold" probes; see, e.g., Seeling et al., "Catalytic Relaxation of Metastable DNA Fuel"; J. Am. Chem. Soc. 2006, Vol. 128(37), pp. 12211–12220).
[0111] At least one attachment region within the barcode domain can be adjacent to at least one side by a double-stranded nucleic acid sequence. At least two attachment regions within the barcode domain can be adjacent to at least one side by a double-stranded nucleic acid sequence. At least three attachment regions within the barcode domain can be adjacent to at least one side by a double-stranded nucleic acid sequence.
[0112] Any attachment region within a barcode domain can be isolated from any adjacent attachment site by a double-stranded nucleic acid sequence called a "pocket oligo." Figure 28 shows an example of a sequencing probe having a barcode domain containing three attachment sites. Attachment site 1 is isolated from the adjacent attachment site 2 by a pocket oligo. Attachment site 2 is further isolated from the adjacent attachment site 3 by another pocket oligo.
[0113] Each nucleic acid contained in a pocket oligo can be a normal base, a modified nucleotide, or a nucleic acid analog. Typical modified nucleotides or nucleic acid analogs useful in pocket oligos are isoguanine and isocytosine. Alternatively, each nucleic acid contained in a pocket oligo can independently be L-DNA. A pocket oligo can contain at least one super T base (5-hydroxybutyl-2'-deoxyuridine). A pocket oligo can have a length of approximately 25 nucleotides.
[0114] In some embodiments, at least one, at least two, or at least three attachment sites within the barcode domain can be flanked by at least one flanking double-stranded polynucleotide. The at least one flanking double-stranded polynucleotide may contain at least one modified nucleotide or nucleic acid analog. The at least one flanking double-stranded polynucleotide may contain L-DNA. The at least one flanking double-stranded polynucleotide may contain at least one super T base (5-hydroxybutyl-2'-deoxyuridine). The at least one flanking double-stranded polynucleotide can be approximately 25 nucleotides in length.
[0115] At least one attachment region within the barcode domain can be flanked on at least one side by any polymer with mechanical properties similar to those of double-stranded DNA. A typical polymer-based spacer is a polyethylene glycol (PEG) type polymer. At least two attachment regions within the barcode domain can be flanked on at least one side by any polymer with mechanical properties similar to those of double-stranded DNA. At least three attachment regions within the barcode domain can be flanked on at least one side by any polymer with mechanical properties similar to those of double-stranded DNA.
[0116] Any attachment region within a barcode domain can be separated from any adjacent attachment site by any polymer with mechanical properties similar to those of double-stranded DNA. A typical polymer spacer is a polyethylene glycol (PEG) type polymer. Figure 29 shows an example of a sequencing probe with three attachment sites. Attachment site 1 is separated from the adjacent attachment site 2 by a PEG linker. Attachment site 2 is further separated from the adjacent attachment site 3 by another PEG linker.
[0117] At least one attachment region within the barcode domain can be adjacent to at least one side by a non-basic single-stranded nucleic acid molecule. A non-basic nucleic acid molecule is one that does not contain purine or pyrimidine bases. At least two attachment regions within the barcode domain can be adjacent to at least one side by a non-basic single-stranded nucleic acid molecule. At least three attachment regions within the barcode domain can be adjacent to at least one side by a non-basic single-stranded nucleic acid molecule.
[0118] Any attachment region within a barcode domain can be separated from any adjacent attachment site by a non-basic single-stranded nucleic acid molecule. Figure 30 shows an example of a sequencing probe having a barcode domain containing three attachment sites. Attachment site 1 is separated from the adjacent attachment site 2 by a non-basic single-stranded nucleic acid molecule. Attachment site 2 is further separated from the adjacent attachment site 3 by a non-basic single-stranded nucleic acid molecule.
[0119] Any attachment region within a barcode domain can be separated from any adjacent attachment site by a 3'-terminal guanosine nucleotide. In some embodiments, at least one of the at least two attachment sites in the barcode domain may contain a 3'-terminal guanosine nucleotide. Figure 53 shows an example of a sequencing probe having a barcode domain containing three attachment sites, each separated by a terminal LG nucleotide. Attachment site 1 is separated from the adjacent attachment site 2 by an LG nucleotide. Attachment site 2 is further separated from the adjacent attachment site 3 by an LG nucleotide. Attachment site 3 ends with an LG nucleotide at its 3' end.
[0120] The sequencing probes of this disclosure may have a total length of approximately 20 nanometers to approximately 50 nanometers (including a target-binding domain, a barcode domain, and an optional domain). The backbone of the sequencing probes may be a polynucleotide molecule containing approximately 120 nucleotides, or approximately 60 nucleotides, or approximately 52 nucleotides, or approximately 27 nucleotides.
[0121] Sequencing probes may include modifications of cleavable linkers. A cleavable linker modification may include at least one, or at least two, or at least three, or at least four, or at least five, or at least six, or at least seven, or at least eight, or at least nine, or at least ten, or any number of cleavable segments. Any cleavable linker modification or cleavable segment known to those skilled in the art may be used. Non-limiting examples of cleavable linker modifications and cleavable segments include UV-cuttable linkers, reducing agent-cuttable linkers, and enzymatically cuttable linkers. An example of an enzymatically cuttable linker is the insertion of deoxyuracil for cleavage by the USER® enzyme. The cleavable linker modification can be located at any position along the length of the sequencing probe, and non-limiting examples of such positions include the region between the target-binding domain and the barcode domain. The right-hand figure of Figure 7 shows a typical cleavable linker modification that can be incorporated into the probe of this disclosure.
[0122] Reporter probe
[0123] The nucleic acid molecule that binds (e.g., hybridizes) to a complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain of the sequencing probe according to this disclosure includes a detectable label (directly or indirectly). This detectable label is referred to herein as the “reporter probe” or “reporter probe complex,” and these terms are used interchangeably herein. The reporter probe may be DNA, RNA, or PNA. The reporter probe is preferably DNA.
[0124] A reporter probe may comprise at least two domains, the first of which can bind to at least one first complementary nucleic acid molecule, and the second domain can bind to a first detectable label and at least a second detectable label. Figure 3 shows a schematic diagram of a typical reporter probe of this disclosure bound to a first attachment site of the barcode domain of a typical sequencing probe. In Figure 3, the first domain of the reporter probe (shown as a chestnut checkerboard pattern) is bound to a complementary nucleic acid sequence within attachment site R1 of the barcode domain, and the second domain of this reporter probe (shown in gray) is bound to two detectable labels (one green and one red).
[0125] Alternatively, the reporter probe may contain at least two domains, the first of which can bind to at least one first complementary nucleic acid molecule, and the second domain can bind to at least one second complementary nucleic acid molecule. These at least one first and second complementary nucleic acid molecule can be different (have different nucleic acid sequences).
[0126] A "primary nucleic acid molecule" is a reporter probe containing at least two domains, the first of which can bind (e.g., hybridize) to a complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain of the sequencing probe, and the second domain can bind (e.g., hybridize) to at least one additional complementary nucleic acid. The primary nucleic acid molecule can bind directly to the complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain of the sequencing probe. The primary nucleic acid molecule can also bind indirectly to the complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain of the sequencing probe via a nucleic acid linker. The nucleic acid linker is called a "connector oligo".
[0127] The connector oligo may contain at least two domains, the first of which can bind (e.g., hybridize) to at least one first complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain, and the second domain can bind (e.g., hybridize) to the first domain of the primary nucleic acid molecule. Figure 31 shows a sequencing probe bound to a reporter probe via the connector oligo.
[0128] The nucleic acids contained in the first or second domain of the connector oligo can be canonical bases, modified nucleotides, or nucleic acid analogs. Typical modified nucleotides or nucleotide analogs useful in the first or second domain of the connector oligo are isoguanine or isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine or isocytosine) can improve the efficiency and accuracy of binding the first domain of the connector oligo to a suitable complementary nucleic acid sequence within at least one attachment region within at least one attachment site of the barcode domain of a sequencing probe, while minimizing binding to other locations (including the target). Using modified nucleotides or nucleotide analogs (such as isoguanine or isocytosine) can improve the efficiency and accuracy of binding the second domain of the connector oligo to a suitable first domain of a reporter probe, while minimizing binding to other locations (including the target). Alternatively, L-DNA can independently be used as the nucleic acid contained in the first or second domain of the connector oligo. In one example of a connector oligo, the first domain contains D-DNA and the second domain contains L-DNA. In another example of a connector oligo, the first domain contains D-DNA and the second domain contains isoguanine and / or isocytosine.
[0129] The first domain of the connector oligo can have a length of approximately 8 to 16 nucleotides. Preferably, the first domain of the connector oligo has a length of 14 nucleotides. The second domain of the connector oligo can have a length of approximately 4 to 12 nucleotides. Preferably, the second domain of the connector oligo has a length of approximately 8 nucleotides.
[0130] In some embodiments including connector oligos, the attachment region can be referred to as a partially double-stranded attachment region. A partially double-stranded attachment region may include a double-stranded region and a single-stranded region. The single-stranded region of a partially double-stranded attachment region may include at least one nucleic acid sequence to which it binds (e.g., hybridizes). A primary nucleic acid molecule can be the at least one complementary nucleic acid sequence to which the single-stranded region of a partially double-stranded attachment region binds (e.g., hybridizes).
[0131] Each nucleic acid contained within the double-stranded region of a partially double-stranded attachment region can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, six, seven, or eight nucleotides within the double-stranded region of a partially double-stranded attachment region can be a modified nucleotide or nucleic acid analog. The typical ratio of modified nucleotides or nucleotide analogs to normal bases within a double-stranded region is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the double-stranded region of a partially double-stranded attachment region are isoguanine and isocytosine. Alternatively, each nucleic acid contained within the double-stranded region of a partially double-stranded attachment region can independently be L-DNA.
[0132] Each nucleic acid contained within the single-stranded region of a partially double-stranded attachment region can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, six, seven, or eight nucleotides within the single-stranded region of a partially double-stranded attachment region can be a modified nucleotide or nucleic acid analog. A typical ratio of modified nucleotides or nucleotide analogs to normal bases within a single-stranded region is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the single-stranded region of a partially double-stranded attachment region are isoguanine and isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine and isocytosine) can improve the efficiency and precision of binding the single-stranded region of the partially double-stranded attachment region to a suitable complementary nucleic acid sequence of a primary nucleic acid molecule, while minimizing binding to other locations (including the target). Alternatively, L-DNA can exist independently as individual nucleic acids contained within the single-stranded regions of the partially double-stranded attachment region.
[0133] The primary nucleic acid molecule may contain a cleavable linker. The cleavable linker may be located between the first and second domains. Preferably, the cleavable linker is photocleavable. The cleavable linker may contain at least one or at least two cleavable moieties. At least one or at least two of these cleavable moieties can be photocleaved.
[0134] The first domain of a primary nucleic acid molecule can have a length of approximately 6 to 16 nucleotides. Preferably, the first domain of the primary nucleic acid molecule has a length of approximately 8 nucleotides.
[0135] Each nucleic acid contained in the first domain of a primary nucleic acid molecule can be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, six, seven, or eight nucleotides within the first domain of the primary nucleic acid molecule can be modified nucleotides or nucleic acid analogs. A typical ratio of modified nucleotides or nucleotide analogs to normal bases within the first domain is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the first domain of a primary nucleic acid molecule are isoguanine and isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine and isocytosine) can improve the efficiency and accuracy of binding the first domain of the primary nucleic acid molecule to a suitable complementary nucleic acid sequence within at least one attachment region within at least one attachment site in the barcode domain of a sequencing probe, while minimizing binding to other locations (including the target). Alternatively, L-DNA can exist independently as each nucleic acid contained within the first domain of the primary nucleic acid molecule.
[0136] In some embodiments, the first domain of the primary nucleic acid molecule may consist entirely of L-DNA, and the second domain of the primary nucleic acid molecule may consist entirely of D-DNA.
[0137] In some embodiments, the first domain of the primary nucleic acid molecule may contain a 3' terminal cytosine nucleotide, which is L-DNA.
[0138] In some embodiments, the first domain of the primary nucleic acid molecule may contain a 5' terminal cytosine nucleotide, which is L-DNA.
[0139] In some embodiments, the first domain of the primary nucleic acid molecule may include at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 3' terminal cytosine nucleotide. In some embodiments, the first domain of the primary nucleic acid molecule may consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 3' terminal cytosine nucleotide. In some embodiments, the first domain of the primary nucleic acid molecule may mainly consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 3' terminal cytosine nucleotide.
[0140] In some embodiments, the first domain of the primary nucleic acid molecule may include at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 5' terminal cytosine nucleotide. In some embodiments, the first domain of the primary nucleic acid molecule may consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 5' terminal cytosine nucleotide. In some embodiments, the first domain of the primary nucleic acid molecule may mainly consist of at least one adenine nucleotide, at least one thymine nucleotide, at least one guanine nucleotide, or any combination thereof, and a 5' terminal cytosine nucleotide.
[0141] In this specification, at least one additional complementary nucleic acid that binds to a primary nucleic acid molecule is referred to as a "secondary nucleic acid molecule." A primary nucleic acid molecule can bind (e.g., hybridize) to at least one, at least two, at least three, at least four, at least five, or more secondary nucleic acid molecules. Preferably, a primary nucleic acid molecule binds (e.g., hybridizes) to four secondary nucleic acid molecules.
[0142] A secondary nucleic acid molecule may contain at least two domains, of which the first domain may bind (e.g., hybridize) to at least one complementary sequence in at least one primary nucleic acid molecule, and the second domain may bind (e.g., hybridize) to (a) a first detectable label and at least a second detectable label; or (b) at least one additional complementary nucleic acid; or (c) a combination thereof. In some embodiments, the first domain of the secondary nucleic acid molecule may consist entirely of L-DNA, and the second domain of the secondary nucleic acid molecule may consist entirely of D-DNA. In some embodiments, both the first and second domains of the secondary nucleic acid molecule may consist entirely of D-DNA.
[0143] The secondary nucleic acid molecule may contain a cleavable linker. The cleavable linker may be located between the first and second domains. Preferably, the cleavable linker can be cleaved by light.
[0144] Each nucleic acid contained in the first domain of a secondary nucleic acid molecule can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, or six nucleotides in the first domain of a secondary nucleic acid molecule can be modified nucleotides or nucleic acid analogs. A typical ratio of modified nucleotides or nucleotide analogs to normal bases in the first domain of a secondary nucleic acid molecule is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the first domain of a secondary nucleic acid molecule are isoguanine and isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine and isocytosine) can improve the efficiency and precision of binding of the first domain of a secondary nucleic acid molecule to a suitable complementary nucleic acid sequence in the second domain of the secondary nucleic acid molecule, while minimizing binding to other locations (including the target).
[0145] In this specification, at least one additional complementary nucleic acid that binds to a secondary nucleic acid molecule is referred to as a "tertiary nucleic acid molecule." A secondary nucleic acid molecule can bind (e.g., hybridize) to at least one, at least two, at least three, at least four, at least five, at least six, at least seven, or more tertiary nucleic acid molecules. Preferably, at least one secondary nucleic acid molecule binds (e.g., hybridizes) to one tertiary nucleic acid molecule.
[0146] The tertiary nucleic acid molecule comprises at least two domains, of which the first domain can bind (e.g., hybridize) to at least one complementary sequence in at least one secondary nucleic acid molecule, and the second domain can bind (e.g., hybridize) to a first detectable label and at least a second detectable label. Alternatively, the second domain can contain a first detectable label and at least a second detectable label, for example, by directly or indirectly attaching these labels during oligonucleotide synthesis using phosphoramidite chemistry or NHS chemistry. In some embodiments, the first domain of the tertiary nucleic acid molecule may consist entirely of L-DNA, and the second domain may consist entirely of D-DNA. In some embodiments, both the first and second domains of the tertiary nucleic acid molecule may consist entirely of D-DNA. The tertiary nucleic acid molecule may include a cleavable linker. The cleavable linker may be located between the first and second domains. Preferably, the cleavable linker can be cleaved by light.
[0147] Each nucleic acid contained in the first domain of a tertiary nucleic acid molecule can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, two, three, four, five, or six nucleotides within the first domain of a tertiary nucleic acid molecule can be modified nucleotides or nucleic acid analogs. A typical ratio of modified nucleotides or nucleotide analogs to normal bases within the first domain of a tertiary nucleic acid molecule is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the first domain of a tertiary nucleic acid molecule are isoguanine and isocytosine. Using modified nucleotides or nucleotide analogs (such as isoguanine and isocytosine) can improve the efficiency and precision of binding of the first domain of a tertiary nucleic acid molecule to a suitable complementary nucleic acid sequence within the second domain of a secondary nucleic acid molecule, for example, while minimizing binding to other locations (including the target).
[0148] The reporter probe is coupled to a first detectable label and at least a second detectable label to produce a two-color combination. This dual combination of fluorescent dyes may include one color overlap (e.g., blue-blue). In this specification, the term “label” includes a single portion capable of generating a detectable signal, or multiple portions capable of generating the same or substantially the same detectable signal. For example, a label may include a single yellow fluorescent dye (e.g., ALEXA FLUOR® 532), or multiple yellow fluorescent dyes (e.g., ALEXA FLUOR® 532).
[0149] The reporter probe can be coupled to a first detectable label and at least a second detectable label, each of which is one of four fluorescent dyes: blue (B), green (G), yellow (Y), and red (R). Using these four dyes, there are 10 possible combinations of two colors (BB; BG; BR; BY; GG; GR; GY; RR; RY; YY). In some embodiments, as shown in Figure 3, the reporter probe of this disclosure is labeled with one of eight possible color combinations, namely BB, BG, BR, BY, GG, GR, GY, and YY. The detectable label and at least the second detectable label may have the same emission spectrum or different emission spectra.
[0150] In embodiments comprising a sequencing probe and a primary nucleic acid molecule, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises any construct listed in Table 1. A typical target-binding domain comprises at least eight nucleotides that can hybridize to a target nucleic acid, where at least six nucleotides within the target-binding domain can identify corresponding (complementary) nucleotides in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule; any of the above at least six nucleotides within the target-binding domain can be modified nucleotides or nucleotide analogs, and the at least two nucleotides within the target-binding domain that do not identify corresponding nucleotides in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain. A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which at least one complementary primary nucleic acid molecule binds, each complementary primary nucleic acid molecule containing a first detectable label and at least one second detectable label, each of the at least three attachment sites corresponds to two nucleotides out of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary primary nucleic acid molecule that binds to each of the at least three attachment sites determine the position and attributes of the two corresponding nucleotides out of the at least six nucleotides in the target nucleic acid to which the target binding domain binds.The at least two nucleotides within the target-binding domain that do not identify the corresponding nucleotides within the target nucleic acid molecule can be any of the four non-target-specific normal bases, universal bases, or degenerate bases specified by the at least six nucleotides within the target-binding domain.
[0151] In some embodiments, at least one nucleotide within the target-binding domain that does not identify the corresponding nucleotide in the target nucleic acid molecule may precede a nucleotide within the target-binding domain that identifies the corresponding nucleotide in the target nucleic acid molecule. In some embodiments, at least one nucleotide within the target-binding domain that does not identify the corresponding nucleotide in the target nucleic acid molecule may follow a nucleotide within the target-binding domain that identifies the corresponding nucleotide in the target nucleic acid molecule.
[0152] In another embodiment, a representative target-binding domain may contain at least six nucleotides capable of hybridizing to a target nucleic acid, and these at least six nucleotides within the target-binding domain may identify the corresponding (complementary) nucleotides within the target nucleic acid molecule; any of these at least six nucleotides within the target-binding domain may not be modified nucleotides or nucleic acid analogs, or any of the at least six nucleotides within the target-binding domain may be modified nucleotides or nucleic acid analogs.
[0153] In embodiments comprising a sequencing probe and a primary nucleic acid molecule, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least 10 nucleotides that can hybridize to a target nucleic acid, at least 6 nucleotides within the target-binding domain that can identify the corresponding (complementary) nucleotides in the target nucleic acid molecule, and at least 4 nucleotides within the target-binding domain that do not identify the corresponding nucleotides in the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least 3 attachment sites, each attachment site comprising at least 1 attachment site containing at least 1 nucleic acid sequence to which at least 1 complementary primary nucleic acid molecule binds. A sequencing probe is also provided, which includes a region and its complementary primary nucleic acid molecule, each containing at least one first detectable label and at least one second detectable label, wherein each of the at least three attachment sites corresponds to two nucleotides of the at least six nucleotides in the target binding domain, and each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary primary nucleic acid molecule bound to each of the at least three attachment sites determine the positions and attributes of the two corresponding nucleotides of the at least six nucleotides in the target nucleic acid to which the target binding domain is bound.
[0154] In embodiments comprising a sequencing probe, a primary nucleic acid molecule, and a secondary nucleic acid molecule, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises any construct listed in Table 1. A typical target-binding domain comprises at least eight nucleotides and can hybridize to a target nucleic acid, wherein at least six nucleotides within the target-binding domain can identify corresponding (complementary) nucleotides in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule; any of the above at least six nucleotides within the target-binding domain can be modified nucleotides or nucleotide analogs, and the at least two nucleotides within the target-binding domain that do not identify corresponding nucleotides in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain. A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which at least one complementary primary nucleic acid molecule binds, and each complementary primary nucleic acid molecule is further bound to at least one complementary primary nucleic acid molecule containing at least one first detectable label and at least one second detectable label, each of the at least three attachment sites corresponds to two nucleotides out of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary secondary nucleic acid molecule that binds to each of the at least three attachment sites determines the position and attributes of the corresponding two nucleotides out of the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0155] In another embodiment, a representative target-binding domain may contain at least six nucleotides capable of hybridizing to a target nucleic acid, and these at least six nucleotides within the target-binding domain may identify the corresponding (complementary) nucleotides within the target nucleic acid molecule; none of these at least six nucleotides are modified nucleotides or nucleotide analogs, or any of these at least six nucleotides are modified nucleotides or nucleotide analogs.
[0156] In embodiments comprising a sequencing probe, a primary nucleic acid molecule, and a secondary nucleic acid molecule, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least 10 nucleotides capable of hybridizing to a target nucleic acid, at least 6 nucleotides within the target-binding domain capable of identifying corresponding (complementary) nucleotides within the target nucleic acid molecule, and at least 4 nucleotides within the target-binding domain not identifying corresponding nucleotides within the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least 3 attachment sites, each attachment site comprising at least 1 attachment region containing at least 1 nucleic acid sequence to which at least 1 complementary primary nucleic acid molecule binds. A sequencing probe is also provided, to which at least one complementary secondary nucleic acid molecule further binds to a complementary primary nucleic acid molecule, the complementary primary nucleic acid molecule having at least one first detectable label and at least one second detectable label, each of the at least three attachment sites corresponds to two nucleotides of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary secondary nucleic acid molecule that binds to each of the at least three attachment sites determines the position and attributes of the two corresponding nucleotides of the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0157] In embodiments comprising a sequencing probe and primary, secondary, and tertiary nucleic acid molecules, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises any construct listed in Table 1. A typical target-binding domain comprises at least eight nucleotides that can hybridize to a target nucleic acid, at least six nucleotides within the target-binding domain can identify corresponding (complementary) nucleotides in the target nucleic acid molecule, and at least two nucleotides within the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule; any of the above at least six nucleotides within the target-binding domain can be modified nucleotides or nucleotide analogs, and the at least two nucleotides within the target-binding domain that do not identify corresponding nucleotides in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain.A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which at least one complementary primary nucleic acid molecule binds, the complementary primary nucleic acid molecule further to which at least one complementary secondary nucleic acid molecule further to which at least one complementary tertiary nucleic acid molecule containing at least one first detectable label and at least one second detectable label, each of the at least three attachment sites corresponds to two nucleotides out of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary tertiary nucleic acid molecule that binds to each of the at least three attachment sites determines the position and attributes of the corresponding two nucleotides out of the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0158] In another embodiment, a representative target-binding domain may contain at least six nucleotides capable of hybridizing to a target nucleic acid, and these at least six nucleotides within the target-binding domain may identify the corresponding (complementary) nucleotides within the target nucleic acid molecule; none of these at least six nucleotides are modified nucleotides or nucleotide analogs, or any of these at least six nucleotides are modified nucleotides or nucleotide analogs.
[0159] In embodiments comprising a sequencing probe and primary nucleic acid molecules, secondary nucleic acid molecules, and tertiary nucleic acid molecules, the present disclosure provides a sequencing probe comprising a target-binding domain and a barcode domain, wherein the target-binding domain comprises at least 10 nucleotides and can hybridize to a target nucleic acid, at least 6 nucleotides in the target-binding domain can identify corresponding (complementary) nucleotides in the target nucleic acid molecule, and at least 4 nucleotides in the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule; the barcode domain comprises a synthetic skeleton and at least 3 attachment sites, each attachment site comprising at least 1 attachment region comprising at least 1 nucleic acid sequence to which at least 1 complementary primary nucleic acid molecule binds, and the complementary primary nucleic acid molecule comprises at least 1 A sequencing probe is also provided, to which a complementary secondary nucleic acid molecule further binds, and to this complementary secondary nucleic acid molecule, at least one complementary tertiary nucleic acid molecule further binds, each of which has at least one first detectable label and at least one second detectable label, and each of the at least three attachment sites corresponds to two nucleotides of the at least six nucleotides in the target binding domain, and each of the at least three attachment sites has a different nucleic acid sequence, and the at least one first detectable label and at least one second detectable label of each complementary tertiary nucleic acid molecule that binds to each of the at least three attachment sites determines the position and attribute of the two corresponding nucleotides of the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0160] This disclosure also provides sequencing probes and reporter probes having detectable labels on both secondary and tertiary nucleic acid molecules. For example, a secondary nucleic acid molecule can bind to a primary nucleic acid molecule, and the secondary nucleic acid molecule can also bind to at least one tertiary nucleic acid molecule containing both a first detectable label and at least a second detectable label, as well as a first detectable label and at least a second detectable label. The first detectable label and at least a second detectable label located on the secondary nucleic acid molecule may have the same emission spectrum or different emission spectra. The first detectable label and at least a second detectable label located on the tertiary nucleic acid molecule may have the same emission spectrum or different emission spectra. The emission spectrum of the detectable label on the secondary nucleic acid molecule may be the same emission spectrum as the emission spectrum of the detectable label on the tertiary nucleic acid molecule or different emission spectrum.
[0161] Figure 4 is a schematic diagram of a typical reporter probe of this disclosure, including a representative primary nucleic acid molecule, a secondary nucleic acid molecule, and a tertiary nucleic acid molecule. The primary nucleic acid molecule contains a first domain at its 3' end, which contains a sequence of 12 nucleotides that hybridizes to a complementary attachment region within the attachment site of the barcode domain of the sequencing probe. At its 5' end, there is a second domain that hybridizes to six secondary nucleic acid molecules. The representative secondary nucleic acid molecule shown contains a first domain at its 5' end that hybridizes to the primary nucleic acid molecule and a domain at its 3' end that hybridizes to five tertiary nucleic acid molecules.
[0162] A tertiary nucleic acid molecule contains at least two domains. The first domain can bind to a secondary nucleic acid molecule. The second domain of the tertiary nucleic acid molecule can bind to a first detectable label and at least a second detectable label. The second domain of the tertiary nucleic acid molecule can be bound to the first detectable label and at least a second detectable label by directly incorporating one or more fluorescently labeled nucleotide monomers into the sequence of the second domain of the tertiary nucleic acid molecule. The second domain of the tertiary nucleic acid molecule can be bound to the first detectable label and at least a second detectable label by hybridizing it with a short polynucleotide that serves as a label for this second domain of the tertiary nucleic acid molecule. These short polynucleotides are called “labeled oligos” and can be labeled by direct incorporation of fluorescently labeled nucleotide monomers or by other nucleic acid labeling methods known to those skilled in the art. The typical tertiary nucleic acid molecules shown in Figure 4 can be considered "labeled oligos," and include a first domain that hybridizes to a secondary nucleic acid molecule, as well as a second domain that is fluorescently labeled, for example, by indirectly attaching the label during oligonucleotide synthesis using NHS chemistry, or by incorporating one or more fluorescently labeled nucleotide monomers during the synthesis of the tertiary nucleic acid molecule. The labeled oligo can be DNA, RNA, or PNA.
[0163] The labeled oligonucleotide may contain a cleavable linker between the fluorescent moiety and the polynucleotide molecule. The cleavable linker is preferably photocatalytically cleavable. The cleavable linker can also be chemically or enzymatically cleaved.
[0164] In another embodiment, a second domain of a secondary nucleic acid can be bound to a first detectable label and at least a second detectable label. The second domain of a secondary nucleic acid can be bound to a first detectable label and at least a second detectable label by directly incorporating one or more fluorescently labeled nucleotide monomers into the second domain of the secondary nucleic acid. The second domain of a secondary nucleic acid molecule can be bound to a first detectable label and at least a second detectable label by hybridizing it with a short polynucleotide that serves as a label for this second domain of the secondary nucleic acid. These short polynucleotides are called labeled oligonucleotides and can be labeled by direct incorporation of fluorescently labeled nucleotide monomers or by other nucleic acid labeling methods known to those skilled in the art.
[0165] A primary nucleic acid molecule can contain approximately 100, 95, 90, 85, 80, or 75 nucleotides. A primary nucleic acid molecule can contain approximately 100 to 80 nucleotides. A primary nucleic acid molecule can contain approximately 90 nucleotides. A secondary nucleic acid molecule can contain approximately 90, 85, 80, 75, or 70 nucleotides. A secondary nucleic acid molecule can contain approximately 90 to 80 nucleotides. A secondary nucleic acid molecule can contain approximately 87 nucleotides. A secondary nucleic acid molecule can contain approximately 25, 20, 15, or 10 nucleotides. A tertiary nucleic acid molecule can contain approximately 20 to 10 nucleotides. A tertiary nucleic acid molecule can contain approximately 15 nucleotides.
[0166] Various designs are possible for the reporter probes of this disclosure. For example, a primary nucleic acid molecule can be hybridized to at least one (e.g., one, two, three, four, five, six, seven, eight, nine, ten, or more) secondary nucleic acid molecules. Each secondary nucleic acid molecule can be hybridized to at least one (e.g., one, two, three, four, five, six, seven, eight, nine, ten, or more) tertiary nucleic acid molecules. To produce a reporter probe labeled with a specific two-color combination, a reporter probe is designed that includes a secondary nucleic acid molecule or a tertiary nucleic acid molecule or a labeled oligo, or any combination of a secondary nucleic acid molecule, a tertiary nucleic acid molecule, and a labeled oligo, each labeled with one of the colors of that specific two-color combination. For example, Figure 4 shows a reporter probe of this disclosure containing a total of 30 dyes (15 dyes for color 1 and 15 dyes for color 2). To prevent color exchange or cross-hybridization between different fluorescent dyes, each tertiary nucleic acid or labeled oligo bound to a specific label or fluorescent dye contains its own unique nucleotide sequence.
[0167] In some embodiments, the present disclosure provides a 5×5 reporter probe. The 5×5 reporter probe comprises a primary nucleic acid molecule, which comprises a first domain consisting of 12 nucleotides. The primary nucleic acid molecule also comprises a second domain, which comprises a nucleotide sequence capable of hybridizing to five secondary nucleic acid molecules. Each secondary nucleic acid molecule comprises a nucleotide sequence capable of hybridizing to its respective secondary nucleic acid with five tertiary nucleic acids bound to a detectable label.
[0168] In some embodiments, the present disclosure provides a 4x3 reporter probe. The 4x3 reporter probe comprises a primary nucleic acid molecule, which comprises a first domain consisting of 12 nucleotides. The primary nucleic acid molecule also comprises a second domain, which comprises a nucleotide sequence capable of hybridizing into four secondary nucleic acid molecules. Each secondary nucleic acid molecule comprises a nucleotide sequence capable of hybridizing into its respective secondary nucleic acid with three tertiary nucleic acids bound to a detectable label.
[0169] In some embodiments, the present disclosure provides a 3×4 reporter probe. The 3×4 reporter probe comprises a primary nucleic acid molecule, which comprises a first domain consisting of 12 nucleotides. The primary nucleic acid molecule also comprises a second domain, which comprises a nucleotide sequence capable of hybridizing to three secondary nucleic acid molecules. Each secondary nucleic acid molecule comprises a nucleotide sequence capable of hybridizing to its respective secondary nucleic acid with four tertiary nucleic acids bound to a detectable label.
[0170] In some embodiments, the present disclosure provides a spacer 3x4 reporter probe. The spacer 3x4 reporter probe comprises a primary nucleic acid molecule, which comprises a first domain consisting of 12 nucleotides. Between the first and second domains of the primary nucleic acid molecule lies a spacer region consisting of 20 to 40 nucleotides. The spacer is identified as having a length of 20 to 40 nucleotides, but the length of the spacer is not limited and can be shorter than 20 nucleotides or longer than 40 nucleotides. The second domain of the primary nucleic acid molecule contains a nucleotide sequence that can hybridize to three secondary nucleic acid molecules. Each secondary nucleic acid contains a nucleotide sequence that allows four tertiary nucleic acids, each bound to a detectable label, to hybridize to the respective secondary nucleic acid.
[0171] In some embodiments, the primary nucleic acid may contain a first domain having a length of 12 nucleotides. However, there is no limit to the length of the first domain of the primary nucleic acid; it can have fewer than 12 nucleotides or 13 or more nucleotides. In one example, the first domain of the primary nucleic acid has 14 nucleotides. In another example, the first domain of the primary nucleic acid has 9 nucleotides. In yet another example, the first domain of the primary nucleic acid has 8 nucleotides. The 9-nucleotide first domain of the primary nucleic acid of the reporter probe includes those shown in Table 15.
[0172] [Table 2]
[0173] Any feature of one particular reporter probe design according to this disclosure can be combined with any feature of another reporter probe design according to this disclosure. For example, a 5×5 reporter probe can be modified to include a spacer region of approximately 20 to 40 nucleotides between the complementary nucleic acid and the primary nucleic acid. In another example, a 4×3 reporter probe can be modified to create a 4×5 reporter probe in which four secondary nucleic acids contain nucleotide sequences that allow five tertiary nucleic acids, bound to a detectable label, to hybridize to each of the secondary nucleic acids.
[0174] If we disregard theory, the fluorescence intensity of a 5x5 reporter probe will be greater because it contains more fluorescent labels (25) than a 4x3 reporter probe (12). The fluorescence detected in a given arbitrary field of view (FOV) is a function of a variety of variables (including the fluorescence intensity of a given reporter probe and the number of arbitrarily bound target molecules within that FOV). The number of arbitrarily bound target molecules per FOV can range from 1 million to 2.5 million targets. Typical values for the number of bound target molecules per FOV are 20,000 to 40,000, or 220,000 to 440,000, or 1 million to 2 million target molecules. A typical FOV is 0.05 mm. 2 ~1 mm 2 A further example of a typical FOV is 0.05 mm. 2 ~0.65 mm 2 That is the case.
[0175] In some embodiments, the present disclosure provides a reporter probe design in which the secondary nucleic acid molecule includes an “extra-handle” located distal to the primary nucleic acid molecule and not hybridized to the tertiary nucleic acid molecule. In some embodiments, the “extra-handle” can be 12 nucleotides long (“dodecamer”), but is not limited to that length, and can be fewer than 12 nucleotides or more than 12 nucleotides. Each “extra-handle” can contain the nucleotide sequence of the first domain of the primary nucleic acid molecule into which the secondary nucleic acid molecule hybridizes. Thus, when the reporter probe includes an “extra-handle”, the reporter probe can hybridize to the sequencing probe via the first domain of the primary nucleic acid molecule or via the “extra-handle”. This increases the probability that the reporter probe will bind to the sequencing probe. This design of the “extra-handle” can also improve the dynamics of hybridization. In theory, the “extra-handle” can increase the effective concentration of the complementary nucleic acid in the reporter probe. A 5×4 "additional handle" reporter probe is expected to produce approximately 4750 fluorescence counts per standard FOV. The 5×3 "additional handle" reporter probe, 4×4 "additional handle" reporter probe, 4×3 "additional handle" reporter probe, and 3×4 "additional handle" reporter probe are all expected to produce approximately 6000 fluorescence counts per standard FOV. Any reporter probe design in this disclosure can be modified to include "additional handles."
[0176] Each secondary nucleic acid molecule in a reporter probe can hybridize to a tertiary nucleic acid molecule labeled with the same detectable label. For example, the left panel of Figure 5 shows a "5×6" reporter probe. The 5×6 reporter probe contains one primary nucleic acid including a second domain, which contains a nucleotide sequence that hybridizes to six secondary nucleic acid molecules. Each secondary nucleic acid contains a nucleotide sequence that allows five tertiary nucleic acid molecules bound to a detectable label to hybridize to that secondary nucleic acid. Each of the five tertiary nucleic acid molecules bound to a particular secondary nucleic acid molecule is labeled with the same detectable label. For example, three secondary nucleic acid molecules are bound to a tertiary nucleic acid molecule labeled with a yellow fluorescent dye, and the other three secondary nucleic acid molecules are bound to a tertiary nucleic acid molecule labeled with a red fluorescent dye.
[0177] Each secondary nucleic acid molecule in a reporter probe can hybridize to tertiary nucleic acid molecules labeled with different detectable labels. For example, the center diagram in Figure 5 shows a "3×2×6" reporter probe design. The 3×2×6 reporter probe contains one primary nucleic acid including a second domain, which contains a nucleotide sequence that hybridizes to six secondary nucleic acid molecules. Each secondary nucleic acid contains a nucleotide sequence that allows five tertiary nucleic acid molecules, bound to a detectable label, to hybridize to that secondary nucleic acid. Each secondary nucleic acid binds to both tertiary nucleic acid molecules labeled with a yellow fluorescent dye and tertiary nucleic acid molecules labeled with a red fluorescent dye. In this specific example, three secondary nucleic acid molecules bind to two red tertiary nucleic acid molecules and three yellow tertiary nucleic acid molecules, while the other three secondary nucleic acid molecules bind to two red tertiary nucleic acid molecules and three yellow tertiary nucleic acid molecules. Each secondary nucleic acid molecule can bind to any number of tertiary nucleic acid molecules, bound to different detectable labels. In the central diagram of Figure 5, the tertiary nucleic acid molecules bound to individual secondary nucleic acid molecules are arranged so that the label colors alternate (i.e., red-yellow-red-yellow-red, or yellow-red-yellow-red-yellow).
[0178] In any of the reporter probe designs described, tertiary nucleic acids labeled with different detectable labels can be arranged in any order along the secondary nucleic acids. For example, the right panel of Figure 5 shows a "FRET-resistant 3×2×6" reporter probe. This is similar to the 3×2×6 reporter probe design, but differs in the arrangement (e.g., linear order or grouping) of the red and yellow tertiary nucleic acid molecules along each secondary nucleic acid molecule.
[0179] Figure 6 shows a more representative reporter probe design of this disclosure, which includes individual secondary nucleic acid molecules that bind to various tertiary nucleic acid molecules. The left figure shows a "6 × 1 × 4.5" reporter probe containing one primary nucleic acid molecule. This primary nucleic acid molecule contains a second domain, which contains a nucleotide sequence that hybridizes to six secondary nucleic acid molecules. Each secondary nucleic acid molecule hybridizes to five tertiary nucleic acid molecules. Four of the five tertiary nucleic acid molecules that hybridize to each secondary nucleic acid molecule are directly labeled with a detectable label of the same color. The fifth tertiary nucleic acid (denoted as branched tertiary nucleic acid) is bound to five labeled oligonucleotides that are the other color of the two-color combination. Three of the six secondary nucleic acids bind to branched tertiary nucleic acids labeled with one color of the two-color combination (red in this example), while the remaining three secondary nucleic acids bind to branched tertiary nucleic acids labeled with the other color of the two-color combination (yellow in this example). In summary, the 6×1×4.5 reporter probe is labeled with a total of 54 dyes (27 dyes for each color). The center figure in Figure 6 shows the "4×1×4.5" reporter probe. This 4×1×4.5 reporter probe shares the same overall architecture as the 6×1×4.5 reporter probe, but differs in that the primary nucleic acid binds to only four secondary nucleic acids, resulting in a total of 36 dyes, with 18 dyes for each color.
[0180] A reporter probe can contain an equal number of dyes for each color in a two-color combination. Alternatively, a reporter probe can contain different numbers of dyes for each color in a two-color combination. The choice of which color's dye to use in greater numbers within the reporter probe is based on the energy levels of light absorbed by the two dyes. For example, the right-hand diagram in Figure 6 shows a "5x5 energy optimized" reporter probe design. This reporter probe design contains 15 yellow dyes (higher energy) and 10 red dyes (lower energy). In this example, the 15 yellow dyes can constitute the first marker, and the 10 red dyes can constitute the second marker.
[0181] The detectable portion, label, or reporter can be attached to a secondary nucleic acid molecule, a tertiary nucleic acid molecule, or a labeled oligo in various ways (direct or indirect attachment of the detectable portion (fluorescent portion, colorimetric portion, etc.)). Those skilled in the art can refer to references on nucleic acid labeling. Non-limiting examples of fluorescent portions include, for example, yellow fluorescent protein (YFP), green fluorescent protein (GFP), cyan fluorescent protein (CFP), red fluorescent protein (RFP), umbelliferone, fluorescein, fluorescein isothiocyanate, rhodamine, dichlorotriazinylamine fluorescein, cyanine, dansilloride, phycocyanin, and phycoerythrin.
[0182] Fluorescent labeling and the attachment of fluorescent labels to nucleotides and / or oligonucleotides are described in numerous publications, including Haugland, *Handbook of Fluorescent Probes and Research Chemicals*, 9th edition (Molecular Probes, Inc., Eugene 2002); Keller and Manak, *DNA Probes*, 2nd edition (Stockton Press, New York, 1993); Eckstein (eds.), *Oligonucleotides and Analogues: A Practical Approach* (IRL Press, Oxford, 1991); and Wetmur, *Critical Reviews in Biochemistry and Molecular Biology*, Vol. 26: pp. 227-259 (1991). Specific methods applicable to this disclosure are described in U.S. Patents 4,757,141, 5,151,507, and 5,091,519, which are sample references. One or more fluorescent dyes can be used as labels for labeled target sequences, and examples of such dyes are disclosed, for example, U.S. Patent No. 5,188,934 (4,7-dichlorofluorescein dye); No. 5,366,860 (spectrally decomposable rhodamine dye); No. 5,847,162 (4,7-dichlororhodamine dye); No. 4,318,846 (ether-substituted fluorescein dye); No. 5,800,996 (energy transfer dye); Lee et al., No. 5,066,580 (xanthine dye); No. 5,688,648 (energy transfer dye), etc. The labeling operation can also be carried out using quantum dots, which are disclosed in the following patents and publications: U.S. Patent Nos. 6,322,901; 6,576,291; 6,423,551; 6,251,303; 6,319,426; 6,426,513; 6,444,143; 5,990,479; 6,207,392; U.S. Patent Application Publication Nos. 2002 / 0045045; 2003 / 0017264.In this specification, the term "fluorescent labeling" includes signaling segments that transmit information through the fluorescent absorption and / or emission properties of one or more molecules. Such fluorescent properties include fluorescence intensity, fluorescence lifetime, emission spectral characteristics, and energy transfer.
[0183] Non-exclusive examples of commercially available fluorescent nucleotide analogs that are easily incorporated into nucleotide and / or oligonucleotide sequences include Cy3-dCTP, Cy3-dUTP, Cy5-dCTP, Cy5-dUTP (Amersham Biosciences, Piscataway, New Jersey), fluorescein-12-dUTP, tetramethylrhodamine-6-dUTP, TEXAS RED™-5-dUTP, CASCADE BLUE™-7-dUTP, BODIPY™FL-14-dUTP, BODIPY™MR-14-dUTP, BODIPY™TR-14-dUTP, RHODAMINE GREEN™-5-dUTP, OREGON GREEN™488-5-dUTP, TEXAS RED™-12-dUTP, and BODIPY™ 630 / 650-14-dUTP, BODIPY™ 650 / 665-14-dUTP, ALEXA FLUOR™ 488-5-dUTP, ALEXA FLUOR™ 532-5-dUTP, ALEXA FLUOR™ 568-5-dUTP, ALEXA FLUOR™ 594-5-dUTP, ALEXA FLUOR™ 546-14-dUTP, Fluorescein-12-UTP, Tetramethylrhodamine-6-UTP, TEXAS RED™-5-UTP, mCherry, CASCADE BLUE™-7-UTP, BODIPY™ FL-14-UTP, BODIPY™ MR-14-UTP, BODIPY™ TR-14-UTP, RHODAMINE Examples include GREEN(trademark)-5-UTP, ALEXA FLUOR(trademark) 488-5-UTP, and LEXA FLUOR(trademark) 546-14-UTP (Molecular Probes, Inc., Eugene, Oregon). Alternatively, the above-mentioned phosphors and those referred to herein can be added during the synthesis of oligonucleotides, for example, using phosphoramidite chemistry or NHS chemistry. Protocols for the custom synthesis of nucleotides with other phosphors are known in this field (see Henegariu et al. (2000) Nature Biotechnol. Vol. 18: p. 345).2-aminopurines are fluorescent bases that can be directly incorporated into oligonucleotide sequences during synthesis. Nucleic acids can also be pre-stained with intercalating dyes (such as DAPI, YOYO-1, ethidium bromide, or cyanine dyes (e.g., SYBR Green)).
[0184] Other phosphors that can be used for post-synthesis attachment include, but are not limited to, ALEXA FLUOR® 350, ALEXA FLUOR® 405, ALEXA FLUOR® 430, ALEXA FLUOR® 532, ALEXA FLUOR® 546, ALEXA FLUOR® 568, ALEXA FLUOR® 594, ALEXA FLUOR® 647, BODIPY 493 / 503, BODIPY FL, BODIPY R6G, BODIPY 530 / 550, BODIPY TMR, BODIPY 558 / 568, BODIPY 558 / 568, BODIPY 564 / 570, BODIPY 576 / 589, BODIPY 581 / 591, BODIPY TR, BODIPY These include 630 / 650, BODIPY 650 / 665, Cascade Blue, Cascade Yellow, Dansyl, Lissamin Rhodamine B, Marina Blue, Oregon Green 488, Oregon Green 514, Pacific Blue, Pacific Orange, Rhodamine 6G, Rhodamine Green, Rhodamine Red, Tetramethylrhodamine, Texas Red (available from Molecular Probes, Inc., Eugene, Oregon), Cy2, Cy3, Cy3.5, Cy5, Cy5.5, Cy7 (Amersham Biosciences, Piscataway, New Jersey), etc. FRET tandem phosphors can also be used, and non-limiting examples include PerCP-Cy5.5, PE-Cy5, PE-Cy5.5, PE-Cy7, PE-Texas Red, APC-Cy7, PE-Alexa dyes (610, 647, 680), and APC-Alexa dyes.
[0185] Metallic silver particles or metallic gold particles can be used to enhance signals from fluorescently labeled nucleotide and / or oligonucleotide sequences (Lakowicz et al. (2003) BioTechniques Vol. 34: p. 62).
[0186] Other labels suitable for oligonucleotide sequences can be included, such as fluorescein (FAM, FITC), digoxigenin, dinitrophenol (DNP), dansyl, biotin, bromooxyuridine (BrdU), hexahistidine (6×His), and phospho-amino acids (e.g., P-tyr, P-ser, P-thr). The following hapten / antibody pairs, namely biotin / α-biotin, digoxigenin / α-digoxigenin, dinitrophenol (DNP) / α-DNP, and 5-carboxyfluorescein (FAM) / α-FAM, can be used for detection, with each antibody being derivatized using a detectable label.
[0187] The detectable labels described herein are spectrally resolvable. “Spectrally resolvable” with respect to multiple fluorescent labels means that the fluorescence emission bands of the multiple labels are sufficiently distinct from each other (i.e., sufficiently non-overlapping) so that the molecular tags to which each label is attached can be identified by a standard photodetector system (e.g., using a system of band-pass filters and photomultiplier tubes) based on the fluorescence signal emitted by each label. Examples of such systems include U.S. Patents 4,230,558 and 4,811,218, or the systems described on pages 21-76 of *Flow Cytometry: Instrumentation and Data Analysis* by Wheeles et al. (Academic Press, New York, 1985). Spectrally resolvable organic dyes (e.g., fluorescein, rhodamine) mean that their maximum emission wavelengths are at least 20 nm apart from each other, or in another embodiment, at least 40 nm apart. Spectrally resolvable with respect to chelated lanthanide compounds, quantum dots, etc., means that their maximum emission wavelengths are at least 10 nm apart from each other, or at least 15 nm apart.
[0188] Based on the existence of three attachment locations within a barcode domain, and the fact that each attachment location can have up to 10 possible two-color combinations, up to 1000 color combinations are possible. When pooling reporter probes in groups of less than 1000 probes per pool, the ability to use parity checks to overcome errors can be utilized. Many potential parity schemes can exist to enable parity checks, and an example of such a scheme is shown in Figure 32. In this example, the actual colors present are not used for the parity check; rather, the presence of single-color (S) reporter probes (e.g., red) and multi-color (M) reporter probes (e.g., red / yellow) at each attachment location within the barcode domain is used. As seen in the parity design, knowledge of the state (S or M) of any two reporter locations allows for the prediction of a third location. In the example shown, if S is observed at any two locations, the unobserved location must be M; if S and M are observed at any two locations, the other location must be S; and if two M reporter probes are observed, the other location must be M. This means that two incorrect readouts are required to obtain the codes for three reporter probes whose detected reporter colors are inaccurate. Figure 32 shows the simulation results when the reporter probe error rate is 5%, and it can be seen that error exclusion increases when parity checks are applied. Numerous applicable parity systems exist, and this is just one example.
[0189] Another error correction routine involves swapping color palettes for each pool of reporter probes. A color palette is the set of reporter probes actually used to measure a pool. If there are many reporter probes not used in any pool, and 500 reporter probes exist in one pool, then only half of the possible color combinations are needed. The simplest way to achieve this is to have two palettes: palette A containing 500 reporter probes and palette B containing another 500 reporter probes. Therefore, if sequencing pools 1, 3, 5, and 7 have palette A, and pools 2, 4, 6, and 8 have palette B, then the running pool order 1, 2, 3, 4, 5, 6, 7, 8 means that each consecutive sequencing pool has a separate palette. Thus, barcodes from pool 2 will not exist in the preceding or following pool (e.g., pools 1 and 3). This allows for easy automation of troubleshooting and limits error detection.
[0190] A reporter probe may include one or more cleavable linker modifications. These one or more cleavable linker modifications can be located anywhere within the reporter probe. The cleavable linker modifications can be located between the first and second domains of the primary nucleic acid molecule of the reporter probe. The cleavable linker modifications can be located between the first and second domains of the secondary nucleic acid molecule of the reporter probe. The cleavable linker modifications can be located between the first and second domains of the primary and secondary nucleic acid molecules of the reporter probe. The left panel of Figure 7 shows a typical reporter probe of this disclosure, which includes cleavable linker modifications between the first and second domains of the primary nucleic acid and cleavable linker modifications between the first and second domains of the secondary nucleic acid. In cases like the one illustrated in the left panel of Figure 7, the cleavable linker modifications may include one or more cleavable portions (as illustrated in the left panel of Figure 7).
[0191] As a modification of a cleavable linker, a compound of formula (I):
Chemical formula
[0192] In one aspect, R1 is C 1-6 Alkyl, preferably C 1-3 Alkyl (methyl, ethyl, propyl, isopropyl, etc.); R2 is NH or N(C) 1-6 R3 is an alkyl group; R3 is a 5- to 6-membered cycloalkyl group, preferably a cyclohexyl group; R4 is C 1-6 Alkylene, preferably C1-3 is an alkylene (such as methylene, ethylene, propylene, isopropylene, etc.); R5 is a 5- to 6-membered heterocyclyl containing one nitrogen atom and optionally 0 or 1 additional heteroatom selected from N, O, S, and this heterocyclyl is optionally substituted with one or two Rs 10 ; R6 is O; R7 is C 1-6 alkylene, preferably C 1-3 alkylene (such as methylene, ethylene, propylene, isopropylene, etc.); R8 is O; R9 is a 5- to 6-membered heterocyclyl containing one nitrogen atom and optionally 0 or 1 additional heteroatom selected from N, O, S, and this heterocyclyl is optionally substituted with one or two Rs 10 ; each R 10 is independently halogen, C 1-6 alkyl, halo C 1-6 alkyl, oxo, -SO2H, -SO3 - or any of them.
[0193] In one embodiment, R3 is cyclohexyl, R4 is methylene, R5 is 1H-pyrrole-2,5-dione, and R9 is pyrrolidine-2,5-dione optionally substituted with SO3 - .
[0194] As a linker compound, [Chemical formula] or its stereoisomers, or its salts are possible.
[0195] As a linker compound, [Chemical formula] or its stereoisomers, or its salts are possible.
[0196] As a linker compound or a modification of the linker, ?[Chemical formula] Either of the following is possible.
[0197] As a linker compound or linker modification, [ka] Either of the following is possible.
[0198] As a modification or severable part of a severable linker, [ka] This is possible.
[0199] The reporter probe can be constructed by mixing three types of storage solutions with water. One storage solution contains primary nucleic acid molecules, one storage solution contains secondary nucleic acid molecules, and the last storage solution contains tertiary nucleic acid molecules. Table 2 shows typical amounts of each storage solution that can be mixed to construct the reporter probe for each individual design.
[0200] [Table 3]
[0201] target nucleic acid
[0202] This disclosure provides a method for sequencing nucleic acids using the sequencing probes disclosed herein. The nucleic acids sequenced using the methods of this disclosure are referred to herein as “target nucleic acids.” The term “target nucleic acids” means nucleic acid molecules (DNA, RNA, PNA) whose sequences are to be determined by the probes, methods, and apparatus of this disclosure. Generally, the terms “target nucleic acids,” “target nucleic acid molecules,” “target nucleic acid sequences,” “target nucleic acid fragments,” “target oligonucleotides,” and “target polynucleotides” are interchangeable and include, in non-limiting examples, nucleotides (deoxyribonucleotides or ribonucleotides) or analogues in polymeric forms of various lengths. Non-limiting examples of nucleic acids include genes, gene fragments, exons, introns, intergenetic DNA (non-limiting examples of which include heterochromatin DNA), messenger RNA (mRNA), transfer RNA, ribosomal RNA, ribozymes, small interfering RNA (siRNA), non-coding RNA (ncRNA), cDNA, recombinant polynucleotides, branched polynucleotides, plasmids, vectors, DNA isolated from a single sequence, RNA isolated from a single sequence, nucleic acid probes, and primers. Prior to sequencing using the methods of this disclosure, the attributes and / or sequence of the target nucleic acid are known, or the attributes and / or sequence are unknown. It is also possible that a portion of the sequence of the target nucleic acid is known prior to sequencing using the methods of this disclosure. For example, this method can be applied to reveal point mutations in a known target nucleic acid molecule.
[0203] The methods of this disclosure directly sequence nucleic acid molecules obtained from a sample (e.g., a sample from a biological organism), and preferably, there is no conversion (or amplification) step. For example, in RNA-based direct sequencing, the methods of this disclosure do not require conversion from RNA molecules to DNA molecules (i.e., through cDNA synthesis) before a sequence can be obtained. Because no amplification or conversion is required, the nucleic acids sequenced in this disclosure retain all unique bases and / or epigenetic markers present in the nucleic acid when it is present in the sample or when it is obtained from the sample. Such unique bases and / or epigenetic markers are lost in sequencing methods known in the art.
[0204] The method disclosed herein enables sequencing with single-molecule resolution. In other words, the method disclosed herein allows users to generate a final sequence based on data recovered from a single target nucleic acid molecule, without the need to combine data from different target nucleic acid molecules, thus preserving all the unique characteristics of that particular target.
[0205] Target nucleic acids can be obtained from any sample or source of nucleic acids (e.g., any cells, tissues, organisms, in vitro, chemical synthesis equipment, etc.). Target nucleic acids can be obtained by any method known in the art. Nucleic acids can be obtained from blood samples of clinical subjects. Nucleic acids can be extracted, isolated, or purified from sources or samples using methods and kits known in the art.
[0206] The target nucleic acid can be fragmented by any means known in the art. Fragmentation is preferably carried out by enzymatic or mechanical means. Mechanical means may include sonication or physical shearing. Enzymatic means may be carried out by digestion using a nuclease (e.g., deoxyribonuclease I (DNase I)) or one or more restriction endonucleases.
[0207] When the nucleic acid molecule containing the target nucleic acid is a single complete chromosome, it is necessary to perform a step of avoiding fragmentation of the chromosome.
[0208] As is well known in the art, the target nucleic acid can contain natural nucleotides or unnatural nucleotides, among which modified nucleotides or nucleic acid analogs are included.
[0209] The target nucleic acid molecule can include DNA molecules, RNA molecules, and PNA molecules with lengths up to several hundred kilobases (e.g., 1 kilobase, 2 kilobases, 3 kilobases, 4 kilobases, 5 kilobases, 10 kilobases, 20 kilobases, 30 kilobases, 40 kilobases, 50 kilobases, 100 kilobases, 200 kilobases, 500 kilobases, or a larger number of kilobases). The target nucleic acid molecule can contain from about 50 to about 400 nucleotides, or from about 90 to about 350 nucleotides.
[0210] Capture probe
[0211] The target nucleic acid molecule can be immobilized on a substrate (e.g., at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more positions).
[0212] Representative examples of useful substrates include substrates containing binding sites, the selection of which is made from a group consisting of ligands, antigens, carbohydrates, nucleic acids, receptors, lectins, and antibodies. Capture probes contain substrate binding sites that can bind to the binding sites on the substrate. Non-limiting examples of useful representative substrates containing reactive sites include surfaces containing any of the following: epoxy, aldehydes, gold, hydrazides, sulfhydryls, NHS esters, amines, alkynes, azides, thiols, carboxylates, maleimides, hydroxymethylphosphines, imide esters, isocyanates, hydroxyls, pentafluorophenyl esters, psoralens, pyridyl disulfide or vinyl sulfones, polyethylene glycol (PEG), hydrogels, or mixtures thereof. Such surfaces can be obtained from commercial suppliers or prepared according to standard techniques. Non-limiting examples of useful representative substrates containing reactive moieties include OptArray-DNA NHS group (Accler8), Nexterion Slide AL (Schott), and Nexterion Slide E (Schott).
[0213] Any solid support known in the art capable of immobilizing the target nucleic acid can be used as the substrate (e.g., coated slides, microfluidic devices). Possible substrates include surfaces, films, beads, porous materials, electrodes, and arrays. Other possible substrates include polymer materials, metals, silicon, glass, and quartz. The target nucleic acid can be immobilized on the surface of any substrate obvious to those skilled in the art.
[0214] When the substrate is an array, it can contain multiple wells. The size and spacing of the wells vary depending on the target nucleic acid molecules to be attached. In one example, the substrate is configured to accommodate an array of target nucleic acids arranged in an ultra-high density. An example of the density of the target nucleic acid array on the substrate is 1 mm 2 500,000 to 10,000,000 target nucleic acid molecules per 1 mm 2 1,000,000 to 4,000,000 target nucleic acid molecules per 1 mm2 Each sample contains approximately 850,000 to 3,500,000 target nucleic acid molecules.
[0215] The wells within the substrate are the sites where target nucleic acid molecules are attached. By functionalizing the surface of the wells using the reactive moieties described above, specific chemical groups present on the surface of the target nucleic acid molecule or on the surface of a capture probe bound to the target nucleic acid molecule can be attracted, immobilized, and bound to it. These functional groups are well known to be able to specifically attract and bind biomolecules through various conjugation chemistry processes.
[0216] To sequence a single nucleic acid molecule on a substrate (such as an array), a universal capture probe, or a universal sequence complementary to the substrate-binding portion of the capture probe, is attached to each well. Then, a single target nucleic acid molecule is bound to the universal capture probe, or to the universal sequence complementary to the substrate-binding portion of the capture probe and bound to the capture probe, thereby initiating sequencing.
[0217] To sequence a single nucleic acid molecule on a substrate (such as an array), a single target nucleic acid molecule can be bound to a capture probe. The substrate-binding portion of the capture probe can then be bound to an adapter oligonucleotide. Finally, by binding the adapter nucleotide to a lone oligonucleotide attached to each well, sequencing can be initiated. Representative sequences for lone oligonucleotides are shown in Table 8.
[0218] [Table 4] [ka]
[0219] Each nucleic acid contained in a lawn oligonucleotide or adapter oligonucleotide can independently be a normal base, a modified nucleotide, or a nucleic acid analog. Typical modified nucleotides or nucleic acid analogs useful in lawn oligonucleotides or adapter oligonucleotides are isoguanine and isocytosine. Alternatively, each nucleic acid contained in a lawn oligonucleotide can independently be L-DNA. In some embodiments, a lawn oligonucleotide can contain L-DNA. A lawn oligonucleotide can consist mainly of L-DNA. Using modified nucleotides or nucleotide analogs (such as isoguanine, isocytosine, and L-DNA) can improve the efficiency and accuracy of binding of the adapter oligonucleotide to a suitable complementary nucleic acid sequence within the lawn oligonucleotide, for example, while minimizing binding to other sites.
[0220] The lone oligonucleotide further contains a 5' amine having a 6-carbon linker (referred to herein as 5AmMC6). 5AmMC6 can be used to attach the lone oligonucleotide to a substrate.
[0221] Figure 33 shows an example of a complex of a capture probe, adapter oligonucleotide, and lone oligonucleotide. In this figure, the sequence of the representative adapter and the sequence of the representative capture probe that hybridize to the target nucleic acid (gene TP53.1 in this example) are shown in green, the sequence of the reverse complement of the representative lone oligo is shown in blue, and the sequence on the capture probe that hybridizes to the target nucleic acid is shown in red. The sequence of the representative capture probe is 3'-CCGGTCAACCGTTTTGTAGAACAACTCCCGTCCCCTCACTCACTAGCCTCCAGTACCGAAAGC-5' (SEQ ID NO: 111). The sequence of the representative adapter is 5'-GAGTGATCGGAG The sequence is GTCATGGCTTTCGAC / iMe-isodC / CTA / iMe-isodC / AAA / iMe-isodC / TCA / iMe-isodC / TA / iMe-isodC / TA / iMe-isodC / CAA / iMe-isodC / AAC / iMe-isodC / TCA / iMe-isodC / CA-3' (SEQ ID NO: 110). A typical sequence of a lone oligonucleotide is TG / iisodG / GAT / iisodG / TTT / iisodG / AGT / iisodG / AT / iisodG / AT / iisodG / GTT / iisodG / TTG / iisodG / AGT / iisodG / GT / 5AmMC6 (SEQ ID NO: 108).
[0222] In some embodiments, the lone oligonucleotide may include at least one affinity moiety, or at least two affinity moieties, or at least three affinity moieties, or at least four affinity moieties, or at least five affinity moieties, or at least six affinity moieties, or at least seven affinity moieties, or at least eight affinity moieties, or at least nine affinity moieties, or at least ten affinity moieties. Biotin can be an affinity moiety. Thus, the lone oligonucleotide may include at least one biotin moiety, or at least two biotin moieties, or at least three biotin moieties, or at least four biotin moieties, or at least five biotin moieties, or at least six biotin moieties, or at least seven biotin moieties, or at least eight biotin moieties, or at least nine biotin moieties, or at least ten biotin moieties.
[0223] In some embodiments, the capture probe of the present disclosure that hybridizes to a target nucleic acid may include at least one affinity moiety (a non-limiting example being a biotin moiety). The capture probe that hybridizes to the target nucleic acid may then hybridize directly or indirectly to at least one lone oligonucleotide on a substrate. This at least one lone oligonucleotide includes at least one affinity moiety (a non-limiting example being a biotin moiety). After the capture probe hybridizes to the lone oligonucleotide, the resulting capture probe-target nucleic acid-lone oligonucleotide complex can be incubated with a second affinity moiety, the second affinity moiety which can bind to a first affinity moiety located on the capture probe and a first affinity moiety located on the lone oligonucleotide. In a non-limiting example, if both the first affinity moiety located on the capture probe and the first affinity moiety located on the lone oligonucleotide are biotin, then neutraavidin can be used as the second affinity moiety. This second affinity moiety binds to the first affinity moiety located on the capture probe and the first affinity moiety located on the lone oligonucleotide, creating a protein crosslink referred to herein as a "protein lock." This protein lock allows for more stable immobilization of the target nucleic acid onto the substrate. Figure 67 shows a schematic diagram of a protein lock using a biotinylated capture probe, a lone oligonucleotide, and neutraavidin.
[0224] One or more capture probes (i.e., two, three, four, five, six, seven, eight, nine, ten, or more) can be bound to the target nucleic acid. Each capture probe contains a domain complementary to a portion of the target nucleic acid and a domain containing a substrate-binding portion. The portion of the target nucleic acid complementary to the capture probe can be one end of the target nucleic acid, or a portion not near the end. The capture probe may contain a cleavable portion between the domain complementary to the target nucleic acid and the domain containing the substrate-binding portion.
[0225] Alternatively, the capture probe may include a first domain complementary to a portion of the target nucleic acid, a second domain containing a substrate-binding portion, and a third domain containing a different substrate-binding portion. The capture probe may also include cleavable portions between any of the domains.
[0226] The capture probe can be phosphorylated at its 5' end. Alternatively, the capture probe can contain at least one phosphorothioate bond. The capture probe can contain at least two phosphorothioate bonds. These at least two phosphorothioate bonds are preferably located at the 5' end of the capture probe.
[0227] Biotin can be used as the substrate binding portion of the capture probe, and avidin (e.g., streptavidin) can be used as the substrate. Useful avidin-containing substrates are commercially available and include TB0200 (Accelr8), SAD6, SAD20, SAD100, SAD500, SAD2000 (Xantec), SuperAvidin (Array-It), streptavidin slide (catalog number MPC 000, Xenopore), and STREPTAVIDINnslide (catalog number 439003, Greiner Bio-one). Avidin (e.g., streptavidin) can be used as the substrate binding portion of the capture probe, and biotin can be used as the substrate. Non-exclusive examples of commercially available useful substrates containing biotin include Optiarray-Biotin (Accler8), BD6, BD20, BD100, BD500, and BD2000 (Xantec).
[0228] A reactive portion that can bind to the substrate by photoactivation is possible as the substrate-binding portion of the capture probe. The substrate may contain a photoactivatable portion, or the first portion of the nanoreporter may contain a photoactivatable portion. Some examples of photoactivatable portions include aryl azides (such as 4-azido-2,3,5,6-tetrafluorobenzoic acid); benzophenone-based reagents (such as 4-benzoylbenzoate succinimidyl); and 5-bromodoxyuridine.
[0229] The substrate binding portion of the capture probe can be a nucleic acid that can hybridize to a complementary binding portion of the substrate. Each nucleic acid contained in the substrate binding portion of the capture probe can independently be a normal base, a modified nucleotide, or a nucleic acid analog. At least one, at least two, at least three, at least four, at least five, or at least six nucleotides in the substrate binding portion of the capture probe can be a modified nucleotide or a nucleic acid analog. The typical ratio of modified nucleotides or nucleotide analogs to normal bases in the substrate binding portion of the capture probe is 1:2 to 1:8. Typical modified nucleotides or nucleotide analogs useful in the substrate binding portion of the capture probe are isoguanine and isocytosine.
[0230] The substrate-binding portion of the capture probe can be immobilized on the substrate via other binding pairs obvious to those skilled in the art. After binding to the substrate, the target nucleic acid can be stretched by applying sufficient force to elongate it (e.g., gravity, hydrodynamic force, electromagnetic force "electrostretching", flow-stretching, receding meniscus techniques, or a combination thereof). The capture probe may include a labelable label (i.e., a reference spot) or be associated with a labelable label.
[0231] A second capture probe can bind to the target nucleic acid, containing a domain complementary to a second portion of the target nucleic acid. The second portion of the target nucleic acid to which the second capture probe binds is different from the first portion to which the first capture probe binds. This portion can be one end of the target nucleic acid, or a portion not near the end. Binding of the second capture probe can occur after the target nucleic acid has elongated, during elongation, or to a target nucleic acid that has not elongated. The second capture probe may have the binding described above.
[0232] A third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth capture probe can be bound to the target nucleic acid, each containing a domain complementary to the third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth portion of the target nucleic acid. This portion can be one end of the target nucleic acid, or a non-terminal portion. Binding of the third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth capture probe can occur after, during, or for unextended target nucleic acids. A third, fourth, fifth, sixth, seventh, eighth, ninth, or tenth capture probe may have the above-described couplings.
[0233] A capture probe can isolate the target nucleic acid from a sample. Here, the capture probe is added to a sample containing the target nucleic acid. The capture probe binds to the target nucleic acid via a region of the probe that is complementary to one region of the target nucleic acid. When the target nucleic acid comes into contact with a substrate that contains the portion of the capture probe that binds to the substrate binding portion, the nucleic acid is immobilized on the surface of the substrate.
[0234] Figure 8 illustrates the capture of a target nucleic acid using the two-probe capture system described herein. Genomic DNA is denatured at 95°C and hybridized into a pool of capture reagents. This pool of capture reagents contains oligonucleotides probe A, probe B, and an antisense block probe. Probe A contains a biotin moiety at its 3' end and a sequence complementary to the 5' end of the target nucleic acid. Probe B contains a purified binding sequence that can be bound to its 5' end by a paramagnetic bead and a nucleotide sequence complementary to the 3' end of the target nucleic acid. The antisense block probe contains a nucleotide sequence complementary to the antisense strand of the target nucleic acid to be sequenced. After hybridization with the capture reagents, a sequencing window is created on the target nucleic acid at a position between the hybridized probes A and B. The target nucleic acid is purified using a paramagnetic bead that binds to the 5' sequence of probe B. After washing away any excess capture reagent or complementary antisense DNA strand, the desired target nucleic acid is purified. Next, the purified target nucleic acid is passed through a flow cell containing a surface capable of binding to the biotin portion (such as streptavidin) on the hybridized probe A. This causes one end of the target nucleic acid to bind to the surface of the flow cell. To capture the other end, the target nucleic acid is flow-stretched, and a biotinylated probe complementary to the purified binding sequence of probe B is added. The biotinylated probe, upon hybridization with the purified binding sequence of probe B, can bind to the surface of the flow cell, resulting in a captured target nucleic acid that is elongated and bound to the surface of the flow cell at both ends.
[0235] To ensure that users can reliably "capture" as many target nucleic acid molecules as possible from highly fragmented samples, it is useful to include multiple capture probes, each complementary to a different region of the target nucleic acid. For example, three pools of capture probes are possible, the first pool complementary to the region near the 5' end of the target nucleic acid, the second pool complementary to the central region, and the third pool complementary to the region near the 3' end. This can be generalized to "n target regions" per target nucleic acid. In this example, each pool of fragmented target nucleic acid is bound to a capture probe containing or bound to a biotin tag. 1 / n (where n is the number of different regions in the target nucleic acid) of the input sample is isolated for each pool chamber. The capture probes bind to the target nucleic acid of interest. The target nucleic acid is then immobilized on an avidin molecule attached to a substrate via the biotin in the capture probe. Optionally, the target nucleic acid is extended, for example, by hydrodynamic or electrostatic force. To simultaneously extend and bind all n pools, or to maximize the number of fully extended molecules, pool 1 (which captures the 5' region) can be extended and bound first, followed by pool 2 (which captures the central region of the target), and finally pool 3.
[0236] The target nucleic acid can be captured using the “Bead-Based Two-Step Purification” system of this disclosure. There are four capture probes, namely probe A, probe B, probe C, and probe D. Probe A contains an OA sequence, which is a nucleic acid sequence complementary to the 5' end of the target nucleic acid, and a nucleic acid sequence attached to the biotin moiety. The OA sequence may contain the nucleotide sequence CGAAAGCCATGACCTCCGATCACTC (SEQ ID NO: 109) and can be bound to a lone oligonucleotide. The nucleic acid sequence attached to the biotin moiety is connected to the nucleic acid sequence complementary to the 5' end of the target nucleic acid via a cleavable linker. Probes B and C contain a nucleic acid sequence complementary to the target nucleic acid and a nucleic acid sequence attached to the biotin moiety. The nucleic acid sequence attached to the biotin moiety is connected to the nucleic acid sequence complementary to the target nucleic acid via a cleavable linker. Probe D contains a nucleic acid sequence complementary to the 3' end of the target nucleic acid (a purified binding sequence named the G sequence) and a biotin moiety. The biotin moiety is connected to the G sequence via a cleavable linker. The four capture probes described above first hybridize to the target nucleic acid. All probes hybridize to non-overlapping positions along the target nucleic acid, and probes B and C hybridize between probes A and D. The target nucleic acid is then purified using streptavidin paramagnetic beads bound to the biotin portion on the capture probes. Excess non-target genomic DNA is washed away from the beads. The target nucleic acid-capture probe complex is then released from the streptavidin magnetic beads by cleavage of the cleavable linkers within each capture probe. The target nucleic acid-capture probe complex is further purified using paramagnetic beads bound to the purified G sequence on probe D. Excess capture probe is washed away, and the target nucleic acid-capture probe complex is eluted from the paramagnetic beads.
[0237] The target nucleic acid can be captured using the "one-step purification system based on beads using λ exonuclease" described herein. There are four capture probes, namely probe A, probe B, probe C, and probe D. Probe A contains a sequence complementary to the 5' end of the target nucleic acid. The 5' end of probe A contains two phosphorothioate bonds. Probes B, C, and D contain a nucleic acid sequence attached to the biotin portion at the 3' end of the probe and a nucleic acid sequence complementary to the target nucleic acid at the 5' end of the probe. The 5' ends of probes B, C, and D are phosphorylated. Probes A, B, C, and D hybridize along the target nucleic acid at non-overlapping positions. After the probes hybridize to the target nucleic acid, the target nucleic acid is purified using streptavidin paramagnetic beads. Excess gDNA and capture probes are washed away. The target nucleic acid-capture probe complex is eluted from the beads. Subsequently, probes B, C, and D are digested using λ exonuclease, which selectively degrades double-stranded DNA with phosphorylated 5' ends.
[0238] The target nucleic acid can be captured using the “One-Step Purification System Based on Beads Using FEN1” described herein. There are four capture probes, namely probe A, probe B, probe C, and probe D. Probe A contains a 3' nucleic acid sequence that does not hybridize to the target nucleic acid, a nucleic acid sequence complementary to the 5' end of the nucleic acid sequence, and a 5' nucleic acid sequence containing a biotin moiety that does not hybridize to the target nucleic acid. Probes B and C contain a 3' nucleic acid sequence that does not hybridize to the target nucleic acid, a nucleic acid sequence complementary to the nucleic acid sequence, and a 5' nucleic acid sequence containing a biotin moiety that does not hybridize to the target nucleic acid. Probe D contains a 3' sequence that does not hybridize to the target nucleic acid and a 5' sequence complementary to the target nucleic acid. When probes A, B, C, and D hybridize to the target nucleic acid, probe A is positioned next to probe B, and the 5' nucleic acid sequence on probe A containing a biotin moiety and not hybridizing to the target nucleic acid sequence, along with the 3' nucleic acid sequence on probe B that does not hybridize to the target nucleic acid sequence, form a branched double-stranded DNA substrate with a 5' DNA flap. When probe B is positioned next to probe C, the 5' nucleic acid sequence on probe B containing a biotin moiety and not hybridizing to the target nucleic acid sequence, along with the 3' nucleic acid sequence on probe C that does not hybridize to the target nucleic acid sequence, form a branched double-stranded DNA substrate with a 5' DNA flap. When probe C is positioned next to probe D, the 5' nucleic acid sequence on probe C containing a biotin moiety and not hybridizing to the target nucleic acid sequence, along with the 3' nucleic acid sequence on probe D that does not hybridize to the target nucleic acid sequence, form a branched double-stranded DNA substrate with a 5' DNA flap. After the probes hybridize to the target nucleic acid sequence, the target nucleic acid is purified using streptavidin paramagnetic beads. Excess genomic DNA and probe are washed away from the beads. The target nucleic acid is eluted from the beads by incubation with heat-stable flap endonuclease 1 (FEN1). FEN1 cleaves the 5' DNA flap, thereby separating the biotin-hybridized capture probe and releasing the target nucleic acid-capture probe complex.
[0239] This disclosure enables users to capture and sequence multiple target nucleic acids simultaneously, and to hybridize multiple capture probes into a sample containing a mixture of multiple target nucleic acids. Multiple target nucleic acids may include groups of two or more nucleic acids, each containing the same sequence, or groups of two or more nucleic acids, each not necessarily containing the same sequence. Similarly, multiple capture probes may include groups of two or more capture probes with the same sequence, or groups of two or more capture probes with different sequences. For example, using multiple capture probes, all containing the same sequence, allows users to capture multiple target nucleic acids, all containing the same sequence. Sequencing these multiple target nucleic acids, all containing the same sequence, can achieve a high level of sequencing accuracy due to data redundancy. In another example, a set of capture probes, each containing a complementary capture probe to a target gene, can be used to capture and simultaneously sequence two or more specific genes of interest. This enables users to perform multiplexed sequencing of specific genes. Figure 9 shows the results from experiments using the method of this disclosure to capture and detect a multi-cancer panel consisting of 100 targets using FFPE samples.
[0240] The capture probe may also include a domain that binds to (hybridizes with) the "multiplexing oligo." The multiplexing oligo may include at least three domains. The first domain may include a nucleic acid sequence that hybridizes with the capture probe. The second domain may include a unique nucleic acid sequence that identifies the sample. The third domain may include a substrate binding portion. Multiple multiplexing oligos can be used in combination with the capture probe of this disclosure to simultaneously sequence multiple target nucleic acids from at least two samples. Multiple target nucleic acids can be sequenced from at least three samples, or at least four samples, or at least five samples, or at least six samples, or at least seven samples, or at least eight samples, or at least nine samples, or at least ten samples, or at least 100 samples, or at least 1000 samples using the multiplexing oligo.
[0241] An example of simultaneously sequencing three target nucleic acid molecules from three samples using a multiplexed oligonucleotide is as follows: Target nucleic acids from each of the three samples (Sample 1, Sample 2, Sample 3) are hybridized into two capture probes (Probe A and Probe B). Probe A contains two domains. The first domain contains a substrate-binding region. The second domain contains a sequence complementary to the 5' end of the target nucleic acid. Probe B contains two domains. The first domain contains a sequence complementary to the 3' end of the target nucleic acid. The second domain contains a sequence complementary to the multiplexed oligonucleotide. After hybridizing these two capture probes into the target nucleic acid, the second domain of Probe B is hybridized into the multiplexed oligonucleotide. The multiplexed oligonucleotide contains three domains. The first domain contains a sequence complementary to the second domain of Probe B. The second domain contains a unique nucleic acid sequence that identifies the sample. The third domain contains a substrate-binding region. After hybridizing the multiplexed oligonucleotides, an endonuclease-based cleavage process is performed to remove all overhang DNA from the target nucleic acid, so that probe A hybridizes to the 5' end of the target nucleic acid and probe B hybridizes to the 3' end of the target nucleic acid. After endonuclease treatment, the multiplexed oligonucleotides are ligated to the 3' end of the target nucleic acid, and then probe B is removed. The target nucleic acid-probe A complex is further purified and then sequenced. Since each target nucleic acid from each sample is ligated to a multiplexed oligonucleotide, the sample from which the target nucleic acid originated can be identified by sequencing of the multiplexed oligonucleotides.
[0242] When sequencing the entire range is desirable, the number of independent capture probes required is inversely related to the size of the target nucleic acid fragment. In other words, highly fragmented target nucleic acids will require more capture probes. For types of samples containing highly fragmented and degraded target nucleic acids (e.g., formalin-fixed paraffin-embedded tissue), including a large pool of capture probes may be useful. Conversely, for samples with long target nucleic acid fragments (e.g., isolated nucleic acids obtained in vitro), one capture probe at the 5' end may suffice.
[0243] In this specification, the region of the target nucleic acid between two capture probes, or the region after one capture probe and before the end of the target nucleic acid, is referred to as the "sequencing window." Figure 8 shows the sequencing window that occurs when one target nucleic acid is captured using two capture probes, with the window labeled. The sequencing window is the portion of the target nucleic acid that can be used for binding to the sequencing probe. The minimum sequencing window is the length of the target binding domain (e.g., 4 to 10 nucleotides), and the maximum sequencing window is most of an entire chromosome.
[0244] When sequencing large target nucleic acid molecules using the method of this disclosure, the size of the sequencing window can be controlled by hybridizing one or more “blocker oligos” along the length of the target nucleic acid. The blocker oligos hybridize to the target nucleic acid at specific locations, thereby preventing sequencing probes from binding to these locations and creating a smaller sequencing window of interest. By creating a smaller sequencing window, the sequencing reaction is limited to specific regions of interest on the target DNA molecule, thereby increasing the speed and accuracy of sequencing. The use of blocker oligos is particularly useful when sequencing specific mutations at known locations within the target nucleic acid, because it is not necessary to sequence the entire target nucleic acid. In a non-limiting example, the method of this disclosure can be used to identify two different haplotypes by sequencing two heterozygous sites.
[0245] The capture probe may include a nucleic acid molecule complex. The nucleic acid molecule complex may include a partially double-stranded nucleic acid molecule. In some embodiments, the partially double-stranded nucleic acid molecule may include a target-specific domain, a double-stranded domain, a single-stranded purified sequence, a cleavable moiety, a single-stranded overhang domain, a sample-specific domain, a substrate-specific domain, or any combination thereof.
[0246] In some embodiments, any one strand of a partially double-stranded nucleic acid molecule may contain approximately 40–150 nucleotides, or approximately 60–135 nucleotides, or approximately 10–90 nucleotides, or approximately 25–75 nucleotides, or approximately 60 nucleotides, or approximately 50–100 nucleotides.
[0247] In some embodiments, any one strand of a partially double-stranded nucleic acid molecule may contain at least one affinity moiety, or at least two affinity moieties, or at least three affinity moieties, or at least four affinity moieties, or at least five affinity moieties, or at least six affinity moieties, or at least seven affinity moieties, or at least eight affinity moieties, or at least nine affinity moieties, or at least ten affinity moieties.
[0248] In some embodiments, any one strand of a partially double-stranded nucleic acid molecule may contain at least one crosslinking site. The crosslinking site can be a chemically crosslinked site or a photoactivated crosslinked site.
[0249] The capture probe may include a single-stranded nucleic acid molecule. In some embodiments, the single-stranded nucleic acid molecule may include a target-specific domain, a double-stranded domain, a single-stranded purified sequence, a cleavable moiety, a single-stranded overhang domain, a sample-specific domain, a substrate-specific domain, or any combination thereof.
[0250] The target-specific domain, double-stranded domain, single-stranded purified sequence, cleavable moiety, single-stranded overhang domain, sample-specific domain, and substrate-specific domain may contain or not contain at least one native base. In some embodiments, the target-specific domain, double-stranded domain, single-stranded purified sequence, cleavable moiety, single-stranded overhang domain, sample-specific domain, and substrate-specific domain may contain or not contain at least one modified nucleotide or nucleic acid analog.
[0251] Target-specific domains, double-stranded domains, single-stranded purified sequences, cleavable regions, single-stranded overhang domains, sample-specific domains, and substrate-specific domains may contain any combination of native bases (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more) and modified nucleotides or nucleic acid analogs (e.g., 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, or more). When present in combination, the native bases and modified nucleotides or nucleic acid analogs may be arranged in any order.
[0252] The target-specific domain may include a nucleic acid sequence that is complementary to and hybridizes with a portion of the target nucleic acid molecule. In some embodiments, the target-specific domain may include about 10 to about 150 nucleotides, or about 25 to about 100 nucleotides, or about 35 to about 100 nucleotides, or about 25 to about 125 nucleotides, or about 15 to about 100 nucleotides.
[0253] In some embodiments, the target-specific domain can hybridize to at least about 100 base pairs from the 3' end of the target nucleic acid molecule. In some embodiments, the target-specific domain can hybridize to at least about 100 base pairs from the 5' end of the target nucleic acid molecule.
[0254] The double-stranded domain may contain nucleic acid sequences that can anneal to another nucleic acid strand to form a partially or completely double-stranded nucleic acid molecule. In some embodiments, the double-stranded domain may contain about 14 to about 45 nucleotides, or about 25 to about 35 nucleotides, or about 30 nucleotides, or about 10 to about 60 nucleotides, or about 30 to about 50 nucleotides.
[0255] A single-stranded purified sequence may contain nucleic acid sequences suitable for purification. A single-stranded purified sequence may contain an F-tag. A single-stranded purified sequence may contain an F-like tag. A single-stranded purified sequence may contain the nucleotide sequence AACATCACACAGACC (SEQ ID NO: 112). A single-stranded purified sequence may contain the nucleotide sequence GCTATCATCACAGC (SEQ ID NO: 113).
[0256] A single-stranded purified sequence may contain at least one affinity moiety, or at least two affinity moieties, or at least three affinity moieties, or at least four affinity moieties, or at least five affinity moieties, or at least six affinity moieties, or at least seven affinity moieties, or at least eight affinity moieties, or at least nine affinity moieties, or at least ten affinity moieties. Biotin can be an affinity moiety. Thus, in some embodiments, a single-stranded purified sequence may contain at least one biotin moiety, or at least two biotin moieties, or at least three biotin moieties, or at least four biotin moieties, or at least five biotin moieties, or at least six biotin moieties, or at least seven biotin moieties, or at least eight biotin moieties, or at least nine biotin moieties, or at least ten biotin moieties.
[0257] A single-stranded purified sequence can contain at least 50 nucleotides, or approximately 15 to 50 nucleotides.
[0258] The cleavable portion may include a portion that can be cleaved by an enzyme. The enzymatically cleavable portion may include a USER sequence for cleavage by the USER enzyme. Alternatively, the cleavable portion may include a portion that can be cleaved by light.
[0259] The single-stranded overhang domain can contain a single-stranded nucleic acid sequence that can combine with a target nucleic acid molecule to form a 5' overhang flap structure.
[0260] The sample-specific domain may include a nucleic acid sequence that identifies the biological sample from which the target nucleic acid molecule was supplied. The sample-specific domain may include L-DNA. The sample-specific domain may include D-DNA. The sample-specific domain may include a combination of L-DNA and D-DNA. The sample-specific domain can be hybridized to any probe of this disclosure. The sample-specific domain may contain approximately 28 nucleotides.
[0261] In some embodiments, the sample-specific domain may include at least one attachment site or at least two attachment sites. In embodiments in which the sample-specific domain includes at least one attachment site or at least two attachment sites, the attachment site may include about 14, about 10, or about 8 nucleotides.
[0262] The substrate-specific domain can include nucleic acid sequences that hybridize to complementary nucleic acid molecules attached to the substrate. An array can be used as the substrate. The substrate-specific domain can also include nucleic acid sequences that hybridize to lone oligonucleotides.
[0263] The substrate-specific domain can contain a poly-A sequence. The substrate-specific domain can contain a poly-T sequence. The substrate-specific domain can contain an L-poly-A sequence, and the nucleotides of the poly-A sequence are L-DNA. The substrate-specific domain can contain an L-poly-T sequence, and the nucleotides of the poly-T sequence are L-DNA. The substrate-specific domain can contain L-DNA. The substrate-specific domain can contain approximately 30 nucleotides.
[0264] Figure 34 shows a schematic diagram of a typical capture probe, which we have named the "c5 probe complex," containing a nucleic acid molecule complex that binds to a target nucleic acid. This c5 probe complex contains a partially double-stranded nucleic acid molecule. One strand of this partially double-stranded nucleic acid molecule contains a target-specific domain that hybridizes to the target nucleic acid, a double-stranded domain that anneals to the other strand of this partially double-stranded nucleic acid molecule, a first single-stranded purified sequence, and a cleavable region located between the target-specific domain and the double-stranded domain. In this non-limiting example, the single-stranded purified sequence contains an F-like tag, and the cleavable region contains an enzymatically cleavable USER sequence. The other strand of the partially double-stranded nucleic acid molecule contains a double-stranded domain that anneals to this other strand of the partially double-stranded nucleic acid molecule, and a single-stranded overhang domain. In this non-limiting example, the single-stranded overhang domain and the target nucleic acid molecule form a 5' overhang flap structure.
[0265] Figure 34 also shows a schematic diagram of a typical capture probe, which we have named a "c3 probe complex," containing a nucleic acid molecule complex that binds to a target nucleic acid. This c3 probe complex contains a partially double-stranded nucleic acid molecule. One strand of this partially double-stranded nucleic acid molecule contains a target-specific domain that hybridizes to the target nucleic acid, a double-stranded domain that anneals to the other strand of this partially double-stranded nucleic acid molecule, and a cleavable region located between the target-specific domain and the double-stranded domain. In this non-limiting example, the cleavable region contains an enzymatically cleavable USER sequence. The other strand of the partially double-stranded nucleic acid molecule contains a double-stranded domain that anneals to this other strand of the partially double-stranded nucleic acid molecule, a sample-specific domain, a substrate-specific domain, a single-stranded purified sequence, and a cleavable region located between the single-stranded purified sequence and the substrate-specific domain. In this non-limiting example, the sample-specific domain contains L-DNA, the substrate-specific domain contains L-DNA, the single-stranded purified sequence contains an F-tag, and the cleavable region contains a photocleavable region.
[0266] Figure 41 shows a schematic diagram of a typical capture probe, which we have named the "c3.2 probe complex," containing a nucleic acid molecule complex that binds to a target nucleic acid. This c3.2 probe complex contains a partially double-stranded nucleic acid molecule. One strand of this partially double-stranded nucleic acid molecule contains a target-specific domain that hybridizes to the target nucleic acid and a double-stranded domain that anneals to the other strand of this partially double-stranded nucleic acid molecule. In some embodiments, this strand optionally contains at least one first affinity moiety. In some embodiments, this strand optionally contains a cleavable moiety located between the target-specific domain and the double-stranded domain. The other strand of the partially double-stranded nucleic acid molecule contains a double-stranded domain that anneals to this other strand of the partially double-stranded nucleic acid molecule and a substrate-specific domain. In some embodiments, this strand optionally contains at least one, at least two, or at least three second affinity moieties.
[0267] Figure 41 also shows a schematic diagram of a typical capture probe, which we have named the "c5.2 probe complex," containing a nucleic acid molecule complex that binds to a target nucleic acid. This c5.2 probe complex contains a partially double-stranded nucleic acid molecule. One strand of this partially double-stranded nucleic acid molecule contains a target-specific domain that hybridizes to the target nucleic acid and a double-stranded domain that anneals to the other strand of this partially double-stranded nucleic acid molecule. In some embodiments, this strand optionally contains a cleavable region located between the target-specific domain and the double-stranded domain. The other strand of the partially double-stranded nucleic acid molecule contains a double-stranded domain that anneals to this other strand of the partially double-stranded nucleic acid molecule, a sample-specific domain, a first single-stranded purified sequence, a first cleavable region located between the double-stranded domain and the sample-specific domain, and a second cleavable region located between the sample-specific domain and the first single-stranded purified sequence. In some embodiments, the first single-stranded purified sequence may contain at least one affinity moiety, for example, at least one biotin moiety. In some embodiments, the first single-stranded purified sequence can be replaced with at least one biotin moiety. Thus, the other strand of a partially double-stranded nucleic acid molecule includes a double-stranded domain that anneals to this other strand of the partially double-stranded nucleic acid molecule, a sample-specific domain, at least one biotin moiety, a first cleavable moiety located between the double-stranded domain and the sample-specific domain, and a second cleavable moiety located between the sample-specific domain and at least one biotin moiety.
[0268] Sample preparation method of this disclosure
[0269] This disclosure provides a sample preparation method that includes immobilizing a target nucleic acid molecule on a substrate.
[0270] The sample preparation method of the present invention may include a CRISPR-based fragmentation step (see, for example, Baker and Mueller, "Isolation of specific megabase pairs of genomic DNA via CRISPR," Nucleic Acids Research 2017, vol. 45(19), p. e165; Tsai et al., "Enrichment and SMRT sequencing of repeat extension disease-causing genomic regions by CRISPR-Cas9 without amplification," bioRxiv 203919; doi: https: / / doi.org / 10.1101 / 203919; Nachmanson et al., "Targeted genomic fragmentation using CRISPR / Cas9 improves hybridization capture of small targets, reduces PCR bias, and enables efficient high-precision sequencing," bioRxiv 207027; doi: https: / / doi.org / 10.1101 / 207027). CRISPR fragmentation allows for in vitro fragmentation of genomic DNA (gDNA) obtained from biological samples by cleaving it proximal to the protospacer adjacent motif (PAM) site located within the gDNA. The PAM site can contain the nucleotide sequence NGG (where N is any nuclear base) or NGA (where N is any nuclear base). Fragments generated by CRISPR-based fragmentation can be purified using a biotinylated CRISPR complex or an anti-CAS9 antibody.
[0271] The method for capturing the target nucleic acid involves (1) fragmenting the gDNA using a CRISPR-based fragmentation process; (2) contacting the fragmented gDNA with at least two capture probes, at least one of which is the c5 probe complex described above, and at least one of which is the c3 probe complex described above, and the c3 probe complex and the c5 probe complex hybridize with the target nucleic acid to form the complex shown in Figure 34; (3) removing the 5' overhang flap structure by contacting this complex with FEN1; and (4) capturing the target nucleic acid. This may include: (5) ligating the end to the 5' end of the chain of the c3 probe complex containing the substrate-specific domain; (6) ligating the single-strand purified sequence of the c5 probe complex to a first substrate; (7) cleaving the cleavable region located between the double-strand domain and the target-specific domain of the c3 probe complex and the c5 probe complex, respectively; (8) ligating the single-strand purified sequence of the c3 probe complex to a second substrate; (9) cleaving the cleavable region located between the ligated single-strand purified sequence of the c3 probe complex and the substrate-specific domain; and (10) hybridizing the substrate-specific domain to a complementary nucleic acid molecule attached to a third substrate.
[0272] In some aspects of the above method, step (9) can be performed before step (8).
[0273] In some embodiments of the above method, steps (3) and (4) can be carried out together.
[0274] In some embodiments, the above method may optionally include one additional step in steps (6) and (7) in which target nucleic acid-capture probe complexes derived from different biological samples are pooled together. In this embodiment, the target nucleic acid-capture probe complexes derived from different biological samples will each contain a c3 probe complex containing its own sample-specific domain, so that the target-specific domain identifies the biological sample from which each acquired target nucleic acid originated.
[0275] An example of the sample preparation method of this disclosure is shown in Figures 34 to 40. In this non-limiting example, gDNA obtained from a biological sample is first fragmented using CRISPR-based fragmentation. After fragmentation, the target nucleic acid hybridizes to two capture probes, as shown in Figure 34. In this non-limiting example, the two capture probes are the c3 probe complex and the c5 probe complex described above. The c3 probe complex and the c5 probe complex hybridize to the target nucleic acid at non-overlapping positions along the target nucleic acid. The c3 probe complex hybridizes to the target nucleic acid via a target-specific domain at a position within eight nucleotides from the 3' end of the target nucleic acid, and the c5 probe complex hybridizes to the target nucleic acid via a target-specific domain, resulting in the c5 probe complex hybridizing to the 5' side of the c3 probe complex. The single-stranded overhang domain of the c5 probe complex and the target nucleic acid molecule form a 5' overhang flap structure. After these two capture probes hybridize, the target nucleic acid-capture probe complex is incubated with FEN1 and ligase. As shown in Figure 35, FEN1 removes the 5' overhang flap structure, and the ligase ligates the 3' end of the target nucleic acid to the strand of the c3 probe complex containing the substrate-specific domain. The resulting complex, shown in Figure 36, binds to an F-like bead that hybridizes to an F-like tag present in the c5 probe complex. The beads are washed, and USER enzyme is added. As shown in Figure 37, the USER enzyme releases the target nucleic acid from the F-like bead by cleaving a cleavable region located between the target-specific domain and the double-stranded domain in both the c3 and c5 probe complexes. As shown in Figure 38, the eluted complex is further purified using SPRI beads. The purified complex then binds to an F-bead that hybridizes to an F-tag present in the c3 probe complex. After washing, as shown in Figure 39, the target nucleic acid is eluted from the F-beads by exposing the F-beads to UV light to cleave the photocatalytically cleavable portion located between the substrate-specific domain and the F-tag within the c3 probe complex.Subsequently, as shown in Figure 40, the substrate-specific domain of the linked c3 probe complex is hybridized with complementary nucleic acids attached to the substrate, thereby binding the resulting complex to the substrate.
[0276] An example of another sample preparation method of this disclosure is shown in Figures 41–46. In this non-limiting example, gDNA obtained from a biological sample is first fragmented using CRISPR-based fragmentation. Following fragmentation, the target nucleic acid is hybridized into two capture probes, as shown in Figure 41. In this non-limiting example, the two capture probes are the c3.2 probe complex and the c5.2 probe complex described above. The c3.2 and c5.2 probe complexes hybridize to the target nucleic acid at non-overlapping positions along the target nucleic acid. The c5.2 probe complex hybridizes to the target nucleic acid via a target-specific domain, resulting in the c5.2 probe complex hybridizing to the 5' side of the c3.2 probe complex. After these two capture probes hybridize, the target nucleic acid is ligated to one strand of the c3.2 probe complex and one strand of the c5.2 probe complex, as shown in Figure 42. The linking can be enzymatic, self-linking, chemical linking, or any combination thereof. In embodiments including enzymatic linking, the enzymatic linking can be performed using a high-fidelity template-dependent nick ligase. The resulting complex, shown in Figure 42, can then be bound to a bead containing at least one oligonucleotide that hybridizes to the single-stranded purified sequence. As shown in Figure 43, the target nucleic acid can be released from the bead by washing and cleaving a cleavable region located between the sample-specific domain and the single-stranded purified sequence. The resulting complex can then be immobilized on the surface of a substrate by hybridizing the substrate-specific domain to an oligonucleotide attached to the substrate, as shown in Figure 44. Any array of the present disclosure is possible as the substrate / oligonucleotide complex.
[0277] The above method may further include hybridizing at least one reporter probe containing a first detectable label and a second detectable label into a sample-specific domain. Subsequently, by identifying the first and second detectable labels, the sample from which the target nucleic acid originates can be identified based on the attributes of the first and second detectable labels.
[0278] Alternatively, the above method may further include hybridizing a first reporter probe containing a first detectable label and a second detectable label into a sample-specific domain. The first and second detectable labels can then be identified. Subsequently, the first and second detectable labels can be removed, and a second reporter probe containing a third and fourth detectable label can be hybridized into the sample-specific domain. By subsequently identifying the third and fourth detectable labels, the sample from which the target nucleic acid originates can be identified based on the attributes of the first, second, third, and fourth detectable labels.
[0279] After identifying the sample from which the target nucleic acid originates, the sample-specific domain can be released by cleaving the cleavable region located between the sample-specific domain and the double-stranded domain, as shown in Figure 45.
[0280] Method of Disclosure
[0281] The sequencing method of this disclosure includes reversibly hybridizing at least one sequencing probe disclosed herein to a target nucleic acid.
[0282] Methods for sequencing nucleic acids include (1) hybridizing a sequencing probe described herein to a target nucleic acid. The target nucleic acid may optionally be immobilized on a substrate at one or more positions. Typical sequencing probes may include a target-binding domain and a barcode domain; the target-binding domain may include any construct listed in Table 1. A typical target-binding domain contains at least eight nucleotides that hybridize to a target nucleic acid, wherein at least six nucleotides within the target-binding domain can identify corresponding nucleotides in the target nucleic acid molecule (for example, these six nucleotides identify six nucleotides complementary to the target molecule into which the target-binding domain hybridizes), and at least two nucleotides within the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule (for example, these at least two nucleotides do not identify two nucleotides complementary to the target molecule into which the target-binding domain hybridizes); any of the above at least six nucleotides within the target-binding domain can be modified nucleotides or nucleotide analogs; the above at least two nucleotides within the target-binding domain that do not identify corresponding nucleotides in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain. A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which a complementary nucleic acid molecule can bind, each of the at least three attachment sites corresponds to two of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the nucleic acid sequence at each of the at least three attachment sites determines the position and attributes of the two corresponding nucleotides among the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0283] In another embodiment, a representative target-binding domain may contain at least six nucleotides that hybridize to a target nucleic acid, and these at least six nucleotides within the target-binding domain can identify the corresponding nucleotides within the target nucleic acid molecule (for example, when the target-binding domain sequence is exactly six nucleotides, these six nucleotides identify six nucleotides complementary to the target molecule they hybridize to); any of these at least six nucleotides within the target-binding domain may not be modified nucleotides or nucleic acid analogs, or any of the at least six nucleotides within the target-binding domain may be modified nucleotides or nucleic acid analogs.
[0284] This method involves hybridizing a sequencing probe to a target nucleic acid, (2) binding a first complementary nucleic acid molecule containing a first detectable label and at least a second detectable label to the first of the at least three attachment sites on the barcode domain, (3) detecting the first detectable label and at least a second detectable label on the bound first complementary nucleic acid molecule, and (4) identifying the location and attributes of at least two nucleotides in the immobilized target nucleic acid. For example, when the first complementary nucleic acid molecule contains two detectable labels, these two detectable labels identify the at least two nucleotides in the immobilized nucleic acid molecule.
[0285] After detecting at least two of the above-mentioned detectable labels, these at least two detectable labels are removed from the first complementary nucleic acid molecule. Thus, this method further includes (5) freeing the first complementary nucleic acid molecule containing the detectable labels by binding the first hybridizing nucleic acid molecule lacking the detectable labels to the first attachment site, or bringing the first complementary nucleic acid molecule containing the detectable labels into contact with a force sufficient to release the first detectable label and at least a second detectable label. Thus, after step (5), no detectable labels are bound to the first attachment site. The method further comprises (6) binding a second complementary nucleic acid molecule containing a third detectable label and at least a fourth detectable label to the second of the at least three attachment sites on the barcode domain; (7) detecting the third detectable label and at least a fourth detectable label on the bound second complementary nucleic acid molecule; (8) identifying the locations and attributes of at least two nucleotides in the optionally immobilized target nucleic acid; (9) identifying a linear sequence of at least six nucleotides for at least a first region of the immobilized target nucleic acid hybridized to the target binding domain of the sequencing probe by repeating steps (5) to (8) until a complementary nucleic acid molecule containing two detectable labels is bound to each of the at least three attachment sites on the barcode domain and these two detectable labels on the bound complementary nucleic acid molecule are detected; and (10) optionally removing the sequencing probe from the immobilized target nucleic acid.
[0286] This method further involves (11) hybridizing a second sequencing probe to a target nucleic acid immobilized on a substrate at one or more locations (the target binding domains of the first sequencing probe and the second sequencing probe are different); (12) binding a first complementary nucleic acid molecule containing a first detectable label and at least a second detectable label to a first attachment site among the at least three attachment sites of the barcode domain; (13) detecting the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule; (14) optionally identifying the locations and attributes of at least two nucleotides in the immobilized target nucleic acid; and (15) freeing the first complementary nucleic acid molecule or complex containing the detectable label by binding a first hybridizing nucleic acid molecule lacking the detectable label to the first attachment site, or releasing the first detectable label and at least a second detectable label from the first complementary nucleic acid molecule or complex containing the detectable label. (16) Apply sufficient force to make contact with the barcode domain; (17) Bind a second complementary nucleic acid molecule containing a third detectable label and at least a fourth detectable label to the second of the at least three attachment sites on the barcode domain; (18) Detect the third detectable label and at least a fourth detectable label on the bound second complementary nucleic acid molecule; (19) Identify the location and attribute of at least two nucleotides in the immobilized target nucleic acid; (10) Identify a linear sequence of at least six nucleotides for at least a second region of the immobilized target nucleic acid hybridized to the target binding domain of the second sequencing probe by repeating steps (15) to (18) until a complementary nucleic acid molecule containing two detectable labels is bound to each of the at least three attachment sites on the barcode domain and these two detectable labels on the bound complementary nucleic acid molecule are detected; (20) Optionally remove the second sequencing probe from the immobilized target nucleic acid.
[0287] This method may further include identifying the sequence of the immobilized target nucleic acid by assembling the identified linear sequences of nucleotides in at least a first region and at least a second region of the immobilized target nucleic acid.
[0288] Steps (5) and (6) may occur sequentially or simultaneously. The first detectable label and at least the second detectable label may have the same emission spectrum or different emission spectra. The third detectable label and at least the fourth detectable label may have the same emission spectrum or different emission spectra.
[0289] The first complementary nucleic acid molecule may contain a cleavable linker. The second complementary nucleic acid molecule may also contain a cleavable linker. The first and second complementary nucleic acid molecules may each contain a cleavable linker. The cleavable linker is preferably cleavable by light. Light can be used as the emission force. UV light is preferred. The light can be provided by a light source selected from the group consisting of arc lamps, lasers, focused UV light sources, and light-emitting diodes.
[0290] The first complementary nucleic acid molecule and the first hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. For example, the first hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence as the portion of the first complementary nucleic acid molecule that binds to the first attachment site among the at least three attachment sites of the barcode domain. The first hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the first attachment site in the barcode domain.
[0291] The second complementary nucleic acid molecule and the second hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. The second hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the second attachment site within the barcode domain.
[0292] This disclosure also provides a method for sequencing nucleic acids, comprising (1) hybridizing a sequencing probe described herein to a target nucleic acid. The target nucleic acid can optionally be immobilized on a substrate at one or more positions. Typical sequencing probes may include a target-binding domain and a barcode domain; the target-binding domain may include any construct listed in Table 1. A typical target-binding domain contains at least eight nucleotides that hybridize to a target nucleic acid, wherein at least six nucleotides within the target-binding domain can identify corresponding nucleotides in the target nucleic acid molecule (for example, these six nucleotides identify six nucleotides complementary to the target molecule they hybridize to), and at least two nucleotides within the target-binding domain do not identify corresponding nucleotides in the target nucleic acid molecule (for example, these at least two nucleotides do not identify two nucleotides complementary to the target molecule they hybridize to); any of the above at least six nucleotides within the target-binding domain can be modified nucleotides or nucleotide analogs, and the above at least two nucleotides within the target-binding domain that do not identify corresponding nucleotides in the target nucleic acid molecule can be any of four non-target-specific canonical bases, universal bases, or degenerate bases specified by the above at least six nucleotides within the target-binding domain. A typical barcode domain includes a synthetic skeleton and at least three attachment sites, each attachment site including at least one attachment region containing at least one nucleic acid sequence to which a complementary nucleic acid molecule can bind, each of the at least three attachment sites corresponds to two of the at least six nucleotides in the target binding domain, each of the at least three attachment sites has a different nucleic acid sequence, and the nucleic acid sequence at each of the at least three attachment sites determines the position and attributes of the two corresponding nucleotides among the at least six nucleotides in the target nucleic acid to which the target binding domain binds.
[0293] In another embodiment, a representative target-binding domain may contain at least six nucleotides that hybridize to a target nucleic acid, and these at least six nucleotides within the target-binding domain can identify the corresponding nucleotides within the target nucleic acid molecule (for example, when the target-binding domain sequence is exactly six nucleotides, these six nucleotides identify six nucleotides complementary to the target molecule they hybridize to); any of the at least six nucleotides within the target-binding domain may not be modified nucleotides or nucleotide analogs, or any of the at least six nucleotides may be modified nucleotides or nucleotide analogs.
[0294] The method includes (2) hybridizing a sequencing probe to a target nucleic acid, then binding a first complementary nucleic acid molecule containing a first detectable label and at least a second detectable label to a first attachment site among the at least three attachment sites of the barcode domain; and (3) detecting and recording the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule.
[0295] After detecting and recording the at least two detectable labels described above, these at least two detectable labels are removed from the first complementary nucleic acid molecule. Thus, this method further includes (4) freeing the first complementary nucleic acid molecule containing the detectable labels by binding the first hybridizing nucleic acid molecule lacking the detectable labels to the first attachment site, or bringing the first complementary nucleic acid molecule containing the detectable labels into contact with a force sufficient to release the first detectable label and at least the second detectable label. Thus, after step (4), no detectable labels are bound to the first attachment site. This method further includes (5) binding a second complementary nucleic acid molecule containing a third detectable label and at least a fourth detectable label to the second of the at least three attachment sites of the barcode domain; (6) detecting and recording the third detectable label and at least a fourth detectable label of the bound second complementary nucleic acid molecule; (7) repeating steps (4) to (6) until a complementary nucleic acid containing two detectable labels is bound to each of the at least three attachment sites in the barcode domain and these two detectable labels of the bound complementary nucleic acid molecule are detected and recorded; (8) using the detectable labels recorded in steps (3), (6), and (7) to identify the positions and attributes of at least six nucleotides in at least a first region of the immobilized target nucleic acid hybridized to the target binding domain of the sequencing probe; and (9) optionally removing the sequencing probe from the immobilized target nucleic acid.
[0296] This method further involves (10) hybridizing a second sequencing probe to a target nucleic acid immobilized on a substrate at one or more arbitrary positions (provided that the target binding domains of the first sequencing probe and the second sequencing probe are different); (11) binding a first complementary nucleic acid molecule containing a first detectable label and at least a second detectable label to a first attachment site among the at least three attachment sites of the barcode domain; (12) detecting and recording the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule; and (13) binding a first hybridizing nucleic acid molecule lacking a detectable label to the first attachment site, thereby freeing the first complementary nucleic acid molecule or complex containing a detectable label, or freeing the first complementary nucleic acid molecule or complex containing a detectable label from the first detectable label. (14) Apply sufficient force to release at least a second detectable label; (15) Bind a second complementary nucleic acid molecule containing a third detectable label and at least a fourth detectable label to the second of the at least three attachment sites of the barcode domain; (16) Detect and record the third detectable label and at least a fourth detectable label of the bound second complementary nucleic acid molecule; (17) Repeat steps (13) to (15) until a complementary nucleic acid molecule containing two detectable labels is bound to each of the at least three attachment sites in the barcode domain and these two detectable labels of the bound complementary nucleic acid molecule are detected and recorded; (18) Use the detectable labels recorded in steps (12), (15), and (16) to determine at least the second of the immobilized target nucleic acid hybridized by the target binding domain of the second sequencing probe. (18) optionally include identifying the positions and attributes of at least six nucleotides in the region of 2, and removing a second sequencing probe from the immobilized target nucleic acid.
[0297] This method may further include identifying the sequence of the immobilized target nucleic acid by assembling the identified linear sequences of nucleotides in at least a first region and at least a second region of the immobilized target nucleic acid.
[0298] Steps (4) and (5) may occur sequentially or simultaneously. The first detectable label and at least the second detectable label may have the same emission spectrum or different emission spectra. The third detectable label and at least the fourth detectable label may have the same emission spectrum or different emission spectra.
[0299] The first complementary nucleic acid molecule may contain a cleavable linker. The second complementary nucleic acid molecule may also contain a cleavable linker. The first and second complementary nucleic acid molecules may each contain a cleavable linker. The cleavable linker is preferably cleavable by light. Light can be used as the emission force. UV light is preferred. The light can be provided by a light source selected from the group consisting of arc lamps, lasers, focused UV light sources, and light-emitting diodes.
[0300] The first complementary nucleic acid molecule and the first hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. For example, the first hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence as the portion of the first complementary nucleic acid molecule that binds to the first attachment site among the at least three attachment sites of the barcode domain. The first hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the first attachment site in the barcode domain.
[0301] The second complementary nucleic acid molecule and the second hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. The second hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the second attachment site within the barcode domain.
[0302] The above method may further include a medium suitable for recording detectable markers. A suitable computer-readable medium could serve as this medium.
[0303] This disclosure further provides a method for sequencing nucleic acids using multiple sequencing probes disclosed herein. For example, by hybridizing a target nucleic acid with two or more sequencing probes, each probe can sequence the portion of the nucleic acid that it has hybridized with.
[0304] This disclosure provides a method for sequencing nucleic acids, comprising: (1) hybridizing a first group of at least one sequencing probe, comprising a plurality of sequencing probes described herein, to a target nucleic acid optionally immobilized at one or more locations on a substrate; (2) binding a first complementary nucleic acid molecule, comprising a first detectable label and at least a second detectable label, to a first attachment site among the at least three attachment sites of a barcode domain; (3) detecting the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule; (4) identifying the locations and attributes of at least two nucleotides in the immobilized target nucleic acid; and (5) freeing the first complementary nucleic acid molecule containing the detectable label, or causing the first complementary nucleic acid molecule containing the detectable label to release the first detectable label and at least a second detectable label, by binding a first hybridizing nucleic acid molecule lacking the detectable label to the first attachment site. A method is also provided which includes: (6) making contact with sufficient force; (7) binding a second complementary nucleic acid molecule containing a third detectable label and at least a fourth detectable label to a second attachment site among the at least three attachment sites of the barcode domain; (8) detecting the third detectable label and at least a fourth detectable label of the bound second complementary nucleic acid molecule; (9) identifying the location and attribute of at least two nucleotides in the optionally immobilized target nucleic acid; (10) identifying the linear sequence of at least one nucleotide of the immobilized target nucleic acid hybridized by the target-binding domain of the sequencing probe by repeating steps (5) to (8) until a complementary nucleic acid containing two detectable labels is bound to each of the at least three attachment sites in the barcode domain and these two detectable labels of the bound complementary nucleic acid molecule are detected; and (11) optionally removing at least one first population of the first sequencing probe from the immobilized target nucleic acid.
[0305] This method further comprises (11) hybridizing a second group of at least one second sequencing probe, comprising a plurality of sequencing probes disclosed herein, to a target nucleic acid optionally immobilized at one or more locations on a substrate (provided that the target binding domains of the first sequencing probe and the second sequencing probe are different); (12) binding a first complementary nucleic acid molecule, comprising a first detectable label and at least a second detectable label, to a first attachment site among the at least three attachment sites of the barcode domain; and (13) binding the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule (14) detect the label; (2) identify the location and attributes of at least two nucleotides in the optionally immobilized target nucleic acid; (3) bind the first hybridizing nucleic acid molecule lacking the detectable label to the first attachment site, thereby freeing the first complementary nucleic acid molecule or complex containing the detectable label, or bringing the first complementary nucleic acid molecule or complex containing the detectable label into contact with sufficient force to release the first detectable label and at least the second detectable label; (4) attach the second complementary nucleic acid molecule containing the third detectable label and at least the fourth detectable label to the second attachment site of the at least three attachment sites in the barcode domain. (17) to bind to; (18) to detect a third detectable label and at least a fourth detectable label on the bound second complementary nucleic acid molecule; (19) to identify the location and attribute of at least two nucleotides in the immobilized target nucleic acid; (11) to identify a linear sequence of at least six nucleotides for at least one second region of the immobilized target nucleic acid hybridized by the target binding domain of the sequencing probe by repeating steps (15) to (18) until the complementary nucleic acid containing two detectable labels binds to each of the attachment sites of the at least three attachment sites in the barcode domain and these two detectable labels on the bound complementary nucleic acid molecule are detected; and (20) optionally to remove at least one second population of sequencing probes from the immobilized target nucleic acid.
[0306] This method may further include identifying the sequence of the immobilized target nucleic acid by assembling the identified linear sequences of nucleotides in at least a first region and at least a second region of the immobilized target nucleic acid.
[0307] Steps (5) and (6) may occur sequentially or simultaneously. The first detectable and at least the second detectable label may have the same emission spectrum or different emission spectra. The third detectable and at least the fourth detectable label may have the same emission spectrum or different emission spectra.
[0308] The first complementary nucleic acid molecule may contain a cleavable linker. The second complementary nucleic acid molecule may also contain a cleavable linker. The first and second complementary nucleic acid molecules may each contain a cleavable linker. The cleavable linker is preferably cleavable by light. Light can be used as the emission force. UV light is preferred. The light can be provided by a light source selected from the group consisting of arc lamps, lasers, focused UV light sources, and light-emitting diodes.
[0309] The first complementary nucleic acid molecule and the first hybridizing nucleic acid molecule lacking a detectable label may contain the same nucleic acid sequence. The first hybridizing nucleic acid molecule lacking a detectable label may contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the first attachment site within the barcode domain.
[0310] The second complementary nucleic acid molecule and the second hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. The second hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the second attachment site within the barcode domain.
[0311] This disclosure provides a method for sequencing nucleic acids, comprising: (1) hybridizing a first group of at least one sequencing probe, comprising a plurality of sequencing probes described herein, to a target nucleic acid optionally immobilized at one or more locations on a substrate; (2) binding a first complementary nucleic acid molecule, comprising a first detectable label and at least a second detectable label, to a first attachment site among the at least three attachment sites in a barcode domain; (3) detecting and recording the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule; (4) freeing the first complementary nucleic acid molecule containing the detectable label by binding a first hybridizing nucleic acid molecule lacking the detectable label to the first attachment site, or bringing the first complementary nucleic acid molecule containing the detectable label into contact with sufficient force to release the first detectable label and at least one second detectable label; and (5) detecting a third detectable label and A method is also provided which includes: (6) binding a second complementary nucleic acid molecule containing at least a fourth detectable label to a second attachment site among the at least three attachment sites of the barcode domain; (7) detecting and recording a third detectable label and at least a fourth detectable label of the bound second complementary nucleic acid molecule; (8) repeating steps (4) to (6) until a complementary nucleic acid molecule containing two detectable labels is bound to each of the at least three attachment sites in the barcode domain and these two detectable labels of the bound complementary nucleic acid molecule are detected and recorded; (9) using the detectable labels recorded in steps (3), (6), and (7), identifying the positions and attributes of at least six nucleotides in at least one first region of the immobilized target nucleic acid hybridized with the target-binding domain of the sequencing probe; and (10) optionally removing at least one first population of the first sequencing probe from the immobilized target nucleic acid.
[0312] This method further comprises: (10) hybridizing a second group of at least one second sequencing probe, comprising a plurality of sequencing probes described herein, to a target nucleic acid optionally immobilized at one or more locations on a substrate (provided that the target binding domains of the first sequencing probe and the second sequencing probe are different); (11) binding a first complementary nucleic acid molecule, comprising a first detectable label and at least a second detectable label, to a first attachment site among the at least three attachment sites of the barcode domain; (12) detecting and recording the first detectable label and at least a second detectable label of the bound first complementary nucleic acid molecule; and (13) recording the first hybridizing nucleic acid molecule lacking a detectable label. (14) A second complementary nucleic acid molecule containing a detectable label is bound to a first attachment site, thereby freeing the first complementary nucleic acid molecule or complex containing a detectable label, or bringing the first complementary nucleic acid molecule or complex containing a detectable label into contact with a force sufficient to release the first detectable label and at least the second detectable label; (15) A second complementary nucleic acid molecule containing a third detectable label and at least the fourth detectable label is bound to the second attachment site of the barcode domain among the at least three attachment sites; (16) A complementary nucleic acid molecule containing two detectable labels is bound to each of the attachment sites of the at least three attachment sites in the barcode domain; (13)–(15) are repeated until these two detectable labels on the bound complementary nucleic acid molecule are detected and recorded; (17) the positions and attributes of the above six nucleotides in at least one second region of the immobilized target nucleic acid hybridized with the target-binding domain of the second sequencing probe are identified using the detectable labels recorded in steps (12), (15), and (16); and (18) the removal of at least one second population of the second sequencing probe from the immobilized target nucleic acid is optionally included.
[0313] This method may further include identifying the sequence of the immobilized target nucleic acid by assembling the identified linear sequences of nucleotides in at least a first region and at least a second region of the immobilized target nucleic acid.
[0314] Steps (4) and (5) may occur sequentially or simultaneously. The first detectable label and at least the second detectable label may have the same emission spectrum or different emission spectra. The third detectable label and at least the fourth detectable label may have the same emission spectrum or different emission spectra.
[0315] The first complementary nucleic acid molecule may contain a cleavable linker. The second complementary nucleic acid molecule may also contain a cleavable linker. The first and second complementary nucleic acid molecules may each contain a cleavable linker. The cleavable linker is preferably cleavable by light. Light can be used as the emission force. UV light is preferred. The light can be provided by a light source selected from the group consisting of arc lamps, lasers, focused UV light sources, and light-emitting diodes.
[0316] The first complementary nucleic acid molecule and the first hybridizing nucleic acid molecule lacking a detectable label may contain the same nucleic acid sequence. The first hybridizing nucleic acid molecule lacking a detectable label may contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the first attachment site within the barcode domain.
[0317] The second complementary nucleic acid molecule and the second hybridizing nucleic acid molecule lacking a detectable label can contain the same nucleic acid sequence. The second hybridizing nucleic acid molecule lacking a detectable label can contain a nucleic acid sequence complementary to the single-stranded nucleotide adjacent to the second attachment site within the barcode domain.
[0318] The above method may further include a medium suitable for recording detectable markers. A suitable computer-readable medium could serve as this medium.
[0319] This sequencing method will be explained further here.
[0320] Figure 10 shows a schematic diagram of a typical sequencing cycle of the present disclosure. While the methods of the present disclosure do not require the immobilization of the target nucleic acid before sequencing, in this example, the method begins with the target nucleic acid captured using a capture probe and bound to the surface of the flow cell, as shown in the upper left figure. Next, a pool of sequencing probes is introduced into the flow cell, allowing the sequencing probes to hybridize to the target nucleic acid. In this example, the sequencing probes are shown in Figure 1. These sequencing probes contain a hexameric sequence within the target-binding domain that hybridizes to the target nucleic acid. Each hexameric has an (N) base adjacent to it. The (N) bases can be universal / degenerate bases, or any of the four non-specific canonical bases not specified by the bases b1- b2- b3- b4- b5- b6. Using the hexameric sequence, 4096(4 6 A set of 4096 sequencing probes makes it possible to sequence any target nucleic acid. In this example, a set of 4096 sequencing probes hybridizes to target nucleic acids in eight pools, each containing 512 sequencing probes. The hexameric sequence within the target-binding domain of the sequencing probe hybridizes along the length of the target nucleic acid at a position where the hexameric sequence perfectly complements the target nucleic acid, as shown in the upper center diagram of Figure 10. In this example, a single sequencing probe hybridizes to the target nucleic acid. Any sequencing probes that do not bind are washed away from the flow cell.
[0321] These sequencing probes also include a barcode domain with three attachment sites R1, R2, and R3, as described above. The attachment region at attachment site R1 contains one or more nucleotide sequences corresponding to the first dinucleotide of the hexamer of the sequencing probe. Therefore, only reporter probes containing complementary nucleic acids corresponding to the attributes of the first dinucleotide present in the target-binding domain of the sequencing probe will hybridize to attachment site R1. Similarly, the attachment region at attachment site R2 of the sequencing probe corresponds to the second dinucleotide present in the target-binding domain, and the attachment region at attachment site R3 of the sequencing probe corresponds to the third dinucleotide present in the target-binding domain.
[0322] This method follows the diagram at the top right of Figure 10. A pool of reporter probes is introduced into a flow cell. Each reporter probe in the pool contains a detectable label in the form of a two-color combination and a complementary nucleic acid that can hybridize to the corresponding attachment region within the attachment site R1 of the sequencing probe. The two-color combination and the complementary nucleic acid of a particular reporter probe correspond to one of the 16 possible dinucleotides, as described above. Each pool of reporter probes is designed so that a two-color combination corresponding to a specific dinucleotide is established before sequencing. For example, in the sequencing experiment shown in Figure 10, for the first pool of reporter probes hybridizing to attachment site R1, the yellow-red two-color combination can be associated with the adenine-thymine dinucleotide. As shown in the upper right diagram of Figure 10, after the reporter probe hybridizes to the attachment site R1, any unbound reporter probes are washed away from the flow cell, and the detectable labeling of the bound reporter probes is recorded to reveal the attributes of the first dinucleotide of the hexamer.
[0323] Remove the detectable label belonging to the reporter probe hybridized at attachment site R1. To remove the detectable label, the reporter probe may be made to include a cleavable linker and an appropriate cleavage reagent may be added. Alternatively, a complementary nucleic acid lacking the detectable label may be hybridized at attachment site R1 of the sequencing probe, replacing the reporter probe with the detectable label. Whatever method is used to remove the detectable label, attachment site R1 no longer generates a detectable signal. The method of making an attachment site of a barcode domain that previously generated a detectable signal no longer generate a detectable signal is referred to herein as "darkening."
[0324] A second pool of reporter probes is introduced into the flow cell. Each reporter probe in the pool contains a detectable label in the form of a two-color combination and a complementary nucleic acid that can hybridize to the corresponding attachment region in the attachment site R2 of the sequencing probe. The two-color combination and the complementary nucleic acid of a particular reporter probe correspond to one of 16 possible dinucleotides. A particular two-color combination may correspond to one dinucleotide in the context of the first pool of reporter probes and to a different dinucleotide in the context of the second pool of reporter probes. As shown in the lower right diagram of Figure 10, after the reporter probes have hybridized to the attachment site R2, any unbound reporter probes are washed out of the flow cell, and the detectable labels of the bound reporter probes are recorded to reveal the attributes of the hexameric second dinucleotide present in the sequencing probe.
[0325] To remove the detectable label at position R2, the reporter probe can be made to include a cleavable linker, and an appropriate cleavage reagent can be added. Alternatively, a complementary nucleic acid lacking the detectable label can be hybridized to the attachment position R2 of the sequencing probe, replacing the reporter probe with the detectable label. Whatever method is used to remove the detectable label, the attachment position R2 will no longer generate a detectable signal.
[0326] Next, a third pool of reporter probes is introduced into the flow cell. Each reporter probe in the third pool contains a detectable label in a two-color combination form and a complementary nucleic acid that can hybridize to the corresponding attachment region within the attachment site R3 of the sequencing probe. The two-color combination and the complementary nucleic acid of a particular reporter probe correspond to one of 16 possible dinucleotides. After the reporter probes hybridize to site R3, as shown in the lower center diagram of Figure 10, any unbound reporter probes are washed out of the flow cell, and the detectable labels of the bound reporter probes are recorded to reveal the attributes of the third dinucleotide of the hexameric group present in the sequencing probe. In this way, all three dinucleotides of the target-binding domain are identified, and they can be assembled together to reveal the sequence of the target-binding domain, and therefore the sequence of the target nucleic acid.
[0327] To continue sequencing of the target nucleic acid, any bound sequencing probes can be removed from the target nucleic acid. Even if a reporter probe remains hybridized at position R3 of the barcode domain, the sequencing probe can be removed from the target nucleic acid. Alternatively, a reporter probe hybridized at position R3 can be removed from the barcode domain before removing the sequencing probe from the target binding domain, for example, by using the darkening procedure described above for reporters at positions R1 and R2.
[0328] The sequencing cycle shown in Figure 10 can be repeated any number of times, and each sequencing cycle can begin by hybridizing the same pool of sequencing probes to the target nucleic acid molecule, or by hybridizing different pools of sequencing probes to the target nucleic acid molecule. It is possible for a second pool of sequencing probes to bind to the target nucleic acid at a position that overlaps with the position where the first sequencing probe, or the first pool of sequencing probes, bound to the target nucleic acid during the first sequencing cycle. In this way, it is possible to sequence some nucleotides in the target nucleic acid two or more times and to use two or more sequencing probes.
[0329] Figure 11 shows a schematic diagram of one entire cycle of the sequencing method of this disclosure and the corresponding imaging data recovered during this cycle. In this example, the sequencing probe used is shown in Figure 1, and the sequencing process is the same as shown in Figure 10 and described above. After hybridizing the sequencing domain of the sequencing probe to the target nucleic acid, the reporter probe is hybridized to the first attachment site (R1) of the sequencing probe. Next, an image of the first reporter probe is acquired and color dots are recorded. In Figure 11, the color dots are shown as dotted circles. The color dots correspond to a single sequencing probe recorded during one entire cycle. In this example, seven sequencing probes are recorded (1-7). Next, the first attachment site of the barcode domain is darkened, and the bifluorescent reporter probe is hybridized to the second attachment site (R2) of the sequencing probe. Next, an image of the second reporter probe is acquired and the color dots are recorded. Then, the second attachment site of the barcode domain is darkened, and the bifluorescent reporter probe is hybridized to the third attachment site (R3) of the sequencing probe. Next, an image of the third reporter probe is acquired and the color dots are recorded. Next, the three color dots from each sequencing probe 1 to 7 are arranged in order. Then, each color dot is mapped to a specific dinucleotide using a decoding matrix to determine the sequence of the target binding domain of sequencing probes 1 to 7.
[0330] The number of reporter probes required to sequence the target-binding domain of any sequencing probe bound to a target nucleic acid during a single sequencing cycle is equal to the number of attachment sites within the barcode domain. Therefore, for a barcode domain with three attachment sites, three reporter probes would be cycled on the sequencing probe.
[0331] A pool of sequencing probes can contain multiple sequencing probes with identical sequences, or multiple sequencing probes with different sequences. When a pool of sequencing probes contains multiple sequencing probes with different sequences, there can be an equal number of each different sequencing probe, or there can be different numbers of each different sequencing probe.
[0332] Figure 12 shows an example configuration of a sequencing probe pool according to this disclosure, where the sequencing probe contains (a) a target-binding domain containing six nucleotides (hexamers) that specifically bind to the target nucleic acid, and (b) three attachment sites (R1, R2, R3) within the barcode domain. Therefore, eight different pools of sequencing probes are designed using the eight color combinations specifically shown above. There are 4096 possible different hexamer sequences (4×4×4×4×4×4=4096). Since each of the three attachment sites within the barcode domain can hybridize to a complementary nucleic acid bound to one of the eight different color combinations, there are 512 different sets (8×8×8=512) consisting of three possible color combinations. For example, in a probe where R1 hybridizes to a complementary nucleic acid bound to color combination GG, R2 hybridizes to a complementary nucleic acid bound to color combination BG, and R3 hybridizes to a complementary nucleic acid bound to color combination YR, the set of three color combinations is GG-BG-YR. Within one pool of sequencing probes, three different sets of color combinations would correspond to different hexamers within the target binding domain. Each pool contains 512 different hexamers, resulting in a total of 4096 possible hexamers. Therefore, eight pools are needed to sequence all possible hexamers (4096 / 512=8). The specific sequencing probes to be placed in each of the eight pools are determined so that each sequencing probe optimally hybridizes to the target nucleic acid. There are several considerations to ensure optimal hybridization. The precautions include (a) separating the complete complements of the hexamers into different pools; (b) separating hexamers with high Tm and low Tm into different pools; and (c) separating hexamers into different pools based on empirically learned hybridization patterns.
[0333] Figure 13 illustrates the difference between the sequencing probe described in U.S. Patent Application Publication 20160194701 and the sequencing probe in this disclosure. As shown in the left panel of Figure 13, U.S. Patent Application Publication 2016 / 0194701 describes a sequencing probe having a barcode domain with six attachment sites that hybridize to complementary nucleic acids. Each complementary nucleic acid binds to one of four different fluorescent dyes. In this configuration, each color (red, blue, green, yellow) corresponds to one nucleotide (A, T, C, G) in the target binding domain. By designing the probe in this way, 4096 different probes (4 6 ) is generated. As shown in the right panel of Figure 13, in one example of this disclosure, the barcode domain of each sequencing probe contains three attachment sites that hybridize to complementary nucleic acids. Unlike U.S. Patent Application Publication 2016 / 0194701, one of eight color combinations (GG, RR, GY, RY, YY, RG, BB, RB) binds to these complementary nucleic acids. Each color combination corresponds to a specific dinucleotide within the target binding domain. In this configuration, 512 combinations (8 3 This generates different probes. To cover all possible hexamer combinations (4096) within the target nucleic acid domain, eight separate pools of these 512 different probes are required to sequence an entire target nucleic acid. Eight color combinations are used to label complementary nucleic acids, but there are 16 possible dinucleotides, so some color combinations will correspond to different dinucleotides depending on which pool of sequencing probes is used. For example, in Figure 13, in the first, second, third, and fourth pools of sequencing probes, color combination BB corresponds to dinucleotide AA, and color combination GG corresponds to dinucleotide AT. In the fifth, sixth, seventh, and eighth pools of sequencing probes, color combination BB corresponds to dinucleotide CA, and color combination CT corresponds to dinucleotide AT.
[0334] Multiple sequencing probes (i.e., two or more sequencing probes) can be hybridized within a sequencing window. During sequencing, the attributes and spatial positions of detectable labels bound to each of the hybridized sequencing probes are recorded. This allows for the subsequent identification of both the positions and attributes of multiple dinucleotides. In other words, hybridizing multiple sequencing probes simultaneously to a single target nucleic acid molecule allows for simultaneous sequencing of multiple locations along this target nucleic acid, thereby improving the sequencing speed.
[0335] In some embodiments, a single sequencing probe can be hybridized to a single captured target nucleic acid. In some embodiments, multiple sequencing probes can be hybridized to a single captured target nucleic acid. The sequencing window between two hybridized 5' and 3' capture probes allows a single or multiple sequencing probes to be hybridized along the length of the target nucleic acid molecule. Hybridizing multiple sequencing probes along the length of the target nucleic acid molecule allows for simultaneous sequencing of two or more locations on the target nucleic acid molecule, thus increasing the sequencing speed. The fluorescence signals from individual probes of multiple probes bound along the length of the target nucleic acid can be spatially separated.
[0336] In some embodiments, sequencing probes can be bound at equal intervals along the length of the target nucleic acid. In some embodiments, the sequencing probes do not need to be bound at equal intervals along the length of the target nucleic acid. By spatially separating the signals from multiple sequencing probes bound along the length of the target nucleic acid, sequencing information can be obtained simultaneously at multiple locations on the target nucleic acid.
[0337] The distribution of probes along the length of the target nucleic acid is crucial for the resolution of the detectable signal. Too many probes in one region can lead to overlapping detectable labels, hindering the separation of two adjacent probes. This can be explained as follows: Since one nucleotide is 0.34 nm long and the lateral (xy) spatial resolution of a sequencing instrument is approximately 200 nm, the resolution limit of the sequencing instrument is approximately 588 base pairs (i.e., 1 nucleotide / 0.34 nm × 200 nm). In other words, when two probes are within approximately 588 base pairs of each other, the above sequencing instrument is unlikely to be able to separate the signals from the two probes hybridized to the target nucleic acid. Therefore, to separate the detectable labels as separate "spots," the two probes should be spaced approximately 600 base pairs apart, depending on the resolution of the sequencing instrument. Thus, the optimal spacing should be one probe every 600 bp of the target nucleic acid. It is preferable that each sequencing probe in a probe population does not bind to each other at a distance of less than 600 nucleotides. Various software approaches (e.g., using fluorescence intensity values and wavelength-dependent ratios) can be used to monitor, limit, and potentially analyze the number of probes hybridizing into a single separable region of the target nucleic acid, allowing for the design of probe populations accordingly. Furthermore, detectable labels (e.g., fluorescent labels) that provide more distinctly different signals can be selected. Additionally, methods described in the literature (Small and Parthasarthy: "Super-Resolution Situation," Annu. Rev. Phys Chem., 2014; Vol. 65: pp. 107-125) describe various super-resolution approaches that reduce the resolution limit of sequencing microscopes to tens of nanometers with structured illumination. Using higher-resolution sequencing equipment allows for the use of probes with shorter target-binding domains.
[0338] As mentioned above, the design of the probe's Tm can affect the number of probes that hybridize to the target nucleic acid. Alternatively, or in addition to this, the concentration of sequencing probes within a single population can be increased to increase the probe coverage within a specific region of the target nucleic acid. Conversely, the concentration of sequencing probes can be decreased to reduce the probe coverage within a specific region of the target nucleic acid, for example, to exceed the resolution limit of the sequencing instrument.
[0339] The resolution limit of the two detectable labels is approximately 600 nucleotides, but this does not hinder the robust sequencing method of this disclosure. In some embodiments, multiple sequencing probes in any population will not be 600 nucleotides apart from each other on the target nucleic acid. However, statistically (following a Poisson distribution), there will be target nucleic acids to which only one sequencing probe is bound, and therefore the sequencing probe can be optically separated. For target nucleic acids that have multiple probes within 600 nucleotides (and therefore cannot be optically separated), the data relating to these inseparable sequencing probes can be discarded. Importantly, the method of this disclosure provides multiple rounds of detection in which multiple sequencing probes are bound. Thus, it is possible to detect signals from all sequencing probes in some rounds, to detect signals from only some of the sequencing probes in some rounds, and to detect no signals from any sequencing probes in some rounds. In some embodiments, the distribution of sequencing probes bound to target nucleic acids can be controlled (for example, by controlling the concentration or dilution) so that only one sequencing probe binds to a single target nucleic acid.
[0340] Randomly, but partly depending on the length of the target-binding domain, the probe's Tm, and the concentration of the probe being applied, two distinctly different sequencing probes within a single population may bind to each other within 600 nucleotides.
[0341] Alternatively, or in addition to that, the concentration of sequencing probes in a single population can be reduced to lower the probe coverage in a specific region of the target nucleic acid, for example, to exceed the resolution limit of the sequencing instrument, thereby enabling single readouts from resolution-limited spots.
[0342] If the sequence of the target nucleic acid, or a portion of its sequence, is known before sequencing the target nucleic acid using the method disclosed herein, the sequencing probes can be designed and selected such that no two sequencing probes bind to each other by more than 600 nucleotides.
[0343] Before hybridizing a sequencing probe to a target nucleic acid, one or more complementary nucleic acid molecules can be bound with a first detectable label, and at least a second detectable label can be hybridized to one or more attachment sites within the barcode domain of the sequencing probe. For example, before hybridizing to the target nucleic acid, one or more complementary nucleic acid molecules bound with the first detectable label and at least a second detectable label can be hybridized to the first attachment site of each sequencing probe. Therefore, since the sequencing probe can generate a detectable signal from the first attachment site upon contact with the target nucleic acid, there is no need to prepare a first pool of complementary nucleic acids or reporter probes directed to the first position on the barcode domain. In another example, one or more complementary nucleic acid molecules bound with the first detectable label and at least a second detectable label can be hybridized to all attachment sites within the barcode domain of the sequencing probe. Therefore, in this example, a sequence consisting of six nucleotides can be read without the need to sequentially exchange the complementary nucleic acids. Using this pre-hybridized sequencing probe-reporter probe complex omits many steps in the described method, thus shortening the time required to acquire sequence information. However, this probe is considered beneficial only when the detectable labels do not overlap. For example, a phosphor is excited by light of non-overlapping wavelengths, or emits light of non-overlapping wavelengths.
[0344] In some embodiments of the methods of this disclosure, target nucleic acids can be accurately sequenced using the signal intensity from recorded color dots. In some embodiments, the probability that a particular color dot corresponds to a color combination where one color overlap (i.e., BB, GG, YY, RR) can be determined using the spot intensity of a particular color within a single color dot.
[0345] Darkening of a specific location within the barcode domain can be achieved by cleaving the strand at the site of a cleavable linker modification present in the reporter probe hybridized to that location. Figure 14 illustrates darkening of a barcode location during a single sequencing cycle using a cleavable linker modification. The first step, shown in the leftmost panel of Figure 14, involves hybridizing the primary nucleic acid of the reporter probe to the first attachment site of the sequencing probe. The primary nucleic acid hybridizes to a specific complementary sequence within the attachment region of the first location of the barcode domain. The first and second domains of the primary nucleic acid are covalently linked by a cleavable linker modification. In the second step, a detectable label is recorded to determine the attribute and location of a specific dinucleotide within the target-binding domain of the sequencing probe. In the third step, the first location of the barcode domain is darkened by cleaving the reporter probe at the site of the cleavable linker modification. This releases the second domain of the primary nucleic acid, thereby releasing a detectable label. The first domain of the primary nucleic acid molecule remains hybridized to the first attachment site of the barcode domain, now lacking any detectable labeling. Therefore, the first position of the barcode domain no longer generates a detectable signal and will not be able to hybridize to any other reporter probe in the subsequent sequencing process. In the final step shown in the rightmost diagram of Figure 14, the reporter probe hybridizes to the second position of the barcode domain, and sequencing continues.
[0346] The attachment site of the barcode domain can be darkened by substituting any secondary or tertiary nucleic acid in the reporter probe to which a detectable label is bound, while the primary nucleic acid of the reporter probe remains hybridized to the sequencing probe. This substitution can be achieved by hybridizing the primary nucleic acid with a secondary or tertiary nucleic acid that is not bound to a detectable label. Figure 15 shows an example of a typical sequencing cycle of this disclosure, in which one site in the barcode domain is darkened by substitution of a secondary nucleic acid with a label. The leftmost panel of Figure 15 shows the start of the sequencing cycle, in which the primary nucleic acid molecule of the reporter probe hybridizes to the first attachment site of the barcode domain of the sequencing probe. Next, the secondary nucleic acid molecule bound to the detectable label hybridizes to the primary nucleic acid molecule, and the detectable label is recorded. To darken the first site in the barcode domain, the secondary nucleic acid molecule bound to the detectable label is substituted with a secondary nucleic acid molecule that does not have a detectable label. In the next step of the sequencing cycle, a reporter probe containing a detectable label is hybridized to a second position in the barcode domain. Darkening of the barcode domain attachment site is possible by hybridizing an unlabeled nucleic acid to the sequencing probe at the corresponding attachment site in the barcode domain, thereby replacing any primary nucleic acid molecule in the reporter probe. If the barcode domain contains at least one single-stranded nucleic acid sequence adjacent to or near at least one attachment site, the unlabeled nucleic acid can replace the primary nucleic acid molecule by hybridizing its flanking sequence to the portion of the barcode domain occupied by that primary nucleic acid molecule. If necessary, the rate of exchange of the detectable label can be increased by incorporating a small single-stranded oligonucleotide that increases the rate of exchange of the detectable label (e.g., "toe-hold" probes; see, e.g., Seeling et al., "Catalytic Relaxation of Metastable DNA Fuel"; J. Am. Chem. Soc. 2006, Vol. 128(37), pp. 12211-12220).
[0347] Complementary nucleic acids containing detectable labels, i.e., reporter probes, can be removed from the attachment site but cannot be replaced with hybridizing nucleic acids lacking detectable labels. This can be achieved, for example, by adding chaotropic agents, and / or increasing the temperature, and / or changing the salt concentration, and / or adjusting the pH, and / or applying hydrodynamic forces. In these examples, fewer reagents (i.e., hybridizing nucleic acids lacking detectable labels) are required.
[0348] The methods of this disclosure can be used to simultaneously capture and sequence RNA and DNA molecules (including mRNA and gDNA) from the same sample. The capture and sequencing of both RNA and DNA molecules from the same sample can be performed within the same flow cell. In some embodiments, the methods of this disclosure can be used to simultaneously capture, detect, and sequence both gDNA and mRNA from an FFPE sample.
[0349] The sequencing method of the present disclosure further includes the step of identifying the sequence of an immobilized target nucleic acid by assembling the linear order of nucleotides identified for each region of the immobilized target nucleic acid, thereby identifying the sequence for the optionally immobilized target nucleic acid. The assembly step utilizes a non-temporary computer-readable storage medium storing an executable program. The nucleic acid sequence is obtained by the program issuing instructions to a microprocessor to assemble the linear order of nucleotides identified for each region of the target nucleic acid. The assembly operation can be performed in "real time," that is, while the data is being retrieved from the sequencing probe, rather than after all the data has been retrieved or after complete data acquisition.
[0350] The raw specificity of the sequencing method of this disclosure is approximately 94%. The accuracy of the sequencing method of this disclosure can be increased to nearly 99% by sequencing the same base in the target nucleic acid using two or more sequencing probes. Figure 16 illustrates how the sequencing method of this disclosure enables sequencing of the same base in the target nucleic acid using different sequencing probes. In this example, the target nucleic acid is a fragment of NRAS exon 2 (SEQ ID NO: 1). The specific base of interest is cytosine (C), which is highlighted in the target nucleic acid. This base of interest will hybridize to two sequencing probes with different hybridization footprints for the target nucleic acid. In this example, sequencing probes 1-4 (barcodes 1-4) bind to the three nucleotides to the left of the base of interest, while sequencing probes 5-8 (barcodes 5-8) bind to the five nucleotides to the left of the base of interest. As a result, the target base is sequenced by two different probes, increasing the amount of base readings at that particular location and thereby improving the overall accuracy at that location. Figure 17 shows how multiple different base readings at a specific nucleotide location on the target nucleotide are recorded by one or more sequencing probes and then combined to create a consensus sequence (SEQ ID NO: 2), thereby improving the accuracy of the final base reading.
[0351] The terms "Hyb & Seq chemistry," "Hyb & Seq sequencing," and "Hyb & Seq" mean the methods described above in this disclosure.
[0352] The array disclosed herein and a method for using this array
[0353] As described in detail herein, this disclosure provides compositions and methods for immobilizing nucleic acid molecules, including arrays and methods utilizing arrays.
[0354] The present disclosure provides a composition comprising: a flat solid support substrate; a first layer on the flat solid support substrate; and a second layer on the first layer; the second layer comprising a plurality of nanowells, each nanowell providing access to an exposed portion of the first layer, and each nanowell comprising a plurality of first oligonucleotides bonded to the exposed portion of the first layer.
[0355] The present disclosure provides a composition comprising: a flat solid support substrate; a first layer on the flat solid support substrate and in contact with a first surface of the flat solid support substrate; a second layer on the first layer and in contact with a second surface of the first layer; wherein the second surface of the first layer is not in contact with the surface of the flat solid support substrate, and the second layer comprises a plurality of nanowells, each nanowell providing access to an exposed portion of the first layer, and each nanowell comprises a plurality of first oligonucleotides bonded to the exposed portion of the first layer.
[0356] The first layer may include a first surface that is in contact with the surface of a flat solid support substrate and a second surface that is in contact with the second layer but not with the surface of the flat solid support substrate.
[0357] The second layer may include a first surface in contact with the surface of the first layer and a second surface exposed to the environment.
[0358] Figure 47 is a schematic cross-sectional view of a typical array of the present invention. This array includes a flat solid support substrate 101, a first layer 102 on the flat solid support substrate 101, and a second layer 103 on the first layer 102. The second layer 103 contains a plurality of nanowells 104. Each nanowell 104 is open on two sides, so that a portion 105 of the first layer is exposed within each nanowell. A plurality of first oligonucleotides 106 are covalently bonded to the exposed first layer 105 within each nanowell.
[0359] In some embodiments, the flat solid support substrate can be any of a surface, a film, beads, a porous material, or an electrode. Non-limiting examples of the flat solid support substrate include, for example, a polymer material, metal, silicon, glass, or quartz.
[0360] In some embodiments, the first layer 102 can include an oxide film, a non-limiting example of which is silicon dioxide.
[0361] In some embodiments, the first layer 102 can have a thickness of about 50 to about 150 nm. The first layer 102 can have a thickness of about 90 nm.
[0362] In some embodiments, the second layer 103 can include, as a non-limiting example, bis(trimethylsilyl)amine (also known as hexamethyldisilazane (HMDS or HDMS)).
[0363] In some embodiments, the second layer 103 can include a material that does not chemically react, so the second layer does not bind to the biopolymer.
[0364] In some embodiments, the second layer 103 can have a thickness of about 1 nm to about 10 nm. The second layer 103 can have a thickness of about 3 nm to about 4 nm.
[0365] In some embodiments, the flat solid support substrate contains silicon, the first layer contains silicon dioxide, and the second layer contains HMDS.
[0366] In some embodiments, the flat solid support substrate contains glass, the first layer contains silicon dioxide, and the second layer contains HMDS.
[0367] In some embodiments, the second layer can include about 0.1×10 5 ~ about 100×10 7 nanowells per square millimeter. The second layer can in...
Claims
1. A method for identifying the presence of a target nucleic acid, the following: (1) Hybridize the target binding domain of the probe to the target nucleic acid molecule, The probe includes a target-binding domain and a barcode domain, The target-binding domain is at least 12 nucleotides in length; The barcode domain comprises L-DNA and includes at least two attachment sites, each attachment site containing at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule; (2) Hybridize the first complementary nucleic acid molecule to the first attachment site of the barcode domain, The first complementary nucleic acid molecule comprises at least one first detectable label and at least one second detectable label; (3) Identify the at least one first detectable label and the at least one second detectable label of the first complementary nucleic acid molecule hybridized to the first attachment site; (4) Remove the at least one first detectable marker and the at least one second detectable marker that have hybridized to the first attachment position; (5) Hybridize a second complementary nucleic acid molecule to the second attachment site of the barcode domain, The second complementary nucleic acid molecule comprises at least one third detectable label and at least one fourth detectable label; (6) Identify the at least one third detectable label and the at least one fourth detectable label of the second complementary nucleic acid molecule hybridized to the second attachment site, This allows for the determination of the presence of the target nucleic acid based on the attributes of the at least one first detectable label, the at least one second detectable label, the at least one third detectable label, and the at least one fourth detectable label. Methods that include...
2. The method according to claim 1, wherein at least one of the at least two attachment sites has a different nucleic acid sequence compared to all the other attachment sites.
3. The method according to claim 1 or 2, wherein each attachment site has a different nucleic acid sequence compared to all other attachment sites.
4. The method according to any one of claims 1 to 3, wherein each of the attachment positions within the barcode domain has a length of 6 to 20 nucleotides.
5. Each attachment position within the aforementioned barcode domain is a) A nucleotide with a length of 9; b) A nucleotide with a length of 12; c) A nucleotide with a length of 14; or d) A nucleotide with a length of 16, The method according to claim 4.
6. The method according to any one of claims 1 to 5, wherein the target-binding domain comprises D-DNA.
7. The method according to any one of claims 1 to 6, wherein steps (4) and (5) are performed sequentially or simultaneously.
8. The method according to any one of claims 1 to 7, wherein the first detectable label and the second detectable label have the same emission spectrum or have different emission spectra.
9. The method according to any one of claims 1 to 8, wherein the third detectable label and the fourth detectable label have the same emission spectrum or have different emission spectra.
10. The method according to any one of claims 1 to 9, wherein the first complementary nucleic acid molecule and / or the second complementary nucleic acid molecule each include a linker capable of cleaving them.
11. The first complementary nucleic acid molecule is a reporter probe, and the reporter probe comprises a primary nucleic acid molecule having at least two domains. The first domain hybridizes to the first attachment position of the barcode domain; and The second domain hybridizes into six secondary nucleic acid molecules, Each of these secondary nucleic acid molecules hybridizes into five tertiary nucleic acid molecules, and each of the tertiary nucleic acid molecules contains a detectable label. The method according to any one of claims 1 to 10.
12. The method according to claim 11, wherein the primary nucleic acid molecule includes a cleavable linker located between the first domain and the second domain.
13. The method according to claim 12, wherein the linker that can be cut is a linker that can be cut by light.
14. Each of the aforementioned secondary nucleic acid molecules comprises at least two domains, The first domain hybridizes to the second domain of the primary nucleic acid molecule. The method according to any one of claims 11 to 13, wherein the second domain hybridizes to the five tertiary nucleic acid molecules.
15. The method according to claim 14, wherein each of the secondary nucleic acid molecules includes a cleavable linker located between the first domain and the second domain.
16. The method according to claim 15, wherein the linker that can be cut is a linker that can be cut by light.
17. Removal of the at least one first detectable marker and the at least one second detectable marker hybridized to the first attachment position, The cleavage of a cleavable linker between the first and second domains of the primary nucleic acid, the cleavage of a cleavable linker between the first and second domains of each secondary nucleic acid, or any combination thereof. The method according to any one of claims 12 to 16, carried out by...
18. A composite, the following: i) probe, The probe includes a target-binding domain and a barcode domain, The target-binding domain is at least 12 nucleotides in length and hybridizes to the target nucleic acid molecule; The barcode domain comprises L-DNA and comprises at least two attachment sites, each attachment site comprising at least one nucleic acid sequence that hybridizes to a complementary nucleic acid molecule; and ii) A complementary nucleic acid hybridized to the first attachment position of the at least two attachment positions of the barcode domain of the probe, The complementary nucleic acid comprises at least one first detectable label and at least one second detectable label. A complex that includes this.
19. The complex according to claim 18, wherein at least one of the at least two attachment sites has a different nucleic acid sequence compared to all the other attachment sites.
20. The complex according to claim 18 or claim 19, wherein each attachment site has a different nucleic acid sequence compared to all other attachment sites.
21. The complex according to any one of claims 18 to 20, wherein each of the attachment sites within the barcode domain has a length of 6 to 20 nucleotides.
22. Each attachment position within the barcode domain is a) A nucleotide with a length of 9; b) A nucleotide with a length of 12; c) A nucleotide with a length of 14; or d) A nucleotide with a length of 16, The composite according to claim 21.
23. The complex according to any one of claims 18 to 22, wherein the target binding domain comprises D-DNA.
24. The composite according to any one of claims 18 to 23, wherein the first detectable label and the second detectable label have the same emission spectrum or have different emission spectra.
25. The complex according to any one of claims 18 to 24, wherein the complementary nucleic acid molecule comprises a cleavable linker.
26. The complementary nucleic acid molecule is a reporter probe, The reporter probe comprises a primary nucleic acid molecule having at least two domains, The first domain hybridizes to the first attachment position of the barcode domain; and The second domain is hybridized into six secondary nucleic acid molecules. Each of these secondary nucleic acid molecules is hybridized into five tertiary nucleic acid molecules, and each of the tertiary nucleic acid molecules contains a detectable label. The composite according to any one of claims 18 to 25.
27. The complex according to claim 26, wherein the primary nucleic acid molecule comprises a cleavable linker located between the first domain and the second domain.
28. The composite according to claim 27, wherein the severable linker is a linker that can be cut by light.
29. Each of the secondary nucleic acid molecules comprises at least two domains, The first domain hybridizes to the second domain of the primary nucleic acid molecule; and The complex according to any one of claims 26 to 28, wherein the second domain hybridizes to the five tertiary nucleic acid molecules.
30. The complex according to any one of claims 26 to 29, wherein the secondary nucleic acid molecule includes a cleavable linker located between the first domain and the second domain.
31. The composite according to claim 30, wherein the severable linker is a linker that can be cut by light.