A method for in vitro transcription and uses thereof
By adding a T7 promoter to the positive strand of single-stranded DNA at the 5' end and performing in vitro transcription, the problems of low amplification efficiency and large bias in existing technologies are solved, achieving efficient and quantitative amplification of low starting amounts and damaged DNA samples, and making it suitable for library construction of various DNA samples.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2025-01-21
- Publication Date
- 2026-07-21
AI Technical Summary
Existing DNA library preparation methods suffer from low amplification efficiency and bias when applied to samples with low starting amounts, damaged DNA, and samples with high requirements for amplification bias. They are particularly ineffective when processing cell-free blood DNA, paraffin-embedded tissue samples, paleontological DNA, and DNA converted from bisulfite.
By employing a method where the 5' end of a single-stranded DNA contains a T7 promoter positive strand, a specific sequence for linear transcription is ligated to the single-stranded DNA using ligase or chemical methods. T7 RNA polymerase is then used for in vitro transcription to generate RNA, achieving unbiased linear amplification.
It achieves efficient and quantitative amplification of low starting amounts and damaged DNA samples, reduces amplification bias, and is suitable for library construction of various DNA samples.
Smart Images

Figure CN122428005A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of RNA in vitro synthesis technology, specifically to a method for obtaining RNA through in vitro transcription using a single-stranded DNA with a T7 promoter at the 5' end as a template, and further applications thereof. Background Technology
[0002] As scientists delve deeper into the study of ribonucleic acid (RNA), RNA has not only become a carrier of genetic information, but has also been developed into drugs for the treatment of diseases, as well as into RNA vaccines and tools such as guide RNA (gRNA) in the CRISPR system and prime editing guide RNA (pegRNA) in gene editing systems.
[0003] Currently, there are two methods for obtaining RNA in vitro: chemical synthesis and enzymatic synthesis. However, chemical synthesis can usually only synthesize relatively short RNAs, and the synthesis cost increases rapidly with the increase in RNA length. Enzymatic synthesis of RNA is a better option. The method based on the T7 promoter and T7 RNA polymerase (T7 RNAP) is the most commonly used method for in vitro RNA synthesis. The principle is as follows: (1) When the template is double-stranded DNA, the ends of the double-stranded DNA contain the T7 promoter sequence. T7 RNAP binds to the T7 promoter region and uses the DNA strand containing the antisense strand of the T7 promoter as a template to synthesize RNA. The sequence of the obtained RNA is consistent with the sequence of the DNA strand containing the sense strand of the T7 promoter in the template strand; (2) When the template is single-stranded DNA, the 3' end of the template DNA needs to contain the antisense strand sequence of the T7 promoter, and the sense strand and antisense strand of the T7 promoter form a local double-stranded DNA sequence at the 3' end of the template strand. Then, under the action of T7 RNAP, RNA is synthesized using DNA as a template, and the obtained RNA sequence is the complementary sequence of the template DNA.
[0004] In addition to the applications mentioned above, researchers have also combined the principle of RNA polymerase transcription to generate RNA by adding promoter sequences to the ends of DNA to produce large amounts of RNA in vitro, thereby achieving linear amplification, quantification, and sequencing of DNA.
[0005] Compared to exponential amplification processes (such as PCR), which suffer from amplification bias and amplified erroneous signals, T7 RNA polymerase-based amplification is limited in scenarios requiring quantification and low amplification bias. T7 RNA polymerase-based amplification is a commonly used linear amplification method. For example, LIATI technology uses the Tn5 transposon to fragment single-cell DNA genomes and add the T7 promoter sequence, further amplifying the template DNA through T7 transcription. However, DNA fragments from sources such as cell-free blood DNA, paraffin-embedded tissue (FFPE) DNA, paleontological DNA, and bisulfite-converted DNA are relatively short, few in quantity, and not entirely complete double-stranded DNA (containing single-stranded DNA and damaged bases). Therefore, the LIATI method is not well-suited for these scenarios. The LABS method adds the T7 promoter to the unconverted double-stranded DNA, followed by bisulfite conversion and T7 transcription of the conversion product, resulting in a large amount of linearly amplified transformed DNA. This allows for the quantification and unbiased amplification of DNA methylation information. However, in LABS, the T7 promoter is added before bisulfite conversion. Bisulfite causes significant damage to DNA, easily breaking it into short fragments. Therefore, this method cannot efficiently utilize all the input DNA information. Considering that the transformation products of bisulfite are broken DNA and incomplete DNA strands, it might be more promising to denature all the bisulfite-converted DNA into single-stranded DNA and then efficiently add transcription adapters and perform linear amplification to achieve efficient library construction using the input DNA. This would further enable library construction with low starting amounts of methylated and hydroxymethylated DNA samples, allowing for quantifiable and unbiased amplification.
[0006] In summary, existing DNA library preparation methods still have limitations when applied to some low-starting-volume DNA samples, damaged DNA samples, and DNA samples with high requirements for amplification bias (such as: cell-free blood DNA, paraffin-embedded tissue (FFPE) DNA, paleontological DNA, and DNA after bisulfite conversion). Summary of the Invention
[0007] Unlike the traditional T7 RNAP principle, this invention reports a novel T7 RNAP transcription method: the 5' end of a single-stranded DNA contains a T7 promoter positive strand, and in the presence of T7 RNA polymerase, this single-stranded DNA can be used as a template for in vitro transcription to generate RNA. Based on this principle, this application provides a method for quantitative and unbiased amplification of template DNA based on linear amplification of single-stranded DNA, the flowchart of which is shown below. Figure 2 As shown. The specific solution is as follows:
[0008] In a first aspect, the present invention provides a method for in vitro transcription, the method comprising: performing in vitro transcription using a single-stranded DNA 5' end containing a specific sequence for linear transcription function as a template.
[0009] Preferably, in vitro transcription is performed using a template with a specific linear transcriptional sequence added to the 5' end of a single-stranded DNA.
[0010] The added linear transcription functional sequence can be directly synthesized or directly or indirectly (e.g., via a nucleic acid functional fragment) linked to single-stranded DNA through a ligation method.
[0011] The ligation method can be any existing ligation method, as long as it can perform the ligation between sequences. Preferred ligation methods include ligase ligation or chemical ligation.
[0012] Specifically, the ligase ligation method involves connecting the 5' end of a single-stranded DNA to the 3' end of a linear transcription functional sequence using a ligase. The addition of the linear transcription functional sequence to the 5' end of the single-stranded DNA forms a single strand, a notched double strand, or a locally formed double strand. For example, through nucleic acid design, the 5' end of the single-stranded DNA can be locally linked to the 3' end of the linear transcription functional sequence to form a single-stranded nucleic acid 1 / single-stranded nucleic acid 2 structure; a nick can be formed between nucleic acid 1 and nucleic acid 2; or a locally formed double-stranded nucleic acid can be constructed.
[0013] More specifically, the ligase is:
[0014] Ligases capable of connecting nicks and / or double-stranded nucleic acids, such as T4 RNA ligase 2, T4 DNA ligase, T3 DNA ligase, SplintR ligase, Taq DNA ligase, 9°N DNA ligase, E. coli ligase, and mutants of the above ligases; or ligases capable of connecting single-stranded nucleic acids, such as T4 RNA ligase 1 (ssRNA ligase), Mth RNA ligase, 5' App DNA / RNA thermostable ligase, TS2126 RNA ligase (TS2126Rnl 1 or CircLigase), Hyper-Thermostable Lysine-Mutatant ssDNA / RNA ligase (HyperLigase), T4 RNA ligase 2, RtcB ligase, and mutants of the above ligases.
[0015] Specifically, the chemical ligation method is a method of ligating the 5' end of a single-stranded DNA to the 3' end of a specific sequence for linear transcription through a chemical reaction;
[0016] More specifically, the chemical reactions include methods such as click chemistry or Michael addition reaction, for example, linkages via oxime bonds, amide bonds, thioether bonds, disulfide bonds, phosphoryl bonds, hydrazone bonds, urea bonds, or ring linkages formed by click chemistry.
[0017] In one specific embodiment of the present invention, the ligation of the linear transcription functional specific sequence is performed using a ligase ligation method, wherein the reaction system contains a ligase, and preferably may also include one or more of the following: polyethylene glycol (PEG), divalent cation, adenine nucleoside triphosphate (ATP), dimethyl sulfoxide (DMSO), buffer (e.g., Tris-HCl buffer), double-distilled water, ribonuclease inhibitor, or bovine serum albumin (BSA).
[0018] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences and RNA ligase.
[0019] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences, RNA ligase, and PEG.
[0020] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences, RNA ligase, PEG, and BSA.
[0021] In one specific embodiment of the present invention, the reaction system for linking the linear transcription functionally specific sequence includes: the linear transcription functionally specific sequence, RNA ligase, PEG, ATP, and buffer (e.g., Tris-HCl buffer).
[0022] In one specific embodiment of the present invention, the reaction system for linking the linear transcription functionally specific sequence includes: the linear transcription functionally specific sequence, RNA ligase, PEG, ATP, RNase inhibitor, buffer (e.g., Tris-HCl buffer), and DMSO.
[0023] The in vitro transcription reaction system includes T7 RNA polymerase, or a mutant thereof.
[0024] Preferably, the method includes using a linear transcription functionally specific sequence linked to the 5' end of a single-stranded DNA as a template, followed by transcription in the presence of T7 RNA polymerase and NTP.
[0025] The final concentration of T7 RNA polymerase in the in vitro transcription reaction system is 1-5000000000 units / ml, preferably 1-50000 units / ml, and more preferably 10-20000 units / ml.
[0026] The in vitro transcription reaction system also includes NTPs, which may be modified or unmodified. Preferably, the NTPs include, but are not limited to, one or more of ATP, UTP, GTP, CTP, modified nucleotides, or non-natural nucleotides.
[0027] The final concentration of NTPs in the in vitro transcription reaction system is 1 nM-100 mM, preferably 100 nM-50 mM. More preferably 1 mM-25 mM.
[0028] In one specific embodiment of the present invention, the in vitro transcription reaction system includes a single-stranded DNA with a 5' end linked to a linear transcription functional sequence as a template, T7 RNA polymerase, and NTP.
[0029] In one specific embodiment of the present invention, the in vitro transcription reaction system includes a single-stranded DNA with a 5' end linked to a specific sequence for linear transcription as a template, T7 RNA polymerase, RNase inhibitor, and NTP.
[0030] The reaction temperature for the in vitro transcription is any value between 10℃ and 65℃, preferably 16-42℃. For example, 10, 15, 20, 25, 30, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 60, 65℃, etc.
[0031] The reaction time for the in vitro transcription is any value from 0.01h to 1000h, preferably any value from 0.1h to 100h, and more preferably any value from 0.5h to 72h. For example, 0.01, 0.1, 0.5, 1, 2, 3, 4, 5, 10, 15, 16, 17, 18, 19, 20, 25, 30, 35, 40, 45, 50, 60, 70, 71, 72, 73, 74, 75, 80, 90, 95, or 100h, etc.
[0032] The single-stranded DNA is phosphorylated before being linked to a specific sequence for linear transcription at its 5' end.
[0033] The phosphorylation reaction system includes T4 polynucleotide kinase (T4 PNK).
[0034] The phosphorylation reaction system also includes ATP.
[0035] Preferably, the final concentration of ATP in the phosphorylation reaction system is 0.01-100 mM, more preferably 1-50 mM, for example 0.01, 0.1, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, or 100. The ATP can be present in a buffer solution.
[0036] In one specific embodiment of the present invention, the phosphorylation reaction system includes ATP, T4 PNK and buffer solution.
[0037] The phosphorylation reaction temperature is any value between 10℃ and 100℃, preferably 16-80℃. For example, 10, 15, 20, 25, 30, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100℃, etc.
[0038] The phosphorylation reaction time is greater than 2 min, for example 5 min, 10 min, 15 min, 20 min, 25 min, 30 min, 35 min, 40 min, 45 min, 50 min, 1 h, 2 h, 3 h, 4 h, 5 h, 6 h, 7 h, 8 h, 9 h, 10 h and above.
[0039] More preferably, the phosphorylation reaction temperature and time are 16-42℃ for 10-50 min and then 70-80℃ for 5-15 min.
[0040] Preferably, the method includes phosphorylating single-stranded DNA, attaching a linear transcription functional specific sequence to the 5' end as a template, and then performing in vitro transcription.
[0041] In one specific embodiment of the present invention, the product of adding a specific sequence for linear transcription to the 5' end of a single-stranded DNA is purified and then transcribed in vitro.
[0042] In one specific embodiment of the present invention, the product of adding a linear transcription-specific sequence to the 5' end of a single-stranded DNA can be directly transcribed in vitro without purification. For example, reagents required for in vitro transcription can be added directly to the reaction system containing the linear transcription-specific sequence.
[0043] The product obtained from the in vitro transcription is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted enriched and / or amplified to obtain a library.
[0044] In one specific embodiment of the present invention, the product obtained by in vitro transcription is purified and then reverse transcribed.
[0045] In one specific embodiment of the present invention, the product obtained by in vitro transcription is directly reverse transcribed without purification.
[0046] Preferably, the method further includes separating the product obtained from the in vitro transcription to obtain template DNA and transcription product RNA; wherein,
[0047] RNA is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted sequencing and / or amplification to obtain a library.
[0048] Template DNA is used for in vitro transcription. Preferably, the template DNA can be left untreated or treated before in vitro transcription.
[0049] Preferably, the special treatment includes DNase digestion, or separation, purification, and denaturation into single strands for direct in vitro transcription as a template, or other treatments, such as transformation with one or more of the following modifications: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), N6-methyladenine (N6-mA), 7-methylguanine (7-mG), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), dihydrouracil (DHU), 5-(β-glucoseoxymethylated)cytosine (5gmC), etc.
[0050] The sequencing includes one or more of the following sequencing methods: whole genome sequencing (WGS), de novo sequencing, whole exome sequencing (WES), metagenomic sequencing, 16S sequencing, methylation sequencing, hydroxymethylation sequencing, single-cell sequencing, chromosome accessibility sequencing, or targeted sequencing.
[0051] Specifically, the methylation sequencing can be whole-genome DNA methylation sequencing (WGBS), simplified genome methylation sequencing (RRBS / dRRBS / XRBS), high-throughput single-cell methylation sequencing (sc-RBS), target gene DNA methylation sequencing (Target-BS / LHC-BS / Capture-BS), precise DNA methylation and hydroxymethylation sequencing (oxBS-seq), single-cell and micro-sample DNA methylation sequencing (Micro DNA-BS), micro-cfDNA genome methylation sequencing (cfDNA-BS), (hydroxy)methylated DNA immunoprecipitation sequencing ((h)MeDIP-seq / 5hmC-Seal), and other sequencing methods.
[0052] Methods for obtaining single-stranded DNA include isolating DNA from a sample. The isolated DNA can be single-stranded or double-stranded. Specifically, the double-stranded DNA can be complete or incomplete double-stranded.
[0053] The 5' end of the double-stranded DNA has more than one free base or nucleotide residue. Preferably, the more than one free base or nucleotide residue at the 5' end of the double-stranded DNA can be DNA directly obtained from a sample, or it can be obtained by physical (e.g., by temperature denaturation, breakage) or chemical (including the use of chemical agents, such as urea or formamide) or biological methods (e.g., enzyme digestion, enzyme ligation, or enzyme addition of bases to the end).
[0054] When the separated DNA is double-stranded, it first denatures into single-stranded DNA.
[0055] Preferably, the denaturation is performed using physical, chemical, or biological methods.
[0056] Specifically, physical methods include temperature denaturation and fracture.
[0057] Specifically, chemical methods include using chemical agents such as urea or formamide.
[0058] Specifically, the chemical method can be alkaline pyrolysis.
[0059] Specifically, the biological method may be enzyme digestion, enzyme ligation, or adding a base to the end of an enzyme, etc.
[0060] Preferably, the nucleic acid chain with a 5' free base or nucleotide residue is cleaved by an endonuclease or USER enzyme (a mixture of uracil DNA glycosylase and DNA endonuclease VIII) or uracil DNA glycosylase (UDG). Alternatively, the nucleic acid chain with a 5' free base or nucleotide residue can be cleaved by an exonuclease. Alternatively, the 5' free nucleic acid chain can be cleaved by a polymerase with 3' end cleavage activity. Alternatively, the 5' free sequence can be extended by a polymerase with 5' end extension activity. Alternatively, the 5' free sequence can be cleaved by Cas enzyme, Fanzor enzyme, or TnpB enzyme (for a detailed introduction to Fanzor enzyme or TnpB enzyme, see the literature: MakotoSaito, et al., Fanzor is a eukaryotic programmable RNA-guided endonuclease, Nature, 2023). Alternatively, the 5' free sequence can be cleaved by a transposase and cleaved by the added 5' free sequence. Alternatively, two or more of the above enzymes can be used.
[0061] AarI, AbsI, Acc36I, Acc65I, AccB1I, AccI, AccIII, AciI, AclI AclWI、AcoI、AcsI、AcyI、AflII、AflIII、AgeI、AgeI-HF、AhlI、AjnI、Alw26I、A lw44I、AlwI、Ama87I、Aor13HI、AoxI、ApaLI、ApeKI、ApoI、ApoI-HF、AscI、AseI 、AsiGI、Asp718I、AspA2I、AspS9I、AsuC2I、AsuII、AsuNHI、AvaI、AvaII、AvrII、 AxyI、BamHI、BamHI-HF、BanI、BauI、BbsI、BbsI-HF、BbvCI、BbvI、BccI、BceAI、 BciT130I, BclI, BclI-HF, BcnI, BcoDI, BcuI, BfaI, BfmI, BfrI, BfuAI, BglII, B isI, BlnI, BlpI, Bme1390I, Bme18I, BmeT110I, BmgT120I, BmrFI, BmsI, BpiI, B pu10I, Bpu1102I, Bpu14I, BpuMI, Bsa29I, BsaHI, BsaI-HFv2, BsaJI, BsaWI, Bse 118I, Bse21I, BseAI, BseBI, BseCI, BseDI, BsePI, BseX3I, BseXI, BseY I、BshNI、BshTI、BshVI、BsiHKCI、BsiSI、BsiWI、BsiWI-HF、BslFI、BsmA I、BsmBI-v2、BsmFI、Bso31I、BsoBI、Bsp119I、Bsp120I、Bsp13I、Bsp140 7I, Bsp143I, Bsp1720I, Bsp19I, BspACI, BspDI, BspEI, BspHI, BspMI, Bs pPI, BspQI, BspT104I, BspT107I, BspTI, BspTNI, BsrFI-v2, BsrGI, Bsr GI-HF、BssAI、BssECI、BssHII、BssMI、BssNI、BssSI-v2、BssT1I、Bst2BI 、Bst2UI、Bst6I、BstACI、BstAFI、BstAUI、BstBI、BstDEI、BstDSI、BstE II、BstEII-HF、BstENI、BstMAI、BstMBI、BstNI、BstPI、BstSCI、BstSFI、BstV1I, BstV2I, BstX2I, BstYI, BstZI, Bsu15I, Bsu36I, BsuTUI, BtgI, Btg ZI、BveI、CciI、CciNI、Cfr10I、Cfr13I、Cfr9I、ClaI、CpoI、CseI、CsiI、Csp 6I、CspAI、CspI、CviAII、CviQI、DdeI、DpnII、EaeI、EagI-HF、Eam1104I、Ea rI、EclXI、Eco130I、Eco31I、Eco47I、Eco52I、Eco81I、Eco88I、Eco91I、EcoN I、EcoO109I、EcoO65I、EcoRI、EcoRI-HF、EcoRII、EcoT14I、ErhI、Esp3I、Fa qI、FatI、FauI、FauNDI、FbaI、FblI、Fnu4HI、FokI、Fsp4HI、FspBI、FspEI、Gl uI、HapII、HgaI、Hin1I、Hin6I、HinP1I、HindIII、HindIII-HF、HinfI、HpaI I、Hpy188III、HpyCH4IV、HpyF3I、HpySE526I、Hsp92I、HspAI、KasI、KflI、Kp n2I, KroI, Ksp22I, Kzo9I, LguI, LpnPI, Lsp1109I, LweI, MabI, MaeI, MaeII 、MaeIII、MauBI、MboI、MfeI、MfeI-HF、MflI、MluCI、MluI、MluI-HF、Mly113 I、MreI、MroI、MroNI、MseI、MspCI、MspI、MspJI、MspR9I、MteI、MunI、MvaI、 NarI, NciI, NcoI, NcoI-HF, NdeI, NdeII, Nfo, NgoMIV, NheI, NheI-HF, NmuCI NotI, NotI-HF, NspV, PaeR7I, PagI, PalAI, PaqCI, PasI, PauI, PciI, PciS I、PfeI、Pfl23II、PflFI、PfoI、PinAI、PleI、PpsI、PpuMI、PscI、PshBI、Psp1 406I、Psp5II、Psp6I、PspEI、PspFI、PspGI、PspLI、PspOMI、PspPI、PspPPI、 PspXI、PsuI、PsyI、PteI、RsaNI、Rsr2I、RsrII、SalI、SalI-HF、SapI、SaqAI、One or more of the following are considered: SatI, Sau3AI, Sau96I, ScrFI, SexAI, SfaNI, SfcI, Sfr274I, SfuI, SgeI, SgrAI, SgrDI, SgsI, SinI, SlaI, SmlI, SmoI, SpeI, SpeI-HF, Sse9I, SsiI, SspDI, SspMI, StyD4I, StyI, StyI-HF, TaqI, TaqI-v2, TasI, TatI, TfiI, Tru1I, Tru9I, TseFI, TseI, Tsp45I, TspMI, Tth111I, Vha464I, VneI, VpaK11BI, VspI, XagI, XapI, XbaI, XhoI, XmaI, XmaJI, XmiI, and XspI.
[0062] The exonucleases mentioned include, but are not limited to, one or more of the following: modified or wild-type T5 exonuclease, exonuclease I from Escherichia coli, exonuclease III from Escherichia coli, exonuclease from bacteriophage λ, or RecJ from thermophilic bacteria.
[0063] The polymerase described herein has exonuclease activity and / or the ability to extend one or more bases into a segment of DNA. It may be a wild-type DNA polymerase or an engineered DNA polymerase, including but not limited to one or more of the following: Phi29 DNA polymerase, Tth DNA polymerase, M2 DNA polymerase, VENT DNA polymerase, T5 DNA polymerase, Bst DNA polymerase, Q5 DNA polymerase, DNA polymerase 1, DNA polymerase 2, DNA polymerase 3, Klenow Fragment, T4 DNA polymerase, T7 DNA polymerase, Bsu DNA polymerase, terminal transferase, Phusion DNA polymerase, Taq DNA polymerase, Pfu DNA polymerase, KOD DNA polymerase, or REPLI-gsc DNA polymerase.
[0064] The Cas enzymes mentioned include, but are not limited to, Cas3, Cas4, Cas8a, Cas8b, Cas8c, Cas9, Cas10, Cas10d, Cas12a, Cas13, Csn2, Csf1, Cmr5, Csm2, Csy1, Cse1, or C2c2, or their mutants, such as Cas9nickase, SpCas9-VQR, SpCas9-VRER, SpCas9-VRQR, xCas9-3.7, SpCas9-NG, SpCas9-NRRH, SpCas9-NRTH, SpCas9-NRCH, iSpyCas9, SpG, spRY, eSpCas9(1.1), SpCas9-HF1, SpCas9-HF2, HypaCas9(Sp), Sniper-Cas9(Sp), HiFi One or more of the following: Cas9, SaCas9, SaCas9-KKH, FnCas9, FnCas9-RHA, St1Cas9, St3Cas9, NMe1Cas9, NMe2Cas9, CjCas9, GeoCas9, ScCas9, ScCas9++, AsCas12a, AsCas12a-RR, AsCas12a-RVR, enAsCas12a, enAsCas12a-HF, LbCas12a, impLbCas12a, BhCas12b-v4, Cas12e, CmCas12f, etc.
[0065] The transposases mentioned include, but are not limited to, one or more of the following: Tn5 transposase or its mutant, TnpA transposase or its mutant, tnpR transposase or its mutant, Tn3 transposase or its mutant, Tn9 transposase or its mutant, and Tn10 transposase or its mutant.
[0066] In one specific embodiment of the present invention, the denaturation is achieved through temperature denaturation and fracture. The temperature denaturation is performed at a temperature greater than 50°C, such as 55-100°C, or even 60-100°C. The time for the temperature denaturation is greater than 1 minute, for example, greater than 1, 2, 3, 4, 5, 6, 7, 8, 9, or 10 minutes.
[0067] The samples may be modified or unmodified.
[0068] The samples can be selected from body fluids, feces, cells, tissues or organs, paraffin-embedded samples, natural environments such as soil and ponds, or fossil samples.
[0069] Specifically, the body fluids mentioned are selected from plasma, tissue fluid, lymph, urine, tears, digestive juices, joint fluid, sweat, cerebrospinal fluid, vaginal secretions, etc.
[0070] Specifically, the digestive juices can be selected from bile, saliva, gastric juice, or pancreatic juice, etc.
[0071] Preferably, the sample is isolated from an organism, a paleontology, or the natural environment.
[0072] Specifically, the organism is selected from humans, non-human animals, plants, or microorganisms.
[0073] More specifically, the microorganisms are selected from prokaryotic microorganisms, eukaryotic microorganisms, viruses, or subviruses.
[0074] The prokaryotic microorganisms include, for example, bacteria, actinomycetes, spirochetes, mycoplasmas, rickettsiae, or chlamydiae.
[0075] The eukaryotic microorganisms include fungi, algae, or protozoa.
[0076] In one specific embodiment of the present invention, the sample includes a paraffin-embedded sample.
[0077] In one specific embodiment of the present invention, when performing methylation sequencing, the sample may be a sample converted by bisulfite or by enzymatic conversion.
[0078] In one specific embodiment of the present invention, when performing chromatin accessibility sequencing, the sample may be enzymatically treated (e.g., ATAC-seq, ChIP-seq, DNase-seq, MNase-seq, and other new enzymatically treated methods for chromatin accessibility studies).
[0079] In one specific embodiment of the present invention, the method includes:
[0080] A) Isolate DNA from the sample. Preferably, the isolated DNA can be single-stranded or double-stranded. When the isolated DNA is double-stranded, it is first denatured into single-stranded DNA.
[0081] B) Phosphorylate the single-stranded DNA obtained in step A);
[0082] C) Use a specific sequence for linear transcription as a template attached to the 5' end of a single-stranded DNA, and then perform in vitro transcription.
[0083] The linear transcription functional sequence includes the T7 promoter sense strand. The 3' end of the T7 promoter sense strand is added to single-stranded DNA.
[0084] The positive chain sequence of the T7 promoter includes TAATACGACTCACTATAG (SEQ ID No. 1), TAATACGACTCACTATAGG (SEQ ID No. 11), or TAATACGACTCACTATAGGG (SEQ ID No. 12).
[0085] The linear transcription functional sequence may also include the antisense strand of the T7 promoter.
[0086] The antisense strand sequence of the T7 promoter includes CTATAGTGAGTCGTATTA (SEQ ID No. 5), CCTATAGTGAGTCGTATTA (SEQ ID No. 3), or CCCTATAGTGAGTCGTATTA (SEQ ID No. 4).
[0087] Preferably, the T7 promoter is added to the 3' end of the positive strand of single-stranded DNA.
[0088] More preferably, the single-stranded DNA added to the 3' end of the T7 promoter's positive strand is a directly synthesized sequence, or the 5' end of the single-stranded DNA is linked to the 3' end of the T7 promoter's positive strand. The linking can be direct or indirect (e.g., via a nucleic acid functional fragment).
[0089] Preferably, the nucleic acid functional fragment can be DNA, RNA, or a fragment containing both DNA and RNA.
[0090] Preferably, the nucleic acid functional fragment can be modified or unmodified.
[0091] In one specific embodiment of the present invention, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment.
[0092] The nucleic acid functional fragment is greater than or equal to 1 megapixel, or of any length designed as required. For example, it can be several, tens, hundreds, thousands, tens of thousands, or even megamers. Preferably, any value from 1 to 10,000 megapixels is preferred; more preferably, any value from 5 to 10,000 megapixels is preferred. Specific examples can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, 2000, 5000, 8000, or 10,000 megapixels or more.
[0093] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0094] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0095] In one specific embodiment of the present invention, the 3' end sequence of the nucleic acid functional fragment comprises a single strand of RNA. The single strand of RNA is used for efficient ligation of the T7 promoter positive strand to single-stranded DNA. Preferably, the 3' end of the single strand of RNA is ligated to the 5' end of the single-stranded DNA (preferably direct ligation).
[0096] Preferably, the nucleotide residues in the RNA single strand can be one or more of the following as bases: adenine (A), guanine (G), cytosine (C), uracil (U), methylcytosine (mC), methyladenine (mA), 5-hydroxymethylcytosine (5-hmC), 5-formylcytosine, 5-carboxylcytosine, hypoxanthine, or xanthine. For example, ribonucleotides, nucleoside pseudouridine (Ψ), dihydrouridine (D), inosine (I), or 7-methylguanosine (m7G) can be used.
[0097] Preferably, the length of the RNA single strand is greater than or equal to 1 mer, or any length designed as required. For example, several, tens, hundreds, thousands, tens of thousands, or even megamers, etc. Preferably greater than or equal to 1 mer, more preferably greater than or equal to 3, 4, 5, 6, 7, or 8 mers, and even more preferably any value from 3 to 10,000 mers or more. For example, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 48, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 6 0, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, or 200, 300, 400, 500, 1000, 2000, 5000, 8000, or 10000mer or above.
[0098] The nucleic acid functional fragment also includes a marker sequence and / or an enzyme cleavage site.
[0099] The marker sequence is used to distinguish different samples or different molecules. Examples include primer sequences, barcodes, UMIs, or indexes.
[0100] The marker sequence is greater than or equal to 1 mer, or of any length designed as required. For example, it can be several, tens, hundreds, thousands, tens of thousands, or even megamers. Preferably, any value from 1 to 10,000 mers is preferred; more preferably, any value from 5 to 10,000 mers is preferred. Specific examples can be 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 40, 50, 60, 70, 80, 90, 100, 200, 300, 400, 500, 1000, 2000, 5000, 8000, or 10,000 mers or more.
[0101] In one specific embodiment of the present invention, the 3' end of the marker sequence or restriction site is linked to the 5' end of the single-stranded DNA.
[0102] In one specific embodiment of the present invention, the nucleic acid functional fragment includes a marker sequence and a single strand of RNA.
[0103] In one specific embodiment of the present invention, the nucleic acid functional fragment includes an enzyme cleavage site and an RNA single strand.
[0104] In one specific embodiment of the present invention, the nucleic acid functional fragment includes a marker sequence, an enzyme cleavage site, and a single strand of RNA.
[0105] The nucleic acid functional fragment includes a 5' end sequence and a 3' end sequence. The 5' end sequence of the nucleic acid functional fragment includes a marker sequence and / or an enzyme cleavage site. The 3' end sequence of the nucleic acid functional fragment includes a single strand of RNA. Preferably, the 3' end of the single strand of RNA is linked to the 5' end of the single-stranded DNA.
[0106] Preferably, the 5' end of the marker sequence or restriction site is linked to the 3' end of the positive strand of the T7 promoter.
[0107] Preferably, the linear transcription functional specific sequence is a single-stranded, double-stranded, or partially double-stranded sequence.
[0108] Preferably, the 3' end sequence of the linear transcription functional specific sequence is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0109] In one specific embodiment of the present invention, the 3' end sequence of the linear transcription functional specific sequence is a single strand of RNA, and the 3' end of the single strand of RNA is connected to the 5' end of the single strand of DNA.
[0110] Preferably, the linear transcription functional specific sequence may also include a hairpin structure.
[0111] The specific sequence for linear transcription can be fixed on a vector.
[0112] Preferably, the carrier includes, but is not limited to, one or more of the following: enzyme-labeled microplates, microparticles, microspheres, affinity membranes, chips, glass slides, test strips, or plastic balls.
[0113] Preferably, the fixation is achieved by linking a nucleotide of a specific functional sequence or a nucleotide on its complementary strand or a modified nucleotide to a vector via linear transcription.
[0114] Preferably, the fixation can be achieved through physical adsorption or covalent coupling.
[0115] The physical adsorption mentioned above includes, for example, hydrophobic or electrostatic forces.
[0116] The covalent coupling method mentioned above includes one or more of the following: carbodiimide method, Huisgen reaction, mixed acid anhydride method, maleimide method, glutaraldehyde method, or biotin-streptavidin.
[0117] In this context, the biotin-streptavidin sequence is a linear transcription functionally specific sequence or vector, one labeled with biotin and the other modified with streptavidin. The Huisgen reaction involves a linear transcription functionally specific sequence or vector, one modified with an azide group and the other with a cyclooctyne terminal group or its derivative. The cyclooctyne terminal group or its derivative preferably includes, but is not limited to, any one of DIFO, BCN, DIBAC, DIBO, ADIBO, or DBCO. Another example is a linear transcription functionally specific sequence or vector, one modified with a cysteine residue and the other with a maleimide terminal group.
[0118] Preferably, the linear transcription functional specific sequence can be an unmodified nucleic acid or a modified nucleic acid.
[0119] Specifically, the modification is selected from one or more of the following: 5' end modification, 3' end modification, introduction of non-natural nucleotides, base modification, phosphate backbone modification, sugar ring modification, or introduction of substances that can be linked to nucleic acids.
[0120] Preferably, the base modification is selected from one or more of the following: 5-position pyrimidine modification, 8-position purine modification, or 5-bromouracil substitution.
[0121] Preferably, the sugar ring modification is selected from one or more groups selected from H, OZ, Z, halo, SH, SZ, NH2, NHZ, NZ2 or CN, wherein Z is an alkyl group.
[0122] Preferably, the phosphate skeleton modification includes thiophosphate modification.
[0123] Preferably, the non-natural nucleotide is one or more of the following: non-natural bases, non-natural sugar moieties, or non-natural backbones.
[0124] Preferably, the non-natural base is selected from 2-aminoadenine-9-yl, 5-hydroxymethyluracil, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymidine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-amino-adenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazoguanine, 3-deazoadenine, 4-thiouracil, 4-thiothymidine, uracil-5-yl, hypoxanthine-9-yl (I), 5-methylcytosine, 5-hydroxymethylcytosine Pyridine, xanthine, hypoxanthine, 5-bromouracil, 5-trifluoromethyluracil, 5-bromocytosine, 5-trifluoromethylcytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, cyclocytosine, cytarabine, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, 5-iodouracil 6-alkyl derivatives of pyrimidine, adenine, and guanine; 6-azyuracil; 6-azouracil; 6-azocytosine; azocytosine; 6-azothymidine; 6-thioguanine; 7-methylguanine; 7-methyladenine; 7-deazoguanine; 7-deazoguanosine; 7-deazoguanosine; 7-deazo-8-azyguanine; 8-azyguanine; 8-azyadenine; 8-aminoadenine; 8-aminoguanine; 8-thiol adenine; 8-thiol guanine; 8-thioalkyladenine; 8-thioalkylguanine; 8-hydroxyadenine; 8-hydroxyguanine; N4-ethylcytosine; N-2-substituted purines; N -6-substituted purines, O-6-substituted purines, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazincytidine ([5,4-b][l,4]benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido[5,4-b][l,4]benzothiazin-2(3H)-one), G-clamps, phenoxazincytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylpiperidine, inosine, N6-isopentene adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyl The following are included in the list of one or more of the following: 2-aminouracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosyl uracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentene adenine, uracil-5-oxyacetic acid, uracil-5-oxyacetic acid, pseudouracil, uracil, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, methyl uracil-5-oxyacetic acid, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, or 2,6-diaminopurine.
[0125] Preferably, the non-natural sugar moiety is selected from the following group of modifications at the 2' position: OH; substituted lower alkyl, alkylaryl, aralkyl, O-alkylaryl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl, O-alkenyl, S-alkenyl, N-alkenyl, O-ynyl, S-ynyl, N-ynyl, O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, wherein the alkyl, alkenyl, and ynyl groups can be substituted or unsubstituted C1-C10 alkyl, C2-C10 alkenyl, C2-C10 ynyl, -O[(CH2)] n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -ONH2 and -O(CH2) n ON[(CH2) n CH3)]2, wherein n and m are from 1 to 10; and / or one or more of the following group of modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), and at the 4' position: 4'-S, heterocyclic alkyl, heterocyclic aryl, aminoalkylamino, polyalkylamino, or substituted silyl.
[0126] Preferably, the substance that can be linked to nucleic acids can be one or more of amino acids, polypeptides, carbohydrates, proteins, or lipids.
[0127] Preferably, the amino acid can be a natural amino acid or a non-natural amino acid. More preferably, the non-natural amino acids include one or more of the following: 2-aminoisobutyric acid (Aib), imidazole-4-acetate (IA), imidazole propionic acid (IPA), α-aminobutyric acid (Abu), tert-butylglycine (Tle), 3-aminomethylbenzoic acid, anthranilic acid, deaminohistidine, β-alanine, 2-aminohistidine, β-hydroxyhistidine, homohistidine, Nα-acetylhistidine, α-fluoro-methylhistidine, α-methylhistidine, α,α-dimethylglutamic acid, m-CF3-phenylalanine, α,β-diaminopropionic acid, 3-pyridylalanine, 2-pyridylalanine, 4-pyridylalanine, (1-aminocyclopropyl)carboxylic acid, (1-aminocyclobutyl)carboxylic acid, (1-aminocyclopentyl)carboxylic acid, (1-aminocyclohexyl)carboxylic acid, (1-aminocycloheptyl)carboxylic acid, or (1-aminocyclooctyl)carboxylic acid.
[0128] In one specific embodiment of the present invention, the linear transcription functional specific sequence may be SEQ ID No. 2 or SEQ ID No. 7.
[0129] TAATACGACTCACTATAGG / rG / / rG / / rA / / rC / (SEQ ID No.7)
[0130] The in vitro transcription method described herein is applicable to single-stranded DNA or double-stranded nucleic acids containing a single-stranded DNA portion.
[0131] In a second aspect, the present invention provides RNA obtained by the above method.
[0132] A third aspect of the present invention provides an application of the above-described in vitro transcription method or the RNA obtained by the above-described method, the application including:
[0133] i) Used as an RNA vaccine;
[0134] ii) Used as an RNA drug;
[0135] iii) Used to construct RNA databases;
[0136] iv) Used for guide RNA or other RNAs derived from guide RNA, such as pegRNA, etc.
[0137] v) Used for linear amplification of target DNA; or
[0138] vi) is used to construct sequencing libraries.
[0139] In a fourth aspect, the present invention provides a method for preparing RNA, the method comprising using a single-stranded DNA containing a linear transcription functionally specific sequence at its 5' end as a template to perform in vitro transcription to obtain RNA; wherein the linear transcription functionally specific sequence includes a T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA.
[0140] Preferably, the single-stranded DNA added to the 3' end of the T7 promoter's positive strand is a directly synthesized sequence, or the 5' end of the single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand.
[0141] More preferably, the linear transcription functional-specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment.
[0142] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0143] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0144] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0145] Preferably, the preparation method includes: after determining the RNA sequence, transcribing DNA capable of being transcribed into the RNA.
[0146] Preferably, the RNA can be linear or circular.
[0147] Preferably, the RNA can be interfering RNA, pegRNA, mRNA, etc.
[0148] A fifth aspect of the present invention provides a library construction method, the method comprising constructing a library using the product obtained from the above-described in vitro transcription.
[0149] Preferably, the product obtained from the in vitro transcription is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted enriched and / or amplified to obtain a library.
[0150] Preferably, the product obtained from the in vitro transcription is separated to obtain template DNA and transcription product RNA; wherein,
[0151] RNA is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted enrichment and / or amplification are performed to obtain a library.
[0152] Template DNA is used for in vitro transcription. Preferably, the template DNA can be used directly as a template for in vitro transcription without any treatment, or it can be specially treated before in vitro transcription.
[0153] Preferably, the special treatment includes DNase digestion, or separation, purification, and denaturation into single strands for direct in vitro transcription as a template, or other treatments, such as transformation with one or more of the following modifications: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), N6-methyladenine (N6-mA), 7-methylguanine (7-mG), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), dihydrouracil (DHU), 5-(β-glucoseoxymethylated)cytosine (5gmC), etc.
[0154] Preferably, the sequencing includes one or more of the following sequencing methods: whole genome sequencing, metagenomic sequencing, de novo sequencing, whole exome sequencing, 16S sequencing, methylation sequencing, hydroxymethylation sequencing, single-cell sequencing, chromosome accessibility sequencing, or targeted sequencing.
[0155] Specifically, the methylation sequencing can be whole-genome DNA methylation sequencing (WGBS), simplified genome methylation sequencing (RRBS / dRRBS / XRBS), high-throughput single-cell methylation sequencing (sc-RBS), target gene DNA methylation sequencing (Target-BS / LHC-BS / Capture-BS), precise DNA methylation and hydroxymethylation sequencing (oxBS-seq), single-cell and micro-sample DNA methylation sequencing (Micro DNA-BS), micro-cfDNA genome methylation sequencing (cfDNA-BS), (hydroxy)methylated DNA immunoprecipitation sequencing ((h)MeDIP-seq / 5hmC-Seal), and other sequencing methods.
[0156] In a sixth aspect, the present invention provides a library obtained by the above-described library construction method.
[0157] A seventh aspect of the present invention provides a sequencing method comprising sequencing a library obtained using the library preparation method described above.
[0158] An eighth aspect of the present invention provides a combined sequencing method for performing two or more sequencing operations on the same sample, the combined sequencing method comprising:
[0159] A) Obtain DNA, preferably by isolating DNA from a sample;
[0160] B) Transformed into a single strand;
[0161] C) Using the 5' end of the single-stranded DNA obtained in step B) containing a linear transcription function-specific sequence as a template, in vitro transcription is performed; preferably, the linear transcription function-specific sequence includes the T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA.
[0162] D) Separate the products obtained from in vitro transcription to obtain template DNA and transcription product RNA;
[0163] E1) The RNA obtained in D) is further reverse transcribed into cDNA, ligated with the adapters required for sequencing, and then sequenced directly, or after ligating the adapters required for sequencing, it is amplified and then sequenced, or it is targeted enriched and then library constructed and sequenced, or library constructed and then targeted enriched and sequenced.
[0164] E2) The template DNA obtained in D) is processed in steps B)-C), and then the product obtained from in vitro transcription is further reverse transcribed into cDNA, ligated with the adapter required for sequencing, and sequenced directly, or after ligating the adapter required for sequencing, it is amplified and sequenced, or targeted enriched and then library constructed and sequenced, or library constructed and then targeted enriched and sequenced.
[0165] E3) After performing steps B), C), and D) on the template DNA obtained in D), proceed to steps E1) and / or E2).
[0166] The flowchart for simultaneously constructing and sequencing WGS and WGBS libraries can be seen in the figure. Figure 33 The flowchart for simultaneously constructing and sequencing dual and multiple libraries can be seen. Figure 34 .
[0167] Preferably, the combined sequencing method further includes phosphorylation of the 5' end of the DNA.
[0168] More preferably, the phosphorylation can be performed before step A), or before denaturation in step B), or after denaturation in step B).
[0169] Preferably, the template DNA in step E2) may not require processing, or it may undergo special processing before proceeding to steps B)-C).
[0170] Preferably, the special treatment includes DNase digestion, or separation, purification, and denaturation into single strands for direct in vitro transcription as a template, or other treatments, such as transformation with one or more of the following modifications: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), N6-methyladenine (N6-mA), 7-methylguanine (7-mG), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), dihydrouracil (DHU), 5-(β-glucoseoxymethylated)cytosine (5gmC), etc.
[0171] Preferably, the single-stranded DNA added to the 3' end of the T7 promoter's positive strand is a directly synthesized sequence, or the 5' end of the single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand.
[0172] In one specific embodiment of the present invention, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment.
[0173] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0174] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0175] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0176] Preferably, the separation in step D) can be screening for antigens and antibodies, screening for biotin-streptomycin magnetic beads, or a method used in the prior art to separate DNA from RNA.
[0177] Preferably, the sequencing in steps E1) and E2) can be the same sequencing method or different sequencing methods.
[0178] Preferably, the sequencing is selected from one or more sequencing methods such as whole genome sequencing, de novo sequencing, whole exome sequencing, metagenomic sequencing, 16S sequencing, methylation sequencing, hydroxymethylation sequencing, single-cell sequencing, chromosome accessibility sequencing, or targeted sequencing.
[0179] A ninth aspect of the present invention provides a linear transcription function-specific sequence for attachment to the 5' end of a template single-stranded DNA for in vitro transcription, the linear transcription function-specific sequence comprising a T7 promoter positive strand and a nucleic acid functional fragment.
[0180] The positive chain sequence of the T7 promoter includes SEQ ID No. 1, SEQ ID No. 11, or SEQ ID No. 12.
[0181] The linear transcription functional sequence may also include the antisense strand of the T7 promoter.
[0182] The antisense strand sequence of the T7 promoter includes SEQ ID No. 5, SEQ ID No. 3, or SEQ ID No. 4.
[0183] The 5' end of the single-stranded DNA is directly or indirectly connected to the 3' end of the positive strand of the T7 promoter.
[0184] Preferably, the 5' end of the nucleic acid functional fragment is connected to the 3' end of the T7 promoter positive strand.
[0185] Preferably, the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0186] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0187] In one specific embodiment of the present invention, the 3' end sequence of the nucleic acid functional fragment comprises a single strand of RNA. This single strand of RNA is used for the efficient ligation of the T7 promoter's positive strand to single-stranded DNA.
[0188] Preferably, the length of the RNA single strand is greater than or equal to 1 mer, more preferably greater than or equal to 3, 4, 5, 6, 7 or 8 mer, and even more preferably 3-10000 mer or more.
[0189] The nucleic acid functional fragment also includes a marker sequence and / or restriction enzyme sites. The marker sequence is used to distinguish different samples or different molecules. Examples include primer sequences, barcodes, UMIs, or indexes.
[0190] In one specific embodiment of the present invention, the 3' end of the marker sequence or restriction site is linked to the 5' end of the single-stranded DNA.
[0191] In one specific embodiment of the present invention, the nucleic acid functional fragment includes a marker sequence (and / or restriction enzyme site) and a single strand of RNA.
[0192] The nucleic acid functional fragment includes a 5' end sequence and a 3' end sequence. The 5' end sequence of the nucleic acid functional fragment includes a marker sequence (and / or an enzyme cleavage site). The 3' end sequence of the nucleic acid functional fragment includes a single strand of RNA.
[0193] Preferably, the 3' end of the RNA single strand is connected to the 5' end of the single strand DNA.
[0194] Preferably, the 5' end of the marker sequence and / or restriction site is linked to the 3' end of the positive strand of the T7 promoter.
[0195] In one specific embodiment of the present invention, the linear transcription functional specific sequence may be SEQ ID No. 2 or SEQ ID No. 7.
[0196] In a tenth aspect of the present invention, there is an application of a single-stranded DNA containing a linear transcription function-specific sequence at its 5' end as a template in in vitro transcription, wherein the linear transcription function-specific sequence includes a T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA.
[0197] Preferably, the single-stranded DNA added to the 3' end of the T7 promoter's positive strand is a directly synthesized sequence, or the 5' end of the single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand.
[0198] In one specific embodiment of the present invention, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment;
[0199] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0200] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0201] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0202] In an eleventh aspect, the present invention provides a reaction system for linking specific sequences of linear transcription function, the reaction system comprising:
[0203] The single-stranded DNA serves as a template, a linear transcription functionally specific sequence, and a linker component; wherein the linear transcription functionally specific sequence includes the T7 promoter sense strand, and the linker component links the 3' end of the T7 promoter sense strand to the 5' end of the single-stranded DNA.
[0204] Preferably, the linker connects the 3' end of the positive strand of the T7 promoter to the 5' end of the single-stranded DNA directly or indirectly (e.g., through a nucleic acid functional fragment).
[0205] The linear transcription functional specific sequences include the T7 promoter positive strand and nucleic acid functional fragments;
[0206] The linker connects the 5' end of the nucleic acid functional fragment to the 3' end of the T7 promoter positive strand, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0207] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0208] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0209] Preferably, the ligation component includes ligases or reagents required for chemical ligation methods.
[0210] In one specific embodiment of the present invention, the ligation component includes a ligase.
[0211] Preferably, the ligase can be a ligase capable of connecting nicks and / or double-stranded nucleic acids, such as T4 RNA ligase 2, T4 DNA ligase, T3 DNA ligase, SplintR ligase, Taq DNA ligase, 9°N DNA ligase, E. coli ligase, and mutants of the above ligases; or it can be a ligase capable of connecting single-stranded nucleic acids, such as T4 RNA ligase 1 (ssRNA ligase), Mth RNA ligase, 5'App DNA / RNA thermostable ligase, TS2126 RNA ligase (TS2126 Rnl 1 or CircLigase), Hyper-Thermostable Lysine-Mutatants DNA / RNA ligase (HyperLigase), T4 RNA ligase 2, RtcB ligase, and mutants of the above ligases.
[0212] In one specific embodiment of the present invention, the reaction system for linking specific sequences of linear transcription function contains one or more of the following: PEG, divalent cations, ATP, dimethyl sulfoxide (DMSO), buffer (e.g., Tris-HCl buffer), double-distilled water, RNase inhibitors, or bovine serum albumin (BSA).
[0213] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences and RNA ligase.
[0214] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences, RNA ligase, and PEG.
[0215] In one specific embodiment of the present invention, the reaction system for linking linear transcription functionally specific sequences includes linear transcription functionally specific sequences, RNA ligase, PEG, and BSA.
[0216] In one specific embodiment of the present invention, the reaction system for linking the linear transcription functionally specific sequence includes: the linear transcription functionally specific sequence, RNA ligase, PEG, ATP, and buffer (e.g., Tris-HCl buffer).
[0217] In one specific embodiment of the present invention, the reaction system for linking the linear transcription functionally specific sequence includes: the linear transcription functionally specific sequence, RNA ligase, PEG, ATP, RNase inhibitor, buffer (e.g., Tris-HCl buffer), and DMSO.
[0218] A twelfth aspect of the present invention provides an in vitro transcription reaction system, said reaction system comprising:
[0219] A linear transcription functionally specific sequence is added to the 5' end of a single-stranded DNA as a template, along with an NTP and a T7 RNA polymerase; wherein the linear transcription functionally specific sequence includes the positive strand of the T7 promoter, and the 3' end of the positive strand of the T7 promoter is added to the single-stranded DNA.
[0220] Preferably, the single-stranded DNA added to the 3' end of the T7 promoter's positive strand is a directly synthesized sequence, or the 5' end of the single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand.
[0221] In one specific embodiment of the present invention, the linear transcription functional specific sequence includes a T7 promoter positive strand and a nucleic acid functional fragment;
[0222] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0223] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0224] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0225] Preferably, the NTP includes, but is not limited to, one or more of ATP, UTP, GTP, CTP, modified nucleotides or non-natural nucleotides.
[0226] In one specific embodiment of the present invention, the in vitro transcription reaction system includes a single-stranded DNA with a 5' end linked to a specific sequence for linear transcription as a template, T7 RNA polymerase, RNase inhibitor, and NTP.
[0227] In a thirteenth aspect of the present invention, a method for in vitro transcription of low-frequency mutant samples is provided, the method comprising using a single-stranded DNA containing a linear transcription function-specific sequence at the 5' end as a template, and then performing in vitro transcription;
[0228] The linear transcription functional specific sequence includes the T7 promoter positive strand and nucleic acid functional fragments;
[0229] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0230] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0231] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0232] In a fourteenth aspect of the present invention, an in vitro transcription method for low-frequency mutant samples is provided, the in vitro transcription method comprising using a single-stranded DNA containing a linear transcription function-specific sequence at the 5' end as a template, and then performing in vitro transcription.
[0233] The linear transcription functional specific sequence includes the T7 promoter sense strand, the 3' end of which is linked to the 5' end of the single-stranded DNA.
[0234] Preferably, the 3' end of the positive strand of the T7 promoter is directly or indirectly (e.g., through a nucleic acid functional fragment) connected to the 5' end of the single-stranded DNA.
[0235] In one specific embodiment of the present invention, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment;
[0236] The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA.
[0237] Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA.
[0238] More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
[0239] In a fifteenth aspect, the present invention provides a method for detecting low-frequency mutations, the method comprising employing the above-described in vitro transcription method or library preparation method, followed by sequencing.
[0240] A sixteenth aspect of the present invention provides a method for unbiased amplification of single-stranded DNA, the method comprising using a single-stranded DNA containing a linear transcription functionally specific sequence at its 5' end as a template.
[0241] In a seventeenth aspect of the present invention, a kit is provided comprising the above-described reaction system, such as a reaction system for linking linear transcription functionally specific sequences, an in vitro transcription reaction system, or a phosphorylation reaction system.
[0242] The kit is an in vitro transcription kit and a library preparation kit.
[0243] The unit "mer" in this invention stands for monomeric unit. It can represent the length of a nucleotide in a single strand or the number of base pairs in a double strand. In some cases, it can be used interchangeably or in combination with nt (nucleotide) or bp (base pair). That is, when describing a single strand or a single-stranded portion, mer represents nt; when describing a double strand or a double-stranded portion, mer represents bp.
[0244] The sequencing method of this invention can be either second-generation sequencing or third-generation sequencing. "Second-generation sequencing" includes, but is not limited to, massively parallel sequencing (MPS), nanoball sequencing, and other second-generation sequencing technologies based on cyclic microarrays. Platforms include, but are not limited to, Roche / 454GS FLX, Illumina / Sol-exa Genome Analyzer, and HeliScope from Helicos BioSciences. TM Single Molecule Sequencer, Polonator from Danaher Motion, and sequencing by ligation (which uses primers to locate nucleic acid information) are all available technologies, with platforms including Applied Biosystems / SOLiD. TM The "third-generation sequencing" mentioned refers to single-molecule sequencing technologies, such as nanopore sequencing, single-molecule real-time sequencing, etc.
[0245] The beneficial effects of this application are:
[0246] 1. The 5' end of a single-stranded DNA cell linked to the positive strand of a T7 promoter allows for in vitro transcription, which differs from the traditional principle. A schematic diagram comparing the traditional T7 transcription principle with the T7 RNAP transcription principle based on single-stranded DNA involved in this application is shown below. Figure 1 As shown. Specifically, in the traditional principle, transcription requires (1) the T7 promoter to be a complete double strand, and the template DNA to also be double-stranded DNA. Figure 1 (1) or (2) the T7 promoter is a complete double-stranded DNA and the 5' end of the antisense strand of the T7 promoter is connected to the 3' end of the template single-stranded DNA. Figure 1 Mode 2). This invention reports a novel transcription mode in which the 3' end of the T7 promoter on the positive strand is located at the 5' end of a single-stranded DNA, allowing this sequence to serve as a template for transcription in the presence of NTPs and T7 RNA polymerase, generating an RNA sequence complementary to the template. Figure 1 (The bottommost diagram).
[0247] 2. Traditional methods of ligating T7 adapters to the 3' end of single-stranded DNA are relatively inefficient. In this invention, the RNA modification contained in the 3' end of the linear transcription functional specific sequence achieves a ligation efficiency of approximately 90% or more, and even up to 96% at the 5' end of single-stranded DNA.
[0248] 3. The operation is simple and there are no additional DNA loss steps. The entire process, from DNA sample input to completion of in vitro transcription, is carried out in the same tube, eliminating the need for additional purification steps and significantly reducing DNA sample loss.
[0249] 4. Wide range of applications: The in vitro transcription method is applicable to various samples of different lengths, including damaged DNA, trace DNA samples, such as cell-free blood DNA, paraffin-embedded tissue sample (FFPE) DNA, paleontological DNA, and DNA after bisulfite conversion.
[0250] 5. This application can detect low-frequency mutant DNA samples, such as DNA samples with a mutation frequency of 0.1%.
[0251] 6. The same template DNA sample can be repeatedly isolated and used, which is expected to enable duplex, triplet, and higher sequencing. Attached Figure Description
[0252] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings, wherein:
[0253] Figure 1 Schematic diagram of T7 RNAP transcription mechanism. Traditional T7 RNAP transcription mechanism: Mode 1: When the template is double-stranded DNA, the ends of the double-stranded DNA contain the T7 promoter sequence. T7 RNAP binds to the T7 promoter region and uses the DNA strand containing the antisense strand of the T7 promoter as a template to synthesize RNA. The sequence of the resulting RNA is consistent with the sequence of the DNA strand containing the sense strand of the T7 promoter in the template strand. Mode 2: When the template is single-stranded DNA, the 3' end of the template DNA needs to contain the antisense strand sequence of the T7 promoter, and the sense and antisense strands of the T7 promoter form a local double-stranded DNA sequence at the 3' end of the template strand. Then, under the action of T7 RNAP, RNA is synthesized using the DNA as a template, and the resulting RNA sequence is the complementary sequence of the template DNA. The T7 RNAP transcription mechanism of this invention: When the 3' end of the sense strand of the T7 promoter is at the 5' end of a single-stranded DNA, this sequence can serve as a template for transcription in the presence of NTPs and T7 RNA polymerase, generating an RNA sequence complementary to the template.
[0254] Figure 2 Linear amplification and application process of single-stranded DNA. Linear amplification of single-stranded DNA ligated to the positive strand of the T7 promoter yields the transcription product RNA, which is then further utilized.
[0255] Figure 3 A schematic diagram comparing the 5' end of a single-stranded DNA strand connected to the T7 promoter positive strand and the double strand connected to the T7 promoter negative strand.
[0256] Figure 4The image shows the separation electrophoresis results on an agarose gel. Lane 1 is the marker, lanes 2, 3, and 4 are the dsT7 group, lanes 5, 6, and 7 are the ssT7 group, lanes 8, 9, and 10 are the ligated unpurified group, lanes 2, 5, and 8 are the transcription products, lanes 3, 6, and 9 are the results after DNase 1 digestion of the transcription products, and lanes 4, 7, and 10 are the results after DNase 1 and RNase A co-digestion of the transcription products.
[0257] Figure 5 : Schematic diagram of ss_T7D. The 5' end of ss_T7D contains the T7 promoter justice chain.
[0258] Figure 6 : Number of reads corresponding to the inserted fragment length in the RNA library. The horizontal axis represents the length of the inserted fragment in the RNA library, in bases. The vertical axis represents the number of reads corresponding to the inserted fragment length as counted from sequencing results.
[0259] Figure 7 Distribution of lambda genomic DNA fragment lengths interrupted by ultrasound.
[0260] Figure 8 : Percentage of each base content in the aligned reads. The horizontal axis represents the reads in bp; the vertical axis represents the percentage of each base content.
[0261] Figure 9 Library insert length statistics. The horizontal axis represents the length of the inserted fragment, in bp; the vertical axis represents the number of corresponding sequencing reads.
[0262] Figure 10 Sequencing reads aligned to the reference genome. The outer circle shows the coordinates of the reference sequence; the middle circle shows the sequencing depth; and the inner gradient layer shows the sequencing coverage.
[0263] Figure 11 : Length of unbroken genomic DNA (0s group).
[0264] Figure 12 The length of genomic DNA was broken down by ultrasound for 30 seconds (30s group).
[0265] Figure 13 The length of genomic DNA was broken down by ultrasound for 45 seconds (45s group).
[0266] Figure 14 The length of genomic DNA was broken down by ultrasound for 60 seconds (60s group).
[0267] Figure 15Statistical analysis of insert fragment lengths in sequencing libraries constructed from template DNA transcripts fragmented at different ultrasound time points. The horizontal axis represents the insert fragment length in bp; the vertical axis represents the proportion of the corresponding sequencing reads.
[0268] Figure 16 Statistical analysis of the repetition rate of sequencing results for libraries constructed from template DNA transcripts fragmented at different ultrasound times. The horizontal axis represents the grouping of template DNA fragments fragmented at different ultrasound times; the vertical axis represents the repetition rate.
[0269] Figure 17 Unique alignment rate statistics of sequencing reads from libraries constructed from template DNA transcripts fragmented at different ultrasound times. The horizontal axis represents the grouping of template DNA fragments fragmented at different ultrasound times; the vertical axis represents the percentage of reads uniquely aligned to the reference genome.
[0270] Figure 18 : Statistics on the number of reads obtained from sequencing libraries constructed from template DNA transcripts fragmented at different sonication times. The horizontal axis represents the grouping of template DNA fragments at different sonication times; the vertical axis represents the number of reads obtained from sequencing.
[0271] Figure 19 The coverage depth of sequencing results of libraries constructed from template DNA transcripts fragmented by sonication at different times on the reference genome. The horizontal axis represents the coordinates of the reference genome; the vertical axis represents the coverage depth of the sequencing reads; the lines of different colors represent the sequencing results of libraries constructed from template DNA transcripts fragmented by sonication at 0s, 30s, 45s, and 60s.
[0272] Figure 20 : HeLa cell genomic DNA disrupted by ultrasound.
[0273] Figure 21 : Number of reads obtained from sequencing libraries in the experimental and control groups under different input levels. The horizontal axis represents the experimental and control groups with different input levels; the vertical axis represents the number of reads obtained from sequencing.
[0274] Figure 22 The unique alignment rate of sequencing data from experimental and control groups under different input levels is statistically analyzed. The horizontal axis represents the experimental and control groups with different input levels; the vertical axis represents the percentage of uniquely aligned reads out of the total reads.
[0275] Figure 23 GC bias results were analyzed using lowercase library sequencing data from a control group with an input of 100 ng.
[0276] Figure 24 GC bias results were analyzed using library sequencing data from a control group with an input of 10 ng.
[0277] Figure 25 Analysis of GC bias results using 5pg input in the experimental group and lowercase library sequencing data.
[0278] Figure 26 Analysis of GC bias results using lowercase sequencing data from a 10pg input in the experimental group.
[0279] Figure 27 Analysis of GC bias results using 100pg input of the experimental group and lower library sequencing data.
[0280] Figure 28 Analysis of GC bias results using 1ng input in the experimental group and lowercase library sequencing data.
[0281] Figure 29 Analysis of GC bias results using library sequencing data from the experimental group with an input of 10 ng.
[0282] Figure 30 : GC content statistics of library sequencing data in experimental and control groups under different input levels. The horizontal axis represents the experimental and control groups with different input levels; the vertical axis represents the percentage of GC base content.
[0283] Figure 31 : GC content statistics of sequencing data from experimental and control groups under different input amounts of FFPE sample DNA. The horizontal axis represents the groups with different input amounts; the vertical axis represents the percentage of GC base content.
[0284] Figure 32 GIV statistics of sequencing data from libraries in the experimental and control groups under different input amounts of FFPE sample DNA. The vertical axis represents GIV, and the horizontal axis represents the grouping of 12 nucleic acid mutations.
[0285] Figure 33 The same template DNA sample is repeatedly isolated and used to simultaneously achieve WGS and WGBS library construction and sequencing.
[0286] Figure 34 The same template DNA sample can be repeatedly isolated and used to simultaneously achieve the construction and sequencing of dual and multiple libraries. Detailed Implementation
[0287] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0288] Example 1: T7 RNA polymerase can perform in vitro transcription using single-stranded DNA with the 5' end of the T7 promoter-positive strand as a template. 1. Sequence Information
[0289] 1) Linear transcription functional specific sequence (DR_T7 sequence for short)
[0290] TAATACGACTCACTATAGGATACNNNNNNNNNNCT / rG / / rG / / rA / / rC / (SEQ ID No.2)
[0291] Among them, TAATACGACTCACTATAGG (SEQ ID No. 1) is the positive strand sequence of the T7 promoter; ATACNNNNNNNNNNCT is the marker sequence; NNNNNNNNNN is the UMI; and / rG / / rG / / rA / / rC / is the single strand of RNA at the 3' end of the linear transcription functional specific sequence.
[0292] 2) T7 promoter antisense sequence: CCTATAGTGAGTCGTATTA (SEQ ID No. 3)
[0293] 3) circ-129nt single-stranded DNA
[0294] GTTTGCATGGTTTGTTGAAAACCGGACATGGCACTCCAGTCGCCTTCCCGTTCCGCTATCGGCTGAATTTGATTGCGAGTGAG ATATTTATGCCAGCCAGCCAGACGCAGACGCGCCGAGACAGAACTT(SEQ ID No. 6)
[0295] 2. Experimental Grouping
[0296] 1) ssT7 group: DR_T7 was purified after ligation with circ-129nt and then transcribed in vitro.
[0297] 2) dsT7 group: DR_T7 was ligated into circ-129nt and purified, then the antisense strand sequence of the T7 promoter was added, annealed into double strand, and transcribed in vitro.
[0298] 3) Unpurified ligation group: DR_T7 was ligated to circ-129nt without purification and directly transcribed in vitro.
[0299] 4) Group D: Add 2ul DNase 1 to digest DNA for 1 hour.
[0300] 5) D+R group: Add 2ul DNase 1, 2ul RNase A, and 2ul RNase H to digest DNA, and react at 37°C for 1 hour.
[0301] 3. Experimental Procedure
[0302] 1) Steps for SST7 group:
[0303] Step 1: Treat circ-129nt single-stranded DNA with T4 PNK to phosphorylate its 5' end.
[0304] The reaction system is shown in Table 1.
[0305] Table 1
[0306] Components volume T4 PNK buffer (M0201V, NEB) 0.36ul ATP (P0756S, NEB) 0.36ul T4 PNK(M0201V, NEB) 0.1ul circ-129nt (10uM, Sangon Biotech) 2ul DEPC Water (W915679-500ml, McLean) 0.78ul
[0307] The reaction conditions were: 37℃ for 30 min, 75℃ for 20 min.
[0308] Step 2: DR_T7 sequence concatenation
[0309] The reaction system is shown in Table 2.
[0310] Table 2
[0311]
[0312]
[0313] The reaction conditions were: 37℃ for 30 min, 75℃ for 10 min.
[0314] Step 3: Purify the ligation product
[0315] The 129nt DNA ligation product was purified using 3X DNA purification magnetic beads (N411-02, Novizan) (for specific operating procedures, refer to the DNA purification magnetic beads (N411-02, Novizan) instruction manual), and eluted with 17ul DEPC water (W915679-500ml, Maclean).
[0316] Step 4: T7 RNA polymerase transcribes the ligation product from the previous step.
[0317] The reaction system is shown in Table 3.
[0318] Table 3
[0319] Components volume Previous purification product 17ul NTP mix (D7387-1ml, Beyotime) 10ul T7 RNA polymerase (M0251S, NEB) 3ul
[0320] The reaction conditions were: 37°C for 20 hours.
[0321] 2) dsT7 group steps:
[0322] Steps one through three are the same as the “ssT7 group steps”. In step four, 2 μL of 10 μM T7 promoter antisense strand is annealed with the ligation product obtained in step three to form a double strand. Then, in vitro transcription is performed using the reaction system described in Table 4. The in vitro transcription reaction conditions are: 37°C for 20 hours.
[0323] Table 4
[0324] Components volume The double chain obtained by annealing 17ul NTP mix (D7387-1ml, Beyotime) 10ul T7 RNA polymerase (M0251S, NEB) 3ul
[0325] 3) Steps for ligating the unpurified group:
[0326] The first and second steps are the same as the "ssT7 group steps". After the second step is completed, in vitro transcription is performed directly using the reaction system shown in Table 5. The in vitro transcription reaction conditions are: 37℃ for 20h.
[0327] Table 5
[0328] Components volume The products of the second step reaction system 20ul DEPC Water (W915679-500ml, McLean) 31ul NTP mix (D7387-1ml, Beyotime) 30ul T7 RNA polymerase (M0251S, NEB) 9ul
[0329] 4) Steps for Group D:
[0330] Group D, which includes the products of "ssT7 group", "dsT7 group" and "ligation unpurified group", had 2 μL of DNase 1 (M0303S, NEB) added to digest the DNA for 1 hour.
[0331] 5) D+R group steps:
[0332] The D+R group includes the products of the "ssT7 group", "dsT7 group", and "ligation-unpurified group", to which 2 μL of DNase 1 (M0303S) is added.
[0333] DNA was digested for 1 hour with 2 μL of RNase A (R8021-25 mg, 100 mg / ml, Solarbio) and RNase H (M0297S, NEB).
[0334] 4. Experimental Results
[0335] Figure 3 This is a schematic diagram showing the contrasting sequences of the 5' end of a single-stranded DNA strand with the T7 promoter attached to it and the double-stranded strand with the T7 promoter attached. Separation results using 1.2% agarose gel are shown below. Figure 4 As shown. By comparing lanes 2, 3, and 4, and lanes 5, 6, and 7, it is shown that both the dsT7 group and the ssT7 group can successfully transcribe RNA. This indicates that the DR_T7 sequence containing the T7 promoter, when linked to the 5' end of single-stranded DNA, can be transcribed under in vitro conditions, proving that the novel transcription method provided in this application, namely, transcription of single-stranded DNA based on the positive strand of the T7 promoter, can be achieved.
[0336] Furthermore, the ligation and transcription steps were tested to see if they could be performed in the same tube, thus avoiding intermediate purification. Therefore, after ligating the 3' end of the T7 promoter's positive strand to the 5' end of the single-stranded DNA, the ligation system was not purified; the transformation substrate and T7 RNA polymerase were directly added to the original ligation system for transcription (named the "ligation-unpurified" group). Comparing lanes 5, 6, and 7, and lanes 8, 9, and 10, it was observed that the products in lanes 8, 9, and 10 were not significantly different from those in lanes 5, 6, and 7. This indicates that the system does not inhibit subsequent transcription under unpurified ligation conditions. This also demonstrates that the ligation step of the DR_T7 sequence to single-stranded DNA and the transcription step of T7 RNA polymerase can be performed in the same tube without additional purification. This will simplify the further development of linear amplification methods.
[0337] Furthermore, the efficiency of DR_T7 sequence ligation to single-stranded DNA was increased by introducing RNA modification (i.e., a single-stranded RNA at the 3' end of a linear transcription functional sequence). The ligation efficiency of the linear transcription functional sequence after introducing RNA modification reached over 90%, and even reached a maximum of 96%. The success of the experiment also proved that T7 RNA polymerase can transcribe templates containing RNA modification.
[0338] Example 2: Validation based on synthesized single-stranded DNA with a T7 promoter at the 5' end of the positive strand.
[0339] 1. Sequence Information
[0340] ss_T7D (117 bases in total, structural diagram as shown) Figure 5 (as shown)
[0341] TAATACGACTCACTATAGGATACAGCCAGGAACGTACTGGTGAAAACACCGCAGCATGTCAAGATCACAGATTTTGGGCTGGCCAAAC TGCTGGGTGCGGAAGAGAAAGAATACCAT (SEQ ID No. 8)
[0342] 2. In vitro transcription is shown in Table 6.
[0343] Table 6
[0344] Components volume ss_T7D (10uM, Bioengineering) 1ul DEPC Water (W915679-500ml, McLean) 16ul NTP mix (D7387-1ml, Beyotime) 10ul RNase inhibitor (R301-03, 40 U / μl, Novizan) 1ul T7 RNA polymerase (M0251S, NEB) 2ul
[0345] React at 37°C for 18 hours. Then add 2 μL of DNase 1 (M0303S, NEB) and react at 37°C for 2 hours to remove the DNA template from the transcription product.
[0346] 3. Purify the transcript.
[0347] The transcript RNA was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) (for specific operating procedures, refer to the instruction manual for RNA purification magnetic beads (N412-03, Novizan), and eluted with 30ul of DEPC water (W915679-500ml, Maclean).
[0348] 4. RNA library construction
[0349] The purified product was used to construct a library using an RNA library construction kit (NR811, Novizan). The procedure was performed according to the kit instructions (NR811, Novizan), as follows:
[0350] 4.1 Connect the 3' end connector
[0351] 4.1.1 Take out the RL3 Adaptor, RL3 Buffer V2, and RL3 Enzyme mix V2, thaw and mix them, briefly centrifuge to collect the contents to the bottom of the tube, and place on ice for later use. All subsequent steps shall be performed on ice. Prepare the reaction system in an RNase-free PCR tube according to the partitioning group in Table 7.
[0352] Table 7
[0353] Components volume RNA 6ul RL3 Adaptor (NR811, Novizan) 1ul
[0354] 4.1.2 Place the PCR tube in a PCR instrument preheated to 70°C for 2 minutes. After the reaction, immediately remove the tube and place it on ice for 2 minutes. The reaction tube contains the components shown in Table 8.
[0355] Table 8
[0356] Components volume Step 4.1.1 Product 7ul RL3 Buffer V2 (NR811, Novizan) 10ul RL3 Enzyme mix V2 (NR811, Novizan) 3ul
[0357] 4.1.3 Mix thoroughly by pipetting 10-15 times, then briefly centrifuge to collect the reaction solution to the bottom of the tube. Place the reaction tube in a PCR instrument with a heated lid and run the program shown in Table 9.
[0358] Table 9
[0359] temperature time 25℃ 1 hour 4℃ maintain
[0360] 4.2 Excess 3' joint sealing
[0361] 4.2.1 Remove the RT Primer (NR811, Novizan), thaw and mix well, briefly centrifuge to collect the contents to the bottom of the tube, and place on ice for later use. All subsequent steps shall be performed on ice. Prepare the reaction system according to Table 10.
[0362] Table 10
[0363] Components volume Step 4.1.3 Product 20ul RT Primer (NR811, Novizan) 1ul DEPC Water (W915679-500ml, McLean) 4.5ul
[0364] 4.2.2 Mix thoroughly by pipetting 10-15 times, then briefly centrifuge to collect the contents at the bottom of the tube. Place the tube in a heated PCR instrument and run the program shown in Table 11.
[0365] Table 11
[0366] temperature time 75℃ 5 minutes 37℃ 15 minutes 25℃ 15 minutes 4℃ Keep
[0367] 4.3 Connect 5' connector
[0368] Remove the RL5 Adaptor (NR811, Novizan), RL5 Buffer (NR811, Novizan), and RL5 Enzyme mix (NR811, Novizan), thaw and mix them, briefly centrifuge to collect the contents to the bottom of the tube, and place them on ice for later use. All subsequent steps shall be performed on ice.
[0369] 4.3.1 Denaturing the 5' adapter: The RL5 Adaptor (NR811, Novizan) may form a secondary structure, which needs to be denatured to open the secondary structure before use. The specific procedure is as follows: Place the RL5 Adaptor (NR811, Novizan) in a PCR instrument at 70°C for 2 minutes, and immediately remove it and place it on ice after the reaction is complete.
[0370] 4.3.2 Prepare the 5' connector reaction system according to Table 12.
[0371] Table 12
[0372] Components volume Step 4.2.2 Product 25.5ul Step 4.3.1 Modified RL5 Adaptor (NR811, Novizan) 1ul RL5 Buffer (NR811, Novizan) 1ul RL5 Enzyme Mix (NR811, Novizan) 2.5ul
[0373] 4.3.3 Use a pipette to mix thoroughly 10-15 times, then briefly centrifuge to collect the reaction solution to the bottom of the tube. Place the reaction tube in a PCR instrument with a heated lid and run the program shown in Table 13.
[0374] Table 13
[0375] temperature time 25℃ 1 hour 4℃ maintain
[0376] 4.4cDNA synthesis
[0377] 4.4.1 Take out the RT Buffer (NR811, Novizan) and RT Enzyme mix V2 (NR811, Novizan), thaw and mix well, briefly centrifuge to collect the contents to the bottom of the tube, and place on ice for later use. Prepare the reverse transcription reaction system according to the components shown in Table 14.
[0378] Table 14
[0379] Components volume Step 4.3.3 Product 30ul RT Buffer (NR811, Novizan) 8ul RT Enzyme mix V2 (NR811, Novizan) 2ul
[0380] 4.4.2 Mix thoroughly by pipetting 10-15 times, then briefly centrifuge to collect the reaction solution to the bottom of the tube. Place the reaction tube in a PCR instrument with a heated lid and run the program shown in Table 15.
[0381] Table 15
[0382] temperature time 50℃ 1 hour 80℃ 5 minutes 4℃ maintain
[0383] 4.5 Library Enrichment
[0384] 4.5.1 Take out Universal Primer (NR811, Novizan), Index Primer (NR811, Novizan), and Amplification mix 3 (NR811, Novizan), thaw and mix well, then place on ice for later use. Prepare the reaction system according to Table 16.
[0385] Table 16
[0386] Components volume Step 4.4.2 Product 40ul Amplification mix 3 (NR811, Novizan) 50ul Universal Primer (NR811, Novizan) 2.5ul Index Primer (NR813, Novizan) 2.5ul ddH2O 5ul
[0387] 4.5.2 Mix the above reaction solution thoroughly and briefly centrifuge to collect the residue at the bottom of the tube. Place the reaction tube in a PCR instrument with a heated lid and run the program shown in Table 17.
[0388] Table 17
[0389]
[0390] 4.6 PCR product purification
[0391] 4.6.1 Remove the VAHTSDNA Clean Beads (N411, Novizan) from 2–8°C 30 minutes in advance and allow them to stand until they reach room temperature.
[0392] 4.6.2 Invert or vortex to thoroughly mix VAHTSDNA Clean Beads (N411, Novizan), add 180 μl (1.8×) to the PCR product, and gently pipette 10 times to mix thoroughly.
[0393] 4.6.3 Incubate at room temperature for 10 minutes to allow DNA to bind to the magnetic beads.
[0394] 4.6.4 Place the sample on the magnetic rack and wait for the solution to clarify (about 10 minutes), then carefully remove the supernatant.
[0395] 4.6.5 Keep the sample on the magnetic rack at all times, add 200 μl of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.
[0396] 4.6.6 Repeat step 4.6.5 once.
[0397] 4.6.7 Keep the sample on the magnetic rack at all times, and open the lid to dry the magnetic beads for about 5 minutes at room temperature.
[0398] 4.6.8 Remove the sample from the magnetic rack, add 20 μl of Nuclease-free ddH2O, gently pipette to mix thoroughly, let stand at room temperature for 2 min, then place on the magnetic rack. After the solution becomes clear (about 5 min), carefully aspirate the supernatant into a new PCR tube.
[0399] 4.7 2100 Bioanalyzer (Agilent) for detecting library length.
[0400] 4.8 The samples were sequenced by Lianchuan Biotechnology Co., Ltd., using an Illumina Novaseq platform PE150.
[0401] 4.9 Bioinformatics was used for analysis and statistics, and the results are as follows: Figure 6 Based on the RNA library construction results, the shortest RNA sequence was 19 bases, which was the complementary sequence to the positive strand of the T7 promoter; the longest RNA sequence was 117 bases, which was an RNA sequence with the same number of bases that was perfectly complementary to the template DNA. This indicates that when the 5' end of a single-stranded DNA is attached to the 3' end of the positive strand of the T7 promoter, the template DNA can be linearly amplified, and the transcription mode is a novel transcription mode different from existing modes.
[0402] Example 3: Library Construction Based on Lambda Genomic DNA Sequencing
[0403] 1. Sample and sequence information
[0404] The linear transcription function-specific sequence is the same as the DR_T7 sequence in Example 1.
[0405] Sample: lambda genomic DNA (D1521, Promega)
[0406] Ultrasonic disruption of lambda genomic DNA resulted in a median fragment length of approximately 313 bp. (See...) Figure 7 .
[0407] 2. In vitro transcription
[0408] 2.1 PNK phosphorylation
[0409] The reaction system is shown in Table 18.
[0410] Table 18
[0411] Components content Broken lambda DNA 8ng T4 PNK buffer (M0201V, NEB) 0.36ul ATP (P0756S, NEB) 0.36ul T4 PNK(M0201V, NEB) 0.1ul DEPC Water (W915679-500ml, McLean) Up to 3.6ul
[0412] The reaction conditions were: 37℃ for 30 minutes; 75℃ for 20 minutes, and the PCR instrument cover temperature was 85℃.
[0413] 2.2 Connecting specific sequences for linear transcription function
[0414] The product from the previous step was heated to 95°C for 3 minutes, then quickly placed on ice for 3 minutes, and then the reaction system shown in Table 19 was prepared.
[0415] Table 19
[0416] Components volume Previous product 3.6ul 10X Reaction Buffer (EL0021, Thermo Scientific) 2ul BSA (EL0021, Thermo Scientific) 2ul T7_DR sequence (100uM, Sangon Biotech) 0.2ul 50% PEG8000 (R0056-2ml, Beyotime) 11.2ul T4 RNA ligase 1 (EL0021, Thermo Scientific) 1ul
[0417] Reaction conditions: 37℃ for 30 minutes; 75℃ for 10 minutes, PCR instrument lid temperature 85℃
[0418] 2.3T7 RNA polymerase transcription
[0419] The reaction system is shown in Table 20.
[0420] Table 20
[0421] Components volume Previous product 20ul DEPC Water (W915679-500ml, McLean) 35ul NTP mix (D7387-1ml, Beyotime) 30ul T7 RNA polymerase (M0251S, NEB) 5ul
[0422] The reaction conditions were: 37℃ for 19 hours.
[0423] 2.4 Digesting DNA
[0424] The reaction system is shown in Table 21
[0425] Table 21
[0426] Components volume Previous product 90ul DNase1 (M0303S, NEB) 10ul
[0427] Reaction conditions: 37℃ for 1 hour.
[0428] 2.5 Purifying RNA
[0429] The transcript RNA was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) (for specific operating procedures, refer to the instruction manual for RNA purification magnetic beads (N412-03, Novizan), and eluted with 18ul of DEPC water (W915679-500ml, Maclean). The RNA concentration was approximately 420ng / ul.
[0430] 3. Database construction
[0431] 3.1 Double-stranded cDNA Synthesis
[0432] Take out the components required for double-stranded cDNA synthesis from -30 to -15°C, thaw on ice, mix by inverting, collect to the bottom of the tube by short centrifugation, and prepare the first-stranded cDNA synthesis reaction system according to Table 22.
[0433] Table 22
[0434] Components volume RNA obtained from the previous purification step 16μl 1st Strand Buffer 6 (NR606, Novizan) 7μl 1st Strand Enzyme Mix 4 (NR606, Novizan) 2μl Total 25μl
[0435] Set the pipette to the 20 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0436] The first-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction temperature and time are shown in Table 23.
[0437] Table 23
[0438] temperature time Heat cap 105℃ On 25℃ 10min 42℃ 15min 70℃ 15min 4℃ Hold
[0439] 3.2 Prepare the second-strand cDNA synthesis reaction system according to Table 24.
[0440] Table 24
[0441] Components volume Previous step 1st Strand cDNA 25μl 2nd Strand Buffer 3 (with dNTP) (NR606, Novizan) 25μl 2nd Strand Enzyme Super Mix 3 (NR606, Novizan) 15μl Total 65μl
[0442] Set the pipette to the 50 μl range and gently aspirate 10 times to mix thoroughly.
[0443] The second-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction time and temperature are shown in Table 25.
[0444] Table 25
[0445] temperature time Heat cap 105℃ On 16℃ 30min 65℃ 15min 4℃ Hold
[0446] The two-chain synthetic product can be temporarily stored at -30 to -15°C for 24 hours.
[0447] 3.3 Connector Connection
[0448] Prepare the connection system according to Table 26.
[0449] Table 26
[0450] Components volume Previous step ds cDNA 65μl Rapid Ligation Buffer 6 (NR606, Novizan) 25μl Rapid DNA Ligase 6 (NR606, Novizan) 5μl Adapter (N321, Novizan) 5μl Total To 100μl
[0451] Set the pipette to the 80 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0452] The ligation reaction was performed in a PCR instrument. The reaction temperature and time are shown in Table 27.
[0453] Table 27
[0454] temperature time Heat cap 105℃ On 20℃ 15min 4℃ Hold
[0455] 3.4 Product Purification
[0456] Purify twice with 1X DNA Clean Beads (N411, Novizan) and elute with 21 μL ddH2O (for specific purification steps for each step, refer to 4.6 PCR product purification in Example 2).
[0457] 3.5 Library Augmentation
[0458] Prepare the PCR reaction system according to Table 28.
[0459] Table 28
[0460] Components volume Purified adapter ligation products 20μl i5 Primer (N321, Novizan) 2.5μl i7 Primer (N321, Novizan) 2.5μl VAHTS HiFi Amplification Mix 3 (NR606, Novizan) 25μl Total 50μl
[0461] Set the pipette to the 30 μl range and gently aspirate 10 times to mix thoroughly.
[0462] The samples were placed in a PCR instrument for library amplification reaction. The reaction conditions are shown in Table 29.
[0463] Table 29
[0464]
[0465] 3.6 Purification of PCR Products
[0466] Purify with 1X DNA Clean Beads (N411, Novizan) and elute with 20 μL ddH2O (for specific operating procedures, refer to 4.6 PCR product purification in Example 2).
[0467] 4. Sequencing and bioinformatics analysis
[0468] The constructed library was sequenced to a length of 2100 byte, and bioinformatics analysis was performed on the sequencing data. The results are shown below. Figure 8-10 The present invention was used to construct a library from 8 ng of lambda genomic DNA to test its applicability to genome library construction. The results showed a total of 3,733,284 reads, with a uniform base distribution in the sequencing data. The sequencing results of the library constructed using the present invention successfully aligned to the lambda reference genome, with a 100% alignment success rate and an average genome coverage depth of 5,342.6014, representing a coverage rate greater than or equal to 95%. Figure 10 This indicates that the developed method for linking single-stranded DNA to the positive strand of the T7 promoter can be further applied to the construction and sequencing of genomic DNA libraries.
[0469] Example 4: Constructing libraries based on mouse genomic DNA of different lengths to test the compatibility of this method with template DNA lengths.
[0470] 1. Sample and sequence information
[0471] The linear transcription function-specific sequence is the same as the DR_T7 sequence in Example 1.
[0472] Preparation of mouse genomic DNA fragment samples: Genomic DNA was extracted from C57BL / 6 mice using a genomic DNA extraction kit (DP304-03, Tiangen Biotech). The extracted genomic DNA was then fragmented into DNA fragments of different lengths using sonication. The lengths were measured at 4200 nm. Results are shown below. Figure 11-14 See Table 30.
[0473] Table 30
[0474] Sample grouping DNA length 0s interrupt group Greater than 1.5kb 30s interrupt group Approximately 1kb 45s interrupt group Approximately 600bp 60s interrupt group Approximately 400bp
[0475] 2. In vitro transcription
[0476] 2.1 PNK phosphorylation
[0477] The reaction system is shown in Table 31.
[0478] Table 31
[0479] Components content Disrupted C57BL / 6 mouse genomic DNA 10ng T4 PNK buffer (M0201V, NEB) 0.36ul ATP (P0756S, NEB) 0.36ul T4 PNK(M0201V, NEB) 0.1ul DEPC Water (W915679-500ml, McLean) Up to 3.6ul
[0480] The reaction conditions were: 37℃ for 30 minutes; 75℃ for 20 minutes, and the PCR instrument cover temperature was 85℃.
[0481] 2.2 Connecting specific sequences in linear transcription
[0482] The product from the previous step was denatured at 95°C for 3 minutes, then quickly placed on ice for 2 minutes, and then the reaction system shown in Table 19 was prepared.
[0483] Reaction conditions: 37℃ for 30 minutes; 75℃ for 10 minutes, PCR instrument lid temperature 85℃
[0484] 2.3T7 RNA polymerase transcription
[0485] The reaction system is shown in Table 32.
[0486] Table 32
[0487] Components volume Previous product 20ul DEPC Water (W915679-500ml, McLean) 16ul NTP mix (D7387-1ml, Beyotime) 48ul RNase inhibitor (R301, Novizan) 3ul T7 RNA polymerase (M0251S, NEB) 3ul
[0488] The reaction conditions were: 37℃ for 16 hours.
[0489] 2.4 Purifying RNA
[0490] Each reaction tube was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) to purify the transcript RNA (refer to the instructions for RNA purification magnetic beads (N412-03, Novizan) for specific operating procedures) and eluted with 21ul DEPC water (W915679-500ml, Maclean).
[0491] 3. Database construction
[0492] 3.1 Double-stranded cDNA Synthesis
[0493] Take out the components required for double-stranded cDNA synthesis from -30 to -15°C, thaw on ice, mix by inverting, collect to the bottom of the tube by short centrifugation, and prepare the first-stranded cDNA synthesis reaction system according to Table 22.
[0494] Set the pipette to the 20 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0495] The first-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction temperature and time are shown in Table 23.
[0496] 3.2 Prepare the second-strand cDNA synthesis reaction system according to Table 24.
[0497] Set the pipette to the 50 μl range and gently aspirate 10 times to mix thoroughly.
[0498] The second-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction time and temperature are shown in Table 25.
[0499] 3.3 Connector Connection
[0500] Prepare the connection system according to Table 26.
[0501] Set the pipette to the 80 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0502] The ligation reaction was performed in a PCR instrument. The reaction temperature and time are shown in Table 27.
[0503] 3.4 Product Purification
[0504] Purification was performed using 2X DNA Clean Beads (N411, Novizan), followed by elution with 21 μL ddH2O (for specific operational steps, refer to 4.6 PCR product purification in Example 2).
[0505] 3.5 Library Augmentation
[0506] Prepare the PCR reaction system according to Table 28. Adjust the pipette to the 30 μl range and gently pipette 10 times to mix thoroughly. Then place the sample in the PCR instrument for library amplification reaction, and the reaction conditions are shown in Table 29.
[0507] 3.6 Purification of PCR Products
[0508] Purify with 1X DNA Clean Beads (N411, Novizan) and elute with 20 μL ddH2O (for specific operating procedures, refer to 4.6 PCR product purification in Example 2).
[0509] 4. Sequencing and bioinformatics analysis
[0510] The bioinformatics analysis results of the ultrasound 0s sequencing data are shown in Table 33 and Figure 15-19 The bioinformatics analysis results of the 30s ultrasound sequencing data are shown in Table 34 and... Figure 15-19 The bioinformatics analysis results of the 45s ultrasound sequencing data are shown in Table 35 and... Figure 15-19 The bioinformatics analysis results of the 60s ultrasound sequencing data are shown in Table 36 and... Figure 15-19 .
[0511] Table 33
[0512] Fragment size 2,728,222,451 number of reads 995,276 Mapped reads 995,276 / 100% Unmapped reads 0 / 0% Mapped paired reads 995,276 / 100%
[0513] Table 34
[0514] Fragment size 2,728,222,451 number of reads 1,559,554 Mapped reads 1,559,554 / 100% Unmapped reads 0 / 0% Mapped paired reads 1,559,554 / 100%
[0515] Table 35
[0516] Fragment size 2,728,222,451 number of reads 820,588 Mapped reads 820,588 / 100% Unmapped reads 0 / 0% Mapped paired reads 820,588 / 100%
[0517] Table 36
[0518]
[0519]
[0520] Due to the different times of sonication fragmentation of template DNA, the length distribution of insert fragments in the sequencing libraries constructed from their transcripts did not differ significantly. Figure 15 The repeatability rates for the 0s, 30s, 45s, and 60s interruption groups were 34.62%, 28.45%, 28.59%, and 24.92%, respectively. Although the repeatability rate decreased with increasing ultrasound interruption time, it remained below 35%. Figure 16 This means that the sequencing results were maintained at a level close to those of existing kits. The unique alignment rates of reads from the 0s, 30s, 45s, and 60s break groups were all relatively high, at 90.7%, 95.1%, 94.6%, and 95%, respectively. Figure 17 Although the number of reads obtained from sequencing different libraries varied, this difference was not regularly related to the breakage time of the template DNA, which may be due to random errors during sequencing. Figure 18 The depth of sequencing data aligned to the reference genome varied among different groups, possibly due to differences in the sequencing data obtained: the 30s and 60s fragmentation groups had higher genome coverage than the 0s and 45s fragmentation groups, consistent with the fact that the 30s and 60s fragmentation groups obtained more reads than the 0s and 45s fragmentation groups. Figure 19 Furthermore, although the genome coverage depth varied among different groups compared to the control group, the overall trend of coverage depth variation across different coordinates was consistent. Figure 19In summary, compared to current next-generation sequencing kits that require template DNA lengths of around 300 bp, the single-stranded DNA linear transcription amplification library preparation method in this application has high compatibility with the starting length of template DNA, and can accommodate template DNA of several hundred to several thousand base pairs.
[0521] Example 5: Construction of libraries based on different amounts of HeLa cell genomic DNA
[0522] (I) Preparation of experimental group library
[0523] 1. Sample and sequence information
[0524] The linear transcription function-specific sequence is the same as the DR_T7 sequence in Example 1.
[0525] The samples were derived from HeLa cells and extracted using a cell genome extraction kit (DP304-03, Tiangen Biotech). The DNA was fragmented into approximately 600 bp segments using sonication. Figure 20 Based on the amount of DNA input, the groups were divided into 5 pg (two replicates), 10 pg (two replicates), 100 pg (two replicates), 1 ng, and 10 ng groups.
[0526] 2. In vitro transcription
[0527] 2.1 PNK phosphorylation
[0528] The reaction system is shown in Table 37.
[0529] Table 37
[0530]
[0531]
[0532] The reaction conditions were: 37℃ for 30 minutes; 75℃ for 20 minutes, and the PCR instrument cover temperature was 85℃.
[0533] 2.2 Connecting specific sequences in linear transcription
[0534] The product from the previous step was denatured at 95°C for 3 minutes, then quickly placed on ice for 2 minutes, and then the reaction system shown in Table 19 was prepared.
[0535] Reaction conditions: 37℃ for 30 minutes; 75℃ for 10 minutes, PCR instrument lid temperature 85℃
[0536] 2.3T7 RNA polymerase transcription
[0537] The reaction system is shown in Table 32.
[0538] The reaction conditions were: 37℃ for 16 hours.
[0539] 2.4 Purifying RNA
[0540] Each reaction tube was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) to purify the transcript RNA (refer to the instructions for RNA purification magnetic beads (N412-03, Novizan) for specific operating procedures) and eluted with 21ul DEPC water (W915679-500ml, Maclean).
[0541] 3. Database construction
[0542] 3.1 Double-stranded cDNA Synthesis
[0543] Take out the components required for double-stranded cDNA synthesis from -30 to -15°C, thaw on ice, mix by inverting, collect to the bottom of the tube by short centrifugation, and prepare the first-stranded cDNA synthesis reaction system according to Table 22.
[0544] Set the pipette to the 20 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0545] The first-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction temperature and time are shown in Table 23.
[0546] 3.2 Prepare the second-strand cDNA synthesis reaction system according to Table 24.
[0547] Set the pipette to the 50 μl range and gently aspirate 10 times to mix thoroughly.
[0548] The second-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction time and temperature are shown in Table 25.
[0549] 3.3 Connector Connection
[0550] Prepare the connection system according to Table 26.
[0551] Set the pipette to the 80 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0552] The ligation reaction was performed in a PCR instrument. The reaction temperature and time are shown in Table 27.
[0553] 3.4 Product Purification
[0554] Purification was performed using 0.6X DNA Clean Beads (N411, Novizan), followed by elution with 22 μL ddH2O (for specific procedures, refer to 4.6 PCR product purification in Example 2).
[0555] 3.5 Library Augmentation
[0556] Prepare the PCR reaction system according to Table 28. Adjust the pipette to the 30 μl range and gently pipette 10 times to mix thoroughly. Then place the sample in the PCR instrument for library amplification reaction, and the reaction conditions are shown in Table 29.
[0557] 3.6 Purification of PCR Products
[0558] Purification was performed using 0.9X DNA Clean Beads (N411, Novizan), followed by elution with 20 μL ddH2O (for specific procedures, refer to 4.6 PCR product purification in Example 2).
[0559] (II) Preparation of the control group library
[0560] 1. Sample and sequence information
[0561] The HeLa cell genomic DNA used is the same as that used in “(I) Preparation of Experimental Library” of this embodiment, which was fragmented by ultrasound. Figure 20 This experiment was divided into two groups based on the amount of DNA input: a 10ng group and a 100ng group.
[0562] 2. End-of-cell DNA repair
[0563] This step involves closing the DNA ends, phosphorylating the 5' end, and adding a dA tail to the 3' end.
[0564] 2.1 Thaw End Prep Buffer 2 (ND610, Novizan), mix it with End Prep Enzyme 2 (ND610, Novizan) by inverting, and then prepare the solution in a sterile PCR tube, as shown in Table 38 (operate on ice):
[0565] Table 38
[0566] Components content HeLa cell genomic DNA disrupted by ultrasound 10ng or 100ng End Prep Buffer 2 (ND610, Novizan) 7ul End Prep Enzyme 2 (ND610, Novizan) 3ul ddH2O Fill up to 60ul
[0567] 2.2 Gently pipette to mix (do not shake to mix), and briefly centrifuge to collect the reaction solution to the bottom of the tube.
[0568] 2.3 Place the PCR tubes in the PCR instrument and perform the following reactions, as shown in Table 39.
[0569] Table 39
[0570] temperature time Heat cap 105℃ On 20℃ 30 minutes 65℃ 30 minutes 4℃ Hold
[0571] 3. Connector connection
[0572] This step involves connecting a connector to the end of the product from the previous step.
[0573] 3.1 Thaw Rapid Ligation Buffer 5 (ND610, Novizan), invert and mix well, then place on ice. Prepare the reaction mixture (Table 40) in the PCR tube from the previous step (operate on ice):
[0574] Table 40
[0575] Components volume Previous step end repair product DNA 60ul Rapid Ligation Buffer 5 (ND610, Novizan) 30ul Rapid DNA Ligase 2 (ND610, Novizan) 10ul DNA Adapter X (N321, Novizan) 5ul ddH2O 5ul total 110ul
[0576] 3.2 Gently pipette to mix (do not shake to mix), and briefly centrifuge to collect the reaction solution to the bottom of the tube.
[0577] 3.3 Place the PCR tubes in the PCR instrument and perform the reaction procedure shown in Table 41:
[0578] Table 41
[0579] temperature time Heat cap 105℃ On 20℃ 15 minutes 4℃ Hold
[0580] 4. Purify DNA
[0581] The reaction product was purified using 0.8X DNA Clean Beads (N411, Novizan) and eluted with 22.5 μL of ddH2O. (For detailed operating procedures, refer to section 4.6 of Example 2 for PCR product purification.)
[0582] 5. Library expansion
[0583] This step involves PCR amplification of the purified adapter ligation product and the addition of complete next-generation sequencing adapters.
[0584] 5.1 Thaw PCR Primer Mix 3 for Illumina (ND610, Novizan) and VAHTS HiFiAmplification Mix 2 (ND610, Novizan), then invert and mix well. Prepare the reaction system in Table 28 in a sterile PCR tube (operate on ice).
[0585] 5.2 Adjust the pipette to the 30 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0586] 5.3 Place the samples in a PCR instrument for library amplification. The reaction program is shown in Table 29. Each sample was amplified for 8 cycles.
[0587] 6. Purify the reaction product using 0.9X DNA Clean Beads (N411, Novizan) and elute with 20 μL of ddH2O. (Refer to section 4.6, PCR product purification, in Example 2 for detailed operating procedures.)
[0588] (III) Sequencing and Bioinformatics Analysis Results
[0589] Bioinformatics analysis of HeLa cell genomic DNA library sequencing data with different input amounts, as follows: Figures 21-30 The number of reads obtained from library sequencing generally increases with the increase of DNA input. Figure 21 Conversely, as the starting amount increased, the percentage of uniquely matched reads gradually decreased: the unique alignment rates for reads with starting amounts of 5 pg, 10 pg, 100 pg, 1 ng, and 10 ng were 99.98%, 99.96%, 99.69%, 97.96%, and 83.6%, respectively. Figure 22 The unique alignment rate of reads in the control group with starting doses of 10 ng and 100 ng was 99.92%, which was lower than that in the experimental groups with lower starting doses of 5 pg and 10 pg. Figure 22 This demonstrates that the library construction method based on linear amplification of single-stranded DNA using the T7 promoter, as used in this invention, has a significant advantage in read alignment rate for constructing DNA libraries with low initial amounts.
[0590] GC bias refers to the tendency for sequences with a GC content closer to 50% to be detected, resulting in more reads and higher coverage. Conversely, sequences with higher or lower GC content are less likely to be detected, leading to fewer reads and lower coverage. GC bias arises from multiple stages, such as the amplification stage during library construction and the cluster amplification stage during sequencing. GC bias affects copy number variation detection, genome assembly, etc., therefore, it is necessary to minimize GC bias as much as possible. (Comparison with control group) Figure 23 and 24 ) and experimental group ( Figure 25-29 In the experimental group, regardless of whether the starting amount was low or relatively high, the sequencing data coverage (circles) was closer to the theoretical distribution (bar chart). The GC contents of the experimental group at different input amounts (5 pg_1, 5 pg_2, 10 pg_1, 10 pg_2, 100 pg_1, 100 pg_2, 1 ng, and 10 ng) were 37.38%, 37.47%, 37.65%, 37.81%, 37.82%, 37.68%, 38.61%, and 39.55%, respectively; while the GC contents of the control group (10 ng and 100 ng) were 41.76% and 41.25%, respectively. It can be seen that the GC contents of the experimental group were all lower than those of the control group. Figure 30 This indicates that the library construction method based on linear amplification of single-stranded DNA using the T7 promoter used in this invention has a lower GC preference than the control group.
[0591] In summary, the library construction method based on T7 promoter linear amplification of single-stranded DNA used in this invention exhibits lower GC bias and higher read uniqueness at low starting amounts. Therefore, the library construction method based on T7 promoter linear amplification of single-stranded DNA reported in this invention has significant advantages over existing double-stranded DNA-based library construction methods in experiments with low starting amounts and GC bias sensitivity.
[0592] Example 6: Construction of DNA Library Based on Mouse FFPE Sample Genome (Part 1) Preparation of FFPE Sample Experimental Group Library
[0593] 1. Sequence and Sample Information
[0594] The linear transcription function-specific sequence is the same as the DR_T7 sequence in Example 1.
[0595] Samples: DNA was extracted from formalin-fixed paraffin-embedded lung tissue samples from C57 / B6J mice 20 months prior using a paraffin-embedded tissue DNA extraction kit (TIANGEN, DP331-02). Experimental groups were divided into four groups based on the amount of DNA input: 10 ng⁻¹, 10 ng⁻², 25 ng⁻¹, 25 ng⁻², 50 ng⁻¹, 50 ng⁻², 100 ng⁻¹, and 100 ng⁻².
[0596] 2. In vitro transcription
[0597] 2.1 PNK phosphorylation
[0598] The reaction system is shown in Table 42.
[0599] Table 42
[0600] Components content Mouse FFPE sample genomic DNA Different input amounts T4 PNK buffer (M0201V, NEB) 0.36ul ATP (P0756S, NEB) 0.36ul T4 PNK(M0201V, NEB) 0.1ul DEPC Water (W915679-500ml, McLean) Up to 3.6ul
[0601] The reaction conditions were: 37℃ for 30 minutes; 75℃ for 20 minutes, and the PCR instrument cover temperature was 85℃.
[0602] 2.2 Connecting specific sequences for linear transcription function
[0603] The product from the previous step was denatured at 95°C for 3 minutes, then quickly placed on ice for 2 minutes, and then the reaction system shown in Table 19 was prepared.
[0604] Reaction conditions: 37℃ for 30 minutes; 75℃ for 10 minutes, PCR instrument lid temperature 85℃
[0605] 2.3T7 RNA polymerase transcription
[0606] The reaction system is shown in Table 32. The reaction conditions were: 37℃ for 16 hours.
[0607] 2.4 Purifying RNA
[0608] Each reaction tube was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) to purify the transcript RNA (refer to the instructions for RNA purification magnetic beads (N412-03, Novizan) for specific operating procedures) and eluted with 21ul DEPC water (W915679-500ml, Maclean).
[0609] 3. Database construction
[0610] 3.1 Double-stranded cDNA Synthesis
[0611] Take out the components required for double-stranded cDNA synthesis from -30 to -15°C, thaw on ice, mix by inverting, collect to the bottom of the tube by short centrifugation, and prepare the first-stranded cDNA synthesis reaction system according to Table 22.
[0612] Set the pipette to the 20 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0613] The first-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction temperature and time are shown in Table 23.
[0614] 3.2 Prepare the second-strand cDNA synthesis reaction system according to Table 24.
[0615] Set the pipette to the 50 μl range and gently aspirate 10 times to mix thoroughly.
[0616] The second-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction time and temperature are shown in Table 25.
[0617] 3.3 Connector Connection
[0618] Prepare the connection system according to Table 26, then adjust the pipette to the 80 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0619] The ligation reaction was performed in a PCR instrument. The reaction temperature and time are shown in Table 27.
[0620] 3.4 Product Purification
[0621] Purification was performed using 0.6X DNA Clean Beads (N411, Novizan), followed by elution with 22 μL ddH2O (for specific procedures, refer to 4.6 PCR product purification in Example 2).
[0622] 3.5 Library Augmentation
[0623] Prepare the PCR reaction system according to Table 28. Adjust the pipette to the 30 μl range and gently pipette 10 times to mix thoroughly. Then place the sample in the PCR instrument for library amplification reaction. The reaction conditions are shown in Table 29. Each sample is amplified for 9 cycles.
[0624] 3.6 Purification of PCR Products
[0625] Purification was performed using 0.9X DNA Clean Beads (N411, Novizan), followed by elution with 20 μL ddH2O (for specific procedures, refer to 4.6 PCR product purification in Example 2).
[0626] (II) Preparation of FFPE sample control library
[0627] 1. Sample Information
[0628] Samples: DNA was extracted from formalin-fixed paraffin-embedded lung tissue samples from C57 / B6J mice 20 months prior using a paraffin-embedded tissue DNA extraction kit (TIANGEN, DP331-02). Control groups were divided into 10ng, 25ng, 50ng, 100ng, and 300ng groups based on the amount of DNA input.
[0629] The DNA used was extracted from formalin-fixed paraffin-embedded samples of C57 / B6J mouse lung tissue obtained 20 months prior using the same paraffin-embedded tissue DNA extraction kit (TIANGEN, DP331-02) used in "(I) Preparation of FFPE Sample Experimental Library" of this embodiment. The control group was divided into 10ng, 25ng, 50ng, 100ng, and 300ng groups based on the amount of DNA used.
[0630] 2. End-of-cell DNA repair
[0631] This step involves closing the DNA ends, phosphorylating the 5' end, and adding a dA tail to the 3' end.
[0632] 2.1 Thaw End Prep Buffer 2 (ND610, Novizan), mix it with End Prep Enzyme 2 (ND610, Novizan) by inverting, and then prepare the solution in a sterile PCR tube, as shown in Table 43 (operate on ice):
[0633] Table 43
[0634] Components content Mouse FFPE sample genomic DNA Different amounts End Prep Buffer 2 (ND610, Novizan) 7ul End Prep Enzyme 2 (ND610, Novizan) 3ul ddH2O Fill up to 60ul
[0635] 2.2 Gently pipette to mix (do not shake to mix), and briefly centrifuge to collect the reaction solution to the bottom of the tube.
[0636] 2.3 Place the PCR tubes in the PCR instrument and perform the reaction according to the procedure in Table 39.
[0637] 3. Connector connection
[0638] This step involves connecting a connector to the end of the product from the previous step.
[0639] 3.1 Thaw Rapid Ligation Buffer 5 (ND610, Novizan), invert and mix well, then place on ice. Prepare the reaction mixture shown in Table 40 in the end-repair PCR tube from the previous step (operate on ice).
[0640] 3.2 Gently pipette to mix (do not shake to mix), and briefly centrifuge to collect the reaction solution to the bottom of the tube.
[0641] 3.3 Place the PCR tubes in the PCR instrument and perform the reaction program in Table 41.
[0642] 4. Purify DNA
[0643] The reaction product was purified using 0.8X DNA Clean Beads (N411, Novizan) and eluted with 22.5 μL of ddH2O. (For detailed operating procedures, refer to section 4.6 of Example 2 for PCR product purification.)
[0644] 5. Library expansion
[0645] This step involves PCR amplification of the purified adapter ligation product and the addition of complete next-generation sequencing adapters.
[0646] 5.1 Thaw PCR Primer Mix 3 for Illumina (ND610, Novizan) and VAHTS HiFiAmplification Mix 2 (ND610, Novizan), then invert and mix well. Prepare the reaction system in Table 28 in a sterile PCR tube (operate on ice).
[0647] 5.2 Adjust the pipette to the 30 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0648] 5.3 Place the samples in a PCR instrument for library amplification. The reaction program is shown in Table 29. Each sample was amplified for 8 cycles.
[0649] 6. Purify the reaction product using 0.9X DNA Clean Beads (N411, Novizan) and elute with 20 μL of ddH2O. (Refer to section 4.6, PCR product purification, in Example 2 for detailed operating procedures.)
[0650] (III) Sequencing and Bioinformatics Analysis Results
[0651] Bioinformatics analysis of genomic DNA library sequencing data from mouse lung tissue FFPE samples with different input amounts, as follows: Figures 31-32As the starting amount increased, the percentage of reads uniquely aligned to the reference genome in both the experimental and control groups gradually decreased. In the experimental group, the unique alignment rates of reads corresponding to starting amounts of 10ng, 25ng, 50ng, and 100ng were close to those in the control group, maintaining a relatively high level (greater than 98%).
[0652] Based on the statistical results of GC content in the sequencing data of the experimental and control groups under different input amounts of FFPE sample DNA ( Figure 31 In the experimental group, the GC content increased with increasing template DNA starting amount, while in the control group, the GC content decreased with increasing template DNA starting amount. This means that the present invention has a significant advantage for FFPE sample DNA that is sensitive to subsequent GC bias and has a low starting amount (less than or equal to 50 ng).
[0653] GIV (Global Imbalance Value) is a statistical method for reflecting DNA damage. A GIV score greater than 1.5 indicates a large number of DNA base substitutions in the sequencing data. Comparing the GIV scores of the experimental and control groups, it can be seen that the GIV scores of the experimental group are all less than 1.5; while in the control group, the GIV scores of G_T and G_A are greater than 1.5 for all input levels, and the GIV score of T_A is also greater than 1.5 even at low input levels.
[0654] In summary, the library construction method based on T7 promoter linear amplification of single-stranded DNA used in this invention maintains a relatively high unique read alignment rate across different starting amounts. Furthermore, it exhibits lower GC bias in library construction results for FFPE sample DNA with low starting amounts (≤50 ng). Additionally, its GIV score is superior to the control group for different starting amounts of DNA. Therefore, for FFPE sample DNA, the library construction method based on T7 promoter linear amplification of single-stranded DNA reported in this invention has significant advantages over existing double-stranded DNA-based library construction methods in constructing libraries for damaged DNA in the 10-300 ng range, and in library construction experiments with low starting amounts and GC bias sensitivity.
[0655] Example 7: Constructing a library based on low-frequency mutation standards
[0656] 1. Sample Information
[0657] EGFR L858R wild-type and mutant primers were synthesized at Sangon Biotech using primer synthesis, and the synthesized primers were annealed into double-stranded DNA. The DNA concentration was measured using a Qubit assay, and a certain proportion of mutant primers was incorporated into the synthesized wild-type DNA to obtain an EGFR L858R DNA standard with a mutation frequency of 0.1%.
[0658] EGFR L858R wild-type sequence:
[0659] AGCCAGGAACGTACTGGTGAAAACACCGCAGCATGTCAAGATCACAGATTTTGGGCTGGCCAAACTGCTGGGTGCGGAAGAGAAAGAA TACCATGCAGAAGGAGGCAAAGTAAGGAGGTGGCTTTAGGTCAGCCAGCATTTTCCTGACACCAGGGACCAGG (SEQ ID No. 9) EGFR L858R mutant sequence:
[0660] AGCCAGGAACGTACTGGTGAAAACACCGCAGCATGTCAAGATCACAGATTTTGGGCGGGCCAAACTGCTGGGTGCGGAAGAGAAAGAA TACCATGCAGAAGGAGGCAAAGTAAGGAGGTGGCTTTAGGTCAGCCAGCATTTTCCTGACACCAGGGACCAGG (SEQ ID No. 10) The T7_DR sequence is the same as in Example 1.
[0661] 2. PNK phosphorylation
[0662] 2.1 PNK phosphorylation
[0663] The reaction system is shown in Table 44.
[0664] Table 44
[0665]
[0666]
[0667] The reaction conditions were: 37℃ for 30 minutes; 75℃ for 20 minutes, and the PCR instrument cover temperature was 85℃.
[0668] 2.2 Connecting linear transcriptional functional sequences
[0669] The product from the previous step was denatured at 95°C for 3 minutes, then quickly placed on ice for 2 minutes, and then the reaction system shown in Table 19 was prepared.
[0670] Reaction conditions: 37℃ for 30 minutes; 75℃ for 10 minutes, PCR instrument lid temperature 85℃
[0671] 2.3T7 RNA polymerase transcription
[0672] The reaction system is shown in Table 32. The reaction conditions were: 37℃ for 16 hours.
[0673] 2.4 Purifying RNA
[0674] Each reaction tube was purified using 1.8X RNA purification magnetic beads (N412-03, Novizan) to purify the transcript RNA (refer to the instructions for RNA purification magnetic beads (N412-03, Novizan) for specific operating procedures) and eluted with 21ul DEPC water (W915679-500ml, Maclean).
[0675] 3. Database construction
[0676] 3.1 Double-stranded cDNA Synthesis
[0677] Take out the components required for double-stranded cDNA synthesis from -30 to -15°C, thaw on ice, mix by inverting, collect to the bottom of the tube by short centrifugation, and prepare the first-stranded cDNA synthesis reaction system according to Table 22.
[0678] Set the pipette to the 20 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0679] The first-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction temperature and time are shown in Table 23.
[0680] 3.2 Prepare the second-strand cDNA synthesis reaction system according to Table 24.
[0681] Set the pipette to the 50 μl range and gently aspirate 10 times to mix thoroughly.
[0682] The second-strand cDNA synthesis reaction was carried out in a PCR instrument. The reaction time and temperature are shown in Table 25.
[0683] 3.3 Connector Connection
[0684] Prepare the connection system according to Table 26, then adjust the pipette to the 80 μl range and gently aspirate and mix 10 times to ensure thorough mixing.
[0685] The ligation reaction was performed in a PCR instrument. The reaction temperature and time are shown in Table 27.
[0686] 3.4 Product Purification
[0687] Purification was performed using 0.6X DNA Clean Beads (N411, Novizan), followed by elution with 22 μL ddH2O (for specific procedures, refer to 4.6 PCR product purification in Example 2).
[0688] 3.5 Library Augmentation
[0689] Prepare the PCR reaction system according to Table 28. Adjust the pipette to the 30 μl range and gently pipette 10 times to mix thoroughly. Then place the sample in the PCR instrument for library amplification reaction. The reaction conditions are shown in Table 29. Each sample is amplified for 9 cycles.
[0690] 3.6 Purification of PCR Products
[0691] 1.2X DNA Clean Beads (N411, Novizan) purification, eluted with 20 μL ddH2O (for specific operating procedures, refer to 4.6 PCR product purification in Example 2).
[0692] 4. Sequencing and Bioinformatics Analysis Results: The sequencing data was 2100 ms long, and bioinformatics analysis was performed on the sequencing data. The mutation frequency statistics are shown in Table 45.
[0693] Table 45
[0694] Grouping 10-3 sets WT Group MUT Group T 1216325 285883 3110 G 1362(Ng3) 165 (Nwt) 257397(Nmut) C 823 142 543 A 210 50 76 total 1218720(Nt3) 286240 261126
[0695] Because there are certain base errors in the synthesis of the WT and MUT groups, a certain proportion of bases will be introduced into the standard. Therefore, by sequencing the synthesized WT and MUT groups and then calculating, the noise caused by the incorrect bases during synthesis can be removed.
[0696] Calculation formula:
[0697] Ng3 / (Nmut+Nwt*(10 n -1))*Nmut / Nt3
[0698] The actual measured mutation result of the standard sample after noise removal is as follows:
[0699] 10-3 groups: 0.68*10 -3
[0700] In summary, considering that this invention is based on T7 linear amplification of template DNA, thus avoiding erroneous amplification during the amplification process, the method described in this invention was used to detect EGFR L858R standards with a mutation frequency of 0.1% to verify whether this method can be used for the detection of low-frequency mutations. Bioinformatics analysis results show that the method described in this invention can detect 0.1% of mutated bases, indicating that this invention is suitable for low-frequency mutation detection, providing a new method for disease-related mutation detection.
[0701] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
[0702] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.
Claims
1. A method for in vitro transcription, characterized in that, The method includes: performing in vitro transcription using a linear transcription functionally specific sequence at the 5' end of a single-stranded DNA as a template; wherein the linear transcription functionally specific sequence includes a T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA.
2. The method according to claim 1, characterized in that, The single-stranded DNA added to the 3' end of the T7 promoter's positive strand is either a directly synthesized sequence or the 5' end of a single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand. Preferably, the nucleic acid functional fragment is DNA, RNA, or a fragment containing both DNA and RNA.
3. The method according to claim 1 or 2, characterized in that, Adding a specific sequence for linear transcription to the 5' end of a single-stranded DNA can form a single strand, a double strand with a notch, or a double strand in a localized area. Preferably, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment; The 5' end of the nucleic acid functional fragment is linked to the 3' end of the T7 promoter positive strand, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA; Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
4. The method according to claim 3, characterized in that, Methods for adding a linear transcriptionally specific sequence to the 5' end of single-stranded DNA include ligase ligation or chemical ligation. Specifically, the ligase ligation method is as follows: a method of ligating the 5' end of a single-stranded DNA to the 3' end of a specific sequence for linear transcription using a ligase; More specifically, the ligase is: Ligases with the function of connecting nicks and / or double-stranded nucleic acids, such as T4 RNA ligase 2, T4 DNA ligase, T3 DNA ligase, SplintR ligase, Taq DNA ligase, and 9... O N DNA ligase, E. coli ligase, and mutants of the above ligases; or ligases with the function of ligating single-stranded nucleic acids, such as T4 RNA ligase 1 (ssRNA ligase), Mth RNA ligase, 5' App DNA / RNA thermostable ligase, TS2126 RNA ligase (TS2126 Rnl 1 or CircLigase), Hyper-Thermostable Lysine-Mutatant ssDNA / RNA ligase (HyperLigase), T4 RNA ligase 2, RtcB ligase, and mutants of the above ligases; Specifically, the chemical ligation method is a method of ligating the 5' end of a single-stranded DNA to the 3' end of a specific sequence for linear transcription through a chemical reaction; More specifically, the chemical reactions include methods such as click chemistry or Michael addition reactions, for example, linkages via oxime bonds, amide bonds, thioether bonds, disulfide bonds, phosphoryl bonds, hydrazone bonds, urea bonds, or ring linkages formed by click chemistry reactions.
5. The method according to any one of claims 1-4, characterized in that, The in vitro transcription reaction system includes T7 RNA polymerase or a mutant thereof; Preferably, the reaction temperature for the in vitro transcription is any value between 10°C and 65°C, and more preferably between 16°C and 42°C. Preferably, the reaction time of the in vitro transcription is any value from 0.01h to 1000h, more preferably any value from 0.1h to 100h, and even more preferably any value from 0.5h to 72h.
6. The method according to any one of claims 3-5, characterized in that, The product of adding a specific sequence for linear transcription to the 5' end of single-stranded DNA can be transcribed in vitro after purification, or it can be transcribed in vitro directly without purification.
7. The method according to any one of claims 1-6, characterized in that, Methods for obtaining single-stranded DNA include isolating DNA from a sample; the isolated DNA can be single-stranded or double-stranded. Specifically, the double strand can be a complete double strand or an incomplete double strand; Preferably, the 5' end of the double-stranded DNA has more than one free base or nucleotide residue. More preferably, the 5' end of the double-stranded DNA contains more than one free base or nucleotide residue, which can be DNA obtained directly from the sample or obtained by physical, chemical or biological methods. Specifically, physical methods include temperature denaturation and fracture; Specifically, chemical methods include the use of chemical agents such as urea or formamide; Specifically, biological methods include enzyme digestion, enzyme ligation, or adding bases to the ends of enzymes.
8. The method according to claim 7, characterized in that, The method includes denaturing double-stranded DNA into single-stranded DNA; Preferably, the denaturation is performed using physical, chemical, or biological methods; Specifically, physical methods include temperature denaturation and fracture; Specifically, chemical methods include the use of chemical agents such as urea or formamide; Specifically, biological methods include enzyme digestion, enzyme ligation, and enzyme addition of bases to the ends.
9. The method according to claim 7, characterized in that, The samples are selected from body fluids, feces, cells, tissues or organs, and may also be selected from paraffin-embedded samples, natural environments such as soil and ponds, or fossil samples, etc. Specifically, the body fluids mentioned are selected from plasma, tissue fluid, lymph, urine, tears, digestive juices, synovial fluid, sweat, cerebrospinal fluid, or vaginal secretions, etc. Specifically, the digestive juices can be selected from bile, saliva, gastric juice, or pancreatic juice, etc. Preferably, the sample is isolated from an organism, a paleontology, or the natural environment; Preferably, the sample includes a paraffin-embedded sample, a sample converted from bisulfite, or a sample treated with enzymes.
10. The method according to any one of claims 1-9, characterized in that, The positive chain sequence of the T7 promoter includes TAATACGACTCACTATAG (SEQ ID No. 1), TAATACGACTCACTATAGG (SEQ ID No. 11), or TAATACGACTCACTATAGGG (SEQ ID No. 12). Preferably, the linear transcription functional specific sequence may further include the antisense strand of the T7 promoter.
11. The method according to any one of claims 2-10, characterized in that, The nucleic acid functional fragment sequence includes a marker sequence and / or an enzyme restriction site, wherein the marker sequence is, for example, a primer sequence, a barcode, a UMI, or an index; Preferably, the 5' end of the marker sequence or restriction enzyme site is linked to the 3' end of the positive strand of the T7 promoter, or the 3' end of the marker sequence or restriction enzyme site is linked to the 5' end of single-stranded DNA.
12. The method according to any one of claims 1-11, characterized in that, The specific sequence for linear transcription function is single-stranded, double-stranded, or contains a partial double strand; Preferably, the 3' end sequence of the linear transcription functional specific sequence is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the linear transcription functional specific sequence is a single strand of RNA, and the 3' end of the single strand of RNA is connected to the 5' end of the single strand of DNA; Preferably, the length of the RNA single strand is greater than or equal to 1 mer, more preferably greater than or equal to 3 mer, and even more preferably 3-10000 mer or more; Preferably, the linear transcription functional specific sequence can be SEQ ID No. 2 or SEQ ID No.
7.
13. The method according to any one of claims 1-12, characterized in that, The specific sequence for linear transcription function can be an unmodified nucleic acid or a modified nucleic acid.
14. The method according to claim 13, characterized in that, The modification is selected from one or more of the following: 5' end modification, 3' end modification, introduction of non-natural nucleotides, base modification, phosphate backbone modification, sugar ring modification, or introduction of a substance that can be linked to nucleic acids; Preferably, the base modification is selected from one or more of the following: 5-position pyrimidine modification, 8-position purine modification, or 5-bromouracil substitution; Preferably, the sugar ring modification is selected from one or more groups selected from H, OZ, Z, halo, SH, SZ, NH2, NHZ, NZ2 or CN, wherein Z is an alkyl group; Preferably, the phosphate skeleton modification includes thiophosphate modification; Preferably, the non-natural nucleotide is one or more of the following: non-natural bases, non-natural sugar moieties, or non-natural backbones. More preferably, the non-natural base is selected from 2-aminoadenine-9-yl, 5-hydroxymethyluracil, 2-aminoadenine, 2-F-adenine, 2-thiouracil, 2-thiothymidine, 2-thiocytosine, 2-propyl and alkyl derivatives of adenine and guanine, 2-amino-adenine, 2-amino-propyl-adenine, 2-aminopyridine, 2-pyridone, 2'-deoxyuridine, 2-amino-2'-deoxyadenosine, 3-deazoguanine, 3-deazoadenine, 4-thiouracil, 4-thiothymidine, uracil-5-yl, hypoxanthine-9-yl (I), 5-methylcytosine, 5-hydroxymethyl Cytosine, xanthine, hypoxanthine, 5-bromouracil, 5-trifluoromethyluracil, 5-bromocytosine, 5-trifluoromethylcytosine, 5-halouracil, 5-halocytosine, 5-propynyluracil, 5-propynylcytosine, 5-uracil, 5-substituted pyrimidine, 5-hydroxycytosine, 5-bromocytosine, 5-bromouracil, 5-chlorocytosine, cyclocytosine, cytarabine, 5-fluorocytosine, fluoropyrimidine, fluorouracil, 5,6-dihydrocytosine, 5-iodocytosine, hydroxyurea, iodouracil, 5-nitrocytosine, 5-bromouracil, 5-chlorouracil, 5-fluorouracil, 5-iodo 6-alkyl derivatives of uracil, adenine, and guanine, including 6-azyuracil, 6-azouracil, 6-azocytosine, azocytosine, 6-azothymidine, 6-thioguanine, 7-methylguanine, 7-methyladenine, 7-deazoguanine, 7-deazoguanosine, 7-deazo-adenine, 7-deazo-8-azyguanine, 8-azyguanine, 8-azyadenine, 8-aminoadenine, 8-aminoguanine, 8-thiol adenine, 8-thiol guanine, 8-thioalkyladenine, 8-thioalkylguanine, 8-hydroxyadenine, 8-hydroxyguanine, N4-ethylcytosine, and N-2-substituted purines. N-6 substituted purines, O-6 substituted purines, fluorinated nucleic acids, tricyclic pyrimidines, phenoxazincytidine ([5,4-b][l,4]benzoxazin-2(3H)-one), phenthiazincytidine (1H-pyrimido[5,4-b][l,4]benzothiazin-2(3H)-one), G-clamps, phenoxazincytidine (9-(2-aminoethoxy)-H-pyrimido[5,4-b][1,4]benzoxazin-2(3H)-one), carbazolecytidine (2H-pyrimido[4,5-b]indole-2-one), pyridoindolecytidine (H-pyrido[3',2':4,5]pyrrolo[2,3-d]pyrimidin-2-one), 5-fluorouracil, 5-bromouracil, 5-chlorouracil, 5-iodouracil, hypoxanthine, xanthine, 4-acetylcytosine, 5-(carboxyhydroxymethyl)uracil, 5-carboxymethylaminomethyl-2-thiouracil, 5-carboxymethylaminomethyluracil, dihydrouracil, β-D-galactosylpiperidine, inosine, N6-isopentene adenine, 1-methylguanine, 1-methylinosine, 2,2-dimethylguanine, 2-methyladenine, 2-methylguanine, 3-methylcytosine, 5-methylcytosine, N6-adenine, 7-methylguanine, 5-methylaminomethyl Uracil, 5-methoxyaminomethyl-2-thiouracil, β-D-mannosyl uracil, 5'-methoxycarboxymethyluracil, 5-methoxyuracil, 2-methylthio-N6-isopentene adenine, uracil-5-oxyacetic acid, weidooxyglycoside, pseudouracil, uracil, 2-thiocytosine, 5-methyl-2-thiouracil, 2-thiouracil, 4-thiouracil, 5-methyluracil, methyl uracil-5-oxyacetic acid, uracil-5-oxyacetic acid, 5-methyl-2-thiouracil, 3-(3-amino-3-N-2-carboxypropyl)uracil, or 2,6-diaminopurine, one or more of these; More preferably, the non-natural sugar moiety is selected from the following group of modifications at the 2' position: OH; substituted lower alkyl, alkylaryl, aralkyl, O-alkylaryl, O-aralkyl, SH, SCH3, OCN, Cl, Br, CN, CF3, OCF3, SOCH3, SO2CH3, ONO2, NO2, N3, NH2F; O-alkyl, S-alkyl, N-alkyl, O-alkenyl, S-alkenyl, N-alkenyl, O-ynyl, S-ynyl, N-ynyl, O-alkyl-O-alkyl, 2'-F, 2'-OCH3, 2'-O(CH2)2OCH3, wherein the alkyl, alkenyl, and ynyl groups can be substituted or unsubstituted C1-C10 alkyl, C2-C10 alkenyl, C2-C10 ynyl, -O[(CH2)] n O] m CH3, -O(CH2) n OCH3, -O(CH2) n NH2, -O(CH2) n CH3, -O(CH2) n -ONH2 and -O(CH2) n ON[(CH2) n CH3)]2, wherein n and m are from 1 to 10; and / or one or more of the following group of modifications at the 5' position: 5'-vinyl, 5'-methyl (R or S), and at the 4' position: 4'-S, heterocyclic alkyl, heterocyclic aryl, aminoalkylamino, polyalkylamino or substituted silyl. More preferably, the substance that can be linked to nucleic acids can be one or more of amino acids, polypeptides, carbohydrates, proteins or lipids; More preferably, the amino acid can be a natural amino acid or a non-natural amino acid; Preferably, the non-natural amino acids include one or more of the following: 2-aminoisobutyric acid (Aib), imidazole-4-acetate (IA), imidazole propionic acid (IPA), α-aminobutyric acid (Abu), tert-butylglycine (Tle), 3-aminomethylbenzoic acid, anthranilic acid, deaminohistidine, β-alanine, 2-aminohistidine, β-hydroxyhistidine, homohistidine, Nα-acetylhistidine, α-fluoro-methylhistidine, α-methylhistidine, α,α-dimethylglutamic acid, m-CF3-phenylalanine, α,β-diaminopropionic acid, 3-pyridylalanine, 2-pyridylalanine, 4-pyridylalanine, (1-aminocyclopropyl)carboxylic acid, (1-aminocyclobutyl)carboxylic acid, (1-aminocyclopentyl)carboxylic acid, (1-aminocyclohexyl)carboxylic acid, (1-aminocycloheptyl)carboxylic acid, or (1-aminocyclooctyl)carboxylic acid.
15. An RNA obtained by any one of the methods described in claims 1-14.
16. An application of the method according to any one of claims 1-14 or the RNA according to claim 15, characterized in that, The applications include: i) Used as an RNA vaccine; ii) Used as an RNA drug; iii) Used to construct RNA databases; iv) Used for guide RNA or other RNAs derived from guide RNA, such as pegRNA, etc. v) Used for linear amplification of target DNA; or vi) is used to construct sequencing libraries.
17. A method for preparing RNA, characterized in that, The preparation method includes using a single-stranded DNA containing a linear transcription functional specific sequence at the 5' end as a template to perform in vitro transcription to obtain RNA; the linear transcription functional specific sequence includes the T7 promoter positive strand, and the 3' end of the T7 promoter positive strand is added to the single-stranded DNA.
18. The preparation method according to claim 17, characterized in that, The single-stranded DNA added to the 3' end of the T7 promoter's positive strand is either a directly synthesized sequence or the 5' end of a single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand. Preferably, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment; The 5' end of the nucleic acid functional fragment is linked to the 3' end of the T7 promoter positive strand, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA; Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
19. The preparation method according to claim 17 or 18, characterized in that, The preparation method includes: after determining the RNA sequence, transcribing DNA that can be transcribed into the RNA; Preferably, the RNA can be linear or circular.
20. A method for building a database, characterized in that, The library construction method includes constructing a library using the product obtained from in vitro transcription according to any one of claims 1-14.
21. The database construction method according to claim 20, characterized in that, The database construction method includes: A) The product obtained from the in vitro transcription is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted enrichment and / or amplification is performed to obtain a library; or... B) The product obtained from the in vitro transcription is separated to obtain template DNA and transcription product RNA; wherein... RNA is further reverse transcribed into cDNA, ligated with adapters required for sequencing, and / or targeted enrichment and / or amplification are performed to obtain a library. Template DNA is transcribed in vitro.
22. The database construction method according to claim 21, characterized in that, The template DNA can be used directly as a template for in vitro transcription without any treatment, or it can be specially treated before in vitro transcription. Preferably, the special treatment includes DNase digestion, or separation, purification, and denaturation into single strands for direct in vitro transcription as a template, or other treatments, such as transformation with one or more of the following modifications: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), N6-methyladenine (N6-mA), 7-methylguanine (7-mG), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), dihydrouracil (DHU), 5-(β-glucoseoxymethylated)cytosine (5gmC), etc.
23. A sequencing method, characterized in that, The sequencing method includes sequencing a library obtained using the library preparation method described in any one of claims 20-22.
24. A combined sequencing method using the same sample for two or more sequencing operations, characterized in that, The combined sequencing method includes: A) Obtain DNA; C) Using the 5' end of the single-stranded DNA obtained in step B) containing a linear transcription function-specific sequence as a template, in vitro transcription is performed, wherein the linear transcription function-specific sequence includes the T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA; D) Separate the products obtained from in vitro transcription to obtain template DNA and transcription product RNA; E1) The RNA obtained in D) is further reverse transcribed into cDNA, ligated with the adapters required for sequencing, and then sequenced directly, or after ligating the adapters required for sequencing, it is amplified and then sequenced, or it is targeted enriched and then library constructed and sequenced, or library constructed and then targeted enriched and sequenced. E2) The template DNA obtained in D) is processed in steps B)-C), and then the product obtained from in vitro transcription is further reverse transcribed into cDNA, ligated with the adapter required for sequencing, and sequenced directly, or after ligating the adapter required for sequencing, it is amplified and sequenced, or targeted enriched and then library constructed and sequenced, or library constructed and then targeted enriched and sequenced. E3) After performing steps B), C), and D) on the template DNA obtained in D), proceed to steps E1) and / or E2); Preferably, the sequencing in steps E1) and E2) can be the same sequencing method or different sequencing methods.
25. The combinatorial sequencing method according to claim 24, characterized in that, The combined sequencing method also includes phosphorylation of the 5' end of the DNA; Preferably, the phosphorylation can be performed before step A), before denaturation in step B), or after denaturation in step B).
26. The combinatorial sequencing method according to claim 24, characterized in that, In step E2), the template DNA may not need to be processed, or it may be specially processed before proceeding to steps B)-C). Preferably, the special treatment includes DNase digestion, or separation, purification, and denaturation into single strands for direct in vitro transcription as a template, or other treatments, such as transformation with one or more of the following modifications: 5-methylcytosine (5mC), 5-hydroxymethylcytosine (5hmC), N6-methyladenine (N6-mA), 7-methylguanine (7-mG), 5-hydroxymethylcytosine (5hmC), 5-formylcytosine (5fC), 5-carboxycytosine (5caC), dihydrouracil (DHU), 5-(β-glucoseoxymethylated)cytosine (5gmC), etc.
27. The library construction method according to claim 21 or the combinatorial sequencing method according to claim 24, characterized in that, The sequencing is selected from one or more of the following sequencing methods: whole genome sequencing, metagenomic sequencing, de novo sequencing, whole exome sequencing, 16S sequencing, methylation sequencing, hydroxymethylation sequencing, single-cell sequencing, chromosome accessibility sequencing, or targeted sequencing.
28. The application of a single-stranded DNA containing a linear transcription function-specific sequence at its 5' end as a template in in vitro transcription, wherein the linear transcription function-specific sequence includes a T7 promoter sense strand, and the 3' end of the T7 promoter sense strand is added to the single-stranded DNA.
29. The combination sequencing method according to claim 24 or the application according to claim 28, characterized in that, The single-stranded DNA added to the 3' end of the T7 promoter's positive strand is either a directly synthesized sequence or the 5' end of a single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand. Preferably, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment; The 5' end of the nucleic acid functional fragment is linked to the 3' end of the T7 promoter positive strand, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA; Preferably, the 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA.
30. A reaction system for linking specific sequences of linear transcription function, characterized in that, The reaction system includes: The single-stranded DNA, a linear transcription functionally specific sequence, and a linker component; wherein the linear transcription functionally specific sequence includes a T7 promoter positive strand, and the linker component links the 3' end of the T7 promoter positive strand to the 5' end of the single-stranded DNA.
31. The reaction system according to claim 30, characterized in that, The linker component links the 3' end of the positive strand of the T7 promoter to the 5' end of the single-stranded DNA directly or indirectly (e.g., through a functional nucleic acid fragment); Preferably, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment; The linker connects the 5' end of the nucleic acid functional fragment to the 3' end of the T7 promoter positive strand, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA. The 3' end sequence of the nucleic acid functional fragment is single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA; Preferably, the ligation component includes ligases or reagents required for chemical ligation methods.
32. An in vitro transcription reaction system, characterized in that, The reaction system includes: The single-stranded DNA contains a linear transcription function-specific sequence at its 5' end as a template, an NTP, and a T7 RNA polymerase; wherein the linear transcription function-specific sequence includes the T7 promoter positive strand, and the 3' end of the T7 promoter positive strand is added to the single-stranded DNA.
33. The reaction system according to claim 32, characterized in that, The single-stranded DNA added to the 3' end of the T7 promoter's positive strand is either a directly synthesized sequence or the 5' end of the single-stranded DNA is directly or indirectly (e.g., through a nucleic acid functional fragment) linked to the 3' end of the T7 promoter's positive strand. Preferably, the linear transcription functional specific sequence includes the T7 promoter positive strand and a nucleic acid functional fragment; The 5' end of the nucleic acid functional fragment is linked to the 3' end of the positive strand of the T7 promoter, and the 3' end sequence of the nucleic acid functional fragment is linked to the 5' end of the single-stranded DNA. The 3' end sequence of the nucleic acid functional fragment is preferably single-stranded or double-stranded, such as double-stranded DNA, single-stranded DNA, single-stranded RNA, DNA containing RNA, or RNA containing DNA. More preferably, the 3' end sequence of the nucleic acid functional fragment is a single-stranded RNA, and the 3' end of the single-stranded RNA is connected to the 5' end of the single-stranded DNA; Preferably, the NTP is selected from one or more of ATP, UTP, GTP, CTP, modified nucleotides, or non-natural nucleotides.