Transposase library building method for reducing linker dimers in library
By modifying the primers and polymerases in the transposase library construction process, the problem of excessive linker dimers in the low-input samples was solved, and the effect of improving sequencing quality and reducing data loss was achieved.
Patent Information
- Application Number
- CN202510098734.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-06-13
AI Technical Summary
During the transposase library construction process, low-input samples are prone to produce linker dimers of specific structures, resulting in a decrease in sequencing quality and loss of data volume.
The primers and polymerase used in the library amplification step are modified so that the amplification step can only amplify the DNA-linker ligation product and not the linker dimer. Specific measures include the use of optimized N5 and N7 primers and binding to polymerases without 3'-5' exonuclease activity.
It effectively reduces the proportion of linker dimers in total library output, improves sequencing quality, significantly reduces the emergence of prickly peak libraries, and improves the integrity and quality of sequencing data.
Smart Images

Figure CN120137962A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biotechnology, and particularly to a transposase library construction method for reducing adapter dimers in a library and related products. Background Art
[0002] DNA library construction is the basis of next-generation sequencing technology, aiming to prepare target DNA into a form compatible with a sequencer. Traditional library construction methods include ultrasonic fragmentation and enzymatic fragmentation. Their basic processes are: (1) Fragmentation: Long-fragment DNA is fragmented using ultrasound or restriction endonucleases; (2) End repair: The ends of DNA fragments are repaired by end repair enzymes to facilitate subsequent ligation reactions; (3) Adapter ligation: DNA fragments are ligated to adapters compatible with the sequencer to form DNA-adapter ligation products; (4) PCR amplification: The DNA-adapter ligation products are amplified by PCR technology to obtain a sufficient library quantity; (5) Library purification: The library is purified by methods such as purification magnetic beads to remove impurities in the system. However, both mechanical and enzymatic methods have inevitable drawbacks. For example, mechanical fragmentation requires a large amount of DNA input and causes serious mechanical damage, and enzymatic methods cause severe bias.
[0003] With the discovery of transposons, the TN5 enzymatic library construction method emerged. With only its terminal core sequence, the transposase can insert and ligate this part of the sequence into the genome (see Figure 1 : Transposon principle; doi:10.3390 / ijms21218329). The transposon terminal sequence can bind to partial sequences of P5 and P7 at both ends of the sequencing adapter to form a coated adapter, and then form a TN5 transposon complex with the transposase. It can fragment genomic DNA while ligating a partial adapter sequence containing the P5 end at one end and a partial adapter sequence containing the P7 end at the other end. Finally, PCR is used to fill the gap between the adapter and the fragmented DNA and add the remaining part of the adapter, thus forming a complete library, greatly reducing the template input amount and shortening the library construction time.
[0004] Although the TN5 enzymatic library construction method greatly improves the library construction efficiency, there are still certain requirements for the template during library construction. For example, it has relatively high requirements for the purity of DNA and the accuracy of DNA concentration. Otherwise, it will affect the fragmentation effect, and adapter dimers are likely to be generated during library construction when the DNA quality is poor or the input amount is too low. When a library with adapter dimers is used for next-generation sequencing, ① the adapter dimers will bind to the anchor sequence on the flow cell of the sequencer and form clusters through bridge PCR amplification, thereby reducing the effective data output of sequencing; ② the adapter dimer sequences are short and will be preferentially amplified during long cluster formation. Moreover, the sequences are fixed, the base complexity is low, and the length is short, which will reduce the Q30 of sequencing and affect the filtering rate of clean reads; ③ as the adapter content increases, the data output drops sharply, resulting in data loss. Summary of the Invention
[0005] The present invention creatively discovers that adapter dimers with a specific structure will appear when constructing a library using a transposase with a low input amount of sample. The purpose of the present invention is to provide a transposase library construction method and related products for reducing adapter dimers in the library. The method of the present invention modifies the primers used in the library amplification step and optimizes the polymerase used in the library amplification step, so that the amplification step can only amplify the DNA-adapter ligation product and cannot amplify the adapter dimer. Using the method of the present invention is beneficial to reducing the proportion of adapter dimers in the total library output and improving the sequencing quality.
[0006] The first aspect of the present invention provides an oligonucleotide, which consists of a 5'-terminal variable region and a 3'-terminal invariant region. Among them, the 5'-terminal variable region contains a sequencing primer sequence, and the 3'-terminal invariant region is composed of a sequence obtained by deleting n nucleotides from the 3' end of the transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA, or TATA, where n is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence starting from the 3' end of the transposase recognition core sequence +1.
[0007] In some embodiments, the sequencing primer sequence is a sequencing primer sequence of the Illumina platform, such as read1 or read2 of the Illumina platform, or for example, the Truseq sequencing primer or Nextera sequencing primer of the Illumina platform, or for example, Truseq read1 or Truseq read2 or Nextera read1 or Nextera read2 of the Illumina platform.
[0008] In some embodiments, the sequencing primer is as shown in SEQ ID NO.5 or 6.
[0009] In some embodiments, the sequencing primer sequence is the sequencing primer sequence of the MGI platform, such as the first-strand sequencing primer or the second-strand sequencing primer of the MGI platform, or for example, the first-strand sequencing primer IP1 or the second-strand sequencing primer IP2 of the MGI platform, or for example, the first-strand sequencing primer Read1 or the second-strand sequencing primer Read2 of the MGI platform.
[0010] In some embodiments, the sequencing primer sequence is the sequencing primer sequence of the Ion Torrent platform.
[0011] In some embodiments, the sequencing primer sequence is the complete or partial sequence of the foregoing sequencing primer sequence.
[0012] In the present invention, a sequencing primer refers to a known nucleotide sequence, which is used for subsequent on-machine sequencing, and its length can also be adjusted adaptively, as long as the complete sequencing primer sequence can be obtained by PCR to achieve sequencing. The sequencing primer can be adjusted according to different sequencing platforms, and the specific sequence is not limited and is within the protection scope of the present invention.
[0013] In some embodiments, n is at least 0. In some embodiments, n is 0, 1 or 2, and preferably, n is 0.
[0014] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase; in some embodiments, the Tn5 transposase recognition core sequence is as shown in SEQ ID NO.7.
[0015] In some embodiments, when the Tn5 transposase recognition core sequence is as shown in SEQ ID NO.7, the restriction sequence closest to the 3'-end of the transposase recognition core sequence is GAGA, the position number of the first nucleotide at the 3'-end of GAGA starting from the 3'-end of the transposase recognition core sequence +1 is 5, and n is at most 5. In some embodiments, the 3'-end invariant region consists of at least 14, 15, 16, 17, 18 and all 19 bases of ME starting from the 5'-end.
[0016] In some embodiments, the oligonucleotide is single-stranded.
[0017] In some embodiments, the oligonucleotide is a completely complementary double-strand or a partially complementary double-strand.
[0018] In some embodiments, at least m bases at the 3'-end of the 3'-end invariant region are stabilized and modified, where m is 1, 2, 3 or more.
[0019] In some embodiments, the stabilization modification refers to a modification that can resist exonuclease degradation while not affecting the extension in the 5'-3' direction, including but not limited to all modification types with the aforementioned functions known to those of ordinary skill in the art.
[0020] In some embodiments, the stabilization modification is selected from thiolation modification and amino modification, preferably thiolation modification.
[0021] In some embodiments, the amino modification is any one of NH2 C7 and NH2 C6.
[0022] In some embodiments, the 5'-terminal variable region further comprises a sequencing solid-phase binding sequence located at the 5'-end of the sequencing primer sequence and optionally a tag sequence located between the sequencing solid-phase binding sequence and the sequencing primer sequence.
[0023] In some embodiments, the sequencing solid-phase binding sequence is the sequencing solid-phase binding sequence of the Illumina platform, such as P5 or P7 or the reverse complementary sequence of P5 or the reverse complementary sequence of P7.
[0024] In some embodiments, the sequencing solid-phase binding sequence is as shown in SEQ ID NO.8 or 9, or for example, as shown by the reverse complementary sequence of SEQ ID NO.8 or 9.
[0025] In some embodiments, the tag sequence is a random sequence with a length of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides. In some embodiments, the tag sequence is color-balanced.
[0026] In some embodiments, the sequencing solid-phase binding sequence is P5 and the sequencing primer sequence is read2, or the sequencing solid-phase binding sequence is P5 and the sequencing primer sequence is read1, or the sequencing solid-phase binding sequence is P7 and the sequencing primer sequence is read1, or the sequencing solid-phase binding sequence is P7 and the sequencing primer sequence is read2.
[0027] In some embodiments, the sequencing solid-phase binding sequence is the reverse complementary sequence of P5 and the sequencing primer sequence is read2, or the sequencing solid-phase binding sequence is the reverse complementary sequence of P5 and the sequencing primer sequence is read1, or the sequencing solid-phase binding sequence is the reverse complementary sequence of P7 and the sequencing primer sequence is read1, or the sequencing solid-phase binding sequence is the reverse complementary sequence of P7 and the sequencing primer sequence is read2.
[0028] In some embodiments, the sequencing solid-phase binding sequence is as shown in SEQ ID NO.8 and the sequencing primer sequence is as shown in SEQ ID NO.5, or the sequencing solid-phase binding sequence is as shown in SEQ ID NO.9 and the sequencing primer sequence is as shown in SEQ ID NO.6, or the sequencing solid-phase binding sequence is as shown in SEQ ID NO.9 and the sequencing primer sequence is as shown in SEQ ID NO.5, or the sequencing solid-phase binding sequence is as shown in SEQ ID NO.8 and the sequencing primer sequence is as shown in SEQ ID NO.6.
[0029] In some embodiments, the sequencing solid-phase binding sequence is the ligation adapter sequence of the MGI platform, such as the upstream ligation adapter sequence of the MGI platform or the upstream ligation adapter sequence of the MGI platform, or for example, the linear sequence of the vesicle adapter of the MGI platform or the vesicular sequence of the vesicle adapter.
[0030] In some embodiments, the sequencing solid-phase binding sequence is as shown in SEQ ID NO.10 or 11, or for example, as shown in the reverse complementary sequence of SEQ ID NO.10 or 11.
[0031] In some embodiments, the sequencing solid-phase binding sequence is the sequencing solid-phase binding sequence of the Ion Torrent platform, such as the A adapter sequence or the P1 adapter sequence of the ion torrent platform.
[0032] The second aspect of the present application provides an oligonucleotide pair, which is composed of an upstream oligonucleotide and a downstream oligonucleotide, wherein,
[0033] (1) The upstream oligonucleotide is composed of an upstream 5'-end variable region and an upstream 3'-end invariant region. Among them, the upstream 5'-end variable region contains an upstream sequencing primer sequence, and the upstream 3'-end invariant region is composed of a sequence obtained by deleting n1 nucleotides from the 3'-end of the transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA or TATA, where n1 is at most the position number of the 3'-end first nucleotide of the restriction sequence closest to the 3'-end of the transposase recognition core sequence starting from the 3'-end of the transposase recognition core sequence +1.
[0034] (2) The downstream oligonucleotide consists of a variable region at the 5'-end of the downstream and a constant region at the 3'-end of the downstream. Among them, the variable region at the 5'-end of the downstream contains a downstream sequencing primer sequence, and the constant region at the 3'-end of the downstream consists of a sequence obtained by deleting n2 nucleotides from the 3'-end of the transposase recognition core sequence, where n2 is at most the position number of the first nucleotide at the 3'-end of the restriction sequence closest to the 3'-end of the transposase recognition core sequence starting from the 3'-end of the transposase recognition core sequence + 1.
[0035] In some embodiments, the upstream sequencing primer sequence and the downstream sequencing primer sequence are the upstream sequencing primer sequence and the downstream sequencing primer sequence of the Illumina platform.
[0036] In some embodiments, the upstream sequencing primer sequence is read1 and the downstream sequencing primer sequence is read2, and vice versa.
[0037] In some embodiments, the upstream sequencing primer sequence is as shown in SEQ ID NO.5 and the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and vice versa.
[0038] In some embodiments, the upstream sequencing primer sequence and the downstream sequencing primer sequence are the upstream sequencing primer sequence and the downstream sequencing primer sequence of the MGI platform. For example, the upstream sequencing primer sequence is a single-strand sequencing primer and the downstream sequencing primer sequence is a double-strand sequencing primer, and vice versa; for another example, the upstream sequencing primer sequence is a single-strand sequencing primer IP1 and the downstream sequencing primer sequence is a double-strand sequencing primer IP2, and vice versa; for another example, the upstream sequencing primer sequence is a single-strand sequencing primer Read1 and the downstream sequencing primer sequence is a double-strand sequencing primer Read2, and vice versa.
[0039] In some embodiments, at least m1 bases at the 3'-end of the upstream constant region at the 3'-end are stabilized and modified, where m1 is 1, 2, 3 or more. In some embodiments, the stabilization modification is selected from thiolation modification and amino modification, preferably thiolation modification.
[0040] In some embodiments, at least m2 bases at the 3'-end of the downstream constant region at the 3'-end are stabilized and modified, where m2 is 1, 2, 3 or more. In some embodiments, the stabilization modification is selected from thiolation modification and amino modification, preferably thiolation modification.
[0041] In some embodiments, the amino modification is any one of NH2 C7 and NH2 C6.
[0042] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase. In some embodiments, the Tn5 transposase recognition core sequence is as shown in SEQ ID NO.7.
[0043] In some embodiments, n1 and n2 are independently 0, 1, or 2. Preferably, both n1 and n2 are 0. In some embodiments, n1 and n2 are the same or different.
[0044] In some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are single-stranded; in some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are fully complementary double-stranded; in some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are partially complementary double-stranded; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is fully complementary double-stranded, and vice versa; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is partially complementary double-stranded, and vice versa; in some embodiments, the upstream oligonucleotide is partially complementary double-stranded and the downstream oligonucleotide is fully complementary double-stranded, and vice versa.
[0045] In some embodiments, the upstream sequencing primer and the downstream sequencing primer are the complete or partial sequences of the aforementioned upstream sequencing primer and downstream sequencing primer sequences.
[0046] In the present invention, the upstream sequencing primer and the downstream sequencing primer refer to a known nucleotide sequence, which is used for subsequent on-machine sequencing. Its length can also be adaptively adjusted as long as the complete sequencing primer sequence can be obtained by PCR to achieve sequencing. This sequencing primer can be adjusted according to different sequencing platforms, and the specific sequence is not limited and is within the protection scope of the present invention.
[0047] The third aspect of the present application provides an oligonucleotide pair, which is composed of an upstream oligonucleotide and a downstream oligonucleotide, wherein,
[0048] (1) The upstream oligonucleotide consists of an upstream 5'-variable region and an upstream 3'-constant region. Among them, the upstream 5'-variable region contains an upstream sequencing primer sequence, and the upstream 5'-variable region also contains an upstream sequencing solid-phase binding sequence located at the 5'-end of the upstream sequencing primer sequence and optionally an upstream tag sequence located between the upstream sequencing solid-phase binding sequence and the upstream sequencing primer sequence. And the upstream 3'-constant region is composed of a sequence obtained by deleting n1 nucleotides from the 3'-end of the transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA or TATA, where n1 is at most the position number of the first nucleotide at the 3'-end of the restriction sequence closest to the 3'-end of the transposase recognition core sequence starting from the 3'-end of the transposase recognition core sequence + 1.
[0049] (2) The downstream oligonucleotide consists of a downstream 5'-variable region and a downstream 3'-constant region. Among them, the downstream 5'-variable region contains a downstream sequencing primer sequence, and the downstream 5'-variable region also contains a downstream sequencing solid-phase binding sequence located at the 5'-end of the downstream sequencing primer sequence and optionally a downstream tag sequence located between the downstream sequencing solid-phase binding sequence and the downstream sequencing primer sequence. And the downstream 3'-constant region is composed of a sequence obtained by deleting n2 nucleotides from the 3'-end of the transposase recognition core sequence, where n2 is at most the position number of the first nucleotide at the 3'-end of the restriction sequence closest to the 3'-end of the transposase recognition core sequence starting from the 3'-end of the transposase recognition core sequence + 1.
[0050] In some embodiments, the upstream sequencing primer sequence and the downstream sequencing primer sequence are the upstream sequencing primer sequence and the downstream sequencing primer sequence of the Illumina platform.
[0051] In some embodiments, the upstream sequencing primer sequence is read2 and the downstream sequencing primer sequence is read1, and vice versa.
[0052] In some embodiments, the upstream sequencing primer sequence is as shown in SEQ ID NO.5 and the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and vice versa.
[0053] In some embodiments, the sequencing primer sequence is the sequencing primer sequence of the MGI platform.
[0054] In some embodiments, the upstream sequencing primer sequence and the downstream sequencing primer sequence are the upstream sequencing primer sequence and the downstream sequencing primer sequence of the MGI platform. For example, the upstream sequencing primer sequence is a first-strand sequencing primer and the downstream sequencing primer sequence is a second-strand sequencing primer, and vice versa; for another example, the upstream sequencing primer sequence is a first-strand sequencing primer IP1 and the downstream sequencing primer sequence is a second-strand sequencing primer IP2, and vice versa; for another example, the upstream sequencing primer sequence is a first-strand sequencing primer Read1 and the downstream sequencing primer sequence is a second-strand sequencing primer Read2, and vice versa.
[0055] In some embodiments, at least m1 bases at the 3'-end of the upstream 3'-end invariant region are stabilized modifications, where m1 is 1, 2, 3 or more. In some embodiments, the stabilized modifications are selected from thiol modifications and amino modifications, preferably thiol modifications.
[0056] In some embodiments, at least m2 bases at the 3'-end of the downstream 3'-end invariant region are stabilized modifications, where m2 is 1, 2, 3 or more. In some embodiments, the stabilized modifications are selected from thiol modifications and amino modifications, preferably thiol modifications.
[0057] In some embodiments, the amino modification is any one of NH2 C7 and NH2 C6.
[0058] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase. In some embodiments, the Tn5 transposase recognition core sequence is as shown in SEQ ID NO.7.
[0059] In some embodiments, n1 and n2 are independently 0, 1 or 2, preferably, both n1 and n2 are 0. In some embodiments, n1 and n2 are the same or different.
[0060] In some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are single-stranded; in some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are fully complementary double-stranded; in some embodiments, both the upstream oligonucleotide and the downstream oligonucleotide are partially complementary double-stranded; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is fully complementary double-stranded, and vice versa; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is partially complementary double-stranded, and vice versa; in some embodiments, the upstream oligonucleotide is partially complementary double-stranded and the downstream oligonucleotide is fully complementary double-stranded, and vice versa.
[0061] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream and downstream sequencing solid-phase binding sequences of the Ion Torrent platform. For example, the upstream sequencing solid-phase binding sequence is the A adapter sequence of the Ion Torrent platform and the downstream sequencing solid-phase binding sequence is the P1 adapter sequence of the Ion Torrent platform, and vice versa.
[0062] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the splint sequences of the MGI platform. For example, the upstream sequencing solid-phase binding sequence is the upstream splint sequence of the MGI platform and the downstream sequencing solid-phase binding sequence is the downstream splint sequence of the MGI platform, and vice versa; or for another example, the upstream sequencing solid-phase binding sequence is the linear sequence of the vesicle adapter of the MGI platform and the downstream sequencing solid-phase binding sequence is the vesicular sequence of the vesicle adapter of the MGI platform, and vice versa.
[0063] In some embodiments, the upstream splint sequence of the MGI platform is as shown in SEQ ID NO.10 and the downstream splint sequence of the MGI platform is as shown in SEQ ID NO.11, and vice versa.
[0064] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream and downstream sequencing solid-phase binding sequences of the Illumina platform.
[0065] In some embodiments, the upstream sequencing solid-phase binding sequence is P5, the downstream sequencing solid-phase binding sequence is P7, and vice versa; in some embodiments, the upstream sequencing solid-phase binding sequence is the reverse complementary sequence of P5, and the downstream sequencing solid-phase binding sequence is the reverse complementary sequence of P7, and vice versa.
[0066] In some embodiments, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8 and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9, and vice versa.
[0067] In some embodiments, the upstream tag sequence and the downstream tag sequence are independently random sequences with a length of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides. In some embodiments, the tag sequences are color-balanced.
[0068] In some embodiments, the upstream sequencing solid-binding sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid-binding sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is P5 and the downstream sequencing primer sequence is read1.
[0069] In some embodiments, the upstream sequencing solid-binding sequence is P5 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing solid-binding sequence is P7 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is P5 and the downstream sequencing primer sequence is read2.
[0070] In some embodiments, the upstream sequencing solid-binding sequence is the reverse complement of P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is the reverse complement of P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid-binding sequence is the reverse complement of P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is the reverse complement of P5 and the downstream sequencing primer sequence is read1.
[0071] In some embodiments, the upstream sequencing solid-binding sequence is the reverse complement of P5 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is the reverse complement of P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing solid-binding sequence is the reverse complement of P7 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is the reverse complement of P5 and the downstream sequencing primer sequence is read2.
[0072] In some embodiments, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8, the upstream sequencing primer sequence is as shown in SEQ ID NO.5, the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9; or, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9, the upstream sequencing primer sequence is as shown in SEQ ID NO.6, the downstream sequencing primer sequence is as shown in SEQ ID NO.5, and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8; or, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9, the upstream sequencing primer sequence is as shown in SEQ ID NO.5, the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8; or, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8, the upstream sequencing primer sequence is as shown in SEQ ID NO.6, the downstream sequencing primer sequence is as shown in SEQ ID NO.5, and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9.
[0073] The fourth aspect of the present invention provides a kit, which comprises the oligonucleotide pair, the extension primer pair and optionally a DNA polymerase with reduced or eliminated proofreading function according to the second aspect of the present invention.
[0074] Wherein, the extension primer pair comprises an upstream extension primer and a downstream extension primer.
[0075] Wherein, the upstream extension primer comprises an upstream sequencing solid-phase binding sequence located at the 5'-end, the upstream sequencing primer sequence located at the 3'-end and optionally an upstream tag sequence located between the upstream sequencing solid-phase binding sequence and the upstream sequencing primer sequence, and the downstream extension primer comprises a downstream sequencing solid-phase binding sequence located at the 5'-end, the downstream sequencing primer sequence located at the 3'-end and optionally the downstream tag sequence located between the downstream sequencing solid-phase binding sequence and the downstream sequencing primer sequence.
[0076] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence of the Ion Torrent platform. For example, the upstream sequencing solid-phase binding sequence is the A adapter sequence of the ion torrent platform and the downstream sequencing solid-phase binding sequence is the P1 adapter sequence of the ion torrent platform, and vice versa.
[0077] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the adapter sequences of the MGI platform. For example, the upstream sequencing solid-phase binding sequence is the upstream adapter sequence of the MGI platform and the downstream sequencing solid-phase binding sequence is the downstream adapter sequence of the MGI platform, and vice versa; for example, the upstream sequencing solid-phase binding sequence is the linear sequence of the vesicle adapter of the MGI platform and the downstream sequencing solid-phase binding sequence is the vesicular sequence of the vesicle adapter of the MGI platform, and vice versa.
[0078] In some embodiments, the upstream adapter sequence of the MGI platform is as shown in SEQ ID NO.10 and the downstream adapter sequence of the MGI platform is as shown in SEQ ID NO.11, and vice versa.
[0079] In some embodiments, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream and downstream sequencing solid-phase binding sequences of the Illumina platform.
[0080] In some embodiments, the upstream sequencing solid-phase binding sequence is P5 and the downstream sequencing solid-phase binding sequence is P7, and vice versa; in some embodiments, the upstream sequencing solid-phase binding sequence is the reverse complementary sequence of P5 and the downstream sequencing solid-phase binding sequence is the reverse complementary sequence of P7, and vice versa.
[0081] In some embodiments, the upstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.8 and the downstream sequencing solid-phase binding sequence is as shown in SEQ ID NO.9, and vice versa.
[0082] In some embodiments, the upstream tag sequence and the downstream tag sequence are independently random sequences with a length of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides. In some embodiments, the tag sequences are color balanced.
[0083] In some embodiments, the upstream sequencing solid-phase binding sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-phase binding sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid-phase binding sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-phase binding sequence is P5 and the downstream sequencing primer sequence is read1.
[0084] In some embodiments, the upstream sequencing solid-binding sequence is P5 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing solid-binding sequence is P7 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is P5 and the downstream sequencing primer sequence is read2.
[0085] In some embodiments, the upstream sequencing solid-binding sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid-binding sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read1.
[0086] In some embodiments, the upstream sequencing solid-binding sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read2, the downstream sequencing solid-binding sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing solid-binding sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read1, the downstream sequencing solid-binding sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read2.
[0087] In some embodiments, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO.8, the upstream sequencing primer sequence is as shown in SEQ ID NO.5, the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO.9; or, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO.9, the upstream sequencing primer sequence is as shown in SEQ ID NO.6, the downstream sequencing primer sequence is as shown in SEQ ID NO.5, and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO.8; or, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO.9, the upstream sequencing primer sequence is as shown in SEQ ID NO.5, the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO.8; or, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO.8, the upstream sequencing primer sequence is as shown in SEQ ID NO.6, the downstream sequencing primer sequence is as shown in SEQ ID NO.5, and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO.9.
[0088] In some embodiments, the DNA polymerase with the proofreading function eliminated includes a polymerase that itself has no proofreading function, such as Class A polymerase, and for another example, Taq DNA polymerase. The DNA polymerase with the proofreading function eliminated also includes a DNA polymerase whose proofreading function has been eliminated through biotechnological modification.
[0089] In some embodiments, the DNA polymerase with weakened or eliminated proofreading function is one or more of Taq polymerase, Bst polymerase, Escherichia coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, TfiDNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I, including but not limited to variants, modified products, and derivatives of the aforementioned polymerases.
[0090] In some embodiments, the kit further comprises one or more of a transposase, a transposase complex, an amplification buffer, amplification primers, purification magnetic beads, and pure water.
[0091] The fifth aspect of the present invention provides a kit comprising the oligonucleotide pair according to the third aspect of the present invention and optionally a DNA polymerase with reduced or eliminated proofreading function.
[0092] In some embodiments, the DNA polymerase with eliminated proofreading function includes a polymerase that itself has no proofreading function, such as class A polymerase, and for example, Taq DNA polymerase. The DNA polymerase with eliminated proofreading function also includes a DNA polymerase that has been biotechnologically modified to eliminate the proofreading function.
[0093] In some embodiments, the DNA polymerase with reduced or eliminated proofreading function is one or more of Taq polymerase, Bst polymerase, Escherichia coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, TfiDNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I, including but not limited to variants, modified products, and derivatives of the aforementioned polymerases.
[0094] In some embodiments, the kit further comprises one or more of a transposase, a transposase complex, an amplification buffer, amplification primers, purification magnetic beads, and pure water.
[0095] The sixth aspect of the present invention provides a method for constructing a sequencing library, which includes: (1) using a transposase complex to fragment a target sequence; (2) performing selective PCR using the oligonucleotide pair according to the second aspect of the present invention and optionally the DNA polymerase with reduced or eliminated proofreading function according to the fourth aspect of the present invention; and (3) performing extension PCR using the extension primer pair according to the fourth aspect of the present invention.
[0096] Wherein, the transposase complex is formed by a transposase dimer embedding a first transposon adaptor and a second transposon adaptor, the transposase recognizes the transposase recognition core sequence, the first transposon contains the upstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end and the second transposon contains the downstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end, and vice versa.
[0097] The seventh aspect of the present invention provides a method for constructing a sequencing library, which includes: (1) using a transposase complex to fragment a target sequence; and (2) performing PCR using the oligonucleotide pair according to the third aspect of the present invention and optionally a DNA polymerase with weakened or eliminated correction function according to the fifth aspect of the present invention.
[0098] Wherein, the transposase complex is formed by a transposase dimer embedding a first transposon adaptor and a second transposon adaptor, the transposase recognizes the transposase recognition core sequence, the first transposon contains the upstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end and the second transposon contains the downstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end, and vice versa.
[0099] In some embodiments, the transposase is Tn5 transposase, including the wild type or mutants of Tn5 transposase.
[0100] In some embodiments, the upstream sequencing primer sequence is read1 and the downstream sequencing primer sequence is read2, and vice versa.
[0101] In some embodiments, the first transposon adaptor is a fully complementary double strand. The first transposon contains a first strand of the upstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end, and also contains a second strand of the first transposon that is fully complementary to the first strand of the first transposon; the second transposon adaptor is a fully complementary double strand. The second transposon contains a first strand of the downstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end, and also contains a second strand of the second transposon that is fully complementary to the first strand of the second transposon.
[0102] In some embodiments, the first transposon adapter is a partially complementary double strand. For example, the first transposon adapter is a Y-shaped adapter. The first transposon first strand of the first transposon adapter contains the upstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end. The first transposon second strand of the first transposon adapter further contains the transposase recognition core sequence at the 5'-end complementary to the first transposon first strand and the downstream sequencing primer sequence that is not complementary to the first transposon first strand. The second transposon adapter is a partially complementary double strand. The second transposon adapter is a Y-shaped adapter. The second transposon first strand of the second transposon adapter contains the downstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end. The second transposon second strand of the second transposon adapter further contains the transposase recognition core sequence at the 5'-end complementary to the second transposon first strand and the upstream sequencing primer sequence that is not complementary to the second transposon first strand.
[0103] For another example, the first transposon adapter is a splint adapter. The first transposon first strand of the first transposon adapter contains the upstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end. The first transposon second strand of the first transposon adapter further contains the transposase recognition core sequence at the 5'-end complementary to the first transposon first strand. The second transposon adapter is a splint-type adapter. The second transposon first strand of the second transposon adapter contains the downstream sequencing primer sequence at the 5'-end and the transposase recognition core sequence at the 3'-end. The second transposon second strand of the second transposon adapter further contains the transposase recognition core sequence at the 5'-end complementary to the second transposon first strand.
[0104] In some embodiments, the method further includes a step of filling in gaps. Preferably, a polymerase is used to fill in the 9bp gaps formed by transposase cleavage. Preferably, the polymerase has strand displacement activity.
[0105] In some embodiments, the PCR amplification is a conventional PCR amplification reaction process known to those of ordinary skill in the art. For example, it includes denaturation, annealing, and extension, and more preferably includes pre-denaturation, denaturation, annealing, extension, and full extension.
[0106] Term Explanation:
[0107] Q30: The percentage of bases with a quality ≥ 30 in the sequencing sequence, mainly used to evaluate the accuracy of sequence sequencing.
[0108] Clean reads: The remaining data after filtering the original second-generation sequencing data to remove low-quality data. Brief Description of the Drawings
[0109] Figure 1:Principle of transposase library construction;
[0110] Figure 2 :Library peak graphs of transposase library construction with input amounts of 1 ng and 10 pg respectively;
[0111] Figure 3A :Complete structure of transposase paired-end library;
[0112] Figure 3B :Main adapter spike-in library structure;
[0113] Figure 3C :Structure of Tn5 transposase complex;
[0114] Figure 3D :Proportion of different adapter spike-in structures;
[0115] Figure 4A :Ligation relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and normal transposase library;
[0116] Figure 4B :Ligation relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and main adapter dimer library;
[0117] Figure 4C :Ligation relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and adapter dimer library with 6 bp missing from the 5′-end ME sequence;
[0118] Figure 5 :Library peak graphs of transposase library construction using conventional N5 / N7 primers and optimized N5 / N7 primers (thio-modified) respectively;
[0119] Figure 6 :Proportion of spike-ins obtained by transposase library construction using conventional N5 / N7 primers and optimized N5 / N7 primers (thio-modified) respectively;
[0120] Figure 7 :Library peak graphs of transposase library construction using conventional N5 / N7 primers with conventional polymerase and optimized N5 / N7 primers with polymerase without 3′-5′ exonuclease activity respectively;
[0121] Figure 8 :Base quality distribution graphs of transposase library construction using conventional N5 / N7 primers with conventional polymerase and optimized N5 / N7 primers with polymerase without 3′-5′ exonuclease activity respectively;
[0122] Figure 9 :Library peak graphs of transposase library construction using conventional N5 / N7 primers and optimized N5 / N7 primers (amino-modified) respectively;
[0123] Figure 10 : Proportion of spike peaks obtained by constructing libraries with transposase using conventional N5 / N7 primers and optimized N5 / N7 primers (amino modification) respectively.
[0124] Specific implementation mode (Example)
[0125] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and through specific implementation modes. However, the following examples are only simple examples of the present invention and do not represent or limit the scope of the protection of the present invention. The scope of protection of the present invention shall be subject to the claims.
[0126] In the following examples, unless otherwise specified, the reagents and consumables used are purchased from conventional reagent manufacturers in the art; unless otherwise specified, the experimental methods and technical means used are conventional methods and means in the art.
[0127] All the sequences included in this application are as follows in the table (direction: 5'-3'):
[0128]
[0129]
[0130] Example 1
[0131] Comparison of library peak maps at 1 ng normal input and 10 pg low input.
[0132] 1. DNA fragmentation
[0133] 1.1 Thaw 5×TTBL (Vazyme#TD504) at room temperature, mix well by inverting up and down, and set aside. Confirm that 6×TSB (Vazyme#TD504) and TWB (Vazyme#TD504) are at room temperature, and flick the tube wall to confirm no precipitation; if there is precipitation, heat at 37°C and vortex to mix well until the precipitation dissolves.
[0134] 1.2 Vortex TNB (Vazyme#TD504) briefly to mix well, and then prepare the following reaction system in a PCR tube:
[0135] Table 1
[0136] Component Volume 1 ng 293 gDNA / 10 pg 293 gDNA 1 μl 5×TTBL 10 μl TNB 10 μl Pure water 39 μl
[0137] 1.3 Mix well by inverting up and down, flick to remove air bubbles, and centrifuge briefly to collect the reaction solution to the bottom of the tube.
[0138] 1.4 Place the reaction tube in a PCR instrument and run the following reaction program:
[0139] Table 2
[0140] Temperature Time 55℃ 15 min 4℃ Hold
[0141] 1.5 Immediately after the reaction is completed, add 10 μl of 6×TSB to the product, gently pipette to mix well, and let it stand at room temperature for 5 min.
[0142] 1.6 Place the reaction tube on the magnetic stand. After the solution becomes clear (about 3 min), carefully remove the supernatant.
[0143] 1.7 Remove the reaction tube from the magnetic stand, add 100 μl of TWB, gently pipette to resuspend the magnetic beads, then place the reaction tube on the magnetic stand again. After the solution becomes clear (about 3 min), remove the supernatant.
[0144] 1.8 Repeat step 7 for a total of two rinses. Discard all the supernatant, cover the tube cap to prevent the TNB magnetic beads from drying out.
[0145] 2. PCR enrichment
[0146] 2.1 Prepare the PCR amplification Mix according to the following reaction system and add it to the TNB magnetic beads after the previous rinse. Gently vortex to mix well.
[0147] Table 3
[0148] Component Volume 2×TAM (Vazyme#TD504) 25 μl N5XX (Vazyme#TD202) 5 μl N7XX (Vazyme#TD202) 5 μl Pure water 15 μl Total 50 μl
[0149] 2.2 Gently flick to remove air bubbles and briefly centrifuge to collect the reaction solution at the bottom of the tube.
[0150] 2.3 Place the reaction tube in the PCR instrument and run the following reaction program:
[0151] Table 4
[0152]
[0153]
[0154] 3. Purification of the amplified product length
[0155] 3.1 After the library amplification is completed, briefly centrifuge the reaction tube. Vortex the VAHTS DNA Clean Beads (Vazyme#N411) to mix well and pipette 60 μl into the 50 μl PCR product. Use the pipette to blow and mix 10 times well, and incubate at room temperature for 5 min.
[0156] 3.2 Briefly centrifuge the reaction tube and place it on the magnetic stand. After the solution becomes clear (about 5 min), carefully remove the supernatant.
[0157] 3.3 Keep the reaction tube on the magnetic stand at all times, add 200 μl of freshly prepared 80% ethanol to wash the magnetic beads, incubate at room temperature for 30 sec, and carefully remove the supernatant.
[0158] 3.4 Repeat step 3.3 for a total of two washes.
[0159] 3.5 Keep the reaction tube on the magnetic stand at all times, open the lid and air-dry the magnetic beads for about 3 min.
[0160] 3.6 Remove the reaction tube from the magnetic stand, add 22 μl of ddH 2 O for elution. Pipette up and down 10 times to mix well, and incubate at room temperature for 5 min.
[0161] 3.7 Briefly centrifuge the reaction tube and place it on the magnetic stand. After the solution becomes clear (about 5 min), carefully pipette 20 μl of the supernatant into a new PCR tube and store it at -20 °C.
[0162] 3.8 Perform length distribution detection using an Agilent 2100 Bioanalyzer, and the results are as Figure 2 shown.
[0163] Result analysis: As Figure 2 shown, the library peak shape is normal at an input amount of 1 ng, and obvious small-fragment spikes appear at a low input amount of 10 pg.
[0164] Example 2
[0165] The sample input amount is 10 pg of 293 gDNA. The library construction steps are the same as in Example 1. After the amplified library is purified, agarose nucleic acid electrophoresis is performed. The spikes are gel-extracted using (Vazyme#DC301) and then cloned and sequenced through (Vazyme#C603) to analyze the spike structure.
[0166] The structure of a normal double-ended transposase library is as follows, as Figure 3A shown, where N7 is the P7 sequencing adapter sequence, index sequence (i7), and sequencing primer sequence (read2), ME is the forward insertion sequence of the transposase recognition core sequence, Insert is the insertion sequence, ME’ is the reverse insertion sequence of the transposase recognition core sequence, and N5 is the P5 sequencing adapter sequence, index sequence (i5), and sequencing primer sequence (read1).
[0167] P5—i5—read1—ME—Insert—ME’—read2—i7—P7
[0168] The main structure of the adapter spike library is as follows, as Figure 3BAs shown, where N7 is the P7 sequencing adapter sequence, index sequence (i7), and sequencing primer sequence (read2), ME is the reverse insertion sequence of the transposase recognition core sequence, ME’-6bp is the reverse insertion sequence of the transposase recognition core sequence lacking 6bp at the 5', and N5 is the P5 sequencing adapter sequence, index sequence (i5), and sequencing primer sequence (read1).
[0169] P5—i5—read1—ME—ME’-6bp—read2—i7—P7
[0170] As Figure 3C shown, the Tn5 transposase complex is formed by embedding a transposase dimer and two pairs of annealed adapters. The adapter structures respectively include the transposase recognition core sequence (ME sequence and ME’ sequence) of the double-stranded part and the sequencing primer sequences (read1 and read2 respectively) of the single-stranded part. After the transposase breaks the target DNA, the ME sequence and the sequencing primer sequence will be ligated to both ends of the insertion sequence (Insert). Subsequently, the complete library structure is introduced to both ends of the target DNA by amplification with N5 primer and N7 primer respectively, obtaining the complete transposase double-end library structure as Figure 3A shown. And through sequencing of the adapter spike library, its main structure is as Figure 3B shown. It does not contain the insertion sequence (Insert), and the proportion of the ME' sequence lacking 6bp near the 3' end is the highest, accounting for more than 85% of the entire spike library structure. The situations of the remaining small amounts of structures are as follows: ① The ME sequence near the 5' end lacks 6bp, ② The ME sequence near the 5' end lacks 4 or 5bp, ③ The ME' sequence near the 3' end lacks 4 or 5bp, etc. The total proportion is statistically shown as Figure 3D shown. Through analysis, since the transposase has a preference for cutting the TCTC and TATABOX regions, and the ME sequence contains this region, at low input amounts, some transposases will use the adapter containing the ME sequence in the system as a template for cutting, and then form the adapter dimer through amplification with N5 and N7 amplification primers.
[0171] Example 3
[0172] The amplification primers were modified to reduce adapter dimers in the library. Among them, the conventional N5 and N7 primer structures are respectively:
[0173] N5: From the 5' end to the 3' end are the P5 sequencing adapter sequence, Index 2 (i5) sequence, and read1 sequence. N7: From the 5' end to the 3' end are the P7 sequencing adapter sequence, Index 1 (i7) sequence, and read2 sequence. The modified N5 and N7 primer structures are respectively:
[0174] Optimized N5 primer: from the 5'-end to the 3'-end are the P5 sequencing adapter sequence, Index 2 (i5) sequence, read1 sequence, and ME sequence respectively, with a thiophosphate modification on the 3 terminal bases of the 3'-end.
[0175] Optimized N7 primer: from the 5'-end to the 3'-end are the P7 sequencing adapter sequence, Index 1 (i7) sequence, read2 sequence, and ME' sequence respectively, with a thiophosphate modification on the 2 terminal bases of the 3'-end.
[0176] The sample input amount is 10 pg of 293 gDNA. The transposase library construction steps for the conventional N5 and N7 primers are the same as in Example 1. The transposase library construction steps for the modified N5 and N7 primers are as follows:
[0177] 1. DNA fragmentation
[0178] 1.1 Thaw 5×TTBL (Vazyme#TD504) at room temperature, invert it up and down to mix well, and set aside. Confirm that 6×TSB (Vazyme#TD504) and TWB (Vazyme#TD504) are at room temperature, and flick the tube wall gently to confirm no precipitation; if there is precipitation, heat it at 37°C and vortex it to mix well until the precipitation dissolves.
[0179] 1.2 Vortex TNB (Vazyme#TD504) briefly to mix well, and then prepare the following reaction system in a PCR tube:
[0180] Table 5
[0181] Component Volume 10 pg 293 gDNA X μl 5×TTBL 10 μl TNB 10 μl Pure water 39 μl
[0182] 1.3 Invert it up and down to mix well, flick to remove air bubbles, and centrifuge briefly to collect the reaction solution to the bottom of the tube.
[0183] 1.4 Place the reaction tube in a PCR instrument and run the following reaction program:
[0184] Table 6
[0185] Temperature Time 55℃ 15 min 4℃ Hold
[0186] 1.5 Immediately after the reaction is completed, add 10 μl of 6×TSB to the product, gently pipette to mix well, and let it stand at room temperature for 5 min.
[0187] 1.6 Place the reaction tube on a magnetic rack, and carefully remove the supernatant after the solution becomes clear (about 3 min).
[0188] 1.7 Remove the reaction tube from the magnetic rack, add 100 μl of TWB, gently pipette to resuspend the magnetic beads, then place the reaction tube on the magnetic rack, and remove the supernatant after the solution becomes clear (about 3 min).
[0189] 1.8 Repeat step 7 for a total of two rinses. Discard all the supernatants completely, cover the tube lid to prevent the TNB magnetic beads from drying out and cracking.
[0190] 2. PCR enrichment
[0191] 2.1 Prepare the PCR amplification Mix according to the following reaction system and add it to the TNB magnetic beads after the previous rinse. Gently vortex to mix well.
[0192] Table 7
[0193] Component Volume 2×TAM (Vazyme#TD504) 25 μl Optimized N5 primer 5 μl Optimized N7 primer 5 μl Pure water 15 μl Total 50 μl
[0194] Specifically, the optimized N5 primer sequence used in this step is (SEQ ID NO.1): 5'-AATG ATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTCAGA TGTGTATAAGAGA*C*A*G-3' (where * represents thiol modification)
[0195] The optimized N7 primer sequence is (SEQ ID NO.2): 5'-CAAGCAGAAGACGGCATA CGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGAC*A*G-3' (where * represents thiol modification)
[0196] 2.2 Gently flick to remove air bubbles and briefly centrifuge to collect the reaction solution at the bottom of the tube.
[0197] 2.3 Place the reaction tube in a PCR instrument and run the following reaction program:
[0198] Table 8
[0199]
[0200] 3. Purification of the amplified product length
[0201] 3.1 After the library amplification is completed, briefly centrifuge the reaction tube. Vortex and mix the VAHTS DNA Clean Beads (Vazyme#N411) well and pipette 60 μl into the 50 μl PCR product. Use a pipette to blow and mix 10 times thoroughly and incubate at room temperature for 5 min.
[0202] 3.2 Briefly centrifuge the reaction tube and place it on a magnetic stand. After the solution becomes clear (about 5 min), carefully remove the supernatant.
[0203] 3.3 Keep the reaction tube on the magnetic stand all the time. Add 200 μl of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 sec, and carefully remove the supernatant.
[0204] 3.4 Repeat step 3.3 for a total of two rinses.
[0205] 3.5 Keep the reaction tube on the magnetic stand at all times, open the lid and air-dry the magnetic beads for about 3 min.
[0206] 3.6 Remove the reaction tube from the magnetic stand, add 22 μl of ddH2O for elution. Use a pipette to pipette up and down 10 times to mix well, and incubate at room temperature for 5 min.
[0207] 3.7 Centrifuge the reaction tube briefly and place it on the magnetic stand. After the solution becomes clear (about 5 min), carefully pipette 20 μl of the supernatant into a new PCR tube and store at -20 °C.
[0208] 3.8 Perform length distribution detection on the Agilent 2100 Bioanalyzer. The results are as Figure 5 shown. Statistically analyze the proportion of the spike peak as Figure 6 shown.
[0209] 3.9 Next-generation sequencing strategy: 10 G / sample, PE150.
[0210] Result analysis:
[0211] The ligation relationships between the primers before and after modification and the normal transposase library and the main adapter dimer library are as Figure 4A and Figure 4B shown. After extending the N5 / N7 primers, they are complementary to the dual-end transposase library sequence and can be amplified normally. However, for the adapter spike library, after primer extension, they cannot be fully complementary to it, so they cannot be amplified. But the polymerase used in the library construction kit usually has 3′-5′ exonuclease activity, and the protruding part of the extended primer will be removed. Therefore, protective modification was carried out on the extended part of the primer, and 2-3 thiol modifications can play a sufficient protective role. Similarly, the modified primer can also effectively prevent the amplification of adapter dimers with a relatively small proportion, such as the ME sequence at the 5′ end lacking 6 bp, as Figure 4C shown. Similarly, for other adapter dimers with a relatively low proportion, such as the ME sequence at the 5′ end lacking 4 bp or 5 bp, and the ME′ sequence at the 3′ end lacking 4 bp or 5 bp, their amplification can be effectively prevented.
[0212] As Figure 5 and Figure 6 shown, the spike peak amount after primer optimization decreased. Statistically analyze the proportion of the spike library before and after modification. The results show that it decreased from 6.32% to 1.66%. Although the spike peak amount after modifying the primer decreased significantly, it could not be completely eliminated. This result indicates that the protective modification cannot protect the modified primer 100%, and its end may still be removed due to the 3′-5′ exonuclease activity of the polymerase.
[0213] Example 4
[0214] On the basis of Example 3, the polymerase for library amplification was further optimized, that is, 2×TAM (Vazyme#TD504) in Table 7 of Example 2 was replaced with Vazyme#PK512 (the 3′-5′ exonuclease activity of the polymerase it contains was greatly weakened), and the remaining steps remained unchanged. The control group was set as the conventional N5 primer and N7 primer (Vazyme#TD202) + the conventional polymerase 2×TAM (Vazyme#TD504), and the test group was the optimized N5 primer (SEQ ID NO.1) and optimized N7 primer (SEQ ID NO.2) + PK512. The sample input amount was 10 pg of 293 gDNA, and the experimental steps were the same as those in Example 2. The library peak diagram was obtained as Figure 7 shown, and the base quality distribution diagram was as Figure 8 shown.
[0215] Result analysis: As Figure 7 , when using the optimized N5 / N7 primers in combination with a polymerase without 3′-5′ exonuclease activity for library construction, the dimer spike peaks were completely eliminated at low input amounts.
[0216] As Figure 8 , the conventional N5 / N7 primers paired with the current conventional polymerase 2×TAM led to an increased base content distribution jitter in read1 bases, reaching 60 bp, and read 2 reaching 30 bp, while the base jitter was maintained within 15 bp after optimization. For PE150 sequencing, with a length of 150 bp sequenced at both ends, theoretically the content of AT and GC in the sequencing results should be equal, but actually there will be a phenomenon of base jitter at the transposase cleavage site. Generally, there will be base jitter at the first 15 bp, resulting in unequal contents of AT and GC. The appearance of the control group's spike library led to more severe base jitter, making the length of the jittered bases longer and reducing the sequencing quality, while the base jitter in the test group returned to normal after eliminating the adapter spike peaks.
[0217] Example 5
[0218] The amplification primers were modified to reduce adapter dimers in the library. Among them, the structures of the conventional N5 and N7 primers were the same as those in Example 3
[0219] The structures of the modified N5 and N7 primers were respectively:
[0220] Optimized N5 primer: From the 5′ end to the 3′ end, it was the P5 sequencing adapter sequence, Index 2 (i5) sequence, read1 sequence, and ME sequence, and the last 2 bases at the 3′ end contained thiol modification
[0221] Optimized N7 primer: from the 5' end to the 3' end are the P7 sequencing adapter sequence, Index 1 (i7) sequence, read2 sequence, and ME' sequence respectively, with an NH2 C7 amino modification on 1 base at the 3' end
[0222] The sample input amount is 10 pg of 293 gDNA. For the transposase library construction steps of the conventional N5 and N7 primers, they are the same as in Example 1. For the transposase library construction steps of the modified N5 and N7 primers, they are the same as in Example 3.
[0223] The specific optimized N5 primer sequence used is (SEQ ID NO.3): 5'-AATGATACGGCGACCACCGAGATCTACACTTCTAGCTTCGTCGGCAGCGTCAGATGTGTATAAGAGAC*A*G-3' (where * is a thiol modification);
[0224] The optimized N7 primer sequence is (SEQ ID NO.4): 5'-CAAGCAGAAGACGGCATACGAGATGTAGAGGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-NH2 C7-3'
[0225] The statistical library peak graph is as Figure 9 shown, and the statistical proportion of spiky peaks is as Figure 10 shown.
[0226] Result description: As Figure 9 - Figure 10 , the spiky peaks of the dimer are significantly reduced after optimization. The proportion of spiky peaks and the output of the spiky peak library are significantly lower than those of the libraries constructed with the conventional N5 / N7 primers, indicating that the NH2C7 modification can achieve results similar to those of the thiol modification.
Claims
1. An oligonucleotide, which consists of a 5' variable region and a 3' constant region, wherein: The 5' variable region comprises a sequencing primer sequence, and optionally, the sequencing primer sequence is a sequencing primer sequence of an Illumina platform or an MGI platform, and further optionally, is read1 or read2 of an Illumina platform, and further optionally, is as shown in SEQ ID NO. 5 or 6, and The 3'-end invariant region is composed of a sequence obtained by deleting n nucleotides from the 3' end of a transposase recognition core sequence comprising at least one restriction sequence selected from TCTC, GAGA or TATA, wherein n is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence + 1 from the 3' end of the transposase recognition core sequence, preferably, n is 0, The transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase, further optionally, as shown in SEQ ID NO.7, preferably, n is 0, 1 or 2, most preferably, n is 0, At least m bases at the 3' end of the 3' end invariant region are stabilizingly modified, wherein m is 1, 2, 3 or more, and further optionally, the stabilizing modification is selected from thio modification and amino modification, and the amino modification is any one of NH2 C7 and NH2 C6, preferably thio modification, Optionally, the oligonucleotide is single stranded.
2. The oligonucleotide according to claim 1, wherein The 5' variable region further comprises a sequencing solid phase binding sequence located at the 5' end of the sequencing primer sequence and an optional tag sequence located between the sequencing solid phase binding sequence and the sequencing primer sequence. Optionally, the sequencing solid phase binding sequence is a sequencing solid phase binding sequence of an Illumina platform, further optionally, it is P5 or P7 of an Illumina platform, further optionally, as shown in SEQ ID NO. 8 or 9, Optionally, the sequencing solid phase binding sequence is a splint sequence of the MGI platform, further optionally, as shown in SEQ ID NO. 10 or 11, Optionally, the tag sequence is a random sequence of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length, Further optionally, the sequencing solid phase binding sequence is P5 and the sequencing primer sequence is read1, or, the sequencing solid phase binding sequence is shown in P7 and the sequencing primer sequence is read2. Further optionally, the sequencing solid phase binding sequence is shown as SEQ ID NO.8 and the sequencing primer sequence is shown as SEQ ID NO.5, or, the sequencing solid phase binding sequence is shown as SEQ ID NO.9 and the sequencing primer sequence is shown as SEQ ID NO.
6.
3. An oligonucleotide pair, which consists of an upstream oligonucleotide and a downstream oligonucleotide, wherein: (1) The upstream oligonucleotide is composed of an upstream 5' end variable region and an upstream 3' end constant region, wherein: The upstream 5' variable region comprises an upstream sequencing primer sequence, and The upstream 3'-end constant region is composed of a sequence obtained by deleting n1 nucleotides from the 3' end of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA or TATA, wherein n1 is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence + 1 from the 3' end of the transposase recognition core sequence, preferably, n1 is 0, at least m1 bases at the 3' end of the upstream 3'-end constant region are stabilizingly modified, wherein m1 is 1, 2, 3 or more, further optionally, the stabilizing modification is selected from thio modification and amino modification, the amino modification is any one of NH2 C7 and NH2 C6, preferably thio modification, and (2) The downstream oligonucleotide consists of a variable region at the downstream 5' end and a constant region at the downstream 3' end, wherein: The downstream 5' end variable region comprises a downstream sequencing primer sequence, and The downstream 3'-end invariant region is composed of a sequence obtained by deleting n2 nucleotides from the 3'-end of the transposase recognition core sequence, wherein n2 is at most the position number of the first nucleotide at the 3'-end of the restriction sequence closest to the 3'-end of the transposase recognition core sequence + 1 from the 3'-end of the transposase recognition core sequence, preferably, n2 is 0. At least m2 bases at the 3' end of the downstream 3' end invariant region are stabilizingly modified, wherein m2 is 1, 2, 3 or more, and further optionally, the stabilizing modification is selected from thio modification and amino modification, and the amino modification is any one of NH2 C7 and NH2 C6, preferably thio modification, Optionally, the upstream sequencing primer sequence and the downstream sequencing primer sequence are the upstream sequencing primer sequence and the downstream sequencing primer sequence of the Illumina platform or the MGI platform, further optionally, the upstream sequencing primer sequence is read1 of the Illumina platform and the downstream sequencing primer sequence is read2 of the Illumina platform, or vice versa, further optionally, the upstream sequencing primer sequence is shown in SEQ ID NO.5 and the downstream sequencing primer sequence is shown in SEQ ID NO.6, or vice versa, The transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase, further optionally, as shown in SEQ ID NO.7, preferably, n1 and n2 are independently 0, 1 or 2, most preferably, n1 and n2 are both 0, Optionally, both the upstream oligonucleotide and the downstream oligonucleotide are single-stranded.
4. The oligonucleotide pair according to claim 3, wherein (1) The upstream 5' end variable region further comprises an upstream sequencing solid phase binding sequence located at the 5' end of the upstream sequencing primer sequence and an optional upstream tag sequence located between the upstream sequencing solid phase binding sequence and the upstream sequencing primer sequence, and (2) The downstream 5' end variable region further comprises a downstream sequencing solid phase binding sequence located at the 5' end of the downstream sequencing primer sequence and an optional downstream tag sequence located between the downstream sequencing solid phase binding sequence and the downstream sequencing primer sequence. in, Optionally, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence of the Illumina platform, further optionally, the upstream sequencing solid-phase binding sequence is P5 of the Illumina platform and the downstream sequencing solid-phase binding sequence is P7 of the Illumina platform, or vice versa, further optionally, the upstream sequencing solid-phase binding sequence is shown in SEQ ID NO.8 and the downstream sequencing solid-phase binding sequence is shown in SEQ ID NO.9, or vice versa, Optionally, the upstream sequencing solid phase binding sequence and the downstream sequencing solid phase binding sequence are the upstream splint sequence and the downstream splint sequence of the MGI platform, further optionally, the upstream splint sequence is as shown in SEQ ID NO.10 and the downstream splint sequence is as shown in SEQ ID NO.11, or vice versa, Optionally, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length, Further optionally, the upstream sequencing solid phase binding sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid phase binding sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid phase binding sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid phase binding sequence is P5 and the downstream sequencing primer sequence is read1, Further optionally, the upstream sequencing solid-phase binding sequence is shown as SEQ ID NO.8, the upstream sequencing primer sequence is shown as SEQ ID NO.5, the downstream sequencing primer sequence is shown as SEQ ID NO.6, and the downstream sequencing solid-phase binding sequence is shown as SEQ ID NO.9; or, the upstream sequencing solid-phase binding sequence is shown as SEQ ID NO.9, the upstream sequencing primer sequence is shown as SEQ ID NO.6, the downstream sequencing primer sequence is shown as SEQ ID NO.5, and the downstream sequencing solid-phase binding sequence is shown as SEQ ID NO.
8.
5. A kit comprising the oligonucleotide pair according to claim 3, an extension primer pair and optionally a DNA polymerase with a weakened or eliminated proofreading function, in, The extension primer pair comprises an upstream extension primer and a downstream extension primer, wherein the upstream extension primer comprises an upstream sequencing solid-phase binding sequence at the 5' end, the upstream sequencing primer sequence at the 3' end, and an optional upstream tag sequence located between the upstream sequencing solid-phase binding sequence and the upstream sequencing primer sequence, and the downstream extension primer comprises a downstream sequencing solid-phase binding sequence at the 5' end, the downstream sequencing primer sequence at the 3' end, and an optional downstream tag sequence located between the downstream sequencing solid-phase binding sequence and the downstream sequencing primer sequence, Optionally, the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence are the upstream sequencing solid-phase binding sequence and the downstream sequencing solid-phase binding sequence of the Illumina platform, further optionally, the upstream sequencing solid-phase binding sequence is P5 of the Illumina platform and the downstream sequencing solid-phase binding sequence is P7 of the Illumina platform, or vice versa, further optionally, the upstream sequencing solid-phase binding sequence is shown in SEQ ID NO.8 and the downstream sequencing solid-phase binding sequence is shown in SEQ ID NO.9, or vice versa, Optionally, the upstream sequencing solid phase binding sequence and the downstream sequencing solid phase binding sequence are the upstream splint sequence and the downstream splint sequence of the MGI platform, further optionally, the upstream splint sequence is as shown in SEQ ID NO.10 and the downstream splint sequence is as shown in SEQ ID NO.11, or vice versa, Optionally, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length, Further optionally, the upstream sequencing solid phase binding sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing solid phase binding sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing solid phase binding sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing solid phase binding sequence is P5 and the downstream sequencing primer sequence is read1, Further optionally, the upstream sequencing solid-phase binding sequence is shown as SEQ ID NO.8, the upstream sequencing primer sequence is shown as SEQ ID NO.5, the downstream sequencing primer sequence is shown as SEQ ID NO.6, and the downstream sequencing solid-phase binding sequence is shown as SEQ ID NO.9; or, the upstream sequencing solid-phase binding sequence is shown as SEQ ID NO.9, the upstream sequencing primer sequence is shown as SEQ ID NO.6, the downstream sequencing primer sequence is shown as SEQ ID NO.5, and the downstream sequencing solid-phase binding sequence is shown as SEQ ID NO.8, Optionally, the DNA polymerase whose proofreading function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, Escherichia coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase and Bsu DNA polymerase I.
6. A kit comprising the oligonucleotide pair according to claim 4 and optionally a DNA polymerase with a weakened or eliminated proofreading function, Optionally, the DNA polymerase whose proofreading function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, Escherichia coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase and Bsu DNA polymerase I.
7. A method for constructing a sequencing library, comprising: (1) Using the transposase complex to disrupt the target sequence; (2) performing selection PCR using the oligonucleotide pair according to claim 3 and optionally the DNA polymerase with weakened or eliminated proofreading function according to claim 5; and (3) performing extension PCR using the extension primer pair according to claim 5, Wherein, the transposase complex is composed of a transposase dimer that encapsulates a first transposon joint and a second transposon joint, the transposase recognizes the transposase recognition core sequence, the first transposon comprises the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon comprises the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.
8. A method for constructing a sequencing library, comprising: (1) Using the transposase complex to disrupt the target sequence; and (2) performing PCR using the oligonucleotide pair according to claim 4 and optionally the DNA polymerase with weakened or eliminated proofreading function according to claim 6, Wherein, the transposase complex is composed of a transposase dimer that encapsulates a first transposon joint and a second transposon joint, the transposase recognizes the transposase recognition core sequence, the first transposon comprises the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon comprises the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.
9. The method according to any one of claims 7-8, wherein the first transposon junction is a completely complementary double-stranded chain, the first transposon comprises the upstream sequencing primer sequence at the 5' end and the first transposon first chain of the transposase recognition core sequence at the 3' end, and further comprises the first transposon second chain that is completely complementary to the first transposon first chain; the second transposon junction is a completely complementary double-stranded chain, the second transposon comprises the downstream sequencing primer sequence at the 5' end and the second transposon first chain of the transposase recognition core sequence at the 3' end, and further comprises the second transposon second chain that is completely complementary to the second transposon first chain; Alternatively, the first transposon adapter is a Y-shaped adapter, comprising a first transposon first chain having the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and further comprising a first transposon second chain having the transposase recognition core sequence at the 5' end complementary to the first transposon first chain and the downstream sequencing primer sequence not complementary to the first transposon first chain; the second transposon adapter is a partially complementary double-stranded chain, the second transposon adapter is a Y-shaped adapter, the second transposon adapter comprises a second transposon first chain having the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and further comprising a second transposon second chain having the transposase recognition core sequence at the 5' end complementary to the second transposon first chain and the upstream sequencing primer sequence not complementary to the second transposon first chain; Alternatively, the first transposon joint is a splint joint, comprising the first transposon first chain with the upstream sequencing primer sequence located at the 5' end and the transposase recognition core sequence located at the 3' end, and further comprising the first transposon second chain with the transposase recognition core sequence located at the 5' end complementary to the first transposon first chain; the second transposon joint is a splint joint, comprising the second transposon first chain with the downstream sequencing primer sequence located at the 5' end and the transposase recognition core sequence located at the 3' end, and further comprising the second transposon second chain with the transposase recognition core sequence located at the 5' end complementary to the second transposon first chain.
Citation Information
Cited By
Single cell transposase library building method
CN121109553A