A transposase library building method that reduces linker dimer in a library

By modifying primers and optimizing polymerases, the problem of adapter dimers in transposase library construction was solved, which improved sequencing quality and data yield, and ensured the effective data yield and Q30 value of the sequencer.

CN122326593APending Publication Date: 2026-07-03NANJING VAZYME BIOTECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING VAZYME BIOTECH CO LTD
Filing Date
2025-01-02
Publication Date
2026-07-03

Smart Images

  • Figure CN122326593A_ABST
    Figure CN122326593A_ABST
Patent Text Reader

Abstract

This invention provides a transposase library construction method and related products for reducing adapter dimers in libraries, belonging to the field of biotechnology. By optimizing the primers and polymerases used in the transposase library construction process, this invention effectively reduces the proportion of adapter dimers in libraries with low input levels, thereby improving sequencing quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of biotechnology, specifically to a method for constructing transposase libraries that reduces linker dimers in a library, and related products. Background Technology

[0002] DNA library construction is the foundation of next-generation sequencing technology, aiming to prepare target DNA in a form compatible with the sequencer. Traditional library construction methods include ultrasonic fragmentation and enzymatic fragmentation. Their basic procedures are: (1) fragmentation: using ultrasound or restriction endonucleases to fragment long DNA fragments; (2) end repair: repairing the ends of DNA fragments with end-repair enzymes to facilitate subsequent ligation reactions; (3) adapter ligation: ligating DNA fragments to adapters compatible with the sequencer to form DNA-adapter ligation products; (4) PCR amplification: amplifying the DNA-adapter ligation products using PCR technology to obtain a sufficient library volume; (5) library purification: purifying the library using methods such as purification beads to remove impurities from the system. However, both mechanical and enzymatic methods have unavoidable shortcomings. For example, mechanical fragmentation requires a large amount of DNA and causes severe mechanical damage, while enzymatic methods suffer from severe bias.

[0003] With the discovery of transposons, the TN5 enzymatic library construction method emerged. Only the terminal core sequence is needed; transposases can insert and ligate this sequence into the genome (see...). Figure 1 Transposition principle (doi:10.3390 / ijms21218329): The terminal sequence of the transposon can bind to the partial sequences of P5 and P7 at both ends of the sequencing adapter to form a coated adapter. Then, it forms the TN5 transposon complex with transposase. It can link the adapter containing the P5 end partial adapter sequence at one end and the adapter containing the P7 end partial adapter sequence at the other end while fragmenting genomic DNA. Finally, the gap between the adapter and the fragmented DNA is filled by PCR and the rest of the adapter is added to form a complete library, which greatly reduces the amount of template input and shortens the library construction time.

[0004] While the TN5 enzyme method significantly improves library construction efficiency, it also places certain demands on the template. For example, it requires high purity and accurate DNA concentration; otherwise, the fragmentation effect will be affected. When the DNA quality is poor or the input amount is too low, adapter dimers are more likely to form during library construction. When libraries containing adapter dimers are sequenced, ① the adapter dimers bind to the anchored sequences on the flow cell of the sequencer, forming clusters through bridge PCR amplification, thus reducing the effective data yield of sequencing; ② the adapter dimer sequences are short, preferentially amplifying long clusters, and their fixed sequence, low base complexity, and short length reduce the sequencing Q30, affecting the filtering rate of clean reads; ③ as the adapter content increases, the data output drops sharply, resulting in data loss. Summary of the Invention

[0005] This invention creatively discovers that adapter dimers with specific structures appear during transposase library construction using low-volume samples. The purpose of this invention is to provide a transposase library construction method and related products that reduce adapter dimers in the library. The method of this invention modifies the primers used in the library amplification step and optimizes the polymerase used in the library amplification step, ensuring that the amplification step only amplifies DNA-adaptor ligation products and not adapter dimers. Using the method of this invention helps reduce the proportion of adapter dimers in the total library output and improves sequencing quality.

[0006] A first aspect of the present invention provides an oligonucleotide comprising a 5' variable region and a 3' invariant region, wherein the 5' variable region contains a sequencing primer sequence, and the 3' invariant region comprises a sequence obtained by deleting n nucleotides from the 3' end of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA, or TATA, wherein n is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence, plus 1.

[0007] In some implementations, the sequencing primer sequences are sequencing primer sequences for the Illumina platform, such as Illumina platform read1 or read2, or Illumina platform Truseq sequencing primers or Nextera sequencing primers, or Illumina platform Truseq read1 or Truseq read2 or Nextera read1 or Nexteraread2.

[0008] In some implementations, the sequencing primers are as shown in SEQ ID NO.5 or 6.

[0009] In some implementations, the sequencing primer sequence is a sequencing primer sequence of the MGI platform, such as a one-stranded sequencing primer or a two-stranded sequencing primer of the MGI platform, or a one-stranded sequencing primer IP1 or a two-stranded sequencing primer IP2 of the MGI platform, or a one-stranded sequencing primer Read1 or a two-stranded sequencing primer Read2 of the MGI platform.

[0010] In some implementations, the sequencing primer sequences are sequencing primer sequences from the Ion Torrent platform.

[0011] In some implementations, the sequencing primer sequence is the complete or partial sequence of the aforementioned sequencing primer sequence.

[0012] In this invention, sequencing primers refer to a known nucleotide sequence used for subsequent sequencing. Their length can be adaptively adjusted, as long as a complete sequencing primer sequence can be obtained through PCR to achieve sequencing. These sequencing primers can be adjusted according to different sequencing platforms, and the specific sequence is not limited; all such adjustments are within the scope of this invention.

[0013] In some implementations, n is at least 0. In some implementations, n is 0, 1, or 2; preferably, n is 0.

[0014] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase; in some embodiments, the Tn5 transposase recognition core sequence is as shown in SEQ ID NO.7.

[0015] In some embodiments, when the Tn5 transposase recognizes the core sequence as shown in SEQ ID NO.7, the restriction sequence closest to the 3' end of the transposase recognizes the core sequence is GAGA, the first nucleotide at the 3' end of GAGA is numbered 5 (+1) from the 3' end of the transposase recognizes the core sequence, and n is at most 5. In some embodiments, the 3' end invariant region consists of at least 14, 15, 16, 17, 18, and all 19 bases of ME starting from the 5' end.

[0016] In some implementations, the oligonucleotide is single-stranded.

[0017] In some embodiments, the oligonucleotide is a fully complementary double strand or a partially complementary double strand.

[0018] In some embodiments, at least m bases at the 3' end of the unchanged 3' region are stabilized, where m is 1, 2, 3 or more.

[0019] In some implementations, the stabilizing modification refers to a modification that can resist exonuclease degradation without affecting the 5'-3' extension, including but not limited to all modification types with the aforementioned functions known to those skilled in the art.

[0020] In some embodiments, the stabilizing modification is selected from thio modification and amino modification, with thio modification being preferred.

[0021] In some embodiments, the amino modification is either NH2C7 or NH2C6.

[0022] In some implementations, the 5' variable region further includes a sequencing immobilization sequence located at the 5' end of the sequencing primer sequence and an optional tag sequence located between the sequencing immobilization sequence and the sequencing primer sequence.

[0023] In some implementations, the sequencing-binding sequence is a sequencing-binding sequence from the Illumina platform, such as P5 or P7 or the reverse complement of P5 or P7.

[0024] In some embodiments, the sequencing-binding sequence is as shown in SEQ ID NO. 8 or 9, or, for example, the reverse complementary sequence of SEQ ID NO. 8 or 9.

[0025] In some embodiments, the tag sequence is a random sequence of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length. In some embodiments, the tag sequence is color-balanced.

[0026] In some implementations, the sequencing primer sequence is P5 and the sequencing primer sequence is read2, or the sequencing primer sequence is P5 and the sequencing primer sequence is read1, or the sequencing primer sequence is P7 and the sequencing primer sequence is read1, or the sequencing primer sequence is P7 and the sequencing primer sequence is read2.

[0027] In some implementations, the sequencing primer sequence is the reverse complementary sequence of P5 and the sequencing primer sequence is read2; or, the sequencing primer sequence is the reverse complementary sequence of P5 and the sequencing primer sequence is read1; or, the sequencing primer sequence is the reverse complementary sequence of P7 and the sequencing primer sequence is read1; or, the sequencing primer sequence is the reverse complementary sequence of P7 and the sequencing primer sequence is read2.

[0028] In some embodiments, the sequencing immobilization sequence is as shown in SEQ ID NO. 8 and the sequencing primer sequence is as shown in SEQ ID NO. 5; or, the sequencing immobilization sequence is as shown in SEQ ID NO. 9 and the sequencing primer sequence is as shown in SEQ ID NO. 6; or, the sequencing immobilization sequence is as shown in SEQ ID NO. 9 and the sequencing primer sequence is as shown in SEQ ID NO. 5; or, the sequencing immobilization sequence is as shown in SEQ ID NO. 8 and the sequencing primer sequence is as shown in SEQ ID NO. 6.

[0029] In some implementations, the sequencing-binding sequence is a clamping adapter sequence of the MGI platform, such as an upstream clamping adapter sequence of the MGI platform or an upstream clamping adapter sequence of the MGI platform, or a linear sequence of a vesicle adapter or a vesicular sequence of a vesicle adapter of the MGI platform.

[0030] In some embodiments, the sequencing-binding sequence is as shown in SEQ ID NO. 10 or 11, or, for example, the reverse complementary sequence of SEQ ID NO. 10 or 11.

[0031] In some implementations, the sequencing-attached sequence is a sequencing-attached sequence from the Ion Torrent platform, such as the A-adaptor sequence or the P1-adaptor sequence from the Ion Torrent platform.

[0032] A second aspect of this application provides an oligonucleotide pair, which comprises an upstream oligonucleotide and a downstream oligonucleotide, wherein...

[0033] (1) The upstream oligonucleotide is composed of an upstream 5' variable region and an upstream 3' invariant region, wherein the upstream 5' variable region contains the upstream sequencing primer sequence, and the upstream 3' invariant region is composed of a sequence obtained by deleting n1 nucleotides from the 3' end of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA or TATA, wherein n1 is at most the position number +1 of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence from the 3' end of the transposase recognition core sequence.

[0034] (2) The downstream oligonucleotide consists of a downstream 5' variable region and a downstream 3' invariant region, wherein the downstream 5' variable region contains the downstream sequencing primer sequence, and the downstream 3' invariant region is a sequence obtained by deleting n2 nucleotides from the 3' end of the transposase recognition core sequence, wherein n2 is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence, plus 1.

[0035] In some implementations, the upstream sequencing primer sequence and the downstream sequencing primer sequence are upstream sequencing primer sequences and downstream sequencing primer sequences of the Illumina platform.

[0036] In some implementations, the upstream sequencing primer sequence is read1 and the downstream sequencing primer sequence is read2, or vice versa.

[0037] In some implementations, the upstream sequencing primer sequence is as shown in SEQ ID NO.5 and the downstream sequencing primer sequence is as shown in SEQ ID NO.6, or vice versa.

[0038] In some implementations, the upstream and downstream sequencing primer sequences are the upstream and downstream sequencing primer sequences of the MGI platform. For example, the upstream sequencing primer sequence is a one-stranded sequencing primer and the downstream sequencing primer sequence is a two-stranded sequencing primer, or vice versa; another example is that the upstream sequencing primer sequence is a one-stranded sequencing primer IP1 and the downstream sequencing primer sequence is a two-stranded sequencing primer IP2, or vice versa; yet another example is that the upstream sequencing primer sequence is a one-stranded sequencing primer Read1 and the downstream sequencing primer sequence is a two-stranded sequencing primer Read2, or vice versa.

[0039] In some embodiments, at least m1 bases at the 3' end of the upstream 3' unchanged region are stabilized, where m1 is 1, 2, 3, or more. In some embodiments, the stabilizing modification is selected from thio modification and amino modification, preferably thio modification.

[0040] In some embodiments, at least m2 bases at the 3' end of the downstream 3' unchanged region are stabilized, where m2 is 1, 2, 3, or more. In some embodiments, the stabilizing modification is selected from thio modification and amino modification, preferably thio modification.

[0041] In some embodiments, the amino modification is either NH2C7 or NH2C6.

[0042] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of the Tn5 transposase. In some embodiments, the Tn5 transposase recognition core sequence is shown in SEQ ID NO.7.

[0043] In some implementations, n1 and n2 are independently 0, 1, or 2; preferably, both n1 and n2 are 0. In some implementations, n1 and n2 may be the same or different.

[0044] In some embodiments, both the upstream and downstream oligonucleotides are single-stranded; in some embodiments, both the upstream and downstream oligonucleotides are fully complementary double-stranded; in some embodiments, both the upstream and downstream oligonucleotides are partially complementary double-stranded; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is fully complementary double-stranded, or vice versa; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is partially complementary double-stranded, or vice versa; in some embodiments, the upstream oligonucleotide is partially complementary double-stranded and the downstream oligonucleotide is fully complementary double-stranded, or vice versa.

[0045] In some implementations, the upstream sequencing primer and the downstream sequencing primer are complete or partial sequences of the aforementioned upstream sequencing primer and downstream sequencing primer sequences.

[0046] In this invention, upstream and downstream sequencing primers refer to a known nucleotide sequence used for subsequent sequencing. Their length can be adaptively adjusted, as long as a complete sequencing primer sequence can be obtained through PCR to achieve sequencing. These sequencing primers can be adjusted according to different sequencing platforms, and the specific sequences are not limited; all are within the scope of protection of this invention.

[0047] A third aspect of this application provides an oligonucleotide pair, which comprises an upstream oligonucleotide and a downstream oligonucleotide, wherein,

[0048] (1) The upstream oligonucleotide is composed of an upstream 5' variable region and an upstream 3' invariant region, wherein the upstream 5' variable region contains the upstream sequencing primer sequence, the upstream 5' variable region also contains an upstream sequencing binding sequence located at the 5' end of the upstream sequencing primer sequence and an optional upstream tag sequence located between the upstream sequencing binding sequence and the upstream sequencing primer sequence, and the upstream 3' invariant region is composed of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA or TATA, wherein n1 is at most the position number + 1 of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence from the 3' end of the transposase recognition core sequence.

[0049] (2) The downstream oligonucleotide consists of a downstream 5' variable region and a downstream 3' invariant region. The downstream 5' variable region contains the downstream sequencing primer sequence. The downstream 5' variable region also contains a downstream sequencing binding sequence located at the 5' end of the downstream sequencing primer sequence and an optional downstream tag sequence located between the downstream sequencing binding sequence and the downstream sequencing primer sequence. The downstream 3' invariant region is composed of a sequence obtained by deleting n2 nucleotides from the 3' end of the transposase recognition core sequence. The n2 is at most the position number + 1 of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence.

[0050] In some implementations, the upstream sequencing primer sequence and the downstream sequencing primer sequence are upstream sequencing primer sequences and downstream sequencing primer sequences of the Illumina platform.

[0051] In some implementations, the upstream sequencing primer sequence is read2 and the downstream sequencing primer sequence is read1, or vice versa.

[0052] In some implementations, the upstream sequencing primer sequence is as shown in SEQ ID NO.5 and the downstream sequencing primer sequence is as shown in SEQ ID NO.6, or vice versa.

[0053] In some implementations, the sequencing primer sequences are sequencing primer sequences from the MGI platform.

[0054] In some implementations, the upstream and downstream sequencing primer sequences are the upstream and downstream sequencing primer sequences of the MGI platform. For example, the upstream sequencing primer sequence is a one-stranded sequencing primer and the downstream sequencing primer sequence is a two-stranded sequencing primer, or vice versa; another example is that the upstream sequencing primer sequence is a one-stranded sequencing primer IP1 and the downstream sequencing primer sequence is a two-stranded sequencing primer IP2, or vice versa; yet another example is that the upstream sequencing primer sequence is a one-stranded sequencing primer Read1 and the downstream sequencing primer sequence is a two-stranded sequencing primer Read2, or vice versa.

[0055] In some embodiments, at least m1 bases at the 3' end of the upstream 3' unchanged region are stabilized, where m1 is 1, 2, 3, or more. In some embodiments, the stabilizing modification is selected from thio modification and amino modification, preferably thio modification.

[0056] In some embodiments, at least m2 bases at the 3' end of the downstream 3' unchanged region are stabilized, where m2 is 1, 2, 3, or more. In some embodiments, the stabilizing modification is selected from thio modification and amino modification, preferably thio modification.

[0057] In some embodiments, the amino modification is either NH2C7 or NH2C6.

[0058] In some embodiments, the transposase recognition core sequence is the transposase recognition core sequence of the Tn5 transposase. In some embodiments, the Tn5 transposase recognition core sequence is shown in SEQ ID NO.7.

[0059] In some implementations, n1 and n2 are independently 0, 1, or 2; preferably, both n1 and n2 are 0. In some implementations, n1 and n2 may be the same or different.

[0060] In some embodiments, both the upstream and downstream oligonucleotides are single-stranded; in some embodiments, both the upstream and downstream oligonucleotides are fully complementary double-stranded; in some embodiments, both the upstream and downstream oligonucleotides are partially complementary double-stranded; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is fully complementary double-stranded, or vice versa; in some embodiments, the upstream oligonucleotide is single-stranded and the downstream oligonucleotide is partially complementary double-stranded, or vice versa; in some embodiments, the upstream oligonucleotide is partially complementary double-stranded and the downstream oligonucleotide is fully complementary double-stranded, or vice versa.

[0061] In some implementations, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are upstream sequencing immobilization sequences and downstream sequencing immobilization sequences of the Ion Torrent platform. For example, the upstream sequencing immobilization sequence is the A adapter sequence of the Ion Torrent platform and the downstream sequencing immobilization sequence is the P1 adapter sequence of the Ion Torrent platform, or vice versa.

[0062] In some implementations, the upstream sequencing-fixed sequence and the downstream sequencing-fixed sequence are clamp sequences of the MGI platform. For example, the upstream sequencing-fixed sequence is an upstream clamp sequence of the MGI platform and the downstream sequencing-fixed sequence is a downstream clamp sequence of the MGI platform, or vice versa. Another example is that the upstream sequencing-fixed sequence is a linear sequence of the vesicle adapter of the MGI platform and the downstream sequencing-fixed sequence is a vesicular sequence of the vesicle adapter of the MGI platform, or vice versa.

[0063] In some implementations, the upstream clamping plate sequence of the MGI platform is as shown in SEQ ID NO.10 and the downstream clamping plate sequence of the MGI platform is as shown in SEQ ID NO.11, and vice versa.

[0064] In some implementations, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are upstream sequencing immobilization sequences and downstream sequencing immobilization sequences of the Illumina platform.

[0065] In some implementations, the upstream sequencing binding sequence is P5, and the downstream sequencing binding sequence is P7, or vice versa; in some implementations, the upstream sequencing binding sequence is the inverse complementary sequence of P5, and the downstream sequencing binding sequence is the inverse complementary sequence of P7, or vice versa.

[0066] In some implementations, the upstream sequencing binding sequence is as shown in SEQ ID NO. 8 and the downstream sequencing binding sequence is as shown in SEQ ID NO. 9, or vice versa.

[0067] In some embodiments, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length. In some embodiments, the tag sequences are color-balanced.

[0068] In some implementations, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read1.

[0069] In some implementations, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read2.

[0070] In some implementations, the upstream sequencing primer sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read2; or, the upstream sequencing primer sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read1.

[0071] In some implementations, the upstream sequencing primer sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read1; or, the upstream sequencing primer sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read2.

[0072] In some embodiments, the upstream sequencing primer sequence is as shown in SEQ ID NO. 8, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5, the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 9; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 9, the upstream sequencing primer sequence is as shown in SEQ ID NO. 6, the downstream sequencing primer sequence is as shown in SEQ ID NO. 5, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 8; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 9, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5, the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 8; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 8, the upstream sequencing primer sequence is as shown in SEQ ID NO. 6, the downstream sequencing primer sequence is as shown in SEQ ID NO. 5, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 9.

[0073] A fourth aspect of the invention provides a kit comprising oligonucleotide pairs, extension primer pairs, and optionally a DNA polymerase with weakened or eliminated proofreading function as described in a second aspect of the invention.

[0074] The extension primer pair includes an upstream extension primer and a downstream extension primer.

[0075] The upstream extension primer comprises an upstream sequencing immobilization sequence at the 5' end, the upstream sequencing primer sequence at the 3' end, and an optional upstream tag sequence located between the upstream sequencing immobilization sequence and the upstream sequencing primer sequence; and the downstream extension primer comprises a downstream sequencing immobilization sequence at the 5' end, the downstream sequencing primer sequence at the 3' end, and an optional downstream tag sequence located between the downstream sequencing immobilization sequence and the downstream sequencing primer sequence.

[0076] In some implementations, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are upstream sequencing immobilization sequences and downstream sequencing immobilization sequences of the Ion Torrent platform. For example, the upstream sequencing immobilization sequence is the A adapter sequence of the Ion Torrent platform and the downstream sequencing immobilization sequence is the P1 adapter sequence of the Ion Torrent platform, or vice versa.

[0077] In some implementations, the upstream sequencing-fixed sequence and the downstream sequencing-fixed sequence are clamp sequences of the MGI platform. For example, the upstream sequencing-fixed sequence is an upstream clamp sequence of the MGI platform and the downstream sequencing-fixed sequence is a downstream clamp sequence of the MGI platform, or vice versa; for example, the upstream sequencing-fixed sequence is a linear sequence of the vesicle adapter of the MGI platform and the downstream sequencing-fixed sequence is a vesicular sequence of the vesicle adapter of the MGI platform, or vice versa.

[0078] In some implementations, the upstream clamping plate sequence of the MGI platform is as shown in SEQ ID NO.10 and the downstream clamping plate sequence of the MGI platform is as shown in SEQ ID NO.11, and vice versa.

[0079] In some implementations, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are upstream sequencing immobilization sequences and downstream sequencing immobilization sequences of the Illumina platform.

[0080] In some implementations, the upstream sequencing binding sequence is P5, and the downstream sequencing binding sequence is P7, or vice versa; in some implementations, the upstream sequencing binding sequence is the inverse complementary sequence of P5, and the downstream sequencing binding sequence is the inverse complementary sequence of P7, or vice versa.

[0081] In some implementations, the upstream sequencing binding sequence is as shown in SEQ ID NO. 8 and the downstream sequencing binding sequence is as shown in SEQ ID NO. 9, or vice versa.

[0082] In some embodiments, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length. In some embodiments, the tag sequences are color-balanced.

[0083] In some implementations, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read2, or the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read1.

[0084] In some implementations, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read1, or the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read2.

[0085] In some implementations, the upstream sequencing primer sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read2; or, the upstream sequencing primer sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read1.

[0086] In some implementations, the upstream sequencing primer sequence is the reverse complementary sequence of P5 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is the reverse complementary sequence of P7 and the downstream sequencing primer sequence is read1; or, the upstream sequencing primer sequence is the reverse complementary sequence of P7 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is the reverse complementary sequence of P5 and the downstream sequencing primer sequence is read2.

[0087] In some embodiments, the upstream sequencing primer sequence is as shown in SEQ ID NO. 8, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5, the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 9; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 9, the upstream sequencing primer sequence is as shown in SEQ ID NO. 6, the downstream sequencing primer sequence is as shown in SEQ ID NO. 5, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 8; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 9, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5, the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 8; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 8, the upstream sequencing primer sequence is as shown in SEQ ID NO. 6, the downstream sequencing primer sequence is as shown in SEQ ID NO. 5, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 9.

[0088] In some embodiments, the DNA polymerase with eliminated correction function includes polymerases that do not have correction function themselves, such as class A polymerases, or Taq DNA polymerases. The DNA polymerase with eliminated correction function also includes DNA polymerases that have been biotechnically modified to eliminate correction function.

[0089] In some embodiments, the DNA polymerase whose correction function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, E. coli DNA polymerase, Klenow Fragment(exo-), Vent(exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I, including but not limited to variants, modified products, and derivatives of the aforementioned polymerases.

[0090] In some embodiments, the kit further comprises one or more of transposase, transposase complex, amplification buffer, amplification primers, purified magnetic beads, and purified water.

[0091] A fifth aspect of the invention provides a kit comprising oligonucleotide pairs according to a third aspect of the invention and optionally a DNA polymerase with weakened or eliminated correction function.

[0092] In some embodiments, the DNA polymerase with eliminated correction function includes polymerases that do not have correction function themselves, such as class A polymerases, or Taq DNA polymerases. The DNA polymerase with eliminated correction function also includes DNA polymerases that have been biotechnically modified to eliminate correction function.

[0093] In some embodiments, the DNA polymerase whose correction function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, E. coli DNA polymerase, Klenow Fragment(exo-), Vent(exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I, including but not limited to variants, modified products, and derivatives of the aforementioned polymerases.

[0094] In some embodiments, the kit further comprises one or more of transposase, transposase complex, amplification buffer, amplification primers, purified magnetic beads, and purified water.

[0095] A sixth aspect of the present invention provides a method for constructing a sequencing library, comprising: (1) fragmenting a target sequence using a transposase complex; (2) performing selective PCR using the oligonucleotide pairs described in the second aspect of the present invention and optionally a DNA polymerase with weakened or eliminated proofreading function described in the fourth aspect of the present invention; and (3) performing extension PCR using the extension primer pairs described in the fourth aspect of the present invention.

[0096] The transposase complex is composed of a transposase dimer that embeds a first transposon linker and a second transposon linker. The transposase recognizes the transposase recognition core sequence. The first transposon contains the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon contains the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.

[0097] A seventh aspect of the present invention provides a method for constructing a sequencing library, comprising: (1) breaking a target sequence using a transposase complex; and (2) performing PCR using oligonucleotide pairs according to a third aspect of the present invention and optionally a DNA polymerase with weakened or eliminated proofreading function according to a fifth aspect of the present invention.

[0098] The transposase complex is composed of a transposase dimer that embeds a first transposon linker and a second transposon linker. The transposase recognizes the transposase recognition core sequence. The first transposon contains the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon contains the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.

[0099] In some embodiments, the transposase is a Tn5 transposase, including wild-type or mutant Tn5 transposases.

[0100] In some implementations, the upstream sequencing primer sequence is read1 and the downstream sequencing primer sequence is read2, or vice versa.

[0101] In some embodiments, the first transposon adapter is a fully complementary double strand, the first transposon comprising the upstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a first transposon second strand that is fully complementary to the first transposon first strand; the second transposon adapter is a fully complementary double strand, the second transposon comprising the downstream sequencing primer sequence at the 5' end and the second transposon first strand that is located at the 3' end, and further comprising a second transposon second strand that is fully complementary to the second transposon first strand.

[0102] In some embodiments, the first transposon adapter is a partially complementary double strand, for example, the first transposon adapter is a Y-type adapter, the first transposon adapter comprising the upstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end as a first transposon first strand, and further comprising the transposon recognition core sequence at the 5' end complementary to the first transposon first strand and the downstream sequencing primer sequence not complementary to the first transposon first strand as a first transposon second strand; the second transposon adapter is a partially complementary double strand, the second transposon adapter is a Y-type adapter, the second transposon adapter comprises the downstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end as a second transposon first strand, and further comprising the transposon recognition core sequence at the 5' end complementary to the second transposon first strand and the upstream sequencing primer sequence not complementary to the second transposon first strand as a second transposon second strand;

[0103] For example, the first transposon adapter is a clip-on adapter, comprising a first transposon first strand with the upstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a first transposon second strand with the transposon recognition core sequence at the 5' end, complementary to the first transposon first strand; the second transposon adapter is a clip-on adapter, comprising a second transposon first strand with the downstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a second transposon second strand with the transposon recognition core sequence at the 5' end, complementary to the second transposon first strand.

[0104] In some embodiments, the method further includes a gap-filling step, preferably, using a polymerase to fill the 9bp gap formed by transposase disruption; preferably, the polymerase has chain displacement activity.

[0105] In some embodiments, the PCR amplification is a conventional PCR amplification reaction process known to those skilled in the art, such as including denaturation, annealing, and extension, more preferably including pre-denaturation, denaturation, annealing, extension, and full extension.

[0106] Terminology Explanation:

[0107] Q30: The percentage of bases with a quality ≥30 in the sequence. Its main purpose is to assess the accuracy of sequence sequencing.

[0108] Clean reads: The remaining data after filtering the raw next-generation sequencing data and removing low-quality data. Attached Figure Description

[0109] Figure 1The principle of transposase library construction;

[0110] Figure 2 : Library peak diagrams for transposase library construction at input levels of 1 ng and 10 pg, respectively;

[0111] Figure 3A Complete transposase paired-end library structure;

[0112] Figure 3B : Main connector spike library structure;

[0113] Figure 3C : Structure of the Tn5 transposase complex;

[0114] Figure 3D The proportion of different connector spike structures;

[0115] Figure 4A The ligation relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and normal transposase libraries;

[0116] Figure 4B The connection relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and the major adapter dimer library;

[0117] Figure 4C The connection relationship between conventional N5 / N7 primers and optimized N5 / N7 primers and the linker dimer library lacking 6bp at the 5′ end of the ME sequence;

[0118] Figure 5 Library peak diagrams obtained by transposase library construction using conventional N5 / N7 primers and optimized N5 / N7 primers (thio-modified);

[0119] Figure 6 The percentage of spike peaks obtained by constructing transposase libraries using conventional N5 / N7 primers and optimized N5 / N7 primers (thio-modified);

[0120] Figure 7 Library peak diagrams obtained by transposase library construction using conventional N5 / N7 primers with conventional polymerase and optimized N5 / N7 primers with polymerases that do not have 3′-5′ exonuclease activity;

[0121] Figure 8 Base mass distribution diagrams obtained by transposase library construction using conventional N5 / N7 primers with conventional polymerase and optimized N5 / N7 primers with polymerases that do not have 3′-5′ exonuclease activity.

[0122] Figure 9 Library peak diagrams obtained by transposase library construction using conventional N5 / N7 primers and optimized N5 / N7 primers (amino modified);

[0123] Figure 10 The percentage of spike peaks obtained by constructing transposase libraries using conventional N5 / N7 primers and optimized N5 / N7 primers (amino-modified) are as follows:

[0124] Detailed Implementation Methods (Examples)

[0125] The technical solution of the present invention will be further described below with reference to the accompanying drawings and specific embodiments. However, the following examples are merely simple examples of the present invention and do not represent or limit the scope of protection of the present invention. The scope of protection of the present invention shall be determined by the claims.

[0126] In the following embodiments, unless otherwise specified, all reagents and consumables used were purchased from conventional reagent manufacturers in the art; unless otherwise specified, all experimental methods and techniques used were conventional methods and techniques in the art.

[0127] The complete sequence contained in this application is shown in the following table (orientation: 5'-3'):

[0128]

[0129]

[0130] Example 1

[0131] Comparison of library peak diagrams at a normal input of 1 ng and a low input of 10 pg.

[0132] 1. DNA fragmentation

[0133] 1.1 Thaw 5×TTBL (Vazyme#TD504) at room temperature, invert and mix well before use. Confirm that 6×TSB (Vazyme#TD504) and TWB (Vazyme#TD504) are at room temperature, and gently tap the tube wall to confirm that there is no precipitate; if there is precipitate, heat at 37°C and vortex to mix until the precipitate dissolves.

[0134] 1.2 Briefly vortex the TNB (Vazyme#TD504) until thoroughly mixed, then prepare the following reaction system in a PCR tube:

[0135] Table 1

[0136] Components volume 1ng 293g DNA / 10pg 293g DNA 1μl 5×TTBL 10μl TNB 10μl pure water 39μl

[0137] 1.3 Invert the tube to mix thoroughly, remove air bubbles, and briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0138] 1.4 Place the reaction tube in the PCR instrument and run the following reaction program:

[0139] Table 2

[0140] temperature time 55℃ 15min 4℃ Hold

[0141] 1.5 Immediately after the reaction is complete, add 10 μl of 6×TSB to the product, gently pipette to mix thoroughly, and let stand at room temperature for 5 min.

[0142] 1.6 Place the reaction tube on the magnetic rack and carefully remove the supernatant after the solution has clarified (about 3 minutes).

[0143] 1.7 Remove the reaction tube from the magnetic rack, add 100 μl of TWB and gently blow to resuspend the magnetic beads. Then place the reaction tube back on the magnetic rack and remove the supernatant after the solution has clarified (about 3 min).

[0144] 1.8 Repeat step 7 for a total of two rinses, discard the supernatant, and cover the tube to prevent the TNB beads from drying out and cracking.

[0145] 2. PCR enrichment

[0146] 2.1 Prepare the PCR amplification mix according to the following reaction system, add it to the TNB magnetic beads that have been rinsed in the previous step, and gently vortex to mix.

[0147] Table 3

[0148] Components volume 2×TAM(Vazyme#TD504) 25μl N5XX(Vazyme#TD202) 5μl N7XX(Vazyme#TD202) 5μl pure water 15μl Total 50μl

[0149] 2.2 Gently tap off the air bubbles, and briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0150] 2.3 Place the reaction tube in the PCR instrument and run the following reaction program:

[0151] Table 4

[0152]

[0153]

[0154] 3. Purification of amplified product length

[0155] 3.1 After the library amplification is completed, briefly centrifuge the reaction tube, vortex to mix the VAHTSDNA Clean Beads (Vazyme#N411), and aspirate 60 μl to 50 μl of the PCR product. Mix thoroughly by pipetting 10 times and incubate at room temperature for 5 min.

[0156] 3.2 Briefly centrifuge the reaction tube and place it on a magnetic rack. After the solution becomes clear (about 5 minutes), carefully remove the supernatant.

[0157] 3.3 Keep the reaction tube on the magnetic rack at all times, add 200 μl of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0158] 3.4 Repeat step 3.3, rinsing a total of two times.

[0159] 3.5 Keep the reaction tube on the magnetic rack at all times, and open the lid to air dry the magnetic beads for about 3 minutes.

[0160] 3.6 Remove the reaction tube from the magnetic rack and add 22 μl of ddH2O to elute. Mix thoroughly by pipetting 10 times and incubate at room temperature for 5 min.

[0161] 3.7 Briefly centrifuge the reaction tube and place it on a magnetic rack. After the solution becomes clear (about 5 min), carefully aspirate 20 μl of supernatant into a new PCR tube and store at -20℃.

[0162] 3.8 Length distribution detection was performed using an Agilent 2100 Bioanalyzer, and the results are as follows: Figure 2 As shown.

[0163] Results analysis: such as Figure 2 As shown, the peak shape is normal at an input of 1 ng, but obvious small-fragment spikes appear at a low input of 10 pg.

[0164] Example 2

[0165] The sample input was 10 pg of DNA (293 g). The library construction steps were the same as in Example 1. After amplification and purification, the library was subjected to agarose gel electrophoresis. The spikes were recovered using a (Vazyme#DC301) gel and then cloned and sequenced using a (Vazyme#C603) gel to analyze the spike structure.

[0166] The normal structure of a biterminal transposase library is as follows: Figure 3A As shown, N7 is the P7 sequencing adapter sequence, index sequence (i7), and sequencing primer sequence (read2); ME is the forward insertion sequence of the transposase recognition core sequence; Insert is the insertion sequence; ME' is the reverse insertion sequence of the transposase recognition core sequence; and N5 is the P5 sequencing adapter sequence, index sequence (i5), and sequencing primer sequence (read1).

[0167] P5—i5—read1—ME—Insert—ME'—read2—i7—P7

[0168] The main structure of the connector spike library is as follows: Figure 3BAs shown, N7 is the P7 sequencing adapter sequence, index sequence (i7), and sequencing primer sequence (read2); ME is the reverse insertion sequence of the transposase recognition core sequence; ME'-6bp is the reverse insertion sequence of the transposase recognition core sequence missing 6bp at 5'; and N5 is the P5 sequencing adapter sequence, index sequence (i5), and sequencing primer sequence (read1).

[0169] P5-i5-read1-ME-ME'-6bp-read2-i7-P7

[0170] like Figure 3C As shown, the Tn5 transposase complex is formed by a transposase dimer and two pairs of annealed adapters. The adapter structures include the transposase recognition core sequences (ME and ME' sequences) in the double-stranded portion and the sequencing primer sequences (read1 and read2, respectively) in the single-stranded portion. After the transposase breaks down the target DNA, the ME sequence and sequencing primer sequences are ligated to both ends of the insert sequence. Subsequently, the complete library structure is introduced to both ends of the target DNA by amplification using primers N5 and N7, respectively, resulting in the complete transposase-terminated library structure as shown. Figure 3A As shown in the figure. Sequencing of the adapter spike library revealed its main structure as follows: Figure 3B As shown, it does not contain insert sequences, and the ME' sequence near the 3' end is missing 6 bp, accounting for more than 85% of the entire spike library structure. The remaining few structural cases are as follows: ① ME sequence near the 5' end is missing 6 bp, ② ME sequence near the 5' end is missing 4 or 5 bp, ③ ME' sequence near the 3' end is missing 4 or 5 bp, etc., and the total proportion is as follows: Figure 3D As shown. Analysis revealed that because transposases prefer to cleave the TCTC and TATABOX regions, and the ME sequence contains these regions, some transposases, at low input levels, will use the adapter containing the ME sequence in the system as a template for cleavage, and then amplify it with N5 and N7 primers to form the adapter dimer.

[0171] Example 3

[0172] The amplification primers were modified to reduce adapter dimers in the library. The conventional N5 and N7 primer structures are as follows:

[0173] N5: From the 5' end to the 3' end are the P5 sequencing adapter sequence, the Index 2 (i5) sequence, and the read1 sequence, respectively.

[0174] N7: From the 5' end to the 3' end are the P7 sequencing adapter sequence, the Index 1 (i7) sequence, and the read2 sequence, respectively.

[0175] The modified N5 and N7 primer structures are as follows:

[0176] Optimized N5 primers: From 5' to 3', the primers consist of the P5 sequencing adapter sequence, Index 2 (i5) sequence, read1 sequence, and ME sequence, respectively, with the 3' end containing a thiolated modification on the three bases.

[0177] Optimized N7 primers: From 5' to 3', the primers consist of the P7 sequencing adapter sequence, Index 1 (i7) sequence, read2 sequence, and ME' sequence, with the two bases at the 3' end containing a thiolated modification.

[0178] The sample input was 10 pg (293 g DNA). The library construction steps for the transposases using the conventional N5 and N7 primers were the same as in Example 1. The library construction steps for the transposases using the modified N5 and N7 primers are as follows:

[0179] 1. DNA fragmentation

[0180] 1.1 Thaw 5×TTBL (Vazyme#TD504) at room temperature, invert and mix well before use. Confirm that 6×TSB (Vazyme#TD504) and TWB (Vazyme#TD504) are at room temperature, and gently tap the tube wall to confirm that there is no precipitate; if there is precipitate, heat at 37°C and vortex to mix until the precipitate dissolves.

[0181] 1.2 Briefly vortex the TNB (Vazyme#TD504) until thoroughly mixed, then prepare the following reaction system in a PCR tube:

[0182] Table 5

[0183] Components volume 10pg 293g DNA Xμl 5×TTBL 10μl TNB 10μl pure water 39μl

[0184] 1.3 Invert the tube to mix thoroughly, remove air bubbles, and briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0185] 1.4 Place the reaction tube in the PCR instrument and run the following reaction program:

[0186] Table 6

[0187] temperature time 55℃ 15min 4℃ Hold

[0188] 1.5 Immediately after the reaction is complete, add 10 μl of 6×TSB to the product, gently pipette to mix thoroughly, and let stand at room temperature for 5 min.

[0189] 1.6 Place the reaction tube on the magnetic rack and carefully remove the supernatant after the solution has clarified (about 3 minutes).

[0190] 1.7 Remove the reaction tube from the magnetic rack, add 100 μl of TWB and gently blow to resuspend the magnetic beads. Then place the reaction tube back on the magnetic rack and remove the supernatant after the solution has clarified (about 3 min).

[0191] 1.8 Repeat step 7 for a total of two rinses, discard the supernatant, and cover the tube to prevent the TNB beads from drying out and cracking.

[0192] 2. PCR enrichment

[0193] 2.1 Prepare the PCR amplification mix according to the following reaction system, add it to the TNB magnetic beads that have been rinsed in the previous step, and gently vortex to mix.

[0194] Table 7

[0195] Components volume 2×TAM(Vazyme#TD504) 25μl Optimized N5 primers 5μl Optimized N7 primers 5μl pure water 15μl Total 50μl

[0196] Specifically, the optimized N5 primer sequence used in this step is (SEQ ID NO.1): 5'-AATGATACGGCGACCACCGAGATCTACACTAGATCGCTCGTCGGCAGCGTCAGATGTGTATAAGAGA*C*A*G-3' (where * indicates thiomodification).

[0197] The optimized N7 primer sequence is (SEQ ID NO.2): 5'-CAAGCAGAAGACGGCATACGAGATTAAGGCGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGAC*A*G-3' (where * indicates thiomodification).

[0198] 2.2 Gently tap off the air bubbles, and briefly centrifuge to collect the reaction solution to the bottom of the tube.

[0199] 2.3 Place the reaction tube in the PCR instrument and run the following reaction program:

[0200] Table 8

[0201]

[0202] 3. Purification of amplified product length

[0203] 3.1 After the library amplification is completed, briefly centrifuge the reaction tube, vortex to mix the VAHTS DNA Clean Beads (Vazyme#N411), and aspirate 60 μl to 50 μl of the PCR product. Mix thoroughly by pipetting 10 times and incubate at room temperature for 5 min.

[0204] 3.2 Briefly centrifuge the reaction tube and place it on a magnetic rack. After the solution becomes clear (about 5 minutes), carefully remove the supernatant.

[0205] 3.3 Keep the reaction tube on the magnetic rack at all times, add 200 μl of freshly prepared 80% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0206] 3.4 Repeat step 3.3, rinsing a total of two times.

[0207] 3.5 Keep the reaction tube on the magnetic rack at all times, and open the lid to air dry the magnetic beads for about 3 minutes.

[0208] 3.6 Remove the reaction tube from the magnetic rack and add 22 μl of ddH2O to elute. Mix thoroughly by pipetting 10 times and incubate at room temperature for 5 min.

[0209] 3.7 Briefly centrifuge the reaction tube and place it on a magnetic rack. After the solution becomes clear (about 5 min), carefully aspirate 20 μl of supernatant into a new PCR tube and store at -20℃.

[0210] 3.8 Length distribution detection was performed using an Agilent 2100 Bioanalyzer, and the results are as follows: Figure 5 As shown, the statistical spike percentage is as follows: Figure 6 As shown.

[0211] 3.9 Next-generation sequencing strategy: 10G / sample, PE150.

[0212] Results analysis:

[0213] The linkage relationships between the primers before and after modification and the normal transposase library and the major adapter dimer library are as follows: Figure 4A and Figure 4B As shown, after extending the N5 / N7 primers, they became complementary to the paired-end transposase library sequence, allowing for normal amplification. However, for the adapter spike library, the extended primers could not achieve complete complementarity, thus preventing amplification. Since the polymerases used in the library construction kits typically have 3′-5′ exonuclease activity, the extended primer protrusions would be removed. Therefore, protective modifications were performed on the primer extension portion; 2-3 thiomodifications were sufficient for adequate protection. Similarly, the modified primers effectively prevented the amplification of adapter dimers, such as those lacking a 6bp ME sequence at the 5′ end. Figure 4C As shown, similarly, the lack of 4 or 5 bp in the ME sequence at the 5′ end and the lack of 4 or 5 bp in the ME′ sequence at the 3′ end can effectively prevent the amplification of other linker dimers with a lower proportion.

[0214] like Figure 5 and Figure 6As shown, the number of spikes decreased after primer optimization. The percentage of spikes in the library before and after modification decreased from 6.32% to 1.66%. Although the number of spikes was greatly reduced after primer modification, it could not be completely eliminated. This result indicates that protective modification cannot protect the modified primers 100%, and their ends may still be cleaved due to the presence of 3′-5′ exonuclease activity of polymerase.

[0215] Example 4

[0216] Based on Example 3, the polymerase for library amplification was further optimized. Specifically, the 2×TAM (Vazyme#TD504) in Table 7 of Example 2 was replaced with Vazyme#PK512 (the 3′-5′ exonuclease activity of the polymerase contained therein was significantly reduced), while the other steps remained unchanged. The control group consisted of conventional N5 and N7 primers (Vazyme#TD202) + conventional polymerase 2×TAM (Vazyme#TD504), and the test group consisted of optimized N5 primers (SEQ ID NO.1) and optimized N7 primers (SEQ ID NO.2) + PK512. The sample input was 10 pg (293 g DNA), and the experimental steps were the same as in Example 2. The resulting library peak chromatogram is shown below. Figure 7 As shown, the base mass distribution diagram is as follows: Figure 8 As shown.

[0217] Results analysis: such as Figure 7 Library construction was performed using optimized N5 / N7 primers in combination with a polymerase that had no 3′-5′ exonuclease activity, and the dimer spikes were completely eliminated at low input levels.

[0218] like Figure 8 The conventional N5 / N7 primers paired with the standard 2×TAM polymerase resulted in increased base jitter in read 1, reaching 60 bp, and read 2, reaching 30 bp. After optimization, the base jitter was maintained within 15 bp. PE150 sequencing was performed, with each end sequenced at 150 bp. Theoretically, the AT and GC content should be equal in the sequencing results. However, in practice, base jitter exists at the transposase breakpoint, typically within the first 15 bp, leading to unequal AT and GC content. The presence of spike peaks in the control group exacerbated the base jitter, increasing the length of the jittered bases and reducing sequencing quality. In the test group, eliminating the adapter spikes restored normal base jitter.

[0219] Example 5

[0220] The amplification primers were modified to reduce adapter dimers in the library. The conventional N5 and N7 primer structures are the same as in Example 3.

[0221] The modified N5 and N7 primer structures are as follows:

[0222] Optimized N5 primers: from 5' to 3', the primers are the P5 sequencing adapter sequence, Index 2 (i5) sequence, read1 sequence, and ME sequence, respectively, with thiomodification on the two bases at the 3' end.

[0223] Optimized N7 primers: from 5' to 3', they are the P7 sequencing adapter sequence, Index 1 (i7) sequence, read2 sequence, and ME' sequence, respectively, with one base at the 3' end containing an NH2 C7 amino group modification.

[0224] The sample input was 10 pg of DNA (293 g). The library construction steps for the transposases of the conventional N5 and N7 primers were the same as in Example 1, and the library construction steps for the transposases of the modified N5 and N7 primers were the same as in Example 3.

[0225] The optimized N5 primer sequence used is (SEQ ID NO.3): 5'-AATGATACGGCGACCACCGAGATCTACACTTCTAGCTTCGTCGGCAGCGTCAGATGTGTATAAGAGAC*A*G-3' (where * represents thiomodification);

[0226] The optimized N7 primer sequence is (SEQ ID NO.4): 5'-CAAGCAGAAGACGGCATACGAGATGTAGAGGAGTCTCGTGGGCTCGGAGATGTGTATAAGAGACAG-NH2 C7-3'

[0227] Statistical library peak diagram as follows Figure 9 As shown, the statistical spike percentage is as follows: Figure 10 As shown.

[0228] Results Explanation: For example Figures 9-10 The optimized dimer spikes were significantly reduced, and the spike ratio and spike library output were significantly lower than those constructed with conventional N5 / N7 primers, indicating that NH2C7 modification can achieve results similar to thiomodification.

Claims

1. An oligonucleotide comprising a 5' variable region and a 3' invariant region, wherein, The 5' variable region contains sequencing primer sequences, optionally, which are sequencing primer sequences for the Illumina platform or the MGI platform, more optionally, are Illumina platform read1 or read2, and more optionally, as shown in SEQ ID NO. 5 or 6. The 3' end invariant region is composed of a sequence obtained by deleting n nucleotides from the 3' end of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA, or TATA. Here, n is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence, starting from the 3' end of the transposase recognition core sequence, plus 1. Preferably, n is 0. Optionally, the transposase recognition core sequence is the transposase recognition core sequence of Tn5 transposase; further optionally, as shown in SEQ ID NO.7; preferably, n is 0, 1, or 2; most preferably, n is 0. Optionally, at least m bases at the 3' end of the unchanged 3' region are stabilized, where m is 1, 2, 3, or more. Further, optionally, the stabilizing modification is selected from thio modification and amino modification, preferably thio modification. Optionally, the oligonucleotide is single-stranded.

2. The oligonucleotide according to claim 1, wherein, The 5' variable region further includes a sequencing immobilization sequence located at the 5' end of the sequencing primer sequence and an optional tag sequence located between the sequencing immobilization sequence and the sequencing primer sequence. Optionally, the sequencing immobilization sequence is a sequencing immobilization sequence from the Illumina platform; more preferably, it is a P5 or P7 sequence from the Illumina platform; and even more preferably, it is as shown in SEQ ID NO. 8 or 9. Optionally, the sequencing-binding sequence is a clamping sequence of the MGI platform, and more preferably, as shown in SEQ ID NO. 10 or 11. Optionally, the tag sequence is a random sequence with a length of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides. Further optionally, the sequencing primer sequence is P5 and the sequencing primer sequence is read1, or the sequencing primer sequence is P7 and the sequencing primer sequence is read2. Further optionally, the sequencing immobilization sequence is as shown in SEQ ID NO.8 and the sequencing primer sequence is as shown in SEQ ID NO.5, or the sequencing immobilization sequence is as shown in SEQ ID NO.9 and the sequencing primer sequence is as shown in SEQ ID NO.

6.

3. An oligonucleotide pair, comprising an upstream oligonucleotide and a downstream oligonucleotide, wherein, (1) The upstream oligonucleotide consists of a variable 5' region and a constant 3' region, wherein, The upstream 5' variable region contains the upstream sequencing primer sequence, and The upstream 3' invariant region is composed of a transposase recognition core sequence containing at least one restriction sequence selected from TCTC, GAGA, or TATA, with n1 nucleotides deleted from its 3' end. Here, n1 is at most the position number of the first nucleotide from the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence, plus 1. Preferably, n1 is 0. Optionally, at least m1 bases at the 3' end of the upstream 3' invariant region are stabilized, where m1 is 1, 2, 3, or more. Further, optionally, the stabilizing modification is selected from thio and amino modifications, preferably thio modification. (2) The downstream oligonucleotide consists of a downstream 5' variable region and a downstream 3' invariant region, wherein, The downstream 5' variable region contains the downstream sequencing primer sequence, and The downstream 3' end unchanged region is composed of a sequence obtained by deleting n2 nucleotides from the 3' end of the transposase recognition core sequence, wherein n2 is at most the position number of the first nucleotide at the 3' end of the restriction sequence closest to the 3' end of the transposase recognition core sequence, plus 1. Preferably, n2 is 0. Optionally, at least m2 bases at the 3' end of the downstream 3' end invariant region are stabilized, where m2 is 1, 2, 3, or more. Further, optionally, the stabilizing modification is selected from thio modification and NH2 C7 modification, preferably thio modification. Optionally, the upstream sequencing primer sequence and the downstream sequencing primer sequence are upstream sequencing primer sequences and downstream sequencing primer sequences for the Illumina platform or the MGI platform. More optionally, the upstream sequencing primer sequence is read1 of the Illumina platform and the downstream sequencing primer sequence is read2 of the Illumina platform, or vice versa. More optionally, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5 and the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, or vice versa. Optionally, the transposase recognition core sequence is the transposase recognition core sequence of the Tn5 transposase. More optionally, as shown in SEQ ID NO.7, preferably, n1 and n2 are independently 0, 1, or 2; most preferably, n1 and n2 are both 0. Optionally, both the upstream oligonucleotide and the downstream oligonucleotide are single-stranded.

4. The oligonucleotide pair according to claim 3, wherein (1) The upstream 5' variable region further includes an upstream sequencing binding sequence located at the 5' end of the upstream sequencing primer sequence and an optional upstream tag sequence located between the upstream sequencing binding sequence and the upstream sequencing primer sequence. (2) The downstream 5' variable region further includes a downstream sequencing immobilization sequence located at the 5' end of the downstream sequencing primer sequence and an optional downstream tag sequence located between the downstream sequencing immobilization sequence and the downstream sequencing primer sequence. in, Optionally, the upstream sequencing solid-binding sequence and the downstream sequencing solid-binding sequence are upstream sequencing solid-binding sequences and downstream sequencing solid-binding sequences of the Illumina platform. More optionally, the upstream sequencing solid-binding sequence is P5 of the Illumina platform and the downstream sequencing solid-binding sequence is P7 of the Illumina platform, or vice versa. More optionally, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO. 8 and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO. 9, or vice versa. Optionally, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are the upstream and downstream clamping sequences of the MGI platform. More optionally, the upstream clamping sequence is as shown in SEQ ID NO.10 and the downstream clamping sequence is as shown in SEQ ID NO.11, or vice versa. Optionally, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length. Further optionally, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read2; or, the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read1. Further optionally, the upstream sequencing primer sequence is as shown in SEQ ID NO.8, the upstream sequencing primer sequence is as shown in SEQ ID NO.5, the downstream sequencing primer sequence is as shown in SEQ ID NO.6, and the downstream sequencing primer sequence is as shown in SEQ ID NO.9; or, the upstream sequencing primer sequence is as shown in SEQ ID NO.9, the upstream sequencing primer sequence is as shown in SEQ ID NO.6, the downstream sequencing primer sequence is as shown in SEQ ID NO.5, and the downstream sequencing primer sequence is as shown in SEQ ID NO.

8.

5. A kit comprising the oligonucleotide pairs, extension primer pairs, and optionally a DNA polymerase with weakened or eliminated proofreading function as described in claim 3. in, The extension primer pair includes an upstream extension primer and a downstream extension primer. The upstream extension primer comprises an upstream sequencing immobilization sequence at the 5' end, the upstream sequencing primer sequence at the 3' end, and an optional upstream tag sequence located between the upstream sequencing immobilization sequence and the upstream sequencing primer sequence; the downstream extension primer comprises a downstream sequencing immobilization sequence at the 5' end, the downstream sequencing primer sequence at the 3' end, and an optional downstream tag sequence located between the downstream sequencing immobilization sequence and the downstream sequencing primer sequence. Optionally, the upstream sequencing solid-binding sequence and the downstream sequencing solid-binding sequence are upstream sequencing solid-binding sequences and downstream sequencing solid-binding sequences of the Illumina platform. More optionally, the upstream sequencing solid-binding sequence is P5 of the Illumina platform and the downstream sequencing solid-binding sequence is P7 of the Illumina platform, or vice versa. More optionally, the upstream sequencing solid-binding sequence is as shown in SEQ ID NO. 8 and the downstream sequencing solid-binding sequence is as shown in SEQ ID NO. 9, or vice versa. Optionally, the upstream sequencing immobilization sequence and the downstream sequencing immobilization sequence are the upstream and downstream clamping sequences of the MGI platform. More optionally, the upstream clamping sequence is as shown in SEQ ID NO.10 and the downstream clamping sequence is as shown in SEQ ID NO.11, or vice versa. Optionally, the upstream tag sequence and the downstream tag sequence are independently random sequences of at least 2, 3, 4, 5, 6, 7, 8 or more nucleotides in length. Further optionally, the upstream sequencing primer sequence is P5 and the upstream sequencing primer sequence is read1, the downstream sequencing primer sequence is P7 and the downstream sequencing primer sequence is read2; or, the upstream sequencing primer sequence is P7 and the upstream sequencing primer sequence is read2, the downstream sequencing primer sequence is P5 and the downstream sequencing primer sequence is read1. Further optionally, the upstream sequencing primer sequence is as shown in SEQ ID NO. 8, the upstream sequencing primer sequence is as shown in SEQ ID NO. 5, the downstream sequencing primer sequence is as shown in SEQ ID NO. 6, and the downstream sequencing primer sequence is as shown in SEQ ID NO. 9; or, the upstream sequencing primer sequence is as shown in SEQ ID NO. 9, the upstream sequencing primer sequence is as shown in SEQ ID NO. 6, the downstream sequencing primer sequence is as shown in SEQ ID NO. 5, and the downstream sequencing primer sequence is as shown in SEQ ID NO.

8. Optionally, the DNA polymerase whose correction function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, E. coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I.

6. A kit comprising the oligonucleotide pairs according to claim 4 and optionally a DNA polymerase with weakened or eliminated correction function. Optionally, the DNA polymerase whose correction function is weakened or eliminated is one or more of Taq polymerase, Bst polymerase, E. coli DNA polymerase, Klenow Fragment (exo-), Vent (exo-) DNA polymerase, Pfu polymerase, Tfi DNA polymerase, Tfl polymerase, KOD polymerase, Hi-Fi polymerase, Tth DNA polymerase, Tth polymerase, TIi polymerase, Sac polymerase, SSo polymerase, Pfutubo polymerase, Pyrobest polymerase, Pwo polymerase, Poc polymerase, Mth polymerase, Pho polymerase, ES4 polymerase, VENT polymerase, DEEPVENT polymerase, Pab polymerase, Expand polymerase, Tbr polymerase, Tfl polymerase, Tru polymerase, Tac polymerase, Tne polymerase, Tma polymerase, Tih polymerase, Bst DNA polymerase, and Bsu DNA polymerase I.

7. A method for constructing a sequencing library, comprising: (1) Use transposase complex to break the target sequence; (2) Selective PCR was performed using the oligonucleotide pairs of claim 3 and optionally the DNA polymerase with weakened or eliminated proofreading function as described in claim 5; and (3) Extension PCR was performed using the extension primer pairs of claim 5. The transposase complex is composed of a transposase dimer that embeds a first transposon linker and a second transposon linker. The transposase recognizes the transposase recognition core sequence. The first transposon contains the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon contains the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.

8. A method for constructing a sequencing library, comprising: (1) Use transposase complex to break the target sequence; (2) PCR is performed using the oligonucleotide pairs according to claim 4 and optionally the DNA polymerase with weakened or eliminated correction function as described in claim 6. The transposase complex is composed of a transposase dimer that embeds a first transposon linker and a second transposon linker. The transposase recognizes the transposase recognition core sequence. The first transposon contains the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and the second transposon contains the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and vice versa.

9. The method according to any one of claims 7-8, wherein the first transposon adapter is a completely complementary double strand, the first transposon comprising the upstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a first transposon second strand completely complementary to the first transposon first strand; the second transposon adapter is a completely complementary double strand, the second transposon comprising the downstream sequencing primer sequence at the 5' end and the second transposon first strand at the 3' end, and further comprising a second transposon second strand completely complementary to the second transposon first strand; Alternatively, the first transposon adapter is a Y-type adapter, comprising a first transposon first strand with the upstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and further comprising the transposase recognition core sequence at the 5' end complementary to the first transposon first strand and a first transposon second strand with the downstream sequencing primer sequence not complementary to the first transposon first strand; the second transposon adapter is a partially complementary double strand, the second transposon adapter is a Y-type adapter, comprising a second transposon first strand with the downstream sequencing primer sequence at the 5' end and the transposase recognition core sequence at the 3' end, and further comprising the transposase recognition core sequence at the 5' end complementary to the second transposon first strand and a second transposon second strand with the upstream sequencing primer sequence not complementary to the second transposon first strand; Alternatively, the first transposon adapter is a clamp adapter, comprising a first transposon first strand with the upstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a first transposon second strand with the transposon recognition core sequence at the 5' end, complementary to the first transposon first strand; the second transposon adapter is a clamp adapter, comprising a second transposon first strand with the downstream sequencing primer sequence at the 5' end and the transposon recognition core sequence at the 3' end, and further comprising a second transposon second strand with the transposon recognition core sequence at the 5' end, complementary to the second transposon first strand.