A high-throughput transcriptome sequencing library construction method and application thereof

By constructing high-throughput transcriptome sequencing libraries and utilizing barcoded reverse transcription primers and Tn5 transposase fragmentation, the problem of sample quantity limitations in conventional methods has been solved, achieving low-cost and high-efficiency transcriptome sequencing suitable for gene expression and genotype analysis of large numbers of samples.

CN114622286BActive Publication Date: 2026-04-24CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2020-12-14
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Conventional transcriptome library construction methods are cumbersome, costly, and difficult to meet the sequencing needs of large numbers of samples, especially in gene expression analysis and genotype analysis, where existing technologies cannot achieve high-throughput and low-cost transcriptome sequencing.

Method used

A high-throughput transcriptome sequencing library construction method was adopted. By using reverse transcription primers with different primary and secondary barcodes, mixed reverse transcription was performed to obtain mixed cDNA. Combined with Tn5 transposase fragmentation and PCR amplification, parallel construction and library isolation of multiple samples were achieved. Library construction and sequencing were performed using a kit.

Benefits of technology

It achieves high-throughput, low-cost transcriptome sequencing, with gene detection efficiency and accuracy comparable to conventional methods, and is suitable for gene expression and genotype analysis of large-scale samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0002833442710000051
    Figure BDA0002833442710000051
  • Figure BDA0002833442710000061
    Figure BDA0002833442710000061
  • Figure BDA0002833442710000071
    Figure BDA0002833442710000071
Patent Text Reader

Abstract

The application discloses a high-throughput transcriptome sequencing library construction method and application thereof. According to the technical scheme of the application, the first level bar code is introduced in the process of reverse transcription to synthesize the first strand of cDNA, so that the different sample libraries can be constructed in parallel in the same experiment. The unique molecular identification code is introduced in the process of reverse transcription to synthesize the first strand of cDNA, so that the repeated sequences generated in the PCR process can be effectively identified and removed. The high-throughput library construction and sequencing are realized by introducing the first level and second level bar codes and by detecting the gene expression level based on the 3' end of the transcript. The application has the gene expression level detection accuracy and expression gene detection efficiency which are equivalent to those of the conventional transcriptome sequencing. Meanwhile, compared with the conventional transcriptome library construction and sequencing, the application has the advantages of high throughput and low cost, and is suitable for large-scale sample gene expression analysis and genotype analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biotechnology and molecular biology, particularly to the field of high-throughput sequencing technology, specifically to a method for constructing high-throughput transcriptome sequencing libraries and its applications. Background Technology

[0002] The transcriptome is the sum of all RNA expressed by a specific species, tissue, or cell under a particular environmental or physiological condition, serving as the link between genomic genetic information and proteomic biological function. Transcriptome sequencing, also known as RNA-seq, refers to the sequencing of transcripts from a biological sample using second-generation high-throughput sequencing technology. Transcriptome sequencing can comprehensively and rapidly obtain information on all transcripts of a specific biological sample under a specific state, and has the advantages of a wide range of gene expression level detection and low background noise, making it the preferred method for transcriptome analysis. Transcriptome sequencing not only provides important tools and methods for studying gene expression and regulation but also facilitates gene function research.

[0003] Conventional transcriptome library construction methods are quite complex, typically involving steps such as mRNA isolation, mRNA fragmentation, reverse transcription to synthesize the first-strand cDNA, second-strand cDNA synthesis, double-strand cDNA end repair and adenosine A addition, adapter ligation, and PCR enrichment. Furthermore, when constructing sequencing libraries for different samples using conventional transcriptome library construction methods, each of the above steps requires independent processing for each sample. Therefore, conventional transcriptome library construction methods are not only costly but also time-consuming and labor-intensive. Additionally, because the full length of transcripts needs to be covered by sequencing reads, the required sequencing data volume is also high. When analyzing the transcriptomes of hundreds or thousands of samples, conventional transcriptome library construction methods become significantly limited in terms of cost and throughput.

[0004] In practice, many scientific studies require transcriptome sequencing analysis of large numbers of samples, including mapping transcriptomes of different tissue samples, dynamic analysis of transcriptomes within the same tissue sample, and analysis of gene expression level variations and regulation in a population. Therefore, there is an urgent need in this field to develop a high-throughput transcriptome library construction and sequencing method for large numbers of samples, building upon conventional transcriptome sequencing library methods and sequencing techniques, while simultaneously achieving low cost. Summary of the Invention

[0005] To overcome the problem that conventional transcriptome library construction and sequencing methods cannot meet the needs of transcriptome sequencing of large numbers of samples, this invention provides the following technical solution:

[0006] One object of the present invention is to provide a high-throughput transcriptome library construction and sequencing method with gene detection efficiency comparable to conventional transcriptome sequencing methods, while simultaneously achieving low cost.

[0007] This invention provides a method for constructing a high-throughput transcriptome sequencing library, comprising the following steps:

[0008] (1) Extract total RNA from N samples to be tested;

[0009] (2) Using the total RNA of each of the test samples as templates, N reverse transcription products are obtained by reverse transcription using N reverse transcription primers, which are the first strands of cDNA of the N test samples.

[0010] Each of the reverse transcription primers includes a sequencing adapter, a primary barcode for distinguishing different samples, polythymine, and a degenerate V at the 3' end;

[0011] The primary barcodes in the N reverse transcription primers are all different;

[0012] The primary barcodes of the N reverse transcription primers are different (there are base differences between the primary barcodes of each reverse transcription primer);

[0013] (3) Mix the reverse transcription products of the N test samples to obtain a mixed cDNA first strand; then use the mixed cDNA first strand to synthesize a cDNA second strand to obtain double-stranded cDNA;

[0014] (4) The double-stranded cDNA was fragmented using a Tn5 transposase assembled with sequencing adapters to obtain fragmented products;

[0015] The sequencing adapters in the Tn5 transposase assembled with sequencing adapters are different from those in the reverse transcription primers and can form a sequencing adapter set with the sequencing adapters in the reverse transcription primers.

[0016] (5) The fragmented product is amplified by PCR to obtain a sequencing library.

[0017] In the above method, the number of differential bases between the primary barcodes in each reverse transcription primer is ≥2.

[0018] In the above method, each of the reverse transcription primers also includes a unique molecular identification code used to distinguish different transcripts of the same gene within the same sample;

[0019] The unique molecular identification code is located between the primary barcode and the polythymine;

[0020] Alternatively, the unique molecular identification code may be located between the sequencing adapter in the reverse transcription primer and the primary barcode.

[0021] The primary barcode is a base sequence with a size greater than or equal to 4 nt;

[0022] The unique molecular identifier is a random base with a size greater than or equal to 4 nt.

[0023] The sequencing adapter in any of the reverse transcription primers described above is selected according to the sequencing platform. Different sequencing platforms use different sequencing adapters, and even for the same sequencing platform, any compatible sequencing adapter can be selected. When the sequencing platform is Illumina, as one feasible option, the sequencing adapter in the reverse transcription primer can specifically be the sequencing adapter shown from position 1 to 34 of the 5' end in Sequence 1 of the sequence listing. When the sequencing adapter selected in the reverse transcription primer is the sequencing adapter shown from position 1 to 34 of the 5' end in Sequence 1 of the sequence listing, as one feasible option, the sequencing adapter carried by the Tn5 transposase can specifically be as shown in Sequence 105 of the sequence listing.

[0024] The specific degree of any of the unique molecular identification codes mentioned above can be 6bp.

[0025] In any of the above-described reverse transcription primers, the unique molecular identification code may or may not be present, and the positions of the unique molecular identification code and the primary barcode can be interchanged. In the embodiments of the present invention, the reverse transcription primer, from the 5' end, consists of a sequencing adapter, a primary barcode, a unique molecular identification code, polythymine, and a degenerate base V. The embodiments of the present invention provide 96 reverse transcription primers, as shown in sequences 1 to 96 of the sequence listing.

[0026] The length of any of the primary barcodes described above is greater than or equal to 4 nt. The length of the primary barcodes in different reverse transcription primers can be the same or different. The number of differential bases between the primary barcodes in each reverse transcription primer is preferably ≥2. In the embodiments of the present invention, the length of the primary barcode is 6 nt. The embodiments of the present invention provide 96 primary barcodes, specifically as shown in sequences 1 to 96 of the sequence listing from position 35-40 of the 5' end.

[0027] In the above method, the polythymidine is a continuous T base of 15-30 nt; the length of any of the above polythymidines can specifically be 18 nt.

[0028] The degenerate base V mentioned above is A, C, or G.

[0029] In the above method, the PCR amplification introduces a secondary barcode for distinguishing different libraries by amplifying at least one primer from the required primer pair.

[0030] In the above method, the primer pair required for amplification is any one of the following 1)-3):

[0031] 1) The primer pair shown consists of M primers A and B;

[0032] Primer A includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer, a secondary barcode for distinguishing different libraries, and a sequencing adapter that binds to the sequencing chip.

[0033] Each primer B includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase and another sequencing adapter that binds to the sequencing chip.

[0034] The secondary barcodes in the M primers A are different;

[0035] 2) The primer pair shown consists of primer C and M primers D;

[0036] Primer C includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer and a sequencing adapter that binds to the sequencing chip.

[0037] Each primer D includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase, a secondary barcode for distinguishing different libraries, and another sequencing adapter that binds to the sequencing chip.

[0038] The secondary barcodes in the M primers D are all different;

[0039] 3) The primer pair shown consists of M primers E and M primers F;

[0040] Each primer E includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer, a secondary barcode for distinguishing different libraries, and a sequencing adapter that binds to the sequencing chip.

[0041] Each primer F includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase, a secondary barcode for distinguishing different libraries, and another sequencing adapter that binds to the sequencing chip.

[0042] The secondary barcodes in the M primers E are all different;

[0043] The secondary barcodes in the M primers F are all different.

[0044] As needed, secondary barcodes can be introduced using forward primers for PCR amplification, which include: sequencing adapters identical or complementary to the sequencing adapters in the reverse transcription primers, secondary barcodes for distinguishing different libraries, and sequencing adapters for binding to the sequencing chip; secondary barcodes can also be introduced using reverse primers for PCR amplification, which include: sequencing adapters identical or complementary to the sequencing adapters carried by the Tn5 transposase, secondary barcodes for distinguishing different libraries, and sequencing adapters for binding to the sequencing chip; secondary barcodes... The sequencing codes can also be introduced by using both the forward and reverse primers used for PCR amplification (the secondary barcodes in the forward and reverse primers can be the same or different). The forward primer includes: a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer, a secondary barcode for distinguishing different libraries, and a sequencing adapter for binding to the sequencing chip. The reverse primer includes: a sequencing adapter that is the same as or complementary to the sequencing adapter carried by the Tn5 transposase, a secondary barcode for distinguishing different libraries, and a sequencing adapter for binding to the sequencing chip. The sequencing adapter is selected according to the sequencing platform; different sequencing platforms use different sequencing adapters, and even within the same sequencing platform, any compatible sequencing adapter can be selected.

[0045] As one feasible approach, the primers in each amplification system of the PCR amplification described above can be composed of any one of the M forward primers containing secondary barcodes (for example, the PCR index primer in Table 2 in the embodiments of this invention) and a universal reverse primer (sequence 104, the PCR common primer, is used in the embodiments of this invention). In the forward primer containing secondary barcodes, sequencing adapters are located on both sides of the secondary barcode; one adapter sequence binds to the sequencing chip, and the other adapter sequence binds to the reverse transcription primer (the sequence is the same as the sequencing adapter sequence in the reverse transcription primer). The embodiments of this invention provide seven types of PCR primers containing secondary barcodes, specifically as shown in sequences 97 to 103 of the sequence listing.

[0046] M is an integer greater than or equal to 1; the length of the secondary barcode is determined according to the read length of the selected sequencing platform, typically 6 bp or 8 bp; the number of differential bases between the secondary barcodes in each PCR primer is preferably ≥2. In the embodiments of the present invention, the length of the secondary barcode is 6 bp. The embodiments of the present invention provide 7 types of secondary barcodes, specifically as shown in sequences 97 to 103 of the sequence listing, from position 35-40 of the 5' end.

[0047] The synthesis of the first and second strands of cDNA also includes a product purification step.

[0048] The process of fragmenting the double-stranded cDNA also includes a product purification step.

[0049] In step (5), the purpose of PCR amplification is to enrich the 3' end fragment of the transcript and add secondary barcodes. Following this step, a PCR amplification product length sorting step is also included. The sorting length can be adjusted according to the actual library state and sequencing strategy (preferably 400-600 bp).

[0050] N or M mentioned above can be any natural number, and can be determined based on the number of samples and the number of documents.

[0051] The library prepared by the above method is also within the scope of protection of this invention.

[0052] Another object of the present invention is to provide a kit for constructing high-throughput transcriptome sequencing libraries.

[0053] The kit provided by this invention includes the above-mentioned N reverse transcription primers.

[0054] The kit also includes the Tn5 transposase assembled with another sequencing adapter and the PCR amplification primers; wherein the forward and / or reverse amplification primers contain secondary barcodes.

[0055] The kit also includes reagents required for routine library construction.

[0056] Another objective of this invention is to provide a method for high-throughput transcriptome sequencing.

[0057] The method provided by the present invention includes the following steps: constructing a sequencing library using the method described in the first objective above, and then sequencing the sequencing library.

[0058] This invention also protects transcriptome sequencing methods, comprising the following steps: constructing a sequencing library using any of the methods described above, and then sequencing the library.

[0059] In the method, sequencing data is allocated to different libraries based on the sequence information of the secondary barcode; and the data of each library is further allocated to the corresponding sequencing sample based on the primary barcode information.

[0060] This invention also protects the application of the transcriptome sequencing method in biological gene expression analysis and genotype analysis.

[0061] The organisms mentioned may specifically include animals, plants, or microorganisms.

[0062] By applying the technical solution of this invention, parallel construction of libraries from different test samples in the same experiment is achieved through primary barcoding of primers used in the reverse transcription synthesis of the first strand of cDNA. The effective identification and removal of repetitive sequences generated during PCR is achieved through unique molecular identification codes of primers used in the reverse transcription synthesis of the first strand of cDNA. High-throughput library construction and sequencing are achieved by introducing primary and secondary barcoding, and by detecting gene expression levels based on the 3' end of the transcript. This invention achieves gene expression level detection accuracy and gene detection efficiency comparable to conventional transcriptome sequencing. Furthermore, compared to conventional transcriptome library construction and sequencing, this invention offers advantages of high throughput and low cost, making it suitable for gene expression and genotype analysis of large-scale samples. Attached Figure Description

[0063] Figure 1 Build a workflow for the document library.

[0064] Figure 2 This is the library quality control result in Example 3.

[0065] Figure 3 The results show the quality assessment of the sequencing reads in Example 3.

[0066] Figure 4 The results show the evaluation of sequencing read coverage in Example 3.

[0067] Figure 5 The results show the gene expression level analysis between two biological duplicates of different samples in Example 3.

[0068] Figure 6 This is a comparison of the number of expressed genes obtained from the high-throughput sequencing library provided by the present invention in Example 3 with the number of expressed genes obtained from the conventional transcriptome sequencing library.

[0069] Figure 7 This is a comparison of gene expression levels obtained from the high-throughput sequencing library provided by the present invention in Example 3 with gene expression levels obtained from conventional transcriptome sequencing libraries.

[0070] Figure 8 The results show the gene expression level analysis between two biological duplicates of different samples in Example 4.

[0071] Figure 9 This is the genetic map constructed in Example 5.

[0072] Figure 10 This is an analysis of eQTL in Example 5. Detailed Implementation

[0073] The following examples are provided to better understand the present invention, but are not intended to limit the invention. Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the experimental materials used in the following examples were purchased from conventional biochemical reagent suppliers.

[0074] Example 1: Primer Design

[0075] I. Reverse Transcription Primer Design

[0076] Examples of reverse transcription primer information are shown in Table 1.

[0077] Table 1 Information on reverse transcription primers

[0078]

[0079]

[0080]

[0081] In Table 1, for each primer, positions 1 to 34 at the 5' end are Illumina sequencing adapters (in practical applications, these can be replaced with different types of sequencing adapters for the Illumina platform, or with sequencing adapters required by other sequencing platforms). Positions 35 to 40 are 6 bp primary barcodes (in practical applications, the length can be adjusted arbitrarily as needed, but 4-20 bp is generally preferred. The lengths of barcodes corresponding to different reverse transcription primers can also differ as needed. To reduce the impact of sequencing errors on barcode recognition, the number of base differences between any two different barcodes should ideally be greater than or equal to two). Positions 41 to 46 are 6 bp unique molecular identification codes (in practical applications, the length can be adjusted arbitrarily as needed, but 4-20 bp is generally preferred. The lengths of unique molecular identification codes corresponding to different reverse transcription primers can also differ as needed). Positions 47 to 64 are Oligo(dT) codes (in practical applications, the length can be adjusted as needed, but 15-30 bp is generally preferred). The 65th degenerate base can be A, C, or G, which helps to improve the binding efficiency between the reverse transcription primer and the polyA of the adjacent 3'UTR.

[0082] In practical applications, the positions of the primary barcode and the unique molecule identification code can also be swapped.

[0083] Primary barcodes can be used to distinguish data generated from different test samples within the same high-throughput transcriptome library; unique molecular identification codes (IMC codes) distinguish different transcripts of the same gene in the same sample and can be used to effectively identify and remove repetitive sequences generated during the PCR process. IMC codes may be omitted if needed.

[0084] In practical applications, each sample to be tested corresponds to a reverse transcription primer in Table 1 (corresponding to a unique primary barcode), and different samples can be distinguished based on different primary barcodes.

[0085] II. PCR enrichment primer design

[0086] PCR common primers (sequence 104):

[0087] 5'- AATGATACGGCGACCACCGAGATCTACAC TCGTCGGCAGCGTCAG-3'

[0088] The PCR common primers are Illumina sequencing adapter sequences. The underlined sequence binds to the sequencing chip, and its right side can bind to the sequencing adapter carried by the Tn5 transposase.

[0089] Examples of PCR index primer information are shown in Table 2.

[0090] Table 2 PCR index primer information

[0091] serial number Sequence (5'-3') PCR index P1 <![CDATA[caagcagaagacggcatacgagat CGTGAT gtgactggagttcagacgtgtgc (Sequence 97)]]> PCR index P2 <![CDATA[caagcagaagacggcatacgagat ACATCG gtgactggagttcagacgtgtgc (Sequence 98)]]> PCR index P3 <![CDATA[caagcagaagacggcatacgagat GCCTAA gtgactggagttcagacgtgtgc (Sequence 99)]]> PCR index P4 <![CDATA[caagcagaagacggcatacgagat TGGTCA gtgactggagttcagacgtgtgc (Sequence 100)]]> PCR index P5 <![CDATA[caagcagaagacggcatacgagat CACTGT gtgactggagttcagacgtgtgc(Sequence 101)]]> PCR index P6 <![CDATA[caagcagaagacggcatacgagat ATTGGC gtgactggagttcagacgtgtgc (Sequence 102)]]> PCR index P7 <![CDATA[caagcagaagacggcatacgagat GATCTG gtgactggagttcagacgtgtgc (Sequence 103)]]>

[0092] In Table 2, the underlined sections represent secondary barcodes. The sequences on either side of the underline are Illumina sequencing adapters; the left sequence binds to the sequencing chip, and the right sequence binds to the reverse transcription primer's sequencing adapter. (In practical applications, these can be replaced with different types of sequencing adapters for the Illumina platform, or with sequencing adapters required by other sequencing platforms.)

[0093] In practical applications, the length of secondary barcodes can vary, and the specific length needs to be determined based on the read length of the selected sequencing platform. To reduce the impact of sequencing errors on barcode recognition, the number of base differences between any two different barcodes should ideally be greater than or equal to two bases. Secondary barcodes can be used to distinguish data generated from different high-throughput transcriptome libraries.

[0094] In practical applications, each library corresponds to a PCR index primer (with a unique secondary barcode), and different libraries can be distinguished based on their different secondary barcodes.

[0095] Example 2: Library Construction and Sequencing Methods

[0096] See the document library construction flowchart. Figure 1 The specific steps are as follows (taking 96 samples as an example):

[0097] I. Extraction and quantification of total RNA from the sample to be tested

[0098] 1. Extract total RNA from the sample to be tested.

[0099] 2. Quantify the total RNA extracted in step 1 and adjust the concentration to 1000 ng / μl.

[0100] 3. Add 0.5 μl of RQ1 RNase-Free DNase 10X Reaction Buffer and 0.5 μl of RQ1 RNase-Free DNase (Promega, catalog number: M6101) to 4 μl of total RNA quantified in step 2, mix well, briefly centrifuge, incubate at 37°C for 30 minutes, add 0.5 μl of RQ1 DNase Stop Solution (Promega, catalog number: M6101), and incubate at 65°C for 10 minutes to inactivate the DNase.

[0101] The reagent in step 3 is from Promega, product number: M6101.

[0102] II. Synthesis of the first strand of cDNA

[0103] 1. After completing step one, add 1 μl of the reverse transcription primers from Table 1 (10 μM / μl), incubate at 70°C for 5 minutes, and immediately place on ice. Add one reverse transcription primer from Table 1 to each sample, and add different reverse transcription primers to different samples.

[0104] 2. After completing step 1, add 4 μl of reverse transcription reagent mixture, mix well, and centrifuge briefly.

[0105] Reverse transcription reagent mixture: 5×Superscript III buffer 2μl, RNasin Ribonuclease Inhibitors 0.3μl (Promega, catalog number: N2111), dNTPs (10mM) 0.5μl (NEB, catalog number: N0447), DTT (0.1M) 0.5μl, Superscript III 0.3μl (Invitrogen, 18080051), ddH2O 0.4μl.

[0106] The 5×Superscript III buffer, DTT, and Superscript III are from Invitrogen, product number: 18080051.

[0107] 3. After completing step 2, reverse transcription is performed.

[0108] Reverse transcription program: 25℃, 10 min; 50℃, 50 min; 70℃, 15 min; 4℃, Hold.

[0109] After reverse transcription, the first strand of cDNA from different samples was obtained.

[0110] Take 3 μl of the first strand of cDNA from each of the above different samples and mix them into the same reaction tube, for a total of 288 μl (at this point, 96 samples with different primary barcodes are mixed together).

[0111] III. Synthesis of the second strand of cDNA

[0112] 1. RNA digestion: After completing step 2, add 2 μl of RNase A (ThermoScientific, catalog number: EN0531) to the obtained 288 μl sample and incubate at 37°C for 30 minutes.

[0113] 2. Product purification:

[0114] (1) Add an equal volume of VAHTS DNA Clean Beads (Vazyme, catalog number: N411-03), mix thoroughly by pipetting 10 times, and incubate at room temperature for 5 minutes.

[0115] (2) Place the reaction tube on the magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear, carefully remove the supernatant.

[0116] (3) Keep the reaction tube on the magnetic rack at all times, add 1 ml of freshly prepared 75% (volume percentage) ethanol aqueous solution to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0117] (4) Repeat step (3) for a total of two rinses.

[0118] (5) Keep the reaction tube on the magnetic rack at all times and open the lid to dry it for about 5 minutes.

[0119] (6) Remove the reaction tube from the magnetic rack and add 39 μl of sterile ultrapure water to elute. Mix thoroughly by pipetting 10 times and incubate at room temperature for 2 min.

[0120] (7) Briefly centrifuge the reaction tube and place it on a magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear (about 5 min), carefully aspirate 36.5 μl of supernatant into a new sterile PCR tube to obtain the purified cDNA first strand.

[0121] 3. cDNA second-strand synthesis:

[0122] Add 13.5 μl of the second-strand synthesis reagent mixture to the first strand of the purified cDNA obtained in step 2 above, mix well, and centrifuge briefly.

[0123] Two-strand synthesis reagent mixture: NEBuffer 2 (10×) (NEB, catalog number: B7002) 5μl, dNTP Mix (10mM) (NEB, catalog number: N0447) 2.5μl, DNA Polymerase I (Invitrogen, catalog number: 18010025, 10U / μL) 5μl, RNase H (Thermo Scientific, catalog number: EN0201, 5U / μL) 1μl.

[0124] Reaction procedure: 16℃, 2.5 hours; 4℃, Hold, to obtain double-stranded cDNA.

[0125] 4. Product purification:

[0126] (1) Add an equal volume of VAHTS DNA Clean Beads to the double-stranded cDNA obtained in step 3 above, mix thoroughly by pipetting 10 times, and incubate at room temperature for 5 minutes.

[0127] (2) Place the reaction tube on the magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear, carefully remove the supernatant.

[0128] (3) Keep the reaction tube on the magnetic rack at all times, add 1 ml of freshly prepared 75% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0129] (4) Repeat step (3) for a total of two rinses.

[0130] (5) Keep the reaction tube on the magnetic rack at all times and open the lid to dry it for about 5 minutes.

[0131] (6) Remove the reaction tube from the magnetic rack and add 27 μl of sterile ultrapure water to elute. Mix thoroughly by pipetting 10 times and incubate at room temperature for 2 min.

[0132] (7) Briefly centrifuge the reaction tube and place it on a magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear (about 5 min), carefully aspirate 25 μl of the supernatant into a new sterile PCR tube to obtain the purified product.

[0133] 5. Quantify the purified product (double-stranded cDNA) obtained in step 4.

[0134] IV. Double-stranded cDNA fragmentation

[0135] The sequencing adapters were assembled using the Tn5 transposase included in the Novizan DNA Library Construction Kit (Vazyme, TD501). (The sequencing adapters carried by the Tn5 transposase are different from those in the reverse transcription primers in Table 1.)

[0136] The adapter sequence it carries is “TCGTCGGCAGCGTCAGATGTGTATAAGAGACAG (Sequence 105)”, which, together with the sequencing adapter in the reverse transcription primer, forms the sequencing adapters required for sequencing, connecting the two ends of the target fragment. The double-stranded cDNA is fragmented as follows:

[0137] (1) Prepare the reaction system in a sterile RCR tube: 10 μl of 5×TTBL (Vazyme, catalog number: TD501), 50 ng of the double-stranded cDNA obtained above, 5 μl of TTE Mix V50 (Vazyme, catalog number: TD501), and ddH2O to 50 μl.

[0138] (2) Use a pipette to gently blow and mix 20 times to thoroughly mix. Place the RCR tube in the PCR instrument and run the reaction program: 105℃, hot cap; 55℃, 10 min; 10℃, Hold.

[0139] (3) Add an equal volume of VAHTS DNA Clean Beads, mix thoroughly by pipetting 10 times, and incubate at room temperature for 5 minutes.

[0140] (4) Place the PCR tube on a magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear, carefully remove the supernatant.

[0141] (5) Keep the PCR tube on the magnetic rack at all times, add 1 ml of freshly prepared 75% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0142] (6) Repeat step (5) for a total of two rinses.

[0143] (7) Keep the PCR tube on the magnetic rack and let it air dry for about 5 minutes.

[0144] (8) Remove the PCR tube from the magnetic rack and add 26 μl of sterile ultrapure water to wash it off. Mix thoroughly by pipetting 10 times and incubate at room temperature for 2 min.

[0145] (9) Briefly centrifuge the PCR tube and place it on a magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear (about 5 min), carefully aspirate 24 μl of the supernatant into a new sterile PCR tube to obtain the fragmented product.

[0146] V. PCR enrichment

[0147] 1. Add 26 μl of reagent mixture to the fragmented product obtained in step 4 above, mix well, and centrifuge briefly.

[0148] Reagent mixture: 5×TAB 10μl, PPM 5μl, PCR common primer 5μl, PCR index P1 primer 5μl, TAE 1μl.

[0149] The reagents were obtained from Novizan's DNA library construction kit (Vazyme, TD501).

[0150] 2. PCR enrichment: 105℃ hot cap; 72℃ for 3 min; 98℃ for 30 sec, 98℃ for 15 sec, 60℃ for 30 sec, for a total of 12 cycles; 72℃ for 3 min, 72℃ for 5 min, 4℃ Hold.

[0151] PCR products were obtained.

[0152] One PCR index primer from Table 2 is added to each library. If more than one library needs to be mixed for sequencing, different PCR index primers are added to each library for PCR enrichment, and then the libraries are mixed according to the required data volume before subsequent operations. For the above 96 samples, they can be mixed and treated as one library, with the PCR index primers added being PCR index P1 from Table 2.

[0153] VI. PCR amplification product length sorting

[0154] The preferred sorting length is 400-600 bp. The sorting length can be adjusted according to the actual library status and sequencing strategy.

[0155] 1. Add 0.2 times the volume (10 μl) of VAHTS DNA Clean Beads to 50 μl of the PCR product obtained in step 5 above, mix thoroughly by pipetting 10 times, and incubate at room temperature for 5 min.

[0156] 2. Place the PCR tube on the magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear, carefully transfer the supernatant to a new sterile PCR tube and discard the magnetic beads.

[0157] This step removes long PCR products. The smaller the volume of VAHTS DNA Clean Beads added in step 1, the longer the long PCR products removed will be.

[0158] 3. Add 0.6 times the volume (30 μl) of VAHTS DNA Clean Beads to the supernatant, mix thoroughly by pipetting 10 times, and incubate at room temperature for 5 min.

[0159] This step removes short PCR products. The larger the total volume of VAHTS DNA Clean Beads added in steps 1 and 3, the shorter the length of the short PCR products removed.

[0160] 4. Keep the PCR tube on the magnetic rack at all times, add 200 μl of freshly prepared 75% ethanol to rinse the magnetic beads, incubate at room temperature for 30 seconds, and carefully remove the supernatant.

[0161] 5. Repeat step 4, rinsing a total of two times.

[0162] 6. Keep the PCR tube on the magnetic rack at all times and allow it to air dry for about 5 minutes.

[0163] 7. Remove the PCR tube from the magnetic rack and add 22 μl of sterile ultrapure water to wash it off. Mix thoroughly by pipetting 10 times and incubate at room temperature for 2 min.

[0164] 8. Briefly centrifuge the PCR tube and place it on a magnetic rack to separate the magnetic beads from the liquid. After the solution becomes clear (about 5 minutes), carefully aspirate 20 μl of the supernatant into a new sterile PCR tube and store at -20°C.

[0165] VII. Library Quality Control and Sequencing

[0166] The product obtained in step six was diluted to 1 ng / μl. 1 μl was used to analyze the length distribution of the inserted fragments in the library using an Agilent 2100 (Agilent Technologies, USA). Another 1 μl was used for real-time quantitative PCR detection; the concentration for sequencing was determined based on the results. After diluting the library to the required level (2 nM), paired-end sequencing was performed on the Illumina Xten sequencing platform.

[0167] 8. Splitting Sequencing Data from Different Samples

[0168] 1. Based on the sequence information of the secondary barcodes corresponding to the PCR index primers used in step five PCR enrichment, the sequencing data obtained in step seven are allocated to different libraries. Each library contains data from 96 sequencing samples.

[0169] 2. Based on the primary barcode information corresponding to the reverse transcription primers used in the synthesis of the first strand of cDNA in step two, the data of each library is further allocated to the corresponding 96 sequencing samples.

[0170] Example 3: Construction and sequencing of transcriptome libraries from Arabidopsis thaliana and maize samples

[0171] Experimental materials: maize inbred line PHBA6: stems and leaves; maize inbred line Chang7-2: stems and leaves; Arabidopsis thaliana Columbia-0: stems; all maize and Arabidopsis materials were obtained from the National Maize Improvement Center of China Agricultural University.

[0172] Stems of maize inbred line PHBA6, leaves of maize inbred line PHBA6, stems of maize inbred line Chang7-2, leaves of maize inbred line Chang7-2, and stems of Arabidopsis thaliana Columbia-0 were used as five test samples. Two biological replicates were set up for each sample, i.e., two libraries were constructed.

[0173] Transcriptome library construction and sequencing were performed according to the method in Example 2, as follows:

[0174] I. Extraction and quantification of total RNA from the sample to be tested: Same as in Example 2;

[0175] II. Synthesis of the first strand of cDNA: Same as in Example 2.

[0176] Biological replicate 1: The reverse transcription primers RTP1, RTP2, RTP3, RTP4, and RTP5 listed in Table 1 were used for the biological replicate 1 samples of maize inbred line PHBA6 stem, maize inbred line PHBA6 leaf, maize inbred line Chang7-2 stem, maize inbred line Chang7-2 leaf, and Arabidopsis thaliana Columbia-0 stem, respectively.

[0177] Biological replicate 2: The reverse transcription primers RTP6, RTP7, RTP8, RTP9, and RTP10 listed in Table 1 were used for the samples of the technical replicate 2 of maize inbred line PHBA6 stem, maize inbred line PHBA6 leaf, maize inbred line Chang7-2 stem, maize inbred line Chang7-2 leaf, and Arabidopsis thaliana Columbia-0 stem.

[0178] First strand of cDNA was obtained from five samples of biological repeat 1 and biological repeat 2, respectively.

[0179] III. Synthesis of cDNA second strand: The cDNA first strands of the five samples of biological repeat 1 and biological repeat 2 above were mixed and the cDNA second strands were synthesized according to the method in Example 2.

[0180] IV. Double-stranded cDNA fragmentation: Same as in Example 2;

[0181] V. PCR enrichment: Same as in Example 2: The PCR index primer for biological repeat 1 is PCR index P1, and the PCR index primer for biological repeat 2 is PCR index P2.

[0182] VI. PCR amplification product length sorting: Same as in Example 2;

[0183] VII. Library quality control and sequencing: Same as in Example 2;

[0184] 8. Splitting of sequencing data from different samples: Same as in Example 2;

[0185] 1. Based on the sequence information of the secondary barcodes corresponding to the PCR index P1 and PCR index P2 primers used in the PCR enrichment in step five, the sequencing data obtained in step seven is split to obtain the sequencing data of biological repeat 1 library and the sequencing data of biological repeat 2 library, each containing data from 5 sequencing samples.

[0186] 2. Based on the primary barcode information corresponding to the reverse transcription primers used in the synthesis of the first strand of cDNA in step two, the data of biological repeat 1 library is further allocated to the corresponding 5 sequencing samples, and the data of biological repeat 2 library is allocated to the corresponding 5 sequencing samples.

[0187] Evaluate the length distribution of inserted fragments in the library, such as Figure 2 As shown (the top image corresponds to the biological repeat 1 library, and the bottom image corresponds to the biological repeat 2 library), the total length of the constructed technical repeat 1 and 2 libraries is between 400-700 bp.

[0188] The quality of the reads produced by sequencing is evaluated, such as... Figure 3 As shown (the top left image shows the quality assessment results of reads 1 of the biological repeat 1 library obtained from sequencing, the top right image shows the quality assessment results of reads 2 of the biological repeat 1 library obtained from sequencing, the bottom left image shows the quality assessment results of reads 1 of the biological repeat 2 library obtained from sequencing, and the bottom right image shows the quality assessment results of reads 2 of the biological repeat 2 library obtained from sequencing), the percentage of low-quality bases (<20) in reads 1 is relatively low, and the percentage of low-quality bases (<20) in the first 12 bases of reads 2 is also relatively low.

[0189] Sequencing reads2 are used to obtain primary barcode information and unique molecular identification code information, while sequencing reads1 are used to analyze gene expression levels.

[0190] The coverage of sequencing reads on the transcripts was evaluated, using the Chang7-2 stem sample as an example. Figure 4As shown in the figure (left image is Chang7-2 stem biological repeat 1, right image is Chang7-2 stem biological repeat 2), the results show that the sequencing reads are mainly concentrated at the 3' end of the transcript, indicating that the constructed high-throughput sequencing library can effectively enrich the 3' end fragment of the transcript.

[0191] Analyzing gene expression levels between two biological duplicates from different samples, such as... Figure 5 As shown ( Figure 5 The correlation of gene expression levels between two biological replicates of five materials (stem of maize inbred line PHBA6, leaf of maize inbred line PHBA6, stem of maize inbred line Chang7-2, leaf of maize inbred line Chang7-2, and stem of Arabidopsis thaliana Columbia-0) was studied. The results showed that the average correlation r value of technical replicates was 0.94, indicating that the constructed high-throughput sequencing library has high reproducibility in detecting gene expression levels.

[0192] Five types of test samples were constructed and sequenced using conventional transcriptome sequencing libraries as controls.

[0193] The number of expressed genes detected by the high-throughput sequencing library constructed by the method of this invention is basically the same as the number of expressed genes detected by the control conventional transcriptome sequencing library. Figure 6 )

[0194] Taking stem samples from maize Chang7-2 and Arabidopsis thaliana as examples, the correlation r values ​​between the gene expression levels detected by the high-throughput sequencing library constructed by the method of this invention (biological replication 1) and the gene expression levels detected by the conventional transcriptome sequencing library were 0.89 and 0.91, respectively, indicating that the constructed high-throughput sequencing library can effectively detect gene expression levels in the samples. Figure 7 The left figure compares the gene expression levels obtained from maize Chang7-2 stem samples using the method of this invention and conventional transcriptome sequencing methods. The right figure compares the gene expression levels obtained from Arabidopsis stem samples using the method of this invention and conventional transcriptome sequencing methods.

[0195] Example 4: Construction and sequencing of transcriptome libraries from human and mouse samples

[0196] Experimental materials: Mice: heart and liver (Experimental Animal Center, Institute of Genetics and Developmental Biology).

[0197] Human: HeLacellline (ATCC-CCL2).

[0198] Mouse heart, mouse liver, and HeLacellline cells were used as three test samples, and two biological replicates were set up for each sample, i.e., two libraries were constructed.

[0199] Transcriptome library construction and sequencing were performed according to the method in Example 2, as follows:

[0200] I. Extraction and quantification of total RNA from the sample to be tested: Same as in Example 2;

[0201] II. Synthesis of the first strand of cDNA: Same as in Example 2.

[0202] Biological replicate 1: The biological replicate 1 samples of mouse heart, mouse liver and HeLacellline cells were prepared using the reverse transcription primers RTP1, RTP2 and RTP3 in Table 1, respectively.

[0203] Biological replicate 2: The biological replicate 2 samples of mouse heart, mouse liver and HeLacellline cells were prepared using the reverse transcription primers RTP:4, RTP5 and RTP6 in Table 1.

[0204] The first strand of cDNA was obtained from three samples of biological repeat 1 and biological repeat 2, respectively.

[0205] III. Synthesis of cDNA second strand: The cDNA first strands of the three samples of biological repeat 1 and biological repeat 2 above were mixed and the cDNA second strands were synthesized according to the method in Example 2.

[0206] IV. Double-stranded cDNA fragmentation: Same as in Example 2;

[0207] V. PCR enrichment: Same as in Example 2: The PCR index primer for biological repeat 1 is PCR index P1, and the PCR index primer for biological repeat 2 is PCR index P2.

[0208] VI. PCR amplification product length sorting: Same as in Example 2;

[0209] VII. Library quality control and sequencing: Same as in Example 2;

[0210] 8. Splitting of sequencing data from different samples: Same as in Example 2;

[0211] 1. Based on the sequence information of the secondary barcodes corresponding to the PCR index P1 and PCR index P2 primers used in the PCR enrichment in step five, the sequencing data obtained in step seven is split to obtain the sequencing data of biological repeat 1 library and the sequencing data of biological repeat 2 library, each containing data from 3 sequencing samples.

[0212] 2. Based on the primary barcode information corresponding to the reverse transcription primers used in the synthesis of the first strand of cDNA in step two, the data of biological repeat 1 library is further allocated to the corresponding 3 sequencing samples, and the data of biological repeat 2 library is allocated to the corresponding 3 sequencing samples.

[0213] Analyzing gene expression levels between two biological duplicates from different samples, such as... Figure 8 As shown, the results indicate that the technical repeatability correlation r values ​​for mouse heart and liver samples were 0.96 and 0.99, respectively, while the technical repeatability correlation r value for HeLa cell samples was 0.87, indicating that the constructed high-throughput sequencing library has high reproducibility in detecting gene expression levels.

[0214] Example 5: Analysis of genotypes and expression patterns of 477 DH (double haploid) line samples.

[0215] Experimental materials: 477 stem tissue samples of the DH line induced from the F1 generation of the Chang7-2 and PHBA6 hybrids. The Chang7-2, PHBA6, and DH line samples were obtained from the National Maize Improvement Center of China Agricultural University.

[0216] Forty-seventy stem tissues from the DH line were used as 477 test samples. They were divided into 10 groups, with groups 1 to 9 each containing 48 test samples and group 10 containing 45 test samples. Libraries were constructed from groups 1 to 10, resulting in a total of 10 libraries.

[0217] Transcriptome library construction and sequencing were performed according to the method in Example 2, as follows:

[0218] I. Extraction and quantification of total RNA from the sample to be tested: Same as in Example 2;

[0219] II. Synthesis of the first strand of cDNA: Same as in Example 2. The 48 DH lines in groups 1 to 9 used RTP1 to RTP48 as reverse transcription primers in Table 1, and the 45 DH lines in group 10 used RTP1 to RTP45 as reverse transcription primers in Table 1.

[0220] First strand of cDNA was obtained from 477 samples.

[0221] III. Synthesis of cDNA second strand: The cDNA first strands of the samples corresponding to groups 1 to 10 above were mixed, and then the cDNA second strands of groups 1 to 10 were synthesized according to the method in Example 2.

[0222] IV. Double-stranded cDNA fragmentation: Same as in Example 2;

[0223] V. PCR enrichment: Same as in Example 2: The PCR index primers for samples in groups 1 to 5 are PCR index P1 to P5; the PCR index primers for samples in groups 6 to 10 are PCR index P1 to P5.

[0224] VI. PCR amplification product length sorting: Same as in Example 2;

[0225] VII. Library quality control and sequencing: Same as in Example 2;

[0226] 8. Splitting of sequencing data from different samples: Same as in Example 2;

[0227] 1. Based on the sequence information of the secondary barcodes corresponding to the PCR index primers used in step 5 PCR enrichment, the sequencing data obtained in step 7 are allocated to different libraries. Libraries 1 to 9 each contain data from 48 sequencing samples, and library 10 contains data from 45 sequencing samples.

[0228] 2. Based on the primary barcode information corresponding to the reverse transcription primers used in the synthesis of the first strand of cDNA in step two, the data of each library is further allocated to the corresponding sequencing samples.

[0229] Based on the high-throughput transcriptome library construction and sequencing of this invention, a total of 20,584 genes were detected to be expressed in at least 50% of the DH lines. The average number of expressed genes identified per line was 19,453. This indicates that the constructed high-throughput sequencing library can effectively analyze gene expression in large-scale samples.

[0230] Based on the high-throughput transcriptome library construction and sequencing of this invention, a total of 35,836 SNPs were detected, of which 85% were distributed in gene exons. This indicates that the sequencing data of the constructed high-throughput library can be used for SNP mining.

[0231] Based on the SNPs mined from the high-throughput transcriptome sequencing libraries of this invention, such as Figure 9 As shown, a genetic map with a total length of 857 cM was successfully constructed. This indicates that the sequencing data from the constructed high-throughput library can be used for genotype analysis and genetic map construction.

[0232] Gene expression and genotype information obtained from the high-throughput transcriptome sequencing library of this invention, such as Figure 10 As shown, eQTL analysis was successfully performed. A total of 24,994 eQTL sites were identified, regulating 15,652 genes. SEQUENCE LISTING <110> China Agricultural University <120> A high-throughput transcriptome sequencing library construction method and its application <160> 105 <170> PatentIn version 3.5 <210> 1 <211> 65 <212> DNA <213> Artificial sequence <220> <221> misc_feature <222> (41) (46) <223> n is a, c, g, or t <400> 1 gtgactggagttcagacgtgtgctcttccgatctagactcnnnnnntttttttttttttt 60 ttttv 65 <210> 2 <211> 65 <212> DNA <213> Artificial sequence <220> <221> misc_feature <222> (41) (46) <223> n is a, c, g, or t <400> 2 gtgactggagttcagacgtgtgctcttccgatctagctagnnnnnntttttttttttttt 60 ttttv 65 <210> 3 <211> 65 <212> DNA <213> Artificial sequence <220> <221> misc_feature <222> (41) (46) <223> n is a, c, g, or t <400> 3 gtgactggagttcagacgtgtgctcttccgatctagctcannnnnntttttttttttttt 60 ttttv 65 <210> 4 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 4 gtgactggagttcagacgtgtgctcttccgatctagcttcnnnnnntttttttttttttt 60 ttttv 65 <210> 5 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 5 gtgactggagttcagacgtgtgctcttccgatctcatgagnnnnnntttttttttttttt 60 ttttv 65 <210> 6 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 6 gtgactggagttcagacgtgtgctcttccgatctcatgcannnnnntttttttttttttt 60 ttttv 65 <210> 7 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 7 gtgactggagttcagacgtgtgctcttccgatctcatgtcnnnnnntttttttttttttt 60 ttttv 65 <210> 8 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 8 gtgactggagttcagacgtgtgctcttccgatctcactagnnnnnntttttttttttttt 60 ttttv 65 <210> 9 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 9 gtgactggagttcagacgtgtgctcttccgatctcagatcnnnnnntttttttttttttt 60 ttttv 65 <210> 10 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 10 gtgactggagttcagacgtgtgctcttccgatcttcacagnnnnnntttttttttttttt 60 ttttv 65 <210> 11 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 11 gtgactggagttcagacgtgtgctcttccgatctaggatcnnnnnntttttttttttttt 60 ttttv 65 <210> 12 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 12 gtgactggagttcagacgtgtgctcttccgatctagtgcannnnnntttttttttttttt 60 ttttv 65 <210> 13 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 13 gtgactggagttcagacgtgtgctcttccgatctagtgtcnnnnnntttttttttttttt 60 ttttv 65 <210> 14 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 14 gtgactggagttcagacgtgtgctcttccgatcttcctagnnnnnntttttttttttttt 60 ttttv 65 <210> 15 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 15 gtgactggagttcagacgtgtgctcttccgatcttctgagnnnnnntttttttttttttt 60 ttttv 65 <210> 16 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 16 gtgactggagttcagacgtgtgctcttccgatcttctgcannnnnntttttttttttttt 60 ttttv 65 <210> 17 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 17 gtgactggagttcagacgtgtgctcttccgatcttcgaagnnnnnntttttttttttttt 60 ttttv 65 <210> 18 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 18 gtgactggagttcagacgtgtgctcttccgatcttcgacannnnnntttttttttttttt 60 ttttv 65 <210> 19 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 19 gtgactggagttcagacgtgtgctcttccgatcttcgatcnnnnnntttttttttttttt 60 ttttv 65 <210> 20 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 20 gtgactggagttcagacgtgtgctcttccgatctgtacagnnnnnntttttttttttttt 60 ttttv 65 <210> 21 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 21 gtgactggagttcagacgtgtgctcttccgatctgtaccannnnnntttttttttttttt 60 ttttv 65 <210> 22 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 22 gtgactggagttcagacgtgtgctcttccgatctgtactcnnnnnntttttttttttttt 60 ttttv 65 <210> 23 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 23 gtgactggagttcagacgtgtgctcttccgatctgtctagnnnnnntttttttttttttt 60 ttttv 65 <210> 24 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 24 gtgactggagttcagacgtgtgctcttccgatctgtctcannnnnntttttttttttttt 60 ttttv 65 <210> 25 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 25 gtgactggagttcagacgtgtgctcttccgatctgttgcannnnnntttttttttttttt 60 ttttv 65 <210> 26 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 26 gtgactggagttcagacgtgtgctcttccgatctgtgacannnnnntttttttttttttt 60 ttttv 65 <210> 27 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 27 gtgactggagttcagacgtgtgctcttccgatctgtgatcnnnnnntttttttttttttt 60 ttttv 65 <210> 28 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 28 gtgactggagttcagacgtgtgctcttccgatctacagtgnnnnnntttttttttttttt 60 ttttv 65 <210> 29 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 29 gtgactggagttcagacgtgtgctcttccgatctaccatgnnnnnntttttttttttttt 60 ttttv 65 <210> 30 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 30 gtgactggagttcagacgtgtgctcttccgatctactctgnnnnnntttttttttttttt 60 ttttv 65 <210> 31 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 31 gtgactggagttcagacgtgtgctcttccgatctactcgannnnnntttttttttttttt 60 ttttv 65 <210> 32 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 32 gtgactggagttcagacgtgtgctcttccgatctacgtacnnnnnntttttttttttttt 60 ttttv 65 <210> 33 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 33 gtgactggagttcagacgtgtgctcttccgatctacgttgnnnnnntttttttttttttt 60 ttttv 65 <210> 34 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 34 gtgactggagttcagacgtgtgctcttccgatctacgtgannnnnntttttttttttttt 60 ttttv 65 <210> 35 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 35 gtgactggagttcagacgtgtgctcttccgatctctagacnnnnnntttttttttttttt 60 ttttv 65 <210> 36 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 36 gtgactggagttcagacgtgtgctcttccgatctctagtgnnnnnntttttttttttttt 60 ttttv 65 <210> 37 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 37 gtgactggagttcagacgtgtgctcttccgatctctaggannnnnntttttttttttttt 60 ttttv 65 <210> 38 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 38 gtgactggagttcagacgtgtgctcttccgatctctcatgnnnnnntttttttttttttt 60 ttttv 65 <210> 39 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 39 gtgactggagttcagacgtgtgctcttccgatctctcagannnnnntttttttttttttt 60 ttttv 65 <210> 40 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 40 gtgactggagttcagacgtgtgctcttccgatctcttcgannnnnntttttttttttttt 60 ttttv 65 <210> 41 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 41 gtgactggagttcagacgtgtgctcttccgatctctgtacnnnnnntttttttttttttt 60 ttttv 65 <210> 42 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 42 gtgactggagttcagacgtgtgctcttccgatctctgtgannnnnntttttttttttttt 60 ttttv 65 <210> 43 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 43 gtgactggagttcagacgtgtgctcttccgatcttgagacnnnnnntttttttttttttt 60 ttttv 65 <210> 44 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 44 gtgactggagttcagacgtgtgctcttccgatcttgcaacnnnnnntttttttttttttt 60 ttttv 65 <210> 45 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 45 gtgactggagttcagacgtgtgctcttccgatcttgcatgnnnnnntttttttttttttt 60 ttttv 65 <210> 46 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 46 gtgactggagttcagacgtgtgctcttccgatcttgcagannnnnntttttttttttttt 60 ttttv 65 <210> 47 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 47 gtgactggagttcagacgtgtgctcttccgatcttgtcacnnnnnntttttttttttttt 60 ttttv 65 <210> 48 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 48 gtgactggagttcagacgtgtgctcttccgatcttgtcgannnnnntttttttttttttt 60 ttttv 65 <210> 49 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 49 gtgactggagttcagacgtgtgctcttccgatcttggtacnnnnnntttttttttttttt 60 ttttv 65 <210> 50 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 50 gtgactggagttcagacgtgtgctcttccgatctgacatgnnnnnntttttttttttttt 60 ttttv 65 <210> 51 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 51 gtgactggagttcagacgtgtgctcttccgatctgatcacnnnnnntttttttttttttt 60 ttttv 65 <210> 52 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 52 gtgactggagttcagacgtgtgctcttccgatctgatctgnnnnnntttttttttttttt 60 ttttv 65 <210> 53 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 53 gtgactggagttcagacgtgtgctcttccgatctgatcgannnnnntttttttttttttt 60 ttttv 65 <210> 54 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 54 gtgactggagttcagacgtgtgctcttccgatctgagtacnnnnnntttttttttttttt 60 ttttv 65 <210> 55 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 55 gtgactggagttcagacgtgtgctcttccgatctagacagnnnnnntttttttttttttt 60 ttttv 65 <210> 56 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 56 gtgactggagttcagacgtgtgctcttccgatctagaccannnnnntttttttttttttt 60 ttttv 65 <210> 57 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 57 gtgactggagttcagacgtgtgctcttccgatctagtgagnnnnnntttttttttttttt 60 ttttv 65 <210> 58 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 58 gtgactggagttcagacgtgtgctcttccgatctaggaagnnnnnntttttttttttttt 60 ttttv 65 <210> 59 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 59 gtgactggagttcagacgtgtgctcttccgatctaggacannnnnntttttttttttttt 60 ttttv 65 <210> 60 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 60 gtgactggagttcagacgtgtgctcttccgatctcaacagnnnnnntttttttttttttt 60 ttttv 65 <210> 61 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 61 gtgactggagttcagacgtgtgctcttccgatctcaaccannnnnntttttttttttttt 60 ttttv 65 <210> 62 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 62 gtgactggagttcagacgtgtgctcttccgatctcaactcnnnnnntttttttttttttt 60 ttttv 65 <210> 63 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 63 gtgactggagttcagacgtgtgctcttccgatctcactcannnnnntttttttttttttt 60 ttttv 65 <210> 64 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 64 gtgactggagttcagacgtgtgctcttccgatctcacttcnnnnnntttttttttttttt 60 ttttv 65 <210> 65 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 65 gtgactggagttcagacgtgtgctcttccgatctcagaagnnnnnntttttttttttttt 60 ttttv 65 <210> 66 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 66 gtgactggagttcagacgtgtgctcttccgatctcagacannnnnntttttttttttttt 60 ttttv 65 <210> 67 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 67 gtgactggagttcagacgtgtgctcttccgatcttcaccannnnnntttttttttttttt 60 ttttv 65 <210> 68 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 68 gtgactggagttcagacgtgtgctcttccgatcttcactcnnnnnntttttttttttttt 60 ttttv 65 <210> 69 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 69 gtgactggagttcagacgtgtgctcttccgatcttcctcannnnnntttttttttttttt 60 ttttv 65 <210> 70 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 70 gtgactggagttcagacgtgtgctcttccgatcttccttcnnnnnntttttttttttttt 60 ttttv 65 <210> 71 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 71 gtgactggagttcagacgtgtgctcttccgatcttctgtcnnnnnntttttttttttttt 60 ttttv 65 <210> 72 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 72 gtgactggagttcagacgtgtgctcttccgatctgtcttcnnnnnntttttttttttttt 60 ttttv 65 <210> 73 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 73 gtgactggagttcagacgtgtgctcttccgatctgttgagnnnnnntttttttttttttt 60 ttttv 65 <210> 74 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 74 gtgactggagttcagacgtgtgctcttccgatctgttgtcnnnnnntttttttttttttt 60 ttttv 65 <210> 75 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 75 gtgactggagttcagacgtgtgctcttccgatctgtgaagnnnnnntttttttttttttt 60 ttttv 65 <210> 76 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 76 gtgactggagttcagacgtgtgctcttccgatctacagacnnnnnntttttttttttttt 60 ttttv 65 <210> 77 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 77 gtgactggagttcagacgtgtgctcttccgatctacaggannnnnntttttttttttttt 60 ttttv 65 <210> 78 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 78 gtgactggagttcagacgtgtgctcttccgatctaccaacnnnnnntttttttttttttt 60 ttttv 65 <210> 79 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 79 gtgactggagttcagacgtgtgctcttccgatctaccagannnnnntttttttttttttt 60 ttttv 65 <210> 80 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 80 gtgactggagttcagacgtgtgctcttccgatctactcacnnnnnntttttttttttttt 60 ttttv 65 <210> 81 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 81 gtgactggagttcagacgtgtgctcttccgatctctcaacnnnnnntttttttttttttt 60 ttttv 65 <210> 82 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 82 gtgactggagttcagacgtgtgctcttccgatctcttcacnnnnnntttttttttttttt 60 ttttv 65 <210> 83 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 83 gtgactggagttcagacgtgtgctcttccgatctcttctgnnnnnntttttttttttttt 60 ttttv 65 <210> 84 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 84 gtgactggagttcagacgtgtgctcttccgatctctgttgnnnnnntttttttttttttt 60 ttttv 65 <210> 85 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 85 gtgactggagttcagacgtgtgctcttccgatcttgagtgnnnnnntttttttttttttt 60 ttttv 65 <210> 86 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 86 gtgactggagttcagacgtgtgctcttccgatcttgaggannnnnntttttttttttttt 60 ttttv 65 <210> 87 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 87 gtgactggagttcagacgtgtgctcttccgatcttgtctgnnnnnntttttttttttttt 60 ttttv 65 <210> 88 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 88 gtgactggagttcagacgtgtgctcttccgatcttggttgnnnnnntttttttttttttt 60 ttttv 65 <210> 89 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 89 gtgactggagttcagacgtgtgctcttccgatcttggtgannnnnntttttttttttttt 60 ttttv 65 <210> 90 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 90 gtgactggagttcagacgtgtgctcttccgatctgaagacnnnnnntttttttttttttt 60 ttttv 65 <210> 91 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 91 gtgactggagttcagacgtgtgctcttccgatctgaagtgnnnnnntttttttttttttt 60 ttttv 65 <210> 92 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 92 gtgactggagttcagacgtgtgctcttccgatctgaaggannnnnntttttttttttttt 60 ttttv 65 <210> 93 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 93 gtgactggagttcagacgtgtgctcttccgatctgacaacnnnnnntttttttttttttt 60 ttttv 65 <210> 94 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 94 gtgactggagttcagacgtgtgctcttccgatctgacagannnnnntttttttttttttt 60 ttttv 65 <210> 95 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 95 gtgactggagttcagacgtgtgctcttccgatctgagttgnnnnnntttttttttttttt 60 ttttv 65 <210> 96 <211> 65 <212> DNA <213>Artificial sequence <220> <221>misc_feature <222> (41)..(46) <223> n is a, c, g, or t <400> 96 gtgactggagttcagacgtgtgctcttccgatctgagtgannnnnntttttttttttttt 60 ttttv 65 <210> 97 <211> 53 <212> DNA <213>Artificial sequence <400> 97 caagcagaagacggcatacgagatcgtgatgtgactggagttcagacgtgtgc 53 <210> 98 <211> 53 <212> DNA <213>Artificial sequence <400> 98 caagcagaagacggcatacgagatacatcggtgactggagttcagacgtg tgc 53 <210> 99 <211> 53 <212> DNA <213>Artificial sequence <400> 99 caagcagaagacggcatacgagatgcctaagtgactggagttcagacgtgtgc 53 <210> 100 <211> 53 <212> DNA <213>Artificial sequence <400> 100 caagcagaagacggcatacgagattggtcagtgactggagttcagacgtgtgc 53 <210> 101 <211> 53 <212> DNA <213>Artificial sequence <400> 101 caagcagaagacggcatacgagatcactgtgtgactggagttcagacgtgtgc 53 <210> 102 <211> 53 <212> DNA <213>Artificial sequence <400> 102 caagcagaagacggcatacgagatattggcgtgactggagttcagacgtgtgc 53 <210> 103 <211> 53 <212> DNA <213>Artificial sequence <400> 103 caagcagaagacggcatacgagatgatctggtgactggagttcagacgtgtgc 53 <210> 104 <211> 45 <212> DNA <213>Artificial sequence <400> 104 aatgatacggcgaccaccgagatctacactcgtcggcagcgtcag 45 <210> 105 <211> 33 <212> DNA <213>Artificial sequence <400> 105 tcgtcggcagcgtcagatgtgtataagaga cag 33

Claims

1. A method for constructing a high-throughput transcriptome sequencing library, comprising the following steps: (1) Extract total RNA from N samples to be tested; (2) Using the total RNA of each of the test samples as templates, N reverse transcription primers are used to reverse transcribe N reverse transcription products, which are the first strands of cDNA of the N test samples. Each of the reverse transcription primers includes a sequencing adapter, a primary barcode for distinguishing different samples, polythymine, and a degenerate V at the 3' end; the polythymine is a continuous T base of 15-30 nt; the degenerate V base is A, C, or G. Each sample to be tested corresponds to one of the aforementioned reverse transcription primers; The primary barcodes in the N reverse transcription primers are all different; Each of the reverse transcription primers also includes a unique molecular identification code used to distinguish different transcripts of the same gene within the same sample; The unique molecular identification code is located between the primary barcode and the polythymine; Alternatively, the unique molecular identification code may be located between the sequencing adapter in the reverse transcription primer and the primary barcode; (3) Mix the N reverse transcription products to obtain a first strand of mixed cDNA; then use the first strand of mixed cDNA to synthesize a second strand of cDNA to obtain double-stranded cDNA; (4) The double-stranded cDNA was fragmented using Tn5 transposase with sequencing adapters assembled to obtain fragmented products; The sequencing adapter in the Tn5 transposase assembled with the sequencing adapter is different from the sequencing adapter in the reverse transcription primer and can form a sequencing adapter set with the sequencing adapter in the reverse transcription primer. (5) The fragmented product is amplified by PCR to obtain a sequencing library; The PCR amplification introduces a secondary barcode used to distinguish different libraries by amplifying at least one primer from the required primer pair; The primer pair required for the amplification is any one of the following 1)-3): 1) The primer pair shown consists of M primers A and B; Primer A includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer, a secondary barcode for distinguishing different libraries, and a sequencing adapter that binds to the sequencing chip. Each primer B includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase and another sequencing adapter that binds to the sequencing chip. The secondary barcodes in the M primers A are different; 2) The primer pair shown consists of primer C and M primers D; Primer C includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer and a sequencing adapter that binds to the sequencing chip. Each primer D includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase, a secondary barcode for distinguishing different libraries, and another sequencing adapter that binds to the sequencing chip. The secondary barcodes in the M primers D are all different; 3) The primer pair shown consists of M primers E and M primers F; Each primer E includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the reverse transcription primer, a secondary barcode for distinguishing different libraries, and a sequencing adapter that binds to the sequencing chip. Each primer F includes a sequencing adapter that is the same as or complementary to the sequencing adapter in the Tn5 transposase, a secondary barcode for distinguishing different libraries, and another sequencing adapter that binds to the sequencing chip. The secondary barcodes in the M primers E are all different; The secondary barcodes in the M primers F are all different.

2. The library prepared by the method of claim 1.

3. A method for high-throughput transcriptome sequencing, comprising the following steps: constructing a sequencing library using the method of claim 1, and then sequencing the sequencing library.

4. The application of the transcriptome sequencing method according to claim 3 in biological gene expression analysis and genotype analysis.

Citation Information

Patent Citations

  • Systems and methods for pooling samples from multi-well devices

    CN108449933A

  • Sequencing library construction method and kit for pathogenic microorganism detection

    CN111188094A