A multiplex method for preparing a sequencing library
By using multiple amplification methods of target capture primer pairs and tagged primer pairs in a single closed tube, combined with the autonormalization step, the problems of nonspecific binding and cross-contamination in targeted genotyping are solved, and efficient and simple uniform amplification of target sequences and sample pooling are achieved, suitable for high-throughput sequencing.
Patent Information
- Application Number
- CN202080064097.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-19
- Filing Date
- 2020-09-11
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-09-11
AI Technical Summary
Existing targeted genotyping methods have problems such as nonspecific binding, cross-contamination and sample processing during multiple amplification, especially in high-throughput sequencing, which is difficult to achieve efficient and simple uniform amplification of target sequences and sample pooling.
Using a single blocking tube method, multiple amplification was performed in the same reaction vessel using a target capture primer pair and a tagged primer pair, combined with a self-normalization step, and a complete construct containing the purified tag was generated by limited concentrations of tagged primers and purification tags, avoiding cross contamination between samples and additional purification steps.
Efficient and simplified targeted high-throughput sequencing in a single reaction vessel is achieved, reducing the risk of cross-contamination, improving the uniformity and sequencing efficiency of the target sequence, and is suitable for multiple amplification of thousands of samples.
Smart Images

Figure CN114364812B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multiplex amplification method for preparing a library for targeted next generations sequencing. Methods of targeted next generations sequencing comprising multiplex amplification are also contemplated. Background Art
[0002] Targeted genotyping (GBS) by sequencing involves merging the multiple amplification products of index and adapter connection into a single library sample for the operation of high-throughput next-generation sequencing. From this sequencing data, individual genotype profiles are extracted and compiled by bioinformatics tools involving demultiplexing and allele counting. There are several modifications of this main method, open source and commercial sources (Eureka and Agriseq Thermofisher, Nugen Allegro, KeygeneSNPselect, LGC SeqSNP, AgriPlex Genomics PlexSeq). Open source methods usually rely on enzyme amplification in enrichment and actual library preparation.
[0003] The development of an in-house workflow for a targeted GBS approach can be divided into the following steps:
[0004] 1. Target-specific multiplex amplification
[0005] 2. Sample indexing, sequencing primers and adapter addition
[0006] 3. Normalization and size selection
[0007] 4. QC and final library merging
[0008] 5. Next-Generation Sequencing
[0009] 6. Data Analysis
[0010] In all of the above methods, target amplification and library preparation are performed in multiple steps (steps 1 to 4). The first reaction is a multiplex amplification of all targets in the panel. Having a high number of targets in the same amplification reaction can cause problems through non-specific binding and the formation of primer dimers. Therefore, in order to achieve a panel that works efficiently, the multiplex nature of the initial reaction needs to be considered in the primer design step to avoid problematic sequences. The final panel will be fine-tuned through a small-scale test iteration process. With amplification-based methods (GT-Seq, MonsterPlex, Hi-plex, PlexSeq), multiplex levels of up to about 500 are typically reported. When run for the first time, MonsterPlex (referenced in the literature as similar to Hi-Plex) expects approximately 85-90% of the markers to work. These libraries are sequenced at a high depth, with approximately 500K reads per sample, to obtain good coverage. The Hi-Plex method, also based on amplification, claims overlap levels of up to 1000 (Nguyen-Dumont et al. 2013, Nguyen-Dumont et al. 2015, Pope et al. 2018, Hammet et al. 2019). Hi-Plex is the only method designed to combine two amplification steps; after the initial amplification, primers that introduce indexes / adapters are incorporated into the multiplex reaction. Therefore, this method still requires opening and resealing the sample plate. The key points listed by the authors in the principle process are high-fidelity Taq enzyme, permissive initial amplification conditions that allow a series of targets to be amplified simultaneously, followed by large-scale amplification of the tails added for the target-specific primers, and relatively long cycles that allow complete priming and extension. The method originally published for the 60-plex set was further developed to achieve 1003-plex by adding another set of so-called short bridge primers to the multiplex. The target-specific amplification mixture here has two shorter primers that are partially complementary to the tails added by the target-specific primers. Large-scale amplification is performed in the first part using target-specific and bridge primers. The tagging primers are not added to the mix until the last four cycles. This differs from the original 60-plex method, in which short bridge primers are not used, in which the majority of the amplification occurs after the addition of the index primers. A recent publication by Hammet et al. (2019) describes further updates to the Hi-Plex method to improve on-target amplification. In the updated method, the initial two-step amplification using target-specific and tagging primers has fewer cycles, followed by size selection before another amplification to amplify the constructed library, but the sample plate still needs to be opened and resealed. Universal short bridge primers are no longer used. Hammet's comparison of the previous and current methods showed similar performance at lower overlap levels below 300, but the updated method had improved capacity at higher target numbers.
[0011] Consistent annealing between target primers is crucial for more uniform amplification of multiple targets simultaneously. Rather than cycling for longer at a single temperature common to all Hi-Plex versions, the GT-Seq method achieves this by including a slow temperature ramp from 95°C to 57°C, resulting in more accurate annealing over the 10 cycles of amplification. Chen et al. (2016) described another approach to balancing amplification by using limited primer availability; multiplexed amplification products are used as targets in a second amplification step without adding more primers. This reaction is allowed to run to exhaust all available primers before indexing and produce a more normalized output (the authors refer to this approach as three-round amplification). Introducing a third amplification reaction is not ideal, but it can be helpful to verify correct primer availability in multiplexed reactions. Like Hi-Plex, GT-Seq and MonsterPlex methods use high-quality Taq enzymes.
[0012] In order to pool samples into one sequencing run, each sample requires a unique tag. This can be accomplished by a unique combination of short sequences (e.g., 6- or 8-mers) at both ends or only one end of the DNA sample. In order to enable thousands of individuals to be sequenced in the same index run, a systematic indexing method is needed; such as one index representing the well position and another representing the plate. This high-density sample pool can be achieved by reusing 96-well or 384-well specific indexes in combination with different plate indexes. In the amplification-based targeted GBS method, indexing is performed by target-specific multiple primers carrying universal tail sequences; the forward primers share the same tail sequence, which is different from the tail shared by the reverse primers. These tails are targeted in a second amplification step performed by adding primers with unique indexes and sequencing adapters. In addition, the universal tails introduced by the target-specific primers are used as sequencing primer targets in the sequencing reaction. After the second amplification step, a sample-specific library containing multiple targets is obtained.
[0013] The process of linking these two reactions can vary. As already mentioned by Chen et al. (2016), an amplification step can be included between the two to deplete all target-specific primers. In GT-Seq, the amplification products from the multiplex reaction are diluted and then used as targets in the amplification with the second marker / adapter on a new plate. The Hi-Plex method uses a different approach. The index primers are added directly to the initial amplification plate without an intermediate dilution step. This reaction uses the enzyme mix added in the first amplification setup, which saves both reagents and plates. However, the plate still needs to be pierced or peeled and further resealed. Due to the risk of sample contamination, especially when using water bath thermal cycling, this last step is more difficult to implement in high-throughput applications.
[0014] GB2536446 describes a single closed tube PCR method for producing an amplicon construct of a target sequence by reverse complementary PCR (RC-PCR). The method uses (a) an oligonucleotide probe comprising a universal sequence and a target-specific sequence at or near its 5' end, wherein the target-specific sequence is capable of hybridizing with the reverse complementary sequence of one of the 3' ends of the target sequence or the sequence flanking it, and (b) a universal primer comprising a sequence capable of hybridizing with the universal sequence of the oligonucleotide probe at its 3' end. A key element of the process is that it does not include a normalization step, but rather uses a blocking group in the RC-PCR method in the initial step when producing full-length oligonucleotides for target amplification+adapter addition, preventing the target-specific oligonucleotide from combining until full-length oligonucleotides are produced. Although this may contribute to target uniformity, the lack of normalization in the method can cause some samples to drown other samples during sequencing and exhaust a large amount of data. This severely limits the number of targets that the method can accurately process. Referring to the Nimagen website (https: / / www.nimagen.com / applications / greenotype), this method (GREENOTYPE RC-PCR) is used for genotyping of no more than 100 targets, while another method (GREENOTYPE MIP CAPTURE) is commercially available for higher throughput (up to 5000 targets). Therefore, the RC-PCR method disclosed in GB2536446 is not suitable for processing a large number of targets.
[0015] Meldgaard et al. describe target amplification and library preparation methods in Clinica Chimica Acta, Vol. 413; No. 19 (2012) and WO 2017 / 044100 (Insilixa Inc.), which use asymmetric PCR to produce single-stranded products captured on printed arrays or beads and subsequently analyzed on a flow cytometer. When one of the two primers in each group is in excess and the limiting primer is used up, single-stranded products are produced. Single-stranded products are amplified in a linear rather than exponential form. In both examples, the ultimate goal is to produce these single-stranded products, which can then be captured by hybridization. Both methods are not suitable for targeted high-throughput (NGS) sequencing or sample pooling.
[0016] Bybee et al. (Genome Biology and Evolution, Vol. 3, January 1, 2011) described a method for generating sequence libraries by a multiplex PCR approach that uses single-column barcodes with R℃he 454 sequencing, which limits the number of samples that can be run at one time. This method requires performing a first PCR reaction in a first container before aliquoting the first reaction product into a second container for the second PCR step. The two-step PCR approach is laborious and time-consuming, and transferring products between containers introduces the risk of cross-contamination between samples. The method includes a bead immobilization step as part of the emulsion PCR step required for R℃he 454 sequencing. In this method, each individual sample is quantified using pico-green and then normalized using a liquid handling robot, which is both time-consuming and expensive.
[0017] It is an object of the present invention to overcome at least one of the above problems. Summary of the Invention
[0018] The present invention provides a single closed tube method for preparing library constructs for targeted high-throughput sequencing by multiplex amplification, without the need to reopen the tube to add additional components. The present invention also has built-in normalization, eliminating the need for a separate normalization step. Figure 1 , the method uses a target capture primer pair specific for each target in the sample, and a tagging primer pair specific for the sample. Each target capture primer pair comprises a forward and a reverse primer, the forward and reverse primers comprising a target-specific sequence and a read sequence, wherein the target-specific sequence is configured to bind to the template nucleic acid at a position flanking the target (step 1). The tagging primer pair comprises a forward and a reverse tagging primer, each primer having an adapter sequence (P5, P7), an index sequence (xxxxx) and a sequence that binds to the read sequence (read sequence primer site) (step 2). The initial amplification step ( Figure 1 (1)) produces an intermediate construct containing the target sequence flanked by read sequences (R1 and R2). In the subsequent amplification step ( Figure 1 (2)), the tagging primer binds to the intermediate construct via the read sequence primer sequence that anneals to the read sequence. In the method of the present invention, one of the tagging primers is provided at a limited concentration and a purification tag (B) is also included, which produces a reaction product containing both partial and complete constructs ( Figure 1(3)). The partial construct contains the target sequence, but only one adapter sequence (P5) and one index sequence (xxxxx). The complete construct, on the other hand, contains the target sequence, the first and second adapter sequences (P5, P7) and the first and second index sequences (xxxxx) as well as a purification tag. Because a limited amount of one tagging primer (in the case shown, the reverse tagging primer) is used, the reaction products contain an excess of the partial construct relative to the complete construct, with the result that each individual reaction produces a complete sequencing construct that has approximately the same amount of each targeted gene locus in relatively equal abundance across a wide range of template DNA (i.e., self-normalized). It will be appreciated that in Figure 1 In FIG, read sequences R1, R2 and adapter sequences P5, P7 are used for exemplary purposes only, and any read sequence, index sequence, and adapter sequence known to those skilled in the art may be used.
[0019] By using the methods of the present invention, target capture, indexing, and adapter addition can be performed in the same reaction vessel without opening the reaction vessel, followed by library capture using a purification tag. After amplification, samples from thousands of individual reactions can be pooled and purified using a purification tag ( Figure 1 The completed library construct is captured and purified using a single primer (biotin in the example). The unique combination of the first and second index sequences (e.g., i5 and i7 barcode sequences) in each reaction allows all samples to be pooled after multiplex amplification without the risk of cross-contamination. When the number of target-specific primers can vary depending on the desired assay set, only two unique tagging primers are required for each reaction. Initial amplification is performed using target capture primers that amplify the desired genomic DNA site. After the first two rounds of amplification, partial binding sites for the tagging primers are generated and they can begin to anneal with their complements. In the next few cycles, the full-length complementary sequences of the tagging primers are generated, and once this occurs, the remainder of the reaction is performed at a higher annealing temperature, reflecting the complete priming site of the tagging primers. The method of the present invention is suitable for multiplex amplification of multiple samples (e.g., up to 200,000), each of which contains thousands of target sequences (e.g., SNPs), all of which can be performed in a single closed tube. For a sample with, for example, 50 target sequences, the method uses 50 target-specific primer pairs and one tagging primer pair (specific for the sample). When the method is applied to multiple samples, for example 500 samples, each with 50 target sequences, the method uses 50 x 500 target-specific primer pairs and 50 tagging primer pairs (one for each sample). Amplification can be performed by polymerase chain reaction (PCR) or by other enzymatic nucleic acid amplification techniques.
[0020] Compared to the method of Bybee et al., the method of the present invention allows the target capture primer pair and the tagging primer pair to be included in the same reaction vessel and incubated together simultaneously from the beginning and during the amplification step. This speeds up and simplifies the process and reduces the risk of cross-contamination, which is an inherent problem of two-step PCR methods such as Bybee, in which the reaction products of the first PCR step must be washed before being aliquoted into the second reaction vessel of the second PCR step. In addition, the incorporation of a self-normalization step in the method of the present invention eliminates the time-consuming and expensive requirement to quantify each sample individually and then normalize it using liquid handling robots.
[0021] Compared with the method of GB 2536446, the introduction of a self-normalization step in the method of the present invention overcomes the problem of low throughput and allows multiplex amplification of hundreds or thousands of targets from multiple samples in a single container in a single run without affecting target uniformity, regardless of different samples having different concentrations or masses of target DNA.
[0022] In a first aspect, the present invention provides a method for preparing a library construct for targeted next generation sequencing by multiplex amplification, comprising the following steps:
[0023] (a) providing a first reaction vessel comprising:
[0024] (i) at least one sample, wherein the sample comprises a plurality of different target sequences;
[0025] (ii) a plurality of target capture primer pairs for the at least one sample, wherein each target capture primer pair comprises:
[0026] a forward primer comprising the first read sequence and the target-specific sequence in the 5' to 3' direction; and
[0027] a reverse primer comprising the second read sequence and the target-specific sequence in the 5' to 3' direction;
[0028] (iii) a tagging primer pair for the at least one sample, comprising:
[0029] a forward tagging primer comprising, in the 5' to 3' direction, a first adapter sequence, a first index sequence, and a first read sequence primer site; and
[0030] a reverse tagging primer comprising, in the 5' to 3' direction, a second adapter sequence, a second index sequence, and a second read sequence primer site,
[0031] (b) performing sequential rounds of amplification at sequential annealing temperatures configured to amplify the target sequence, generate the target sequence comprising the first and second read sequences, and provide a reaction product comprising a library of adaptor-ligated constructs in a sequential manner; and
[0032] (c) capturing the adaptor-ligated library construct from the reaction products,
[0033] Characterized in that one of the forward and reverse tagging primers comprises a purification tag at the 5' end and is provided in a limited concentration, whereby the library of adaptor-ligated constructs comprises:
[0034] a partial construct containing only one of the first and second index sequences and one of the first and second adapter sequences; and
[0035] a complete construct comprising the first and second index sequences, the first and second adaptor sequences, and a purification tag,
[0036] wherein the reaction product comprises an excess of the partial construct relative to the complete construct, and wherein step (c) comprises capturing only the complete construct comprising the purification tag.
[0037] Typically, separation (capture) of the intact library construct from the reaction products involves contacting the reaction products with a support comprising a ligand for a purification tag.
[0038] The sample comprises or consists of a nucleic acid containing a target sequence. Typically, the sample is DNA. The target sequence is typically a sequence variation, such as a single nucleotide polymorphism (SNP) or a short insertion / deletion (indel).
[0039] The capture step typically involves reacting the reaction product with a support comprising a ligand for a purification tag. The purification tag can be biotin, and the support can comprise streptavidin (e.g., streptavidin beads). Other purification tags and capture ligands can be used, the details of which are known to those skilled in the art.
[0040] The complete library construct can be released from the support for subsequent high-throughput sequencing, or it can be amplified in a multiplex amplification step while still attached to the support to provide amplification products that can then be sequenced using high-throughput sequencing. In one embodiment, multiple amplification cycles are performed and the amplification products are typically separated from the separation support.
[0041] Typically, steps (a) and (b) are performed in a closed container (eg, a tube), which is typically not opened until these steps are completed.
[0042] The thermal cycling of the amplification step is configured to amplify the target sequence in a sequential manner, generate a target sequence comprising a first and a second read sequence, and provide a reaction product comprising a library construct ligated to the adapter in a sequential manner. In the described embodiment, Illumina R1 and R2 read sequences, and i5 and i7 index sequences are used, and the thermal cycling employs initial, intermediate, and final rounds of amplification with increasing annealing temperatures. It should be understood that different thermal cycling methods can be used when using read sequences and index sequences from different sources. The use of Illumina reads and index sequences is exemplary and is not intended to limit the scope of this application.
[0043] Thus, for example, sequential rounds of amplification at sequential annealing temperatures may include:
[0044] (i) performing one or more initial rounds of amplification in the reaction vessel at a first annealing temperature;
[0045] (ii) performing one or more intermediate rounds of amplification in the reaction vessel at a second annealing temperature, the second annealing temperature being configured to generate a target sequence comprising the first and second read sequences; and
[0046] (iii) performing one or more final rounds of amplification in the reaction vessel at a third annealing temperature, wherein the third annealing temperature is configured to provide a reaction product comprising a library of library constructs, including partial constructs and complete constructs.
[0047] Typically, the second annealing temperature is higher than the first annealing temperature, and the third annealing temperature is higher than the second annealing temperature. However, depending on the context, a different order of annealing temperatures may be employed.
[0048] In one embodiment, the first annealing temperature is 62°C ± 5°C, the second annealing temperature is 67°C ± 5°C, and the third annealing temperature is 72°C ± 5°C.
[0049] In one embodiment, the sequential rounds of amplification include 1-5 initial rounds of amplification (preferably 1-3 or 2 rounds), 1-5 intermediate rounds of amplification (preferably 2-4 or 3 rounds), and 10-20 final rounds of amplification (e.g., 10-50).
[0050] In one embodiment, the purification tag is biotin and the isolating step comprises reacting the reaction product with streptavidin beads.
[0051] In one embodiment, steps (a) and (b) are carried out on multiple samples, wherein the reaction products of different samples are merged, and capture / separation step (c) is carried out on the merged reaction products. In one embodiment, 2 to 200,000 samples, for example, at least 100, 1000, 10,000, 20,000, 50,000, 100,000, 150,000 or 200,000 samples are carried out steps (a) and (b), and then merged. In one embodiment, steps (a) and (b) are carried out in the wells of a microtiter plate. In one embodiment, each sample comprises 1-10,000 target sequences, for example, at least 10, 50, 100, 500, 1000, 2000, 3000, 4000, 5000, 6000, 7000, 8000, 9000 or 10,000 target sequences.
[0052] Amplification is usually performed by polymerase chain reaction (PCR). It should be understood that in order to perform PCR on a target sequence, additional reagents are required and these reagents will depend on the type of PCR employed. For example, in order to perform PCR, the method of the present invention uses a thermostable DNA polymerase (e.g., Taq polymerase) and all four deoxyribonucleotides (dATP, dTTP, cCTP, cGTP).
[0053] In another aspect, the present invention provides a method for targeted high-throughput sequencing, comprising the following steps:
[0054] The method according to the invention provides a library of complete constructs; and
[0055] A library of complete constructs was subjected to high-throughput sequencing.
[0056] In a preferred embodiment, the high-throughput sequencing is a next generation sequencing technology. One example is luminescent dye sequencing, in which the adapter sequences (and typically the target primer pair and the tagging primer pair) are configured for luminescent dye sequencing.
[0057] In one embodiment, the high throughput sequencing has an on-target rate of at least 30%, 40%, 50% or 60%.
[0058] The method of the present invention can be used to genotype a sample for medical, diagnostic or commercial purposes. In one embodiment, the method of the present invention can be used to genetically confirm the origin or identity of a food, such as meat, fish, game, vegetable, legume, grain or fruit product.
[0059] In another aspect, the present invention provides a kit for preparing a library construct by multiplex amplification, which is suitable for targeted high-throughput sequencing of one or more samples, the kit comprising a tagging primer pair for each sample, each tagging primer pair comprising:
[0060] a forward tagging primer comprising, in the 5' to 3' direction, a first adapter sequence, a first index sequence, and a first read sequence primer site; and
[0061] A reverse tagging primer comprising a second adapter sequence, a second index sequence, and a second read sequence primer site in the 5' to 3' direction
[0062] Only one of the forward or reverse tagging primers contains a purification tag at the 5' end.
[0063] In one embodiment, the kit comprises:
[0064] Multiple target-specific primer pairs (one per target sequence per sample), wherein each primer pair comprises:
[0065] a forward primer comprising the first read sequence and the target-specific sequence in the 5' to 3' direction; and
[0066] a reverse primer comprising the second read sequence and the target-specific sequence in the 5' to 3' direction;
[0067] In one embodiment, the purification tag is biotin.
[0068] In one embodiment, the kit comprises a support comprising streptavidin, such as beads or magnetic beads coated with streptavidin.
[0069] In one embodiment, the target primer pair and the tagging primer pair are configured for luminescent dye sequencing.
[0070] Further aspects and preferred embodiments of the present invention are defined and described in the detailed description set out below. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] Figure 1 . Closed tube multiplex library construction.
[0072] Phase 1: Set up an initial amplification sequence containing template DNA and the target SNP site marked with an asterisk. The forward and reverse target capture primers are marked to anneal to the template DNA at positions flanking the target sequence via their respective Illumina target-specific sequences. The tagging primers are also marked:
[0073] the forward tagging primer comprises, in the 5' to 3' direction, a first adapter sequence, a first index sequence, and a read sequence primer site; and
[0074] The reverse tagging primer contains a biotin purification tag (B), an adapter sequence, an index sequence, and a read sequence primer site in the 5' to 3' direction.
[0075] Phase 2: After an initial low annealing temperature amplification cycle, the partially formed DNA construct is combined with the merged read sequences R1 and R2.
[0076] Stage 3: In the final stage, due to the limited amount of biotin-labeled reverse tagging primers, amplification is allowed to proceed for many cycles, generating a large number of partial constructs (containing the target sequence, R1 and R2 read sequences, one adapter sequence (P5) and one index sequence), while also generating a certain amount of complete library constructs (containing the target sequence, R1 and R2 read sequences, two adapter sequences (P5 and P7) and two index sequences).
[0077] Figure 2 .Distribution of reads between loci and the percentage of forward primers containing target sequence.
[0078] Figure 3 .Read distribution between samples and summary data.
[0079] Figure 4 Example of a single-point scatter plot. Each data point is a sample. DETAILED DESCRIPTION
[0080] All publications, patents, patent applications, and other references mentioned herein are incorporated by reference in their entirety for all purposes to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference and was set forth in its entirety.
[0081] Definition and general preferences
[0082] As used herein, unless expressly stated otherwise, the following terms are intended to have the following meanings, as well as any broader (or narrower) meanings that these terms may have in the art:
[0083] Unless the context requires otherwise, as used herein, the singular should be understood to include the plural, and vice versa. The term "a" or "an" in relation to an entity should be understood to refer to one or more of that entity. Therefore, the terms "a," "one," "one or more," and "at least one" are used interchangeably herein.
[0084] As used herein, the term "comprise" or variations such as "include" will be understood to indicate the inclusion of any recited integer (e.g., features, elements, characteristics, properties, method / process steps, or limitations) or group of integers (e.g., features, elements, characteristics, properties, method / process steps, or limitations), but not the exclusion of any other integer or group of integers. Thus, as used herein, the term "comprise" is inclusive or open-ended and does not exclude additional, unrecited integers or method / process steps.
[0085] "Target sequence" refers to a sequence to be amplified, such as a variation of a nucleic acid sequence, such as a polymorphism (e.g., a single nucleotide polymorphism (SNP) or a short indel). The nucleic acid sequence can be any part of a genomic nucleic acid, such as the coding or non-coding part of a genome.
[0086] "Sample" refers to a composition comprising a nucleic acid containing a target sequence. A sample is typically a nucleic acid (e.g., DNA) purified from a sample of biological material (e.g., tissue). Other amplifiable nucleic acids include cDNA. The DNA to be amplified can be obtained from any part of the genome of any organism. For example, DNA can be obtained from a human or other animal genome, a plant genome, a fungal genome, a bacterial genome, a viral genome, or any other DNA molecule. Methods for isolating nucleic acids from biological materials and tissue samples are well known to those skilled in the art. The method of the present invention can be performed on a single sample comprising multiple targets, multiple samples comprising a single target, or multiple samples each comprising multiple targets. The method of the present invention is particularly suitable for use with multiple targets, for example, greater than 100, 200, 300, 400, or 500 targets from one or more samples.
[0087] "High-throughput sequencing" refers to the sequencing of multiple different target sequences in parallel in the same reaction. The term includes "next-generation sequencing" or "NGS" (also known as "deep sequencing" or "massively parallel sequencing"), which is a DNA sequencing technology that can sequence millions of small DNA fragments in parallel. There are various NGS technologies, including synthesis sequencing (Illumina), pyrosequencing (454), ion semiconductor (ion torrent sequencing), and combined probe anchor synthesis (cPAS-BGI-MGI). Targeted NGS is a next-generation sequencing technology that focuses on amplicons and specific genes. The desired gene or amplicon is amplified by enzymatic amplification and then sequenced on the NGS platform. It was reported in Bybee et al. ("Targeted Amplicon Sequencing (TAS): A Scalable Next-Gen Approach to Multiplication, Multitaxa Phylogenetics". Genome Biology and Evolution. 3: 1312-1323. doi: 10.1093 / gbe / evr106. PMC 3236605) and Masser et al. ("Targeted DNA Methylation Analysis by Next-generation Sequencing".Journal of Visualized Experiments.96:52488.doi:10.3791 / 52488.PMC 4354667).
[0088] "Library construct" refers to the product of the multiplex amplification reaction of the present invention, which comprises a target sequence and a first and a second adapter sequence to allow the construct to be sequenced by next generation or high throughput sequencing (NGS) technology. The construct typically also includes a first and a second read sequence (e.g., Illumina R1 and R2 sequences) flanking the target sequence, and two index sequences (e.g., Illumina i5 and i7 sequences) located between the read sequence and the adapter sequence. The library typically comprises a partial construct (the partial construct comprises only one adapter sequence without a purification tag) and a complete (intact) construct (the complete construct comprises the adapter sequence, index sequence, and purification tag required for NGS). Capture technology using purification tags allows the complete construct to be captured and used in NGS. Complete (intact) and partial adapter library constructs such as Figure 1 (Phase 3) shown.
[0089] An "adapter sequence" is a DNA sequence configured to allow the adapter-connected library construct to be sequenced by a next-generation sequencing technology, such as any of the NGS technologies described herein, including Illumina dye sequencing in a flow cell. The adapter sequence can be configured to allow the adapter-connected construct to be sequenced by any NGS technology. Illumina's adapters are described in https: / / support.illumina.com / downloads / illumina-adapter-sequences-document-1000000002694.html. The adapter sequences for all major NGS platforms are provided by IDT (Integrated DNA Technologies) in Coralville, Ohio, USA (xGenTM UDI-UMI Adapters, xGenTM Stubby Adapters). IDT also provides a custom adapter configurator tool that provides guidance for the design of custom NGS adapters.
[0090] A "read sequence" is a portion of a target-specific primer. Both primers in a target-specific primer pair will contain a read sequence. When the target DNA is amplified and forms a sequence to which the tagging primer can anneal through the tagging primer read sequence primer site, the read sequence is incorporated into the amplification product. Examples of read sequences can include Illumina R1 and R2 read sequences. Read sequences can also be obtained from Thermo Fisher as part of its Ion AmpliSeq TMPart of the DNA panel and library kit (https: / / www.thermofisher.com / ie / en / home / life-science / sequencing / next-generation-sequencing / ion-torrent-next-generation-sequencing-products-services / sequencing-reagents.html). The "read sequence primer site" is the part of the tagging primer that is complementary to one of the read sequences and allows the tagging primer to anneal to the target DNA containing one of the read sequences.
[0091] The "index sequence" is part of the tagging primer that is introduced into the final library construct and acts as a barcode to identify the construct when analyzing sequencing data. Typically, the index sequence is 6-8 nucleotides in length and is located between the read sequence primer site and the adapter sequence in the tagging primer. The tagging primer pair will include two different index sequences, both of which will be incorporated into the complete (intact) library construct. An example of an index sequence is the Illumina i5 and i7 index sequences and the index oligonucleotides described in US2018334712. Suitable index sequences can be created using commercial products, such as:
[0092] https: / / rdrr.io / bioc / DNABarcodes /
[0093] https: / / omictools.com / dnabarcodes-tool
[0094] In the methods of the present invention, tagging primers containing a purification tag are provided at a "limited concentration". This means that the concentration provided results in insufficient tagging primers to bind to partially formed DNA constructs, resulting in an excess of partial constructs (lacking the purification tag) in the final product, while the complete (intact) construct containing the purification tag is limited. This, together with uniform initial target amplification, results in equal yields across samples, thereby avoiding the need for sample-specific downstream normalization. Figure 1 In the examples provided, reverse tagging primers containing the P7 adaptor sequence are provided in limited concentrations.
[0095] A "purification tag" refers to a tag that can be incorporated into the end of a primer sequence and can be used to purify the products of multiplex amplification. Typically, the tag is an affinity tag configured for purification of the amplified products using a support containing an affinity tag ligand (e.g., streptavidin beads). In one embodiment, the tag is biotin, the details of which are known to those skilled in the art. Other purification tags that can be used in the methods and products of the present invention include oligonucleotide tags (Acridite).
[0096] "On-target rate" refers to the percentage of sequence data returned that is specific for the target region. The higher the on-target rate, the more efficiently the sequencing run is utilized.
[0097] Example
[0098] The present invention will now be described with reference to specific embodiments. These are exemplary only and are for illustrative purposes only: they are not intended to limit the scope of the claimed patent or the invention described in any way. These embodiments constitute the best mode currently contemplated for carrying out the invention.
[0099] The present invention provides a single-tube method for preparing targeted libraries for next-generation sequencing, including a normalization step using a purification tag, such as biotin. Typically, multiplex library preparation involves two steps: target capture, followed by tagging and adapter addition. In many prior art methods, these two steps are typically achieved through amplification. Typically, the two amplification steps are performed separately from the purification step. Hi-plexing is a method designed to combine these two amplification steps, but in Hi-plexing, index / adapter introduction primers are incorporated into the multiplex reaction after the initial amplification. This means that the method still requires opening and resealing the sample container, which is not ideal for high-throughput processes (Nguyen-Dumont et al. 2013, Nguyen-Dumont et al. 2015, Pope et al. 2018, Hammet et al. 2019). The present invention relates to a true single-tube method, in which target capture, indexing, and adapter addition are performed in the same reaction vessel, followed by library capture using a purification tag, eliminating the need for opening and resealing the sample container and requiring no purification steps before tagging and adapter addition.
[0100] The idea of the method of the present invention is that each individual reaction will produce a complete sequencing construct with approximately the same number of each targeted locus and equal relative abundance over a wide range of template DNA inputs. The unique combination of index barcode sequences in each reaction allows all samples to be pooled after multiplex amplification without the risk of cross-contamination. Tagging primers labeled with a low concentration of purification tag allow multiplex amplification to produce an excess of constructs containing only one adapter and index sequence while limiting the amount of completed library constructs. This, together with uniform initial target amplification, achieves identical output across samples, avoids the need for sample-specific downstream normalization, and allows multiplex amplification of a large number of targets, ranging from hundreds to thousands. After the amplification reaction, samples from tens of thousands of individual reactions can be pooled together to facilitate capture and purification of the complete library construct using a purification tag, for example by using biotin as a purification tag and streptavidin beads to capture the complete construct with a biotin tag. Figure 1 An overview of the single-tube approach is shown. Unlike other approaches, both the target capture primer and the tagging primer are contained in the same reaction vessel, and the order of their activity is guided by thermal cycling. As the number of target-specific primers can vary depending on the desired set, only two tagging primers are required for each reaction. An initial amplification is performed using the target capture primer that amplifies the desired genomic DNA site. After the first two rounds of amplification, partial binding sites for the tagging primers from the read sequence are generated, and they can begin to anneal to their complementary sequences. Over the next few cycles, the full-length complementary sequences of the tagging primers are generated, and once this occurs, the remainder of the reaction is performed at a higher annealing temperature that reflects the complete binding site (priming site) of the tagging primer.
[0101] Current indicative cycling conditions are shown in Table 1.
[0102]
[0103] Table 1. Cycling conditions
[0104] Primer design
[0105] The target capture primers are designed so that the annealing temperature (Tm) of the target-specific 3' end is between 59°C and 61°C. Primers are designed using pipelines generated for multiplex amplification to avoid most sources of downstream primer heterodimer artifacts. The target capture primers carry a full-length read sequence, such as an Illumina R1 or R2 read sequence, that matches the 3' end of the tagging primer. Each tagging primer contains a 6-base index barcode sequence sandwiched between an adapter sequence (such as an Illumina p5 or p7 capture sequence) and a read sequence primer site. This allows the use of dual-index sequencing, such as sequencing from Illumina, and allows demultiplexing of sequences from genetic data (such as fastq data) for downstream genotyping of a single sample. For example, at the p5 end of the construct, the tagging primer can consist of the p5 sequence, a unique 6-base barcode sequence, and an Illumina read 1 sequencing primer site. For the p7 end of the construct, the tagging primers consist of a biotinylated 5' end, the p7 sequence, a unique 6-base barcode sequence, and an Illumina Read 2 sequencing primer site. The overlapping portions of the primers typically have an annealing temperature of approximately 60°C. Once the full complementary sequence is generated, the Tm for the complete primer sites at each end of the construct is approximately 74°C.
[0106] Normalization
[0107] This method does not require a separate normalization step after library construction. Instead, identical presentation of different samples is achieved by cycling conditions and concentration control of the tagging primers that limit the purification tag. The tagging primers are depleted in the later stages of thermal cycling, leaving equal amounts of the complete construct of the library bound to the support via the added purification tag; the total amount of construction may vary between samples, but the fraction containing the complete tag should be similar. Therefore, all samples can be pooled together after library construction and all downstream steps are performed on a single sample; library capture using a purification support, release of the library from the support (e.g. with short thermal cycles), final library cleanup (e.g. using magnetic beads) and library quantification. Suitable tagging primers are the i7 tagging primers from Illumina. A suitable purification tag is a biotin tag that can be captured using streptavidin beads.
[0108] Material
[0109] DNA used in the experiments was extracted from pig ear tag tissue samples using magnetic beads, with an average yield of 24 ng / μl (SD 11 ng / μl). All target capture primers were ordered from IDT, synthesized at a scale of 25 nmoles at a concentration of 200 μM, and purified using standard desalting in Tris-EDTA pH 8.0 buffer. Biotinylated i7 tagging primers were freeze-dried in tubes and resuspended in nuclease-free water to a stock concentration of 10 μM.
[0110] Multiplex amplification
[0111] The target capture primers were pooled and diluted to a concentration of 0.5 uM per primer as a working stock solution. For the test library, a primer pool was generated using 51 target capture primers. A single reaction was set up in a total volume of 7 uL (3.5 uL Qiagen Plus multiplex premix, 2 uL DNA extract, 0.5 uL target capture primer mixture, 1 uL 10 uM i5 tagging primer, ~0.01 uL biotinylated i7 tagging primer). The number of i5 tagging primers needed to vary depending on the number of target capture primers present in the library, so the molar amount of available primer sites was roughly equal to the molar amount of i5 primers. In this case, under the final reaction conditions, the final concentration of all i5 tagging primer sites was ~1.8 uM, and the concentration of the added i5 tagging primer used was ~1.4 uM. Amplification was set up in a 96-well plate by first preparing a master mix by mixing 52uL of pooled target primer mixture, 1uL of 10uM biotinylated i7 tagging primer, and 371uL of Qiagen plus multiplex premix (Qiagen, UK). 4uL of the master mix was dispensed into each well, followed by 1uL of 10uM i5 tagging primer and 2uL of template DNA using a multichannel pipette. The plate was heat-sealed, vortexed gently, and briefly centrifuged. Amplification was performed under the following cycling conditions (95°C-15m (1x hot start); 95°C-30s, 61°C-30s (slow ramp 0.2°C / s), 72°C-30s (5x); 95°C-30s, 67°C-30s (x3); 95°C-30s, 72°C-30s (x22); 4°C-hold).
[0112] Library capture and normalization
[0113] After amplification, 4uL of each reaction was combined into one pool using a multichannel pipette. 500uL aliquots were combined with 500uL 2x binding buffer (10mM Tris-HCl (pH 7.5), 1mM EDTA, 200mM NaCl, and 0.02% Tween20 buffer). 1 microliter of streptavidin beads was washed in 1mL 1x binding buffer in a 1.5mL test tube and captured using a magnetic stand. The supernatant was discarded, the test tube was removed from the magnetic stand, and the magnetic beads were resuspended using 1mL of the combined amplification mixture. The beads were incubated at room temperature for 15 minutes and then placed back on the magnetic stand for approximately 3 minutes until the supernatant was clear. The supernatant was removed from the magnetic stand and the beads were washed once with 1mL 1x binding buffer. The supernatant was discarded again and the beads were resuspended in 20uL elution buffer (10mM Tris-HCl pH 8.0).
[0114] In order to release the captured library construct from the beads, another amplification is performed using resuspended streptavidin beads. Single tube reaction is set up by mixing 20uL Qiagen plus master mix (any PCR master mix can be used), 4uL of 10uM Illumina P5 primers, 4uL of 10uM Illumina P7 primers and 12uL of resuspended streptavidin beads. Amplification is performed using the following thermal cycling conditions (94°C-15m (hot start); 94°C-30s, 60°C-30s, 72°C-30s (5x); 4°C-holding). After 5 amplification cycles, the mixture is transferred to a new 1.5mL tube and placed on a magnetic rack. 25uL of clarified supernatant is transferred to a new 1.5mL test tube and mixed with 15uL CleanNGS (CleanNA, the Netherlands) beads. The mixture is incubated at room temperature for 5 minutes, and the test tube is then placed back on the magnetic rack for 3 minutes. The clear supernatant was then transferred to a new 1.5 mL tube and mixed with another 13 uL of CleanNGS beads. The mixture was incubated again at room temperature for 5 minutes and then placed on a magnetic rack for 3 minutes. The supernatant was removed and, while still on the magnetic rack, the beads were washed twice with 200 uL of fresh 70% ethanol. The ethanol was removed and the tube was air-dried for 10 minutes to evaporate any residual ethanol. The beads were then resuspended in 15 uL of elution buffer and the beads were captured on the magnetic rack. The clear supernatant was collected and transferred to a new 1.5 mL tube and 1.5 uL of elution buffer containing 1% Tween 20 was added. This was the finalized library.
[0115] Library quantification
[0116] The final library concentration was determined using the Illumina Library Quantification Kit from Kapa Biosystems (Roche Sequencing, USA) according to the manufacturer's instructions on an Applied Biosystems StepOnePlus instrument (thermofisher scientific, UK).
[0117] Sequencing and data analysis
[0118] The completed libraries were sequenced using an Illumina MiSeq instrument using a paired-end 75-cycle kit. After sequencing, the data were converted to fastq format using an Illumina conversion script, which allows the addition of i7 and i5 index sequences to the header line of each sequence. Sequencing reads were split into separate fastq files for each individual sample using the expected barcode combinations provided as an input file to the python script GTseq_BarcodeSplit_MP.py. Each individual fastq file was then used for genotyping using the perl script GTseq_Genotyper_v3.pl. Summary files containing genotypes, allele ratios, and other sequence read data were then analyzed to assess the efficiency of the multiplex amplification library preparation method.
[0119] result
[0120] Optimal performing conditions were tested on a set of 24 samples. Four samples were removed from analysis because these samples contained no template, the reagents evaporated during PCR, or performed poorly under all conditions tested. For the subset of samples representing the proposed conditions, 835,017 raw reads containing the expected 6-base barcode combinations were returned after sequencing. Evidence of evaporation holes due to poor heat sealing was noted prior to sequencing, and as expected, these holes returned poor read counts. Because these samples did not represent typical reaction conditions, they were removed from further analysis. The raw read count from each of the remaining individual samples (n=20) was relatively uniform, averaging 41,751 reads with a standard deviation of 18,068. The number of raw reads returned ranged from 26,890 to 93,038. Across the 20 analyzed samples, the percentage of targeted sequences (as opposed to non-target or artifactual sequences—also known as the "on-target fraction") averaged 31%. Even with the low number of on-target reads, all 20 analyzed samples had sufficient reads to perform genotyping well, with an average call rate of 98.9%, and all call rates above 98%. The uniformity of reads across loci means that a small number of reads are required to achieve high call rates, and only 2,400 on-target reads are needed to meet the 90% threshold ( Figure 3 The selected target loci behaved as expected and produced clean allele ratios that were easily scored using the genotyping pipeline ( Figure 4 ).
[0121] Equivalent plan
[0122] The foregoing description describes in detail the presently preferred embodiments of the present invention. In view of this description, it is expected that those skilled in the art will make numerous modifications and variations in their practice. Such modifications and variations are intended to be encompassed by the appended claims. For the sake of brevity, the embodiments are described as a single embodiment, however, it should be understood that various combinations of these embodiments are within the scope of the present invention.
[0123] References
[0124] Hammet, F., Mahmood, K., Green, TR, Nguyen-Dumont, T., Southey, MC, Buchanan, DD, Lonie, A., Nathanson, K., L., Couch, FJ, Pope, BJ and Park, DJ2019: Hi-Plex2: a simple and robust approach to targeted sequencing-based geneticscreening.BioTechniques.67(3):00-00(September 2019).10.2144 / btn-2019-0026.
[0125] Nguyen-Dumont, Tu, Pope, B., Hammet, F., Southey, M. and Park, D. 2013: A high-plex PCR approach for massively parallel sequencing. Biotechniques. 55: 69-74.
[0126] Nguyen-Dumont, T., Hammet, F., Mahmoodi, M., Pope, B., Giles, G., Hopper, J., Southey, M., and Park, D. 2015: Abridged adapterprimers increase the target scope of Hi-Plex. BioTechniques. 58:33-36.
[0127] Pope,B.,Hammet,F.,Nguyen-Dumont,T.and Park,D.2018:Hi-Plex for simple,accurate,and cost-effective amplicon-based targetedDNA sequencing.Chapter 5in Steven R.Head et al.(eds),NextGeneration Sequencing:Methods andProtocols.vol.1712。
Claims
1. A method for preparing a library construct for targeted next generation sequencing by multiplex amplification, comprising the following steps: (a) providing a first reaction vessel comprising: (i) at least one sample, wherein the sample comprises a plurality of different target sequences; (ii) a plurality of target-specific primer pairs for the at least one sample, wherein each primer pair comprises: a forward primer comprising the first read sequence and the target-specific sequence in the 5' to 3' direction; and a reverse primer comprising the second read sequence and the target-specific sequence in the 5' to 3' direction; (iii) a tagging primer pair for the at least one sample, comprising: a forward tagging primer comprising, in the 5' to 3' direction, a first adapter sequence, a first index sequence, and a first read sequence primer site; and a reverse tagging primer comprising, in the 5' to 3' direction, a second adapter sequence, a second index sequence, and a second read sequence primer site, (b) performing sequential rounds of amplification in a first reaction vessel at a sequential annealing temperature, the temperature being configured to amplify the target sequence, generate the target sequence comprising the first and second read sequences, and provide a reaction product comprising a library of adaptor-ligated constructs in a sequential manner; and (c) capturing a library of adaptor-ligated constructs from said reaction products, Characterized in that one of the forward and reverse tagging primers comprises a purification tag at the 5' end and is provided in a limited concentration, whereby the library of adaptor-ligated constructs comprises: a partial construct comprising a first index sequence and a first adapter sequence or a partial construct comprising a second index sequence and a second adapter sequence; and a complete construct comprising the first and second index sequences, the first and second adaptor sequences, and a purification tag, wherein the reaction product comprises an excess of the partial construct relative to the complete construct, and wherein step (c) comprises capturing only the complete construct comprising the purification tag.
2. The process of claim 1, wherein the first reaction vessel is closed during step (b).
3. The method of claim 1 or 2, wherein the sequential rounds of amplification at sequential annealing temperatures comprise: (i) performing one or more initial rounds of amplification in the reaction vessel at a first annealing temperature; (ii) performing one or more intermediate rounds of amplification in the reaction vessel at a second annealing temperature, wherein the second annealing temperature is configured to generate a target sequence comprising the first or second read sequence; as well as (iii) performing one or more final rounds of amplification in the reaction vessel at a third annealing temperature, wherein the third annealing temperature is configured to provide a reaction product comprising a uniform amount of the complete library construct. The method of claim 3 , wherein the second annealing temperature is higher than the first annealing temperature, and the third annealing temperature is higher than the second annealing temperature. 5 . The method of claim 4 , wherein the first annealing temperature is 61° C.±5° C., the second annealing temperature is 67° C.±5° C., and the third annealing temperature is 72° C.±5° C. The method according to claim 3 , comprising 1-5 initial rounds of amplification, 1-5 intermediate rounds of amplification and 10-20 final rounds of amplification.
7. The method of claim 1 or 2, wherein the purification tag is biotin, and the capturing step comprises reacting the reaction product with streptavidin beads.
8. A method according to claim 1 or 2, wherein steps (a) and (b) are performed on a first sample in the first reaction vessel to produce a first reaction product comprising a first library of adaptor-ligated constructs, and steps (a) and (b) are performed on a second sample in a second reaction vessel to produce a second reaction product comprising a second library of adaptor-ligated constructs, wherein the first and second reaction products are combined and the capture step (c) is performed on the combined reaction products.
9. The method of claim 8, wherein steps (a) and (b) are performed on each of more than 100 samples.
10. The method of claim 8, wherein steps (a) and (b) are performed on each of more than 1000 samples.
11. The method of claim 1 or 2, wherein each sample comprises at least 10 target sequences.
12. The method of claim 1 or 2, wherein each sample comprises at least 50 target sequences.
13. The method of claim 1 or 2, wherein the capturing step comprises capturing the complete construct on a support and subsequently amplifying the complete construct attached to the support.
14. A method for targeted next generation sequencing, comprising the following steps: Providing a library of complete constructs according to the method of any preceding claim; as well as A library of complete constructs was subjected to high-throughput sequencing.
15. The method of claim 14, wherein the next generation sequencing is luminescent dye sequencing.
16. A library of adaptor-ligated constructs obtainable by the method according to any one of claims 1 to 15.
Citation Information
Patent Citations
Pcr method
GB2536446A
Universal short adapters for indexing of polynucleotide samples
US20180334712A1
Methods and systems for multiplex quantitative nucleic acid amplification
WO2017044100A1
Methods and compositions for dna profiling
CN106164298A