Multiplex PCR amplification library construction method based on single-end sequencing and use thereof
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2026-08-13
Smart Images

Figure CN2025076576_13082026_PF_FP_ABST
Abstract
Description
A Multiplex PCR Amplification Library Preparation Method Based on Single-End Sequencing and Its Application Technical Field
[0001] This invention relates to the fields of biomedical technology and molecular diagnostic technology, and in particular to a method for library preparation based on single-end sequencing multiplex PCR amplification and its application. Background Technology
[0002] Cancer is one of the leading diseases threatening human health, and its complexity and diversity pose enormous challenges to modern medicine. With advancements in modern medicine and molecular biology, personalized cancer treatment has become a crucial direction in cancer therapy. Cancers are often accompanied by multiple gene mutations, which not only determine the occurrence and development of the tumor but also affect the patient's sensitivity to certain drugs. Therefore, accurately detecting and identifying these gene mutations is essential for developing appropriate treatment plans.
[0003] Currently, targeted cancer therapy and next-generation sequencing (NGS) technology have become important components of precision oncology. Targeted cancer therapy achieves its effects by specifically inhibiting key signaling pathways in tumor cells. Representative molecular markers include EGFR, HER2, KRAS, BRAF, ALK, ROS1, and MET. Mutations or amplifications of these genes have significant clinical implications in various cancer types. For example, EGFR mutations are an important basis for selecting EGFR inhibitors for non-small cell lung cancer (NSCLC) patients; HER2 gene amplification is an indication for trastuzumab treatment in breast cancer patients; NSCLC patients with ALK gene rearrangements (fusions) can benefit from crizotinib treatment; NSCLC patients with ROS1 rearrangements can benefit from crizotinib treatment; NSCLC patients with MET gene exon 14 skipping mutations can benefit from gumetinib and terpoxtinib treatment; and patients carrying BRAF gene V600E mutations can benefit from combination therapy with BRAF inhibitors and MEK inhibitors.
[0004] Next-generation sequencing (NGS) technology can analyze hundreds to thousands of gene loci simultaneously and is widely used in research and clinical practice to comprehensively understand the complexity of tumor genomes. With its high throughput, high sensitivity, and high specificity, this technology has become the mainstream method for gene mutation detection. However, in practical applications, existing NGS kits still face some problems in multiplex PCR amplification and library construction: (1) Poor amplification specificity: Multiplex PCR reactions easily produce non-specific amplification products, which increases the risk of false positives. (2) Strong amplification bias: There may be differences in amplification efficiency between different target fragments, resulting in some low-frequency mutations not being effectively detected. (3) Complex library construction: NGS library construction steps are cumbersome, involving multiple purification and enzyme digestion processes, resulting in long operation times and easy introduction of errors. (4) Long NGS sequencing time makes it difficult for doctors to obtain gene detection results in a timely manner, making it difficult to adjust individualized treatment plans based on the patient's gene characteristics, affecting clinical decision-making and treatment plan adjustments.
[0005] Therefore, developing a novel, efficient, and rapid NGS detection kit and a corresponding multiplex PCR amplification library construction method to improve amplification specificity and homogeneity and simplify the library construction process is particularly urgent and necessary. This will not only improve the accuracy and reliability of detection but also better serve the needs of personalized cancer treatment and promote the development of precision medicine. Summary of the Invention
[0006] The purpose of this invention is to provide a multiplex PCR amplification library preparation method based on single-end sequencing and its application. This amplification library preparation method has high sensitivity and specificity, can detect gene mutations related to tumor targeted therapy, and has a rapid detection process, which can complete nucleic acid extraction, library preparation and sequencing in as little as 16 hours. It can be used for cancer diagnosis, recurrence monitoring and medication guidance.
[0007] To achieve the above-mentioned objectives, the present invention provides the following technical solution:
[0008] This invention provides a method for library preparation using multiplex PCR amplification based on single-end sequencing, comprising the following steps:
[0009] (1) Extract DNA and / or RNA samples from the sample;
[0010] (2) Use primer pools to perform multiplex PCR amplification and magnetic bead purification on cDNA samples obtained by reverse transcription of DNA or RNA samples from step (1).
[0011] (3) Digest and purify the non-specific amplification products in the product of step (2);
[0012] (4) Use Index primers to perform secondary PCR labeling and purification on the product of step (3) to obtain an NGS library for tumor-targeted drugs.
[0013] (5) Perform single-end high-throughput sequencing on the NGS library constructed in step (4) and perform biological analysis on the sequencing data.
[0014] Preferably, in step (2), the primer pool includes one or more specific primer pairs for detecting SNV / Indel, CNV, Fusion, and MSI mutations of proto-oncogenes and tumor suppressor genes.
[0015] Preferably, in step (2), the primer pool includes primers for detecting AKT1, ALK, ARAF, BRAF, CD274, CTNNB1, DDR2, DPYD, EGFR, ERBB2, ESR1, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MET, NRAS, NTRK1, NTRK3, PDGFRA, and PIK3CA. A DNA primer pool consisting of primers for any one or more genes or MSI from POLD1, POLE, PTEN, RET, ROS1, SMAD4, TERT, TP53, UGT1A1, or an RNA primer pool consisting of primers for detecting any one or more genes from ALK, BRAF, CLDN18, EGFR, FGFR1, FGFR2, FGFR3, FGFR4, MET, NRG1, NTRK1, NTRK2, NTRK3, RET, ROS1;
[0016] The DNA primer pool is used to detect any one or more gene mutation types among SNV / Indel, CNV, and MSI mutations, and the RNA primer pool is used to detect Fusion gene mutation types.
[0017] The primer pool contains primers with structures including bridging sequences and specific amplification sequences for building NGS libraries.
[0018] Preferably, the specific amplification sequences in the DNA primer pool include one or more of the sequences shown in SEQ ID NO.1 to SEQ ID NO.354; the specific amplification sequences in the RNA primer pool include one or more of the sequences shown in SEQ ID NO.355 to SEQ ID NO.410; and the bridging sequences used to establish the NGS library are shown in SEQ ID NO.414 and SEQ ID NO.415.
[0019] Preferably, any one of the specific amplification sequences shown in SEQ ID NO.1 to SEQ ID NO.410 can be linked with the bridging sequences shown in SEQ ID NO.414 and SEQ ID NO.415, respectively, to detect any one or more of the SNV / Indel, CNV, Fusion, and MSI mutations of proto-oncogenes and tumor suppressor genes.
[0020] Preferably, in step (4), the Index primer sequence is as follows:
[0021] SEQ ID NO.411:
[0022] CAAGCAGAAGACGGCATACGAGANNNNNNNNCGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC, where NNNNNNNN represents 8 different base sequences used to identify tag sequences for different samples.
[0023] Preferably, in step (5), the biological analysis steps include: performing quality control on the sequencing data and aligning it to a reference genome, and then using variant detection software to identify gene mutations and their frequencies related to tumor-targeted drugs.
[0024] This invention also provides an application of a single-end sequencing-based multiplex PCR amplification library preparation method in screening tumor-targeted drugs, cancer diagnosis, or recurrence monitoring.
[0025] The present invention also provides an NGS detection device or kit that can be used to implement the aforementioned single-end sequencing-based multiplex PCR amplification library preparation method.
[0026] The present invention also provides an NGS detection device or kit that can be used to implement a single-end sequencing-based multiplex PCR amplification library preparation method for screening tumor-targeted drugs, cancer diagnosis or recurrence monitoring.
[0027] The beneficial effects of this invention compared to the prior art are as follows:
[0028] This invention provides a method for library construction based on single-end sequencing and multiplex PCR amplification, comprising: extraction and processing of DNA or RNA, first-round PCR reaction and purification, digestion reaction and purification, and second-round PCR reaction and purification to construct a library for detection. This method significantly improves the sensitivity and specificity of the detection process and simplifies the operation steps by optimizing primer design, improving the amplification system, and streamlining the library construction process. It can be used to detect gene mutations related to tumor targeted therapy, including single nucleotide variants (SNVs), insertions and deletions (Indels), gene fusions, copy number variants, and microsatellite instability (MSI). Based on the detection results, cancer diagnosis, recurrence monitoring, and medication guidance can be performed, providing powerful tool support for scientific research and clinical applications.
[0029] (2) The accuracy and specificity of the single-end sequencing-based multiplex PCR amplification library preparation method of this invention reach 100%, and the minimum detection limit of SNV / Indel is 1%, the minimum detection limit of CNV is 4 copies, the minimum detection limit of Fusion is 100 copies, and the minimum detection limit of MSI is 10%. It can accurately detect tumor-related mutations, is suitable for large-scale clinical testing and scientific research applications, and is expected to play an important role in personalized cancer treatment.
[0030] (3) The single-end sequencing-based multiplex PCR amplification library preparation method of the present invention can simultaneously sequence from upstream and downstream of the primer in single-end sequencing mode, achieving the effect of double-end sequencing, and the sequencing time is shortened to 1 / 2 of that of double-end sequencing. The nucleic acid extraction, library preparation and sequencing can be completed in as little as 16 hours, which can be used for cancer diagnosis, recurrence monitoring and medication guidance. Attached Figure Description
[0031] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0032] Figure 1 illustrates the multiplex PCR amplification library preparation method based on single-end sequencing of this invention. Detailed Implementation
[0033] The embodiments of the present invention are described in detail below. These embodiments are intended to explain the present invention and should not be construed as limiting the present invention. Where specific techniques or conditions are not specified in the embodiments, they are performed according to the techniques or conditions described in the literature in the art or according to the product instructions. Reagents or instruments used, unless otherwise specified, are all commercially available conventional products.
[0034] This invention provides two construction methods for different samples (see Figure 1):
[0035] 1. When the extracted sample is RNA, the library construction steps include:
[0036] (1) RNA samples were extracted from the samples, and the RNA samples were pretreated and then reverse transcribed and purified to obtain cDNA samples;
[0037] (2) The cDNA sample from step (1) was subjected to a first round of multiplex PCR amplification and magnetic bead purification using an RNA primer pool;
[0038] (3) Digest and purify the non-specific amplification products in the product of step (2);
[0039] (4) The product from step (3) was labeled and purified by secondary PCR using index primers to obtain a targeted library for tumor-targeting drugs; the targeted library was quality controlled and quantified using Qubit4.0 (Qubit dsDNA HS Assay Kit, Q32854). When the library concentration was ≥0.4ng / μL, it was considered to be qualified and an NGS library was obtained.
[0040] (5) Perform high-throughput sequencing on the qualified NGS library constructed in step (4) and perform bioinformatics analysis on the sequencing data.
[0041] 2. When the extracted sample is DNA, the library construction steps include:
[0042] (1) Extract DNA samples from the sample;
[0043] (2) The DNA sample from step (1) was subjected to a first round of multiplex PCR amplification and magnetic bead purification using a DNA primer pool;
[0044] (3) Digest and purify the non-specific amplification products in the product of step (2);
[0045] (4) The product from step (3) was labeled and purified by secondary PCR using index primers to obtain a targeted library for tumor-targeting drugs; the targeted library was quality controlled and quantified using Qubit4.0 (Qubit dsDNA HS Assay Kit, Q32854). When the library concentration was ≥0.4ng / μL, it was considered to be qualified and an NGS library was obtained.
[0046] (5) Perform high-throughput sequencing on the qualified NGS library constructed in step (4) and perform bioinformatics analysis on the sequencing data.
[0047] The target gene amplified in this invention is shown in Table 1.
[0048] Table 1 Target Gene
[0049] The DNA primer pool used in the following examples includes:
[0050] The amplification primer sets obtained by ligating the sequences shown in SEQ ID NO.1 to SEQ ID NO.354 with the bridging sequences shown in SEQ ID NO.414 and SEQ ID NO.415, respectively, wherein the concentration of each primer is 75 nM;
[0051] The RNA primer pool used in the following examples includes: amplification primer sets obtained by ligating the sequences shown in SEQ ID NO.355 to SEQ ID NO.410 with the bridging sequences shown in SEQ ID NO.414 and SEQ ID NO.415, respectively, wherein the concentration of each primer is 25 nM.
[0052] Example 1
[0053] Example 1 of this invention detected EGFR mutations in lung cancer samples. The specific steps are as follows:
[0054] (1) DNA extraction:
[0055] DNA was extracted from lung cancer tissue sections using a nucleic acid extraction kit (Zhenyue Biotechnology, YCP-1201) at a concentration of 10 ng / μL.
[0056] (2) Multiplex PCR amplification and purification:
[0057] Prepare a DNA primer pool containing EGFR detection: Use the DNA primer pool to perform multiplex PCR amplification on the DNA isolated in step (1). The PCR amplification system is as follows: use 2 μL of PCR amplification mixture (PARGON GENOMICS, 5X mPCR Mix), 2 μL of DNA primer pool, 2 μL of DNA, and nuclease-free water to make up to 10 μL for amplification.
[0058] The amplification conditions were as follows: 95℃ pre-denaturation for 10 minutes; 98℃ denaturation for 15 seconds; 60℃ annealing / extension for 5 minutes, for a total of 10–12 cycles; hold at 10℃. After amplification, 2 μL of stop buffer (PARGON GENOMICS) was added.
[0059] (3) Purify the amplification products using purification magnetic beads:
[0060] Mix the PCR product from step (2) with the purified magnetic beads at a volume ratio of 1.3:1 and let stand at room temperature for 10 minutes. Collect the magnetic beads using a magnetic rack and discard the supernatant. Wash the magnetic beads twice with 70% ethanol, using 180 μL of ethanol each time, let stand for 30 seconds, and remove the ethanol. Open the tube and dry the magnetic beads at room temperature for 5 minutes. Add 20 μL of elution buffer, mix thoroughly, and proceed directly to the next step of the reaction.
[0061] (4) Digestion and purification:
[0062] Add 2 μL of PARGON GENOMICS (CP Digestion Reagent), 2 μL of PARGON GENOMICS (CP Reagent Bufffer), and 6 μL of nuclease-free water to the elution buffer from step (3) and mix. Then incubate at 37°C for 10 minutes. After incubation, add 2 μL of PARGON GENOMICS (Stop Buffer).
[0063] The product was purified again using the magnetic bead purification step in step (3).
[0064] (5) Index labeling and purification:
[0065] The digestion product from step (4) was subjected to secondary PCR labeling using the primer mixtures shown in SEQ ID NO.412 to SEQ ID NO.413 (all at a concentration of 5 μM). The secondary PCR labeling system was as follows: using amplification mixture 2 (PARGON GENOMICS, 5X 2) nd Add 8 μL of PCR Mix, 4 μL of Index primer mixture, and 18 μL of nuclease-free water to the elution product from step (4).
[0066] The amplification conditions are as follows:
[0067] Pre-denaturation at 95℃ for 10 minutes; denaturation at 98℃ for 15 seconds; annealing / extending at 60℃ for 75 seconds, for a total of 12-15 cycles; hold at 10℃.
[0068] Mix the PCR product with purified magnetic beads at a volume ratio of 1:1 and let stand at room temperature for 10 minutes. Collect the magnetic beads using a magnetic rack and discard the supernatant. Wash the magnetic beads twice with 70% ethanol, 180 μL of ethanol each time, let stand for 30 seconds, and remove the ethanol. Open the tube and dry the magnetic beads at room temperature for 5 minutes. Add 20 μL of elution buffer, mix thoroughly, let stand on a magnetic rack, collect the eluent, and obtain the targeted library.
[0069] (6) DNA library quality control:
[0070] The library was quantified using Qubit 4.0 (Qubit dsDNA HS Assay Kit, Q32854). The library concentration was 1.6 ng / μL, yielding the NGS library.
[0071] (7) NGS sequencing: After constructing the NGS library, sequencing was performed on the Illumina MiniSeq platform.
[0072] Data Analysis: After sequencing, the sequencing data was automatically analyzed in the tumor gene detection data management software (Zhenyue Biotechnology, version 2.0 or higher) to obtain gene mutation results. The results showed a deletion mutation in exon 19 of EGFR in this sample, with a mutation frequency of 15%. The patient from whom this sample was obtained was 55 years old and diagnosed with stage IIIA (T3N1M0) lung adenocarcinoma. He had previously received EGFR inhibitor targeted therapy with icotinib and osimertinib, achieving a beneficial overall survival of 52 months.
[0073] Example 2
[0074] Example 2 of this invention performed ALK fusion mutation detection in lung cancer samples. The specific steps are as follows:
[0075] (1) RNA extraction:
[0076] RNA was extracted from lung cancer sections using a nucleic acid extraction kit (Zhenyue Biotechnology, YCP-1301) at a concentration of 50 ng / μL.
[0077] (2) Preparation of cDNA:
[0078] 4 μL of extracted RNA was mixed with 5 μL of reverse transcription buffer 1 (Abclonal, 2X Frag / Elute Buffer) and nuclease-free water to a final volume of 10 μL, and then pretreated at 65°C for 5 minutes. After pretreatment, 2 μL of reverse transcriptase (Abclonal, First strand Synthesis Enzyme Mix) and 8 μL of reverse transcription buffer 2 (Abclonal, RT Reagent) were added, and cDNA synthesis was performed under the following conditions:
[0079] Temperature and time settings: 25℃ for 10 minutes; 42℃ for 15 minutes; 70℃ for 15 minutes, then maintain at 4℃.
[0080] Purification is performed using purifying magnetic beads following these steps:
[0081] Mix the product with purified magnetic beads at a volume ratio of 2.2:1 and let stand at room temperature for 10 minutes. Collect the magnetic beads using a magnetic rack and discard the supernatant. Wash the magnetic beads twice with 70% ethanol, using 180 μL of ethanol each time, let stand for 30 seconds, and remove the ethanol. Open the tube and dry the magnetic beads at room temperature for 5 minutes. Add 10 μL of elution buffer, mix thoroughly, and proceed directly to the next step of the reaction.
[0082] (3) Multiplex PCR amplification and purification:
[0083] Prepare an RNA primer pool containing ALK fusion detection: Use the RNA primer pool to perform multiplex PCR amplification on the cDNA obtained by reverse transcription in step (2). The PCR amplification system is as follows: 4 μL of PCR amplification mixture (PARGON GENOMICS, 5X mPCR Mix), 2 μL of RNA primer pool, 1 μL of exogenous quality control (non-human plasmid DNA, 100 copies), and 3 μL of nuclease-free water are added to the reverse transcription product in step (2) and mixed before amplification.
[0084] The amplification conditions were as follows: 95℃ pre-denaturation for 10 minutes; 98℃ denaturation for 15 seconds; 60℃ annealing / extension for 5 minutes, for a total of 10–12 cycles; hold at 10℃. After amplification, 2 μL of stop buffer (PARGON GENOMICS, Stop Buffer) was added.
[0085] (4) Purify the amplification products using purification magnetic beads:
[0086] The amplification product was purified using purified magnetic beads, referring to the method in step (3) of Example 1.
[0087] (5) Digestion and purification:
[0088] The digestion and purification were carried out according to the method in step (4) of Example 1.
[0089] (6) Index labeling and purification:
[0090] The digestion product from step (5) was labeled with secondary PCR using the Index primer mixtures shown in SEQ ID NO.412 to SEQ ID NO.413 (all at a concentration of 5 μM). The secondary PCR labeling system was as follows: 8 μL of amplification mixture 2 (PARGON GENOMICS, 5X 2nd PCR Mix), 1.6 μL of Index primer mixture, and 20.4 μL of nuclease-free water were added to the elution product from step (5) and mixed.
[0091] The amplification conditions and purification steps are the same as those in step (5) of Example 1, to obtain the target library.
[0092] (7) RNA library quality control:
[0093] The library was quantified using Qubit4.0 (Qubit dsDNA HS Assay Kit, Q32854) to obtain the NGS library. The library concentration was 0.3 ng / μL.
[0094] (8) NGS sequencing:
[0095] After constructing the NGS library, sequencing was performed on the Illumina MiniSeq platform.
[0096] Data Analysis: After sequencing, the sequencing data was automatically analyzed in the tumor gene detection data management software (Zhenyue Biotechnology, version no lower than 2.0) to obtain gene mutation results. The patient from whom this sample originated was a 34-year-old diagnosed with stage IIIC left lung adenocarcinoma. He benefited from treatment with the ALK-TKI third-generation targeted drug lorlatinib, achieving partial response (PR) after 3 months of treatment, and subsequent efficacy assessments maintained PR. At the last follow-up in June 2024, the patient's progression-free survival (PFS) exceeded 70 months.
[0097] Example 3
[0098] Example 3 of this invention performed microsatellite instability (MSI) detection in colorectal cancer samples. The specific steps are as follows:
[0099] (1) DNA extraction: DNA was extracted from colorectal cancer sections using a nucleic acid extraction kit (Zhenyue Biotechnology, YCP-1201) at a concentration of 10 ng / μL.
[0100] (2) The target library was obtained by referring to steps (1) to (5) of Example (1). The difference from Example (1) is that the DNA used in the multiplex PCR amplification and purification in step (2) is the DNA of step (1) of Example 3.
[0101] (3) DNA library quality control:
[0102] The library was quantified using Qubit 4.0 (Qubit dsDNA HS Assay Kit, Q32854). The library concentration was 2.2 ng / μL, indicating that the library preparation was satisfactory.
[0103] (4) NGS sequencing:
[0104] After constructing the library, sequencing was performed on the Illumina MiniSeq platform.
[0105] Data Analysis: After sequencing, the sequencing data was automatically analyzed in the tumor gene detection data management software (Zhenyue Biotechnology, version no lower than 2.0) to obtain gene mutation results. The test results showed that the colorectal cancer patient had a high level of microsatellite instability (MSI-H). This patient was a 36-year-old metastatic colorectal cancer patient who had failed second-line treatment. Simultaneous testing confirmed that the patient was a dMMR (mismatch repair deficient) patient identified through Lynch syndrome screening and benefited from PD-1 antibody immunotherapy. After four cycles of treatment, the multiple enlarged retroperitoneal lymph nodes essentially disappeared, achieving complete clinical remission. Treatment was discontinued two years later, and the patient remained in disease-free survival (NED) at the most recent follow-up examination (June 2022).
[0106] Example 4
[0107] Example 4 of this invention tested the effectiveness of the multiplex PCR amplification library preparation method based on single-end sequencing of this application. The specific steps are as follows:
[0108] (1) Initial inventory assessment
[0109] To assess the impact of different initial nucleic acid amounts on test results.
[0110] Eighteen lung cancer slides were selected, and DNA was extracted using a nucleic acid extraction kit (Zhenyue Biotechnology, YCP-1201). Quantification was then performed using Qubit4.0 (ThermoFisher, Q32854), with initial DNA amounts of 1 ng, 2.5 ng, 5 ng, 10 ng, 20 ng, and 30 ng used for assessment. Four lung cancer slides were selected, and RNA was extracted using a nucleic acid extraction kit (Zhenyue Biotechnology, YCP-1301). Quantification was then performed using Qubit4.0 (ThermoFisher, Q32852), with initial RNA amounts of 10 ng, 25 ng, 50 ng, 100 ng, and 300 ng used for assessment. Each sample was tested in triplicate. The impact of different initial amounts on the detection results was compared to evaluate whether sample quality control was passed and whether the target variant was detected under different initial amounts.
[0111] Quality control standards:
[0112] DNA library: [1] The library pass rate reaches 100%, and the pass standard is that the library concentration is not less than 0.4 ng / uL; [2] The data quality control pass rate reaches 100%, and the pass standard is that the average sequencing depth is not less than 800x and the uniformity is not less than 70%; [3] All sample mutation results are detected 100%.
[0113] RNA library: [1] Data quality control pass rate reached 100%, and the data pass standard was that the total reads were not less than 20,000 and the number of reads of any housekeeping gene was not less than 20; [2] Mutation results of all samples were detected 100%.
[0114] Table 2. Mean values of quality control parameters under different DNA starting amounts
[0115] Table 3. Mutation detection under different DNA starting amounts
[0116] Table 4. Mean values of quality control parameters under different RNA starting amounts
[0117] Table 5. Mutation detection under different RNA initiation amounts
[0118] The results showed that the sample quality control was passed and all target variants were detected under the conditions of a minimum starting amount of 10ng for DNA and a minimum starting amount of 50ng for RNA. Therefore, the minimum input amount of DNA is 10ng and the minimum input amount of RNA is 50ng.
[0119] (2) Accuracy assessment
[0120] Sixteen cell line standards (Nanjing Kober) were tested, with variant types including SNV, Indel, Fusion, CNV, and MSI. Sample information is detailed in the table below. All mutations were validated using Sanger and ddPCR methods. Each sample was tested three times, and all corresponding mutations should be detected positively.
[0121] Four negative cell line standards were tested, including two wild-type cell line references and two positive mutations outside the detection range of this kit. Each sample was tested three times, and the results should all be negative.
[0122] Table 6. Mutation Information of Cell Line Standards
[0123] Table 7 shows that all target variants were detected, and no target variants were detected in any negative samples, indicating an accuracy of 100%.
[0124] Table 7 Test Results
[0125] (3) Evaluation of the lowest detection limit:
[0126] Fifteen nucleic acid samples derived from clinical lung cancer slides and three cell line standards (Nanjing Kobo) were selected. The mutation types included SNV, Indel, Fusion, CNV, and MSI. Sample preparation information is shown in the table below. Frequency gradients for SNV and Indel were set at 0.6%, 1%, and 2.5%; copy number gradients for CNV were set at 3, 4, and 6 copies; copy number gradients for Fusion were set at 50 copies / 50 ng, 100 copies / 50 ng, and 200 copies / 50 ng; and tumor cell line percentage gradients for MSI were set at 5%, 10%, and 15%. Each mutation frequency was replicated 20 times, and the lowest mutation frequency achieving a 95% positive site detection rate was defined as the optimal detection limit for this kit.
[0127] (1) Preparation of SNV / Indel mutation frequency gradient DNA samples: Eleven clinical lung cancer positive DNA samples containing the nine mutation sites to be tested were taken: DS0530, DS0187, DS0255, DS0244, DS0154, DS0071, DS0223, DS0382, DS0629, DS0574, and DS0551. Each sample contained a single gene mutation site. Based on the initial ddPCR frequency determination results, the mutation frequencies of the nine clinical positive samples were diluted to three gradients of 0.6%, 1%, and 2.5% using negative cell line GM12878D DNA as the sample diluent, and ddPCR detection was performed. The preparation results are shown in Table 10.
[0128] (2) Preparation of ERBB2 and MET gene copy number amplification gradient DNA samples: ERBB2 and MET copy number amplification cell line DNA standards purchased from Kebai were used. Based on the initial ddPCR copy number quantification results, negative cell line GM12878D DNA was used as the sample diluent to dilute the ERBB2 and MET gene copy numbers to three gradients: 6, 4, and 3. ddPCR was then performed for detection. The preparation results are shown in Table 10.
[0129] (3) Preparation of DNA samples with different tumor cell contents in MSI: The MSH cell line DNA standard purchased from Kebai was used with a tumor cell content of 100%. The negative cell line GM12878DDNA was used as the sample diluent to dilute the MSH cell line DNA standard to three gradients with tumor cell contents of 15%, 10%, and 5%. The background purity and mutation frequency unique to the MSH cell line were detected by ddPCR to label the content of incorporated tumor cells. The preparation results are shown in Table 10.
[0130] (4) Preparation of FUSION positive copy gradient RNA samples: Four clinical lung cancer fusion positive RNA samples RS0001, RS0076, RS0047 and RS0054 containing the four fusion sites to be tested were taken. According to the initial ddPCR copy number quantification results, the positive fusion copy number was diluted to three gradients of 50 copies / 50ng, 100 copies / 50ng and 200 copies / 50ng respectively using negative cell line GM12878R FFPE RNA as sample diluent. ddPCR detection was performed. The preparation results are shown in Table 11.
[0131] Note: The mutation frequency of SNV / Indel sites was detected by ddPCR using Bio-Rad QX200, and the verification criteria are as follows:
[0132] Table 8
[0133] ddPCR detection for CNV copy number quantification was performed using Bio-Rad QX200, with the following verification criteria:
[0134] Table 9
[0135] Table 10. ddPCR frequency determination results for DNA dilution samples
[0136] Table 11 Results of ddPCR quantification of RNA diluted samples
[0137] Table 12 shows the results:
[0138] 1) The mutation frequency of SNV and Indel mutation sites is 1%, and the detection rate is 95%, therefore the limit of detection for SNV / Indel is 1%; 2) The copy number of CNV mutation is 4 copies, and the detection rate is 95%, therefore the limit of detection for CNV is 4 copies; 3) The copy number of Fusion mutation is 100 copies, and the detection rate is 95%, therefore the limit of detection for Fusion is 100 copies; 4) The proportion of MSI tumor cells is 10%, and the detection rate is 95%, therefore the limit of detection for MSI is 10%.
[0139] Table 12 Test Results
[0140] (5) Repeatability assessment:
[0141] Ten cell line standards (Nanjing Kober) were selected, and ten intra-batch replicates and ten inter-batch replicates were performed. That is, the same sample was tested 10 times using the same batch of reagents, and all corresponding mutations should be detected; the same sample was tested 10 times (3 times) using three different batches, for a total of 30 tests, and all corresponding mutations should be detected. Reference information is as follows:
[0142] Table 13 Information on Cell Line Standards
[0143] Table 14 shows that the repeatability test results within and between batches of the same sample are consistent.
[0144] Table 14 Test Results
[0145] (6) Interference analysis:
[0146] Two cell line standards (Nanjing Kober) were selected, and sample information is shown in Table 1. Five potential interfering substances from clinical samples were selected for study: ethanol, formalin, melanin, hemoglobin, and proteinase K. Different types and concentrations of interfering substances were added to nucleic acid samples before detection, with blank controls included. Each type was tested three times. The interfering substances and their corresponding concentrations are shown in Table 16.
[0147] Table 15 Information on Cell Line Standards
[0148] Table 16 Gradient of Interference
[0149] Table 17 shows that when the interfering substances do not exceed 8% v / v ethanol, 0.02% v / v formalin, 0.5 mg / L melanin, 0.5 g / L hemoglobin and 20 mg / L proteinase K, they have no effect on the detection results.
[0150] Table 17 Test Results
[0151] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for library construction using multiplex PCR amplification based on single-end sequencing, characterized in that, Includes the following steps: (1) Extract DNA and / or RNA samples from the sample; (2) Use primer pools to perform multiplex PCR amplification and magnetic bead purification on cDNA samples obtained by reverse transcription of DNA or RNA samples from step (1). (3) Digest and purify the non-specific amplification products in the product of step (2); (4) Use Index primers to perform secondary PCR labeling and purification on the product of step (3) to obtain an NGS library for tumor-targeted drugs. (5) Perform single-end high-throughput sequencing on the NGS library constructed in step (4) and perform biological analysis on the sequencing data.
2. The method for library preparation based on single-end sequencing using multiplex PCR amplification according to claim 1, characterized in that, In step (2), the primer pool includes one or more specific primer pairs for detecting any one or more of the following mutations: SNV / Indel, CNV, Fusion, and MSI, which are used to detect proto-oncogenes and tumor suppressor genes.
3. The method for library preparation based on single-end sequencing using multiplex PCR amplification according to claim 2, characterized in that, In step (2), the primer pool includes primers for detecting AKT1, ALK, ARAF, BRAF, CD274, CTNNB1, DDR2, DPYD, EGFR, ERBB2, ESR1, FBXW7, FGFR1, FGFR2, FGFR3, FGFR4, GNAS, HRAS, IDH1, IDH2, KIT, KRAS, MAP2K1, MET, NRAS, NTRK1, NTRK3, PDGFRA, PIK3CA, P DNA primer pools consisting of primers for any one or more genes or MSI from OLD1, POLE, PTEN, RET, ROS1, SMAD4, TERT, TP53, UGT1A1, or RNA primer pools consisting of primers for detecting any one or more genes from ALK, BRAF, CLDN18, EGFR, FGFR1, FGFR2, FGFR3, FGFR4, MET, NRG1, NTRK1, NTRK2, NTRK3, RET, ROS1; The DNA primer pool is used to detect any one or more gene mutation types among SNV / Indel, CNV, and MSI mutations, and the RNA primer pool is used to detect Fusion gene mutation types. The primer pool contains primers with structures including bridging sequences and specific amplification sequences for building NGS libraries.
4. The method for multiplex PCR amplification and library preparation based on single-end sequencing according to claim 3, characterized in that, The specific amplification sequences in the DNA primer pool include one or more of the following sequences: The sequences shown in SEQ ID NO.1 to SEQ ID NO.2 are used to amplify AKT1; The sequences shown in SEQ ID NO.3 to SEQ ID NO.10 are used to amplify ALK; The sequences shown in SEQ ID NO.11 to SEQ ID NO.12 are used to amplify ARAF; The sequences shown in SEQ ID NO.13 to SEQ ID NO.16 are used to amplify BRAF; The sequences shown in SEQ ID NO.17 to SEQ ID NO.36 are used to amplify CD274; The sequences shown in SEQ ID NO.37 to SEQ ID NO.43 are used to amplify CTNNB1; The sequences shown in SEQ ID NO.44 to SEQ ID NO.45 are used to amplify DDR2; The sequences shown in SEQ ID NO.46 to SEQ ID NO.55 are used to amplify DPYD; The sequences shown in SEQ ID NO.56 to SEQ ID NO.67 are used to amplify EGFR; The sequences shown in SEQ ID NO. 68 to SEQ ID NO. 95 are used to amplify ERBB2; The sequences shown in SEQ ID NO.96 to SEQ ID NO.103 are used to amplify ESR1; The sequences shown in SEQ ID NO.104 to SEQ ID NO.105 are used to amplify FBXW7; The sequences shown in SEQ ID NO.106 to SEQ ID NO.113 are used to amplify FGFR1; The sequences shown in SEQ ID NO.114 to SEQ ID NO.121 are used to amplify FGFR2; The sequences shown in SEQ ID NO.122 to SEQ ID NO.125 are used to amplify FGFR3; The sequences shown in SEQ ID NO.126 to SEQ ID NO.132 are used to amplify FGFR4; The sequences shown in SEQ ID NO.133 to SEQ ID NO.136 are used to amplify GNAS; The sequences shown in SEQ ID NO.137 to SEQ ID NO.144 are used to amplify HRAS; The sequences shown in SEQ ID NO.145 to SEQ ID NO.148 are used to amplify IDH1; The sequences shown in SEQ ID NO.149 to SEQ ID NO.152 are used to amplify IDH2; The sequences shown in SEQ ID NO.153 to SEQ ID NO.160 are used to amplify KIT; The sequences shown in SEQ ID NO.161 to SEQ ID NO.166 are used to amplify KRAS; The sequences shown in SEQ ID NO.167 to SEQ ID NO.178 are used to amplify MAP2K1; The sequences shown in SEQ ID NO.179 to SEQ ID NO.190 are used to amplify MET; The sequences shown in SEQ ID NO.191 to SEQ ID NO.224 are used to amplify MSI; The sequences shown in SEQ ID NO.225 to SEQ ID NO.230 are used to amplify NRAS; The sequences shown in SEQ ID NO.231 to SEQ ID NO.234 are used to amplify NTRK1; The sequences shown in SEQ ID NO.235 to SEQ ID NO.238 are used to amplify NTRK3; The sequences shown in SEQ ID NO.239 to SEQ ID NO.244 are used to amplify PDGFRA; The sequences shown in SEQ ID NO.245 to SEQ ID NO.254 are used to amplify PIK3CA; The sequences shown in SEQ ID NO.255 to SEQ ID NO.270 are used to amplify POLD1; The sequences shown in SEQ ID NO.271 to SEQ ID NO.280 are used to amplify POLE; The sequences shown in SEQ ID NO.281 to SEQ ID NO.302 are used to amplify PTEN; The sequences shown in SEQ ID NO.303 to SEQ ID NO.318 are used to amplify RET; The sequences shown in SEQ ID NO.319 to SEQ ID NO.322 are used to amplify ROS1; The sequences shown in SEQ ID NO.323 to SEQ ID NO.332 are used to amplify SMAD4; The sequences shown in SEQ ID NO.333 to SEQ ID NO.334 are used to amplify TERT; The sequences shown in SEQ ID NO.335 to SEQ ID NO.350 are used to amplify TP53; The sequences shown in SEQ ID NO.351 to SEQ ID NO.354 are used to amplify UGT1A1; The specific amplification sequences in the RNA primer pool include one or more of the following sequences: The sequences shown in SEQ ID NO.355 to SEQ ID NO.360 are used to amplify ALK; The sequences shown in SEQ ID NO.361 to SEQ ID NO.364 are used to amplify BRAF; The sequences shown in SEQ ID NO.365 to SEQ ID NO.366 are used to amplify CLDN18; The sequences shown in SEQ ID NO.367 to SEQ ID NO.368 are used to amplify EGFR; The sequences shown in SEQ ID NO.369 to SEQ ID NO.370 are used to amplify FGFR1; The sequences shown in SEQ ID NO.371 to SEQ ID NO.373 are used to amplify FGFR2; The sequences shown in SEQ ID NO.374 to SEQ ID NO.375 are used to amplify FGFR3; The sequences shown in SEQ ID NO.376 to SEQ ID NO.377 are used to amplify FGFR4; The sequences shown in SEQ ID NO.378 to SEQ ID NO.381 are used to amplify MET; The sequences shown in SEQ ID NO.382 to SEQ ID NO.385 are used to amplify NRG1; The sequences shown in SEQ ID NO.386 to SEQ ID NO.388 are used to amplify NTRK1; The sequences shown in SEQ ID NO.389 to SEQ ID NO.390 are used to amplify NTRK2; The sequences shown in SEQ ID NO.391 to SEQ ID NO.394 are used to amplify NTRK3; The sequences shown in SEQ ID NO.395 to SEQ ID NO.399 are used to amplify RET; The sequences shown in SEQ ID NO.400 to SEQ ID NO.410 are used to amplify ROS1; The bridging sequences used to build the NGS library are shown in SEQ ID NO.414 and SEQ ID NO.
415.
5. The method for multiplex PCR amplification and library preparation based on single-end sequencing according to claim 1, characterized in that, Any of the specific amplification sequences shown in SEQ ID NO.1 to SEQ ID NO.410 can be linked with the bridging sequences shown in SEQ ID NO.414 and SEQ ID NO.415, respectively, to detect any one or more of the SNV / Indel, CNV, Fusion, and MSI mutations in proto-oncogenes and tumor suppressor genes.
6. The method for multiplex PCR amplification and library preparation based on single-end sequencing according to claim 1, characterized in that, In step (4), the Index primer sequence is as follows: SEQ ID NO.411: CAAGCAGAAGACGGCATACGAGANNNNNNNNCGTGACTGGAGTTCAGACGTGTGCTCTTCCGATC, where NNNNNNNN represents 8 different base sequences used to identify tag sequences for different samples. SEQ ID NO.412: AATGATACGGCGACCACCGAGATCTACACTCTTTCCCTACACGACGCTCTTCCGATC.
7. The method for library preparation based on single-end sequencing using multiplex PCR amplification according to claim 1, characterized in that, In step (5), the biological analysis includes the following steps: The sequencing data were quality controlled and aligned to a reference genome. Then, variant detection software was used to identify gene mutations and their frequencies associated with targeted cancer drugs.
8. The application of the single-end sequencing-based multiplex PCR amplification library preparation method of claim 1 in screening tumor-targeted drugs, cancer diagnosis, or recurrence monitoring.
9. An NGS detection device or kit for implementing the single-end sequencing-based multiplex PCR amplification library preparation method of claim 1.
10. The application of the NGS detection device or kit of claim 9, which can be used to implement the multiplex PCR amplification library preparation method based on single-end sequencing, in screening tumor-targeted drugs, cancer diagnosis, or recurrence monitoring.