Evaluation and optimization method for ultra-high sensitivity multiplex sampling based on primer depolymerization algorithm

By evaluating and optimizing the complementary regions within the primer set through a primer depolymerization algorithm, the problem of primer dimer formation in multiplex PCR was solved, the detection sensitivity and specificity were improved, and ultra-high sensitivity multiple sampling was achieved.

CN119479773BActive Publication Date: 2025-09-19SUZHOU PALM SPRING BIOTECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411583271.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-09-19
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

The formation of primer dimers in existing multiplex PCR technology leads to reduced detection sensitivity and specificity, and conventional optimization algorithms are difficult to obtain ideal results within limited time and computing resources.

Method used

A primer depolymerization algorithm is used to evaluate and optimize the complementary regions within the primer set through mathematical models to reduce the possibility of dimer formation. The primer dimer problem is given priority when designing primers, and a hash function is used to quickly evaluate and optimize primer combinations.

Benefits of technology

It significantly reduces the formation of primer dimers, improves the detection sensitivity and specificity of multiplex PCR, and achieves ultra-high sensitivity multiple sampling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119479773B_ABST
    Figure CN119479773B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for evaluating the possibility of dimer generation caused by complementary regions between primers in a candidate primer set, wherein the complementary subregion in the primer sequence is the HBJ region, and the possibility of dimer generation caused by the HBJ region is evaluated by the function LHS of the mathematical model (1). m Indicates that LHS m The higher the value, the higher the possibility of inducing dimers, and the more the corresponding primers need to be optimized: the evaluation and optimization method of the present invention can minimize the formation of primer dimers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of biomolecule detection, and in particular relates to an evaluation and optimization method for ultra-high sensitivity multiple sampling based on a primer depolymerization algorithm. Background Art

[0002] After decades of development and clinical application, the detection of trace target substances in biomolecular detection technology has become increasingly important. These target substances include characteristic DNA sequences of pathogenic microorganisms, harmful mutations on DNA sequences, and drug-resistant mutations in microbial genomes. There is a mismatch between the people's growing health needs and the lack of significant progress in the sensitivity of various molecular detection methods.

[0003] Second-generation high-throughput sequencing (NGS) is one of the most sensitive technologies. It can detect critical medical information from massive amounts of DNA sequences, such as cancer-driving mutations, residual disease signatures, drug resistance information, fusion site information, and target pathogen signature sequences. Its detection sensitivity can reach 1 in 100,000 or even 1 in 1,000,000. Today, despite significant reductions in the cost of high-throughput sequencing, the time, management, and financial costs associated with the required high-depth data sets remain unacceptable for clinical applications.

[0004] The fluorescent PCR platform is currently the most widely used molecular detection platform, with the characteristics of simple workflow, short operation time, low reagent cost, and easy-to-understand test results. However, the characteristic of the fluorescent PCR platform with few fluorescent channels means that only a few sites or information can be detected in one test. Even with the help of melting curves, only dozens of sites or information can be detected. In addition, the detection principle based on the Ct value of the fluorescence curve or the melting curve has not made significant progress in detection sensitivity compared with more than ten years ago. A survey study in 2022 found that the analytical performance of 10 multiple nucleic acid detection kits for pathogens with different detection technology principles was evaluated, including 4 respiratory tract infection kits, 3 central nervous system infection kits and 3 bloodstream infection kits. The products came from 7 companies, including Guangzhou Wondfo, Suzhou Chuanglan, BGI, BioMérieux Shanghai, Hunan Shengxiang, Tianjin Baikangxin, and Shanghai Geneo. The median detection limit was 10 3 ~10 6 CFU / mL, the detection limit can reach 1 CFU / mL.

[0005] Based on the characteristics of the second-generation high-throughput sequencing platform and the fluorescent PCR platform, scientists have used the target region enrichment method to overcome the shortcomings of these two technical platforms. The target region enrichment method mainly includes hybridization capture probe technology and multiplex PCR technology. Among them, the multiplex PCR technology process has more time advantages than the hybridization capture probe technology process and requires less template DNA. However, the presence of primer dimers in the multiplex PCR technology and its unpredictable nonlinear increase reduce the positioning speed of multiple primers in the genome, accelerate the consumption of reaction system resources, and finally drag down the detection sensitivity, specificity and repeatability. Indicators, such as, it is difficult for multiplex PCR to cover hundreds or even thousands of genes as easily as hybridization capture probe technology.

[0006] The current methods for removing primer dimers in multiplex PCR technology mainly include using digestive enzymes to remove primer dimers or removing dimers by purification, which are both remedial measures after the generation of dimers.

[0007] There are two main challenges in designing highly multiplexed PCR primer sets. The first challenge is that single-plex PCR contains two primers, while N-plex PCR contains 2N primers, which means that the possibility of each primer forming a dimer with another primer has increased by 2N-1 times, that is, each primer has the potential to cause dimer formation with 2N-1 primers. The second challenge is that considering the complexity of the genome, under the conditions of limiting the target gene, target sequence and amplicon length, each primer usually has more than P>5 candidates. Therefore, for N-plex PCR, the possibility of primer combinations is P 2N When N=100, P=6, 6 100 In our opinion, this is close to infinity. Within limited time and computing resources, it is not easy to obtain ideal results for so many primer combinations using conventional optimization algorithms such as the steepest descent method, Newton's method, conjugate gradient method, penalty function method, feasible direction method, standard convex algorithm, etc. Summary of the Invention

[0008] In view of the problem of easy induction of dimers in the current design of multiplex PCR primer sets, the present invention takes a different approach and puts the problem of primer dimers first in the primer design stage, thereby developing a multiplex primer design algorithm to minimize the formation of primer dimers.

[0009] To achieve the above object, the present invention adopts the following technical solutions:

[0010] The first aspect of the present invention provides an ultra-high sensitivity multiple sampling evaluation and optimization method based on a primer depolymerization algorithm, which is used to evaluate and optimize the possibility of the complementary region between each primer in a candidate primer set to induce dimer production, wherein the complementary subregion in the primer sequence is the HBJ region, and the possibility of the HBJ region inducing dimer production is calculated using the function LHS of the mathematical model (1). m Indicates that LHS m The higher the value, the higher the possibility of inducing dimers, and the more the corresponding primers need to be optimized:

[0011]

[0012]

[0013] in:

[0014] m is the number of cycles to be evaluated, m = m + 1, m ≥ 0;

[0015] PDST represents the probability of a certain HBJ region inducing dimer formation in multiple primers; R represents the density of a certain HBJ region in the primers; n is the number of HBJ regions in each round of optimized primer combination, n ≥ 0;

[0016] a is the base length of the HBJ region, and a is an integer ≥ 3; b is the total number of G bases and C bases in the HBJ region; c is the total number of A bases and T bases in the HBJ region;

[0017] D1, D2, D3 and D4 are all non-zero integers; among them, D1 and D2 are the base distances of the candidate primer cp closest to the 5′ end of the HBJ region. i and cp j If the base number of the distance is zero, D1 or D2 is equal to 1; D3 and D4 are the base distances of the candidate primer cp closest to the 3′ end of the HBJ region. i and cp j The number of bases at the 3′ end of the last base. If the number of bases at the distance is zero, D3 or D4 is equal to 1;

[0018] num(cp i ) is the number of primers containing a specific HBJ region, num(cp j ) is the number of primers containing the HBJ complementary region; i and j are non-zero integers.

[0019] Furthermore, the method needs to be embedded in an evaluation cycle to ultimately achieve the goal of optimizing multiple primers. The process of the evaluation cycle is as follows:

[0020] S1: Generate a global candidate primer pool for X target regions. Generate Y upstream and downstream primers for each target region. The global candidate primer pool contains X×Y primers.

[0021] S2: Select the first-round primers from the global candidate primer pool. For each target region, select a pair of upstream and downstream primers to form the first-round candidate primer set, which is recorded as LHS1.

[0022] S3: Calculate the HBJ region between each primer in the candidate primer set of this round, and obtain n HBJ regions, which are recorded as LHS1 (PDST n ,R n );

[0023] S4: According to the mathematical model (1), a hash function method is used for rapid evaluation: for the first HDJ sequence, the D1 / D3 data set containing the HDJ sequence is used as the hash function value, and the D2 / D4 data set containing the HDJ complementary sequence is used as the key of the hash function to calculate the specific PDST value and obtain LHS1(PDST1,R1); then the PDST and R values ​​of the second to nth HBJ regions are calculated to obtain LHS1(PDST2,R2), ..., LHS1(PDST n-1 ,R n-1 ), LHS1(PDST n ,R n ), and finally get LHS1(ΣPDST n ,ΣR n );

[0024] S5: Sort the n PDST values ​​in LHS1 from large to small, take the top 50%, and then sort them from large to small according to their R values, take the top 60-90%;

[0025] S6: The corresponding primer is randomly selected or replaced, that is, the second primer is randomly selected from the corresponding remaining primers in the corresponding primer pool. After replacement, a new round of candidate primer set LHS2 is formed, and the evaluation process is repeated to obtain LHS2 (Σ′PDST n ,Σ′R n );

[0026] S7: Compare Σ′PDST n and ΣPDST n :

[0027] S71: If Σ′PDST n ≤ΣPDST n , then the candidate primer set was considered to have been optimized once, and the process of PDST value and R value sorting and random replacement of primers was repeated on the basis of LHS2;

[0028] S72: If Σ′PDST n >ΣPDST n , further judgment is needed:

[0029] S721: If Σ′R n >ΣR n , then LHS2 cannot be recognized, and it is necessary to repeat the process of PDST value and R value sorting and random replacement of primers based on LHS1 to obtain LHS2 again;

[0030] S722: If Σ′R n ≤ΣR n , then the probability of LHS2 being recognized is:

[0031] p=(ΣR n / Σ′R n +ΣPDST n / Σ′PDST n ) / 2,

[0032] S8: If LHS2 is not recognized, repeat the process of PDST value and R value sorting and random primer replacement based on LHS1 to obtain LHS2 again until LHS2 is recognized;

[0033] S9: Based on the recognition of LHS2, the process of PDST value and R value sorting and random replacement of primers in S5 to S8 was repeated to further obtain LHS3, ..., LHS m ,LHS m+1 until the LHS appears k times in a row m =LHS m+1 When , the evaluation process stops, and the primer set of round m+1 is the optimal primer combination of this evaluation, where k is a preset value, an integer ≥3, and the value range is 10 to 1000.

[0034] According to the present invention, preferably, the upstream and downstream primer amplification regions of the candidate primers conform to the preset insert fragment length range; the Gibbs free energy of the primers or the specific region of the primers is in the range of -8.0 to -23.0 kcal / mol under a salt concentration of 0.5 mol / L and an environment of 72°C; the GC content of the primers or the specific region of the primers is between 25% and 75%; the primers or the specific region of the primers are compared for sequence similarity in the corresponding reference genome or reference sequence, and primers with an E value less than 0.01 are selected as candidate primers.

[0035] According to the present invention, the preferred principle for randomly selecting or replacing primers in step S6 is: primers with a GC content of 48% to 52% and a Gibbs free energy in the range of -13.0 to -17.0 kcal / mol are preferably selected from the primer pool at each position.

[0036] Furthermore, if there are multiple primers that meet this condition, one is randomly selected; if the number of primers that meet this condition is zero, one is randomly selected from the entire primer pool.

[0037] According to the present invention, in step S5, the R values ​​are sorted from large to small, and the top 80% are selected.

[0038] According to the present invention, the default value of k in step S9 is 10.

[0039] A second aspect of the present invention provides a primer design and evaluation system for amplifying multiple target genomic regions, using the method described above, the primer design and evaluation system comprising:

[0040] one or more physical computer processors;

[0041] An electronic memory configured to store computer-readable instructions, the computer-readable instructions may include physical information of multiple target regions in a reference genome, multiple sequence information and relative positions of the target regions in the sequences, and / or a certain primer optimization combination.

[0042] The readable instructions, when executed by one or more physical computer processors, cause the one or more physical computer processors to:

[0043] a) generating an initial candidate primer set (LHS1) comprising at least 10 primers;

[0044] b) Primer dimerization evaluation function PDST n 、R n , and LHS1(ΣPDST n ,ΣR n )wait;

[0045] c) Further generate LHS2 based on LHS1, and so on m , and the optimization and replacement logic therein; and

[0046] d) An important parameter k at the end of the evaluation process, k is an integer ≥ 3, and the default value is 10.

[0047] According to the present invention, the output of the physical computer processor includes primer sequence, primer specific region or sequence, physical position of primer, sequence and physical position of amplification region, GC content, Tm value, and auxiliary information.

[0048] According to the present invention, the auxiliary information is selected from a molecular tag sequence, a sequencing primer sequence or a combination of these information.

[0049] The present invention has the following beneficial effects:

[0050] To address the problem of dimers that are easily induced when designing multiplex PCR primer sets, the present invention innovatively prioritizes the issue of primer dimers during the primer design stage. The proposed ultra-high-sensitivity multiplex sampling evaluation and optimization method based on a primer depolymerization algorithm can minimize the formation of primer dimers. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 The calculated values ​​of L(Sg) at different generations g in Example 2 are shown, representing the optimization trajectory of the method of the present invention for 167-plex primers.

[0052] Figure 2 The results of dimer ratio analysis after NGS sequencing of the products at the sampling points after library construction are shown.

[0053] Figure 3 The results of a wet experiment using a primer set of 856 primers and a primer set similar to the state of the spot check point D are shown. DETAILED DESCRIPTION

[0054] The present invention will be further described in detail below with reference to specific examples. It should be understood that the following examples are only used to illustrate the present invention and are not intended to limit the scope of the present invention.

[0055] In the context of the present invention, the "specific region" of the primer refers to the region that is identical or complementary to the reference genome. For example, if the primer is designed for humans, the specific region of the primer refers to the region that is identical or complementary to the sequence of the human reference genome (such as hg19 or hg38 versions); similarly, the "non-specific region" of the primer generally refers to sequences such as artificially added tags, barcodes, and indices. Generally speaking, such sequences have no identical or similar sequences in the human reference genome.

[0056] Example 1

[0057] 1.1 Evaluation and Optimization Methods for Ultra-High Sensitivity Multiplex Sampling Based on Primer Depolymerization Algorithm

[0058] This embodiment is an evaluation and optimization method for ultra-high sensitivity multiple sampling based on a primer depolymerization algorithm, which is used to evaluate and optimize the possibility of dimerization between the complementary regions of each primer in a candidate primer set. The complementary subregion within the primer sequence is called the "complementary polynuclear region", referred to as the "HBJ region". We use the following mathematical model (1) to describe the possibility of dimerization in the HBJ region, referred to as "dimerization inferiority", and the function LHS m Indicates that the higher the dimerization badness value, the higher the possibility of inducing dimers, and the more the corresponding primers need to be optimized:

[0059]

[0060] in:

[0061] m is the number of cycles to be evaluated, m = m + 1, m ≥ 0;

[0062] PDST represents the probability of a certain HBJ region inducing dimer formation in multiple primers; R represents the density of a certain HBJ region in the primers; n is the number of HBJ regions in each round of optimized primer combination, n ≥ 0;

[0063] a is the base length of the HBJ region, and a is an integer ≥ 3; b is the total number of G bases and C bases in the HBJ region; c is the total number of A bases and T bases in the HBJ region;

[0064] D1, D2, D3 and D4 are all non-zero integers; among them, D1 and D2 are the base distances of the candidate primer cp closest to the 5′ end of the HBJ region. i and cp j If the base number of the distance is zero, D1 or D2 is equal to 1; D3 and D4 are the base distances of the candidate primer cp closest to the 3′ end of the HBJ region. i and cp j The number of bases at the 3′ end of the last base. If the number of bases at the distance is zero, D3 or D4 is equal to 1;

[0065] num(cp i ) is the number of primers containing a specific HBJ region, num(cp j ) is the number of primers containing the HBJ complementary region; i and j are non-zero integers.

[0066] Furthermore, the method needs to be embedded in an evaluation cycle to ultimately achieve the goal of optimizing multiple primers. The process of the evaluation cycle is as follows:

[0067] S1: Generate a global candidate primer pool for X target regions. Generate Y upstream and downstream primers for each target region. The global candidate primer pool contains X×Y primers.

[0068] S2: Select the first-round primers from the global candidate primer pool. For each target region, select a pair of upstream and downstream primers to form the first-round candidate primer set, which is recorded as LHS1.

[0069] S3: Calculate the HBJ region between each primer in the candidate primer set of this round, and obtain n HBJ regions, which are recorded as LHS1 (PDST n ,R n );

[0070] S4: Based on the mathematical model (1), a hash function method is used for rapid evaluation: for the first HDJ sequence, the D1 / D3 data set containing the HDJ sequence is used as the value of the hash function, and the D2 / D4 data set containing the HDJ complementary sequence is used as the key of the hash function. The specific PDST value is calculated to obtain LHS1(PDST1,R1);

[0071] The specific method is shown in Table 1 below: For a certain HDJ sequence, the D1 / D3 data set containing the HDJ sequence is used as the value of the hash function, and the D2 / D4 data set containing the HDJ complementary sequence is used as the key of the hash function to calculate the specific PDST value:

[0072] Table 1

[0073]

[0074] Assume that a=7,b=3,c=4 in HBJ,

[0075] PDST=(87.17+115.33+356.16)×2^(0.2×7+1.2×3+0.5×4)=71508.48,

[0076] Among them, the cp of this HBJ region i Primer and cp j The primers interacted 9 times, i.e., R = 9; after the first HDJ region was evaluated, LHS1 (PDST1, R1) was obtained;

[0077] Then calculate the PDST and R values ​​of the 2nd to nth HBJ regions, and obtain LHS1(PDST2,R2), ..., LHS1(PDST n-1 ,R n-1 ), LHS1(PDST n ,R n ), and finally get LHS1(ΣPDST n ,ΣRn );

[0078] S5: sort the n PDST values ​​in LHS1 from large to small, take the top 50%, and then sort them from large to small according to their R values, take the top 60-90%, preferably take the top 80%;

[0079] S6: The corresponding primer is randomly replaced, that is, the second primer is randomly selected from the corresponding remaining primers in the corresponding primer pool. After replacement, a new round of candidate primer set LHS2 is formed, and the evaluation process is repeated to obtain LHS2 (Σ′PDST n ,Σ′R n );

[0080] S7: Compare Σ′PDST n and ΣPDST n :

[0081] S71: If Σ′PDST n ≤ΣPDST n , then the candidate primer set was considered to have been optimized once, and the process of PDST value and R value sorting and random replacement of primers was repeated on the basis of LHS2;

[0082] S72: If Σ′PDST n >ΣPDST n , further judgment is needed:

[0083] S721: If Σ′R n >ΣR n , then LHS2 cannot be recognized, and it is necessary to repeat the process of PDST value and R value sorting and random replacement of primers based on LHS1 to obtain LHS2 again;

[0084] S722: If Σ′R n ≤ΣR n , then the probability of LHS2 being recognized is:

[0085] p=(ΣR n / Σ′R n +ΣPDST n / Σ′PDST n ) / 2,

[0086] S8: If LHS2 is not recognized, repeat the process of PDST value and R value sorting and random primer replacement based on LHS1 to obtain LHS2 again until LHS2 is recognized;

[0087] S9: Based on the recognition of LHS2, the process of PDST value and R value sorting and random replacement of primers in S5 to S8 was repeated to further obtain LHS3, ..., LHS m,LHS m+1 until the LHS appears k times in a row m =LHS m+1 When , the evaluation process stops, and the primer set of round m+1 is the optimal primer combination of this evaluation;

[0088] k is a preset value, an integer ≥ 3, ranging from 10 to 1000, and with a default value of 10.

[0089] In this embodiment, preferably, the regions amplified by the upstream and downstream primers conform to a predetermined insert length range; the Gibbs free energy of the primers or primer-specific regions is within the range of -8.0 to -23.0 kcal / mol at a salt concentration of 0.5 mol / L and a temperature of 72°C. Preferably, the GC content of the primers or primer-specific regions is between 25% and 75%; the primers or primer-specific regions are compared for sequence similarity with a corresponding reference genome or reference sequence, for example, using a program such as Blast or Blastn. Preferably, primers with an E-value of less than 0.01 are selected as candidate primers.

[0090] In this embodiment, the preferred principle for randomly selecting or replacing primers is: at each position, primers with a GC content of 48% to 52% and a Gibbs free energy within the range of -13.0 to -17.0 kcal / mol are preferentially selected from the primer pool. If multiple primers meet this condition, one is randomly selected. If the number of primers meeting this condition is zero, one is randomly selected from the entire primer pool.

[0091] 1.2. Primer Design and Evaluation System for Amplifying Multiple Target Genomic Regions

[0092] Based on the aforementioned method for evaluating and optimizing the likelihood of dimer formation between complementary regions within a candidate primer set, a primer design and evaluation system for amplifying multiple target genomic regions was further developed to enhance its ease of use. The system includes:

[0093] one or more physical computer processors;

[0094] An electronic memory configured to store computer-readable instructions, the computer-readable instructions may include physical information of multiple target regions in a reference genome, multiple sequence information and relative positions of the target regions in the sequences, and / or a certain primer optimization combination.

[0095] The readable instructions, when executed by one or more physical computer processors, cause the one or more physical computer processors to:

[0096] a) generating an initial candidate primer set (LHS1) comprising at least 10 primers;

[0097] b) Primer dimerization evaluation function PDST n 、R n , and LHS1(ΣPDST n ,ΣR n )wait;

[0098] c) Further generate LHS2 based on LHS1, and so on m , and the optimization and replacement logic therein; and

[0099] d) An important parameter k at the end of the evaluation process, where k is an integer ≥ 3 and the default value is 10;

[0100] In this embodiment, the content output by the physical computer processor may include primer sequence, primer-specific region or sequence, physical position of the primer, sequence and physical position of the amplification region, GC content, Tm value, and various ancillary information, such as molecular tag sequence, sequencing primer sequence, or a combination of such information.

[0101] Example 2: Design and optimization of a single-tube 117-plex primer set and experimental evaluation

[0102] This example uses the method of Example 1 to design and optimize a 167-plex primer set targeting multiple target regions of multiple genes. Figure 1 The calculated values ​​of L(Sg) for different generations g are shown, representing the optimization trajectory of the program for the 167-plex primer set. Furthermore, we used a wet run experiment to observe the optimization effect. We tested the optimized primer sets at four different optimization time points: sampling points A, B, C, and D. Sampling point A was used during the initial unoptimized period, sampling points B and C were used during the intermediate optimization period, and sampling point D was used during the final optimized period.

[0103] The wet experiment conditions are as follows: 20 ng human peripheral blood DNA was selected, the concentration of each primer was 25 nM, 50 μL system, 20 cycles, the obtained products were analyzed by gel electrophoresis, and the products were concentrated at 240 nt. The results are shown in Figure 1 .

[0104] Depend on Figure 1 It can be seen that the degree of optimization of the primer combination of sampling point A is the worst, and almost no specific products are obtained, all of which are primer dimers. The degree of optimization of the primer combination of sampling points B, C, and D is gradually improved, and the proportion of specific products is getting higher and higher.

[0105] We constructed a library for the products of the sampling points and then performed NGS sequencing. After analysis, we found that the dimer proportions of sampling points A, B, C, and D were 96.2%, 55.7%, 34.3%, and 5.6%, respectively. Figure 2 The verification results of the wet experiment show that the primer depolymerization algorithm disclosed in the present invention has an excellent optimization effect.

[0106] Next, we used the primer combination at point D to test 30 clinical samples: 10 non-small cell lung cancer tissue samples, 10 healthy volunteer peripheral blood samples, and 10 breast cancer samples. We logarithmized the number of reads for each sequence type within the sample data, calculated the mean for each sample, and used these mean values ​​to analyze the mean and standard deviation for each group. The results are shown in Table 2 below.

[0107] Table 2

[0108] Statistical items Sample type <![CDATA[Log 10 Average value of (reads)]]> Std value On target part Non-group 3.62 0.64 Breast cancer group 3.91 0.81 Normal group 3.54 0.64 Dimer part Non-group 1.97 0.69 Breast cancer group 2.34 0.59 Normal group 2.4 0.64 Nonspecific fraction Non-group 2.6 0.36 Breast cancer group 2.69 0.38 Normal group 2.91 0.33

[0109] From the results in Table 2, we can see that in the same statistical project, there is no significant difference in the mean and Std values ​​of different sample types, which means that the primer combination of sampling point D is uniform in performance and will not have significant differences due to differences in samples.

[0110] Furthermore, we continued to develop a primer set with a total of 856 primers, and used a primer set similar to the state of the sample D point to perform wet experiment verification. Figure 3 shown.

[0111] Depend on Figure 3 The results show that the dimer ratio is around 3% and the overall on-target rate is 67%, showing a good effect.

[0112] The working conditions of the above tests are as follows:

[0113] 1. Primers: All primers were obtained from a domestic general biological company, premixed by the supplier, and stored at 4°C for a short term and at -20°C for a long term.

[0114] 2. Samples: Store gDNA samples from peripheral blood of volunteers at -20°C. Mix gDNA with plasmid DNA template containing specific mutation information at varying ratios to create samples containing different ratios. Dilute the gDNA sample and synthesized DNA template in 1× TE buffer containing 0.1% Tween 20.

[0115] 3. Multiplex PCR protocol: Multiplex PCR was performed on an ABI7500 PCR instrument. The total volume of each reaction was 50 μL, and the amount of DNA added was 10-100 ng per tube.

[0116] The specific steps are: 95°C, 3 minutes, then 20 cycles: 95°C, 30 seconds, 60°C, 3 minutes, 72°C, 30 seconds, and 72°C, 5 minutes after the cycle ends.

[0117] The reaction system is as follows:

[0118] Final concentration volume PCR buffer 10× 1× 5μL dNTP 10mM 0.4mM 2μL Upstream primer mix, 200 nM 30nM 7.5 μL Downstream primer mix, 200 nM 30nM 7.5 μL gDNA template 10ng PCR amplification enzyme 2unit Add deionized water to 50 μL

[0119] 4. Library construction: use Ultra TM Libraries were constructed from the multiplex PCR products using the ELISA II DNA Library Preparation Kit (E7645L) according to the manufacturer's instructions. After library construction, sequencing was performed using the ILMN platform, and the data were analyzed.

[0120] 5. Primer-dimer analysis: First, we removed the sequencing adapter sequences from the raw FASTQ data and separated all reads by their length: we saved reads >60 nt as long read files and reads shorter than 60 nt as short read files. It is worth noting that when the primer structure lacks other sequences such as molecular tags or common sequences, 60 nt is selected as the filtering threshold. If the primer structure contains other sequence structures in the sequencing direction, the filtering threshold is 60 nt plus the sum of the lengths of these sequence structures.

[0121] 6. Use plug-ins such as Bowtie2 to align reads: The principles include aligning long-read files with reference sequences using Bowtie2, and sequences in long-read files and short-read files that cannot be aligned are considered potential dimer sequences.

[0122] 7. Distinguishing primer dimers from nonspecific amplification products: There are commonly used methods to distinguish between the two. We make a slight change. First, the 6-10 base sequence at the 3′ end of the primer sequence used in the project is used as the primer dimer seed sequence. These seed sequences are aligned with the potential dimer sequences mentioned above. On the aligned reads sequence, starting from the seed sequence, a 3-10 base sequence is taken from the 3′ end and aligned with the primer sequence of this project. If alignment is possible, the potential dimer sequence is marked as a dimer sequence, otherwise it is marked as a nonspecific amplification product. Nonspecific amplification products also include sequences that can be aligned to the genome but are not in the target region.

[0123] Although the preferred embodiments of the present invention are described in detail above, it should be understood that a person skilled in the art may make several improvements and modifications without departing from the principles of the present invention, and these improvements and modifications also fall within the scope of the present invention.

Claims

1. A multiple sampling evaluation and optimization method based on a primer depolymerization algorithm for evaluating and optimizing the possibility of dimerization between complementary regions of primers in a candidate primer set, characterized in that: The complementary subregion within the primer sequence is the HBJ region, and the possibility of the HBJ region initiating dimerization is determined by the function LHS of the mathematical model (1) m Indicates that LHS m The higher the value, the higher the possibility of inducing dimers, and the more the corresponding primers need to be optimized: ; in: m is the number of cycles to be evaluated, m=m+1, m≥0; PDST represents the probability of a certain HBJ region inducing dimer formation in multiple primers; R represents the density of a certain HBJ region in the primers; n is the number of HBJ regions in each round of optimized primer combination, n ≥ 0; a is the base length of the HBJ region, and a is an integer ≥ 3; b is the total number of G bases and C bases in the HBJ region; c is the total number of A bases and T bases in the HBJ region; D1, D2, D3 and D4 are all non-zero integers; among them, D1 and D2 are the base distances of the candidate primer cp closest to the 5' end of the HBJ region. i and cp j The number of bases at the 5' end of the primer. If the number of bases at the distance is zero, D1 or D2 is equal to 1; D3 and D4 are the base distances from the candidate primer cp to the base closest to the 3' end of the HBJ region. i and cp j The number of bases at the 3' end of the last base. If the number of bases at the distance is zero, D3 or D4 is equal to 1; num(cp i ) is the number of primers containing a specific HBJ region, num(cp j ) is the number of primers containing a region complementary to the HBJ region; i and j are non-zero integers.

2. The method according to claim 1, characterized in that The method needs to be embedded in an evaluation cycle to ultimately achieve the goal of optimizing multiple primers. The process of the evaluation cycle is as follows: S1: Generate a global candidate primer pool for X target regions. Generate Y upstream and downstream primers for each target region. The global candidate primer pool contains X × Y primers. S2: Select the first-round primers from the global candidate primer pool. For each target region, select a pair of upstream and downstream primers to form the first-round candidate primer set, which is recorded as LHS1. S3: Calculate the HBJ region between each primer in the candidate primer set of this round, and obtain n HBJ regions, which are recorded as LHS1 (PDST n ,R n ); S4: According to the mathematical model (1), a hash function method is used for rapid evaluation: for the first HBJ sequence, the D1 / D3 data set containing the HBJ sequence is used as the value of the hash function, and the D2 / D4 data set containing the HBJ complementary sequence is used as the key of the hash function. The specific PDST value of the HBJ region is calculated to obtain LHS1(PDST1, R1); then the PDST and R values ​​of the second to nth HBJ regions are calculated to obtain LHS1(PDST2, R2), ..., LHS1(PDST n-1 , R n-1 ), LHS1(PDST n ,R n ), and finally get LHS1(ΣPDST n , ΣR n ); S5: Sort the n PDST values ​​in LHS1 from large to small, take the top 50%, and then sort them from large to small according to their R values, take the top 60-90%; S6: The corresponding primer is randomly selected or replaced, that is, the second primer is randomly selected from the corresponding remaining primers in the corresponding primer pool. After replacement, a new round of candidate primer set LHS2 is formed, and the evaluation process is repeated to obtain LHS2 (Σ´PDST n , Σ´R n ); S7: Compare Σ´PDST n and ΣPDST n : S71: If Σ´PDST n ≤ ΣPDST n , then the candidate primer set was considered to have been optimized once, and the process of PDST value and R value sorting and random replacement of primers was repeated on the basis of LHS2; S72: If Σ´PDST n > ΣPDST n , further judgment is needed: S721: If Σ´R n > ΣR n , then LHS2 cannot be recognized, and it is necessary to repeat the process of PDST value and R value sorting and random replacement of primers based on LHS1 to obtain LHS2 again; S722: If Σ´R n ≤ ΣR n , then the probability of LHS2 being recognized is: p=(ΣR n / Σ´R n +ΣPDST n / Σ´PDST n ) / 2, S8: If LHS2 is not recognized, repeat the process of PDST value and R value sorting and random primer replacement based on LHS1 to obtain LHS2 again until LHS2 is recognized; S9: Based on the recognition of LHS2, the process of PDST value and R value sorting and random replacement of primers in S5 to S8 was repeated to further obtain LHS3, ..., LHS m , LHS m+1 until the LHS appears k times in a row m =LHS m+1 When , the evaluation process stops, and the primer set of round m+1 is the optimal primer combination of this evaluation, where k is a preset value, which is an integer in the range of 10 to 1000.

3. The method according to claim 1 or 2, characterized in that: The upstream and downstream primer amplification regions of the candidate primers conform to the preset insert fragment length range; The primer or the specific region of the primer has a Gibbs free energy within the range of -8.0 to -23.0 kcal / mol at a salt concentration of 0.5 mol / L and a temperature of 72°C; The GC content of the primer or the specific region of the primer is between 25% and 75%; The primers or specific regions of the primers are compared for sequence similarity in the corresponding reference genome or reference sequence, and primers with an E value less than 0.01 are selected as candidate primers.

4. The method according to claim 2, characterized in that The conditions for randomly selecting or replacing primers in step S6 are: the GC content in the primer pool at each position is 48% to 52%, and the Gibbs free energy is within the range of -13.0 to -17.0 kcal / mol in a 0.5 mol / L salt concentration and a 72°C environment.

5. The method according to claim 4, characterized in that If there are multiple primers that meet this condition, one is randomly selected; if the number of primers that meet this condition is zero, one is randomly selected from the entire primer pool.

6. The method according to claim 2, characterized in that In step S5, the R values ​​are sorted from large to small, and the top 80% are selected.

7. The method according to claim 2, characterized in that The default value of k in step S9 is 10.

8. A primer design and evaluation system for amplifying multiple target genomic regions, using the method of any one of claims 1 to 7, characterized in that: The primer design and evaluation system includes: one or more physical computer processors, one or more electronic memories storing computer-readable instructions, wherein the computer-readable instructions include physical information of multiple target regions in a reference genome, multiple sequence information and relative positions of the target regions in the sequences, and / or a certain primer optimization combination; The readable instructions, when executed by one or more physical computer processors, cause the one or more physical computer processors to: a) generating an initial candidate primer set LHS1 comprising at least 10 primers, b) Primer dimerization evaluation function PDST n 、R n , and LHS1(ΣPDST n , ΣR n ), c) Further generate LHS2 based on LHS1, and so on m , and the optimization and replacement logic therein, and d) The important parameter k at the end of the evaluation process, with a default value of 10.

9. The primer design and evaluation system according to claim 8, characterized in that The output of the physical computer processor includes primer sequence, primer specific region or sequence, physical position of primer, sequence and physical position of amplification region, GC content, Tm value, and auxiliary information.

10. The primer design and evaluation system according to claim 9, characterized in that: The auxiliary information is selected from a molecular tag sequence, a sequencing primer sequence, or a combination of the two.

Citation Information

Patent Citations

  • Multi-PCR (Polymerase Chain Reaction) primer design method based on Primer 3

    CN107937497A

  • Highly multiplex PCR methods and compositions

    WO2014018080A1