Universal nucleic acid probe design method based on hybridization thermodynamics and dynamics
Through the intelligent nucleic acid hybridization algorithm, the primer and probe design is optimized, and the flux and amplification specificity of multiple detection in the prior art is solved, and efficient and sensitive nucleic acid detection is achieved to adapt to complex detection environments.
Patent Information
- Application Number
- CN202510455896.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-08-05
AI Technical Summary
In multiple detection, existing nucleic acid detection technologies have problems such as limited detection throughput, low amplification specificity, nonlinear growth in thermodynamic interference between probes, exponential increase in computational complexity, and difficulty in dynamic regulation of performance, which cannot meet the detection needs of high throughput and high sensitivity.
An intelligent nucleic acid hybridization algorithm based on hybrid thermodynamics and kinetics is used to combine chemical thermodynamics, kinetic parameters, sequence complementarity and enzymatic reaction characteristics to generate the optimal primer pair and probe combination, and efficient and specific nucleic acid detection is achieved through multi-dimensional constraint conditions and objective function optimization.
It improves the flux and amplification uniformity of nucleic acid detection, enhances the flexibility and adaptability of probe design, reduces system errors, and achieves efficient multi-nucleic acid variation detection.
Smart Images

Figure CN120432005A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of gene detection technology, and more particularly to a universal nucleic acid probe design method based on hybridization thermodynamics and kinetics. Background Art
[0002] As a major contributor to genetic diseases and cancer, accurate detection of gene mutations is crucial for early diagnosis and treatment. However, existing diagnostic technologies face bottlenecks such as low throughput, high costs, and long testing cycles, making it difficult to meet the clinical needs of simultaneous detection of multiple nucleic acid mutations. In the field of tumor genetic testing, the multiplicity and throughput limitations of the detection system directly restrict the depth of biomarker discovery and its clinical application value.
[0003] PCR is a key technology in molecular biology, capable of amplifying DNA by a billion-fold in a short period of time. This powerful amplification ability makes PCR indispensable for detecting small amounts of mutant molecules, and it plays a particularly important role in molecular diagnostics, gene cloning, mutation detection, and pathogen detection. However, the success of PCR technology depends largely on the quality of primer design, and traditional primer design methods often use fixed-length primers, usually 20 nucleotides (20nt). Although this one-size-fits-all approach performs well in simple PCR, it exposes a series of problems in complex multiplex PCR and ultra-high sensitivity detection.
[0004] The difference in binding efficiency of different primers will show an exponential deterioration during the PCR cycle, which will eventually affect the uniformity of the amplified products. This uneven amplification will not only result in some target sequences not being fully amplified, but may also cause potential loss of nucleotides, thereby affecting the capture of variant nucleic acid molecules and leading to false negative diagnostic results. This problem is particularly prominent for application scenarios involving multi-gene testing or multiple primers, such as cancer mutation detection or pathogen screening. Secondly, primers with lower binding efficiency are not fully amplified during the amplification process, which will lead to data deviation in subsequent next-generation sequencing (NGS). In order to compensate for this deviation, it is usually necessary to continuously increase the sequencing depth, which not only increases the cost of testing, but may also lead to a significant increase in the complexity of data analysis.
[0005] Interactions between primers, especially the formation of primer dimers, are another key challenge in PCR reactions. Primer dimers not only reduce the specificity and sensitivity of the reaction, but can also lead to nonspecific amplification, thereby interfering with the detection of the target sequence. In multiplex PCR, the probability of primer dimer formation increases significantly with the increase in the number of primers, which makes the complexity and difficulty of the reaction increase exponentially. Traditional solutions, such as using digestive enzymes to remove or purify dimers, are often remedial measures after the fact and cannot fundamentally solve the problem. Therefore, minimizing the formation of primer dimers during the primer design stage is the key to improving PCR efficiency and sensitivity.
[0006] With the development of bioinformatics and computational biology, algorithm-based primer design methods have gradually become the mainstream direction for solving the above problems. These methods accurately calculate the biophysical properties of primer sequences (such as GC content, melting temperature, secondary structure, etc.) and use mathematical models to predict the interactions between primers, thereby optimizing the reaction conditions at the primer design stage. For example, by introducing the calculation of Gibbs free energy, the stability of primer-template binding can be evaluated, and primer combinations with higher binding efficiency can be screened. In addition, by modeling the formation mechanism of primer dimers, primer sequences that are prone to dimer formation can be eliminated at the design stage, thereby improving the specificity and sensitivity of the reaction. These methods not only perform well in single-plex PCR, but also have significant advantages in multiplex PCR and super-multiplex PCR.
[0007] Although algorithm-based primer design methods are theoretically very advanced, they still face several challenges in practical application. For example, the complexity and diversity of different target sequences can lead to a sharp increase in the computational workload of the optimization algorithm. This is especially true in super-multiplex PCR, where the number of primer combinations increases exponentially, placing higher demands on the algorithm's computational efficiency and accuracy. Furthermore, other factors must be considered during primer design, such as primer specificity, secondary structure stability, and the mismatch rate between primer and template. These factors are intertwined, making primer design a highly complex multi-objective optimization problem.
[0008] The primer design methods CN202411805012.1 and CN202011633512.3 for multiplex detection control Tm, GC content, primer length, etc., and use the primer design tools provided by NCBI to mainly infer the primer sequence based on Tm. This design method patent does not consider the effects of the secondary structure of the primer itself and the interaction between primers on amplification. CN202310092169.3 uses a machine learning model to optimize primer design and amplification efficiency, focusing on predicting the ideal molecular weight concentration ratio of primer pairs and adjusting the primer ratio. CN202411583271.4 evaluates primer design by calculating the PDST and R value cycles of the HBJ region, mainly focusing on the interaction between primers, without considering their own secondary structure.
[0009] However, these published patents still have obvious defects: 1) the detection throughput of existing primer design methods is limited; 2) existing methods do not consider the impact of amplification specificity and yield caused by primer structure; 3) there is a lack of optimization of secondary structure and thermodynamic properties.
[0010] Current mainstream nucleic acid detection technologies exhibit significant differences in their multi-scale detection capabilities. As a core enabling technology for genomic analysis and molecular diagnostics, nucleic acid probes play a key role in cross-scale studies, including transcriptome dynamics tracking, genetic variation screening, infectious pathogen lineage typing, and chromatin modification regulation. Currently, faced with the exponential growth in the complexity of omics data analysis and the demand for simultaneous capture of multi-dimensional targets for clinical precision testing, classic probe construction systems are gradually revealing their multi-dimensional adaptability limitations.
[0011] Non-specific detection systems based on double-stranded DNA binding dyes (such as SYBR Green) rely on melting curve analysis to achieve target identification. Although low-cost, a single-tube reaction can only detect one target sequence, and multiplex detection must be achieved through a split-tube design. To increase the throughput of single-tube detection, several studies have combined sequence-specific probes (such as TaqMan probes) with a fluorophore quenching mechanism to exploit the spectral differences of different fluorescence channels to achieve multiplex detection in a single reaction tube. However, in actual operation, due to spectral overlap between different channels, the number of effective independent detection channels is usually ≤4; and the probe design must strictly meet the orthogonality rule, forcing the actual application to be controlled at 3-4 multiplexes.
[0012] Meanwhile, while detection platforms based on high-throughput sequencing (NGS) technology theoretically enable parallel detection of thousands of targets through hybridization capture, breaking the traditional throughput barrier, their actual capture efficiency is severely constrained by thermodynamic competition within the probe pool. This "preemptive effect" of high-abundance probes due to their hybridization kinetics reduces coverage depth of low-abundance targets, and exponentially amplifies GC bias during PCR amplification. Such systematic errors not only reduce data signal-to-noise ratios and multiply sequencing redundancy, but also significantly reduce the efficiency of resolving critical low-frequency signals in multi-omics association analyses.
[0013] In high-density multiplex probe systems (such as 10,000-level probe pools), thermodynamic interference caused by cross-probe interactions has become a key bottleneck affecting detection uniformity and data fidelity. Although traditional methods optimize hybridization performance by adjusting the GC content and melting temperature of single probes, in large-scale parallel reactions, competitive hybridization will lead to preferential amplification of dominant probes (such as high GC probes), ultimately causing target coverage bias. For example, in multiplex detection of lung cancer-related genes (EGFR / ALK) and tumor suppressor genes (TP53 / RB1), high GC probes trigger increased polymerase occupancy, resulting in signal attenuation in low GC regions.
[0014] Existing probe design methods primarily optimize hybridization performance by fixing primer length, standardizing GC content, controlling single-base repeat sequences, and regulating melting temperature, among other static parameters. However, these optimization strategies have significant limitations: First, traditional design methods struggle to accurately model the dynamic thermodynamic equilibrium during probe-target binding, making it impossible to accurately predict and regulate hybridization behavior in complex environments. Second, multiple parameter constraints significantly compress the available probe combination space. In regions with high single nucleotide variation (SNV) coverage density, the number of candidate probes that simultaneously meet all screening criteria is less than 20% of the target region, greatly limiting the flexibility and adaptability of probe design. Furthermore, existing methods struggle to effectively address the thermodynamic interference between different probes in the probe pool, resulting in a significant reduction in detection uniformity and data fidelity. In complex detection systems, systematic errors in multiple links couple with each other and exhibit a cumulative amplification effect, significantly impacting the credibility of the final detection results. Potential errors introduced in the probe design stage are gradually amplified during hybridization, amplification, and detection, exacerbating target coverage bias. Especially in high-throughput detection scenarios, traditional methods are unable to achieve dynamic balance control of probe performance, resulting in preferential amplification of advantageous probes (such as high GC probes), further exacerbating the systematic error of the detection results.
[0015] As omics research develops towards higher dimensions, the demand for large-scale multiplexed detection systems is becoming increasingly prominent. However, existing technologies face severe challenges in designing large-scale probe pools: first, the thermodynamic interference between probes increases nonlinearly with the number of probes, significantly affecting the stability and reliability of the detection system; second, the computational complexity of probe design increases exponentially with the number of probes, making it difficult for traditional methods to optimize the design of large-scale probe pools within a reasonable time; finally, existing probe design methods struggle to dynamically control and optimize probe performance, making them unable to adapt to complex and changing detection environments. Summary of the Invention
[0016] The purpose of the present invention is to provide a universal nucleic acid probe design method based on hybridization thermodynamics and kinetics, so as to solve the problems of limited detection flux, high amplification specificity, and low amplification yield of primer design methods in the existing technology, as well as the problems of nonlinear growth of thermodynamic interference between probes, exponential increase in computational complexity, and difficulty in dynamically regulating performance to adapt to complex environments when dealing with large-scale probe pool design.
[0017] In order to solve the above technical problems, the present invention adopts the following technical solutions:
[0018] Provided is a universal nucleic acid probe design method based on hybridization thermodynamics and kinetics, comprising the following steps: 1) determining a target amplification region based on a fragment interval, generating a primer pair that meets the length requirement and covers the target fragment, using an intelligent nucleic acid hybridization algorithm to calculate the Gibbs free energy change of nucleic acid hybridization, and combining chemical thermodynamics, kinetic parameters, sequence complementarity, and the characteristics of the enzymatic reaction to generate an optimal primer pair under multidimensional constraints; 2) based on the generated optimal primer pair, generating a probe combination that meets a free energy difference threshold as a candidate pool, screening the probe combination with the highest objective function value from the candidate pool through an iterative process, and gradually updating the candidate pool until the primer pair and probe combination with the highest objective function value are output.
[0019] In step 1), the intelligent nucleic acid hybridization algorithm is based on the following formula: ΔG°=ΔH°-TΔS°ΔG°=ΔH°-TΔS°; Where ΔH° is the enthalpy change, ΔS° is the entropy change, and T is the absolute temperature. The values of ΔH° and ΔS° are calculated using the Nearest-Neighbor Model, which allows the calculated results of ΔG° to be accurately correlated with hybridization stability.
[0020] In step 1), the multidimensional constraints include: primer length is between 18 and 30 bases to ensure hybridization efficiency and specificity; GC content is controlled between 40% and 60% to balance hybridization stability and melting temperature; in the case of homopolymers, it is necessary to evaluate whether the homopolymer length does not exceed 4; when the primer itself has complementary base pairs and causes primer dimers, the number of complementary bases with itself or another primer at the near 3' end position does not exceed 4, and cannot exceed 6 at the middle position.
[0021] In step 1), the multidimensional constraints also include: using NUPACK to verify the secondary structure of the primers themselves. At the reaction temperature, nucleic acid sequences with obvious secondary structure are directly discarded. Primers that pass the test are further screened for nucleic acid sequence specificity in the corresponding target gene library according to design requirements. This assesses the specific binding of the primer or probe to the target sequence and identifies potential cross-reactions with non-target sequences. By accurately calculating the alignment score and mismatch penalty, primers and probes with high specificity are selected.
[0022] In step 2), the multidimensional constraint condition further includes: using the Smith-Waterman algorithm for local alignment, calculating the alignment score of the query sequence with the target sequence and the non-target sequence, with a score of +2 for a matched base pair, a penalty of -3 for a mismatched base pair, a penalty of -5 for introducing a gap, and a penalty of -2 for extending a gap, and obtaining a final score to define a specificity score: Set Specificity according to the amplification specificity requirements s Threshold: If the score of the query sequence and the non-target sequence exceeds the set threshold (Score non-target >0.8×Score target ), then cross-reaction is considered to exist.
[0023] Preferably, in step 2), the objective function is introduced: Objective_Function=Specificity s -λ·∣ΔGprimer-ΔGprobe∣; Among them, Specificity s is the specificity score, ΔGprimer-ΔGprobe is the change in hybridization free energy caused by primer and probe, and λ is an adjustment parameter used to balance amplification specificity and thermodynamic stability, which is quickly calculated using the objective function.
[0024] According to the present invention, in step 2), in each selection step, the candidate combination with the highest objective function value is selected as the current set, the candidate pool is updated, and the selected combinations are excluded. The above process is repeated until the candidate pool is empty or the preset number of iterations is reached. Finally, the primer pair and probe combination with the highest objective function value is output.
[0025] The universal nucleic acid probe design method also includes: introducing functional sequences and chemical modifications into primers and probes to expand their applications in nucleic acid storage, error analysis, error correction coding, and high-precision detection.
[0026] The functional sequences include: adapter sequence, stem-loop structure sequence, molecular tag sequence, redundant error correction sequence, and tag sequence; the chemical modifications include: RNA, LNA, PNA, XNA, dU, Spacer, PEG, fluorescent group, phosphorylation group, reverse dT, and methylated base.
[0027] The present invention is based on an intelligent nucleic acid hybridization algorithm based on hybridization thermodynamics and kinetics. Compared with the existing technology of designing primers based on Tm values, the present invention more comprehensively considers aspects such as primer length, GC content, self-secondary structure and primer dimers. It can optimize primer design under multidimensional constraints and greatly improve nucleic acid detection throughput.
[0028] Although some existing tools (such as Primer3, OligoCalc, and NUPACK) have introduced some automated algorithms to assist in primer and probe design, these tools are usually limited to the optimization of specific parameters (such as Tm value, GC content, dimer formation, etc.) and lack a comprehensive evaluation of specificity and amplification efficiency. The design of some libraries has always been based on manual experience selection or optimization. Due to the contradiction between specificity and amplification efficiency, amplification bias is easily generated, thereby affecting the final throughput.
[0029] Precisely to solve the above problems, the present invention provides a universal nucleic acid probe design method based on hybridization thermodynamics and kinetics. The main inventive point of the present invention is that, first, it provides an intelligent nucleic acid hybridization algorithm, whose core goal is to optimize primer design by accurately predicting and controlling the nucleic acid hybridization process, so as to achieve efficient and specific nucleic acid detection. The algorithm combines chemical thermodynamics, kinetic parameters, sequence complementarity and the characteristics of enzymatic reactions to generate the optimal primer design scheme under multidimensional constraints; secondly, on the basis of the specificity algorithm of the intelligent nucleic acid hybridization algorithm, the objective function is used to evaluate the effect of primers and probes, and the specificity and thermodynamic stability are uniformly quantified. The balance relationship between the inhibition of amplification effect and amplification efficiency of different probe designs is further given, and the parameter λ is used to achieve control of the detection effect, and finally the primer pair and probe combination with the highest objective function value is obtained.
[0030] Compared with the prior art, the present invention has the following beneficial effects:
[0031] 1) This invention introduces intelligent algorithms to achieve automated design and optimization. By combining thermodynamic models, kinetic models, and specificity assessment, it automatically generates optimal primer and amplification inhibition probe combinations and provides an integrated verification solution to ensure the reliability and performance of the design.
[0032] 2) The present invention fully considers the hybridization thermodynamics of primers, probes, and templates, thereby achieving uniform amplification of target sequences and amplification and enrichment of mutant templates;
[0033] 3) The present invention greatly improves the throughput of a single library, thereby enabling a wider range of nucleic acid variation detection in a single tube. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Intelligent nucleic acid hybridization design process;
[0035] Figure 2 Primer design and validation examples;
[0036] Figure 3 Target enrichment probe design workflow. DETAILED DESCRIPTION
[0037] The present invention will be further described below with reference to specific examples. It should be understood that the following examples are intended to illustrate the present invention only and are not intended to limit the scope of the present invention. Unless otherwise specified, the techniques used in the examples are conventional in the art, or according to the experimental methods recommended by the kit and instrument manufacturers. The reagents and materials used in the examples are commercially available unless otherwise specified.
[0038] According to the present invention, a universal nucleic acid probe design method based on hybridization thermodynamics and kinetics is provided, comprising the following steps:
[0039] 1) Intelligent nucleic acid hybridization design to generate optimal primer pairs;
[0040] 2) Based on the generated optimal primer pairs, design a high-throughput nucleic acid probe set and optimize it based on biophysical parameters;
[0041] 3) Synthesize DNA and store it under appropriate conditions;
[0042] 4) Amplify the target sequence by multiplex PCR based on the designed primer set;
[0043] 5) Reading the interaction signal between the probe and the target sequence by sequencing, fluorescence detection, or electrochemical detection;
[0044] like Figure 1The figure shows a flowchart for step 1). This flowchart illustrates the process for determining the optimal primer pair: first, the target amplification region is determined based on the fragment interval, followed by generating a primer pair that meets the required length and covers the target fragment. Next, an intelligent nucleic acid hybridization algorithm is used to check the primer pair for free energy, GC content, melting temperature, homopolymers, dimers, and other parameters. Secondary structure verification and specificity screening are then performed in sequence. If any step fails, the design is changed until the optimal primer pair is finally determined.
[0045] The core goal of intelligent nucleic acid hybridization algorithms is to optimize primer design by accurately predicting and controlling the nucleic acid hybridization process, thereby achieving efficient and specific nucleic acid detection. The algorithm combines chemical thermodynamics, kinetic parameters, sequence complementarity, and the characteristics of enzymatic reactions to generate optimal primer designs under multidimensional constraints.
[0046] This intelligent nucleic acid hybridization algorithm calculates the Gibbs free energy change (ΔG°) of nucleic acid hybridization using chemical thermodynamic parameters. ΔG° is a key indicator for measuring the stability of the hybridization process. The calculation formula is as follows: ΔG°=ΔH°-TΔS°ΔG°=ΔH°-TΔS°
[0047] Among them, ΔH° is the enthalpy change, ΔS° is the entropy change, and T is the absolute temperature. The values of ΔH° and ΔS° are calculated by the Nearest-Neighbor Model, which takes into account the interaction between adjacent base pairs, rather than relying solely on the characteristics of a single base pair, and can accurately predict the stability of the hybridization process. The stability of nucleic acid hybridization is highly dependent on the interaction between adjacent base pairs (such as hydrogen bonds, base stacking forces, etc.). The neighbor model accurately calculates ΔH° and ΔS° by analyzing the combination of adjacent base pairs, fully considering the synergistic effect between bases. Substituting the ΔH° and ΔS° obtained based on this model into the formula can more realistically reflect the energy changes in the nucleic acid hybridization process, so that the calculation results of ΔG° are accurately correlated with the hybridization stability, which enhances the application value of the formula in nucleic acid systems.
[0048] To further optimize the ΔG° value, the intelligent nucleic acid hybridization algorithm introduces the following constraints: Primer length is generally between 18 and 30 bases to ensure hybridization efficiency and specificity. The GC content is controlled between 40% and 60% to balance hybridization stability and melting temperature (Tm). In the case of homopolymers (consecutive identical bases are called homopolymers, such as AAAA, TTTT, CCC, and GGG), it is necessary to evaluate whether the homopolymer length does not exceed 4. When a primer itself has complementary base pairs, resulting in primer dimers, the number of complementary bases with itself or another primer at the near-terminal position should not exceed 4, and at the middle position, it cannot exceed 6. The secondary structure of the primer itself is verified using NUPACK and other methods. At the reaction temperature, nucleic acid sequences with obvious secondary structure should be directly discarded. Primers that pass the test are further screened for nucleic acid sequence specificity in the corresponding target gene library according to design requirements to evaluate the specific binding degree of the primer or probe domain to the target sequence and identify potential cross-reactions with non-target sequences. By accurately calculating the alignment score and mismatch penalty, primers and probes with high specificity are screened. The Smith-Waterman algorithm is used for local alignment to calculate the alignment score between the query sequence and the target sequence and non-target sequences. The score for a matching base pair is +2, the penalty for a mismatched base pair is -3, the penalty for introducing a gap is -5, and the penalty for extending a gap is -2, and the final score is obtained. Define the specificity score:
[0049] Set Specificity according to the amplification specificity requirements s Threshold: If the score of the query sequence and the non-target sequence exceeds the set threshold (Score non-target >0.8×Score target ), then cross-reaction is considered to exist.
[0050] It should be understood that the present invention significantly improves the specificity, stability, and efficiency of primer and probe designs through system integration, innovative scoring methods, objective function design, and algorithmic optimization. Existing related patents do not fully consider all of these conditions, resulting in a failure to achieve high uniformity in amplification efficiency, which in turn limits amplification throughput.
[0051] like Figure 2 As shown in Figure 2, we constructed eight libraries with specified content and used the aforementioned intelligent nucleic acid hybridization algorithm to generate 35,406 pairs of amplification primers for standard NGS library construction. Compared to conventional hybridization algorithms that generate primers based on thermodynamics and kinetics, the sequencing results from our method demonstrated high amplification uniformity, enabling stable read output across all eight libraries at the current sequencing depth.
[0052] This demonstrates that the present invention utilizes a nucleic acid hybridization algorithm to precisely control the efficiency and specificity of primer combination amplification. Compared with randomly designed primers, it more fully considers the thermodynamic and kinetic effects of the primers themselves or with each other, thereby achieving uniform amplification of the target template.
[0053] This technology, which automates the design of targeted detection primer pairs and amplification inhibitor probes on the genome, aims to enrich and amplify interval mutations by leveraging the thermodynamic differences between primers and probes. Based on the primer pair design, the amplification inhibitor probe is a possible nucleic acid sequence near the mutation site, ranging in length from 15-25bp, with a perfect match to the wild-type sequence and one or more mismatches with the mutant sequence. When no mismatches are present, the amplification inhibitor probe has a higher thermodynamic stability than the primer. When mismatches occur, the thermodynamic stability is slightly lower than that of the primer, leading to competitive strand displacement between the probe and primer.
[0054] Based on the above principles, we introduced a targeted enrichment probe design algorithm.
[0055] like Figure 3 As shown in the figure, the targeted enrichment probe design process is shown: the process on the left starts with the positioning of the target region, a single targeted probe is designed, all candidate targeted probes are generated and sorted from low to high by score, and finally a list containing FP, RP, and probes is output. The process on the right is refined. First, a sequence is generated by a random sequence generator, and the score is combined with the GC ratio and repetitive bases (Rep.bases); after screening by the basic quality controller, low-scoring sequences are excluded and high-scoring sequences are retained. Subsequently, a probe is created by a random probe generator, the "inner badness" is evaluated, and the candidate targeted probes are sorted (from low to high), and finally the screening output of the candidate targeted probes is completed.
[0056] First, based on the generated primer pairs, a candidate pool of probe combinations that meet the free energy difference threshold is generated. Next, the advantages and disadvantages of the candidate probe and primer combinations are evaluated: Objective_Function=Specificity s -λ·|ΔGprimer-ΔGprobe|
[0057] Among them, Specificity s is the specificity score, ΔGprimer-ΔGprobe is the change in hybridization free energy caused by the primer and probe, and λ is the adjustment parameter used to balance amplification specificity and thermodynamic stability. The objective function is used for fast calculation.
[0058] At each selection step, the candidate combination with the highest objective function value is selected as the current set. The candidate pool is updated, and the selected combinations are eliminated. This process is repeated until the candidate pool is empty or the preset number of iterations is reached. Finally, the primer pair and probe combination with the highest objective function value is output.
[0059] Based on the specificity algorithm of the intelligent nucleic acid hybridization algorithm, the objective function is used to evaluate the effect of primers on probes, and the specificity and thermodynamic stability are uniformly quantified. The balance between the inhibition of amplification effect and the amplification efficiency of different probe designs is further given, and the parameter λ is used to achieve control of the detection effect.
[0060] Furthermore, in order to achieve the combined design of multiple inhibition amplification probes and primers, the present invention also uses an iterative algorithm to provide the current recursive optimal design in the given multiple targets, thereby maximizing the primer amplification efficiency, inhibition amplification effect and specificity.
[0061] The present invention relates to a scalable method for designing nucleic acid probes and primers. By introducing functional sequences and chemical modifications into primers and probes, their applications in nucleic acid storage, error analysis, error correction coding, and high-precision detection are expanded.
[0062] It should be understood that while the addition of functional sequences and chemical modifications is a common practice, the addition of modified sequences during primer or probe design can affect the hybridization efficiency and specificity between nucleic acid chains, thereby affecting the detection performance of the library. For example, the ΔΔG and λ values used to evaluate the objective function vary depending on the chemical groups modified at the tail of the inhibitory amplification probe. Furthermore, the introduction of different functional sequences or tag sequences can optimize the amplification versatility of the currently defined region, prioritizing the evaluation of amplification efficiency and specificity of functional sequences or tag sequences in subsequent design calculations.
[0063] For example, we can add adapters, molecular tags, or error-correcting sequences to primers or probes to simplify the subsequent library construction process, adapt the library construction and sequencing platform, and locate amplification errors. Furthermore, to further enhance the enrichment of target molecules by amplification suppression probes or primers, we can modify the central region of the primer or probe with multiple dU (deoxyuracil) bases. By binding to uracil-DNA glycosylase (UDG), dU bases are specifically cleaved, reducing primer-dimer formation and improving amplification efficiency and specificity. Nucleic acid analogs such as RNA, LNA (locked nucleic acid), PNA (peptide nucleic acid), or XNA (synthetic nucleic acid) can be introduced into primers or probes to alter their binding ability to the template sequence. This improves single-base resolution, enabling high-precision detection of single-base mismatches. Molecular beacons with fluorescent and quenching groups can be modified at the 5' end of primers or probes. By monitoring changes in the fluorescence signal, real-time monitoring of the PCR amplification process is achieved, improving experimental convenience and reliability.
[0064] The algorithm uses a dynamic modular approach, treating functional sequences and chemical modifications as independently combinable modules. Integrating thermodynamic and kinetic models with specificity assessment ensures efficient and accurate design. Modules can be flexibly selected and combined to achieve customized designs based on specific needs.
[0065] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the present invention. Various modifications are possible. Any simple, equivalent changes and modifications made in accordance with the claims and description of the present invention are within the scope of protection of the patent claims. Anything not fully described in this invention is conventional technology.
Claims
1. A universal nucleic acid probe design method based on hybridization thermodynamics and kinetics, characterized in that: The following steps are involved: 1) Determine the target amplification region based on the fragment interval and generate primer pairs that meet the length requirements and cover the target fragment. Use an intelligent nucleic acid hybridization algorithm to calculate the change in Gibbs free energy of nucleic acid hybridization. Combined with chemical thermodynamics, kinetic parameters, sequence complementarity, and the characteristics of the enzymatic reaction, generate the optimal primer pair under multidimensional constraints. 2) Based on the generated optimal primer pair, a probe combination that meets the free energy difference threshold is generated as a candidate pool. The probe combination with the highest objective function value is screened from the candidate pool through an iterative process, and the candidate pool is gradually updated until the primer pair and probe combination with the highest objective function value is output.
2. The universal nucleic acid probe design method according to claim 1, wherein: In step 1), the intelligent nucleic acid hybridization algorithm is based on the following formula: ΔG°=ΔH°-TΔS°ΔG°=ΔH°-TΔS°; Where ΔH° is the enthalpy change, ΔS° is the entropy change, and T is the absolute temperature. The values of ΔH° and ΔS° are calculated using the Nearest-Neighbor Model, which allows the calculated results of ΔG° to be accurately correlated with hybridization stability.
3. The universal nucleic acid probe design method according to claim 1, wherein: In step 1), the multidimensional constraints include: The primer length is between 18 and 30 bases to ensure hybridization efficiency and specificity; The GC content was controlled between 40% and 60% to balance hybridization stability and melting temperature; In the case of homopolymers, it is necessary to evaluate whether the homopolymer length does not exceed 4; When a primer itself has complementary base pairs and forms a primer dimer, the number of complementary bases with itself or another primer at the 3' end position should not exceed 4, and at the middle position should not exceed 6.
4. The universal nucleic acid probe design method according to claim 3, characterized in that: In step 1), the multidimensional constraint conditions also include: The secondary structure of the primers themselves is verified using NUPACK. Nucleic acid sequences with obvious secondary structure at the reaction temperature are discarded. Approved primers are further screened for nucleic acid sequence specificity in the corresponding target gene library based on design requirements. This assesses the specific binding of the primer or probe to the target sequence and identifies potential cross-reactions with non-target sequences. Highly specific primers and probes are screened by accurately calculating alignment scores and mismatch penalties.
5. The universal nucleic acid probe design method according to claim 4, characterized in that: In step 2), the multidimensional constraint conditions further include: The Smith-Waterman algorithm was used for local alignment to calculate the alignment scores of the query sequence with the target sequence and non-target sequences. The score for a matched base pair was +2, the penalty for a mismatched base pair was -3, the penalty for introducing a gap was -5, and the penalty for extending a gap was -2. The final score was obtained and the specificity score was defined as: Set Specificity according to the amplification specificity requirements s Threshold: If the score of the query sequence and the non-target sequence exceeds the set threshold (Score non-target >0.8×Score target ), then cross-reaction is considered to exist.
6. The universal nucleic acid probe design method according to claim 5, characterized in that: In step 2), the objective function is introduced: Objective_Function=Specificity s -λ·∣ΔGprimer-ΔGprobe∣; Among them, Specificity s is the specificity score, ΔGprimer-ΔGprobe is the change in hybridization free energy caused by primer and probe, and λ is an adjustment parameter used to balance amplification specificity and thermodynamic stability, which is quickly calculated using the objective function.
7. The universal nucleic acid probe design method according to claim 6, characterized in that: In step 2), in each selection step, the candidate combination with the highest objective function value is selected as the current set, the candidate pool is updated, and the selected combinations are excluded. The above process is repeated until the candidate pool is empty or the preset number of iterations is reached. Finally, the primer pair and probe combination with the highest objective function value is output.
8. The universal nucleic acid probe design method according to claim 1, wherein: The universal nucleic acid probe design method also includes: introducing functional sequences and chemical modifications into primers and probes to expand their applications in nucleic acid storage, error analysis, error correction coding, and high-precision detection.
9. The method according to claim 8, characterized in that The functional sequences include: adapter sequence, stem-loop structure sequence, molecular tag sequence, redundant error correction sequence, and tag sequence; the chemical modifications include: RNA, LNA, PNA, XNA, dU, Spacer, PEG, fluorescent group, phosphorylation group, reverse dT, and methylated base.
10. The universal nucleic acid probe design method according to claim 8, characterized in that: A dynamic modular approach is adopted, in which functional sequences and chemical modifications are treated as independently combinable modules, combined with thermodynamic models, kinetic models and specificity assessment to ensure the efficiency and accuracy of the design.
Citation Information
Patent Citations
Super-multiplex primer design method
CN112687337A
Multi-PCR amplification optimization method, system and equipment based on machine learning and medium
CN116092585A
Primer design method of multiple PCR (Polymerase Chain Reaction) targeted sequencing technology
CN119296644A
Ultrahigh-sensitivity multi-sampling evaluation and optimization method based on primer depolymerization algorithm
CN119479773A
Cited By
Gene sequencing and molecular detection method based on multiple fluorescence labeling
CN120758608A
Screening method of multivalent binding probe for bacterial detection
CN122266456A
Screening method for multivalent binding probes for bacterial detection
CN122266456B