A method and device for determining a primer set
By constructing a primer set determination method in multiple PCR amplification, using dimer scoring matrix, Gibbs free energy matrix and corrected free energy matrix to screen primers, the false positive problem caused by primer dimers was solved, and the accuracy of the detection and experimental success rate were improved.
Patent Information
- Application Number
- CN202510421754.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In prior art In multiple PCR amplification, the probability of primer dimers leading to false positive results is high, affecting detection accuracy and diagnostic accuracy.
By designing the primer set determination method, multi-dimensional evaluation and screening is performed using the dimer score matrix, Gibbs free energy matrix and corrected free energy matrix to exclude the combination of primers that are prone to form dimers and reduce the false positive rate.
It effectively reduces the formation of primer dimers, significantly reduces the probability of false positive results, and improves the accuracy of detection and experimental success rate.
Smart Images

Figure CN119920326B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of biomedical technologies, and particularly to a method and device for determining a primer set. Background Art
[0002] Multiplex PCR amplification is a combination of ultra-multiplex PCR amplification and high-throughput sequencing technologies, which can detect dozens to hundreds of known pathogenic microorganisms and their virulence or drug resistance genes in a test sample. When detecting low-concentration pathogenic microorganisms, especially when detecting their virulence or drug resistance genes, this technology has advantages such as a clear pathogenic spectrum, high detection accuracy, and low sequencing cost compared to pathogen metagenomic sequencing (mNGS). However, how to design high-quality and highly specific primers is the key point of this technology, which will have a very important impact on both the operation of wet experiments and the analysis of conventional next-generation sequencing sequences. Good primers can ensure the efficiency of experiments, significantly improve the success rate of experiments, and reduce the cost of repeated trial and error.
[0003] Currently, there are many methods for designing and simulating highly multiplex PCR primers. Among them, the wet experiment method has solved the primer design problem faced by targeted high-throughput sequencing to a certain extent, but it cannot effectively remove a large number of potential primer dimers. Primer dimers or non-specific binding may lead to false positive results, resulting in a relatively high probability of false positives and a relatively high error rate in diagnosis. Summary of the Invention
[0004] This application provides a method for determining a primer set to reduce the occurrence of false positive results caused by primer dimers, thereby reducing the probability of false positive rates. This application also provides a device for determining a primer set.
[0005] In a first aspect, this application provides a method for determining a primer set, including:
[0006] Obtain a target template sequence;
[0007] Design an initial primer set according to the target template sequence;
[0008] Determine the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and use each dimer score to compare with a score threshold to construct a dimer score matrix, and use each Gibbs free energy to compare with a Gibbs free energy threshold to construct a Gibbs free energy matrix. The larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer;
[0009] Calibrate each of the Gibbs free energies using a loss function to obtain calibrated free energies, and construct a calibrated free energy matrix by comparing each of the calibrated free energies with a calibrated free energy threshold. The calibrated free energy incorporates the influence of non-canonical base pairing, and a larger value indicates a higher probability of dimer formation;
[0010] Determine a target primer set using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the calibrated free energy matrix.
[0011] Optionally, the constructing a dimer score matrix by comparing each of the dimer scores with a score threshold includes:
[0012] Compare each of the dimer scores with a score threshold, mark the initial primer pairs with dimer scores greater than the score threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the score threshold as 0;
[0013] Construct a dimer score matrix based on the markings.
[0014] Optionally, the constructing a Gibbs free energy matrix by comparing each of the Gibbs free energies with a Gibbs free energy threshold includes:
[0015] Compare each of the Gibbs free energies with a Gibbs free energy threshold, mark the initial primer pairs with Gibbs free energies greater than the Gibbs free energy threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the Gibbs free energy threshold as 0;
[0016] Construct a Gibbs free energy matrix based on the markings.
[0017] Optionally, the calibrating each of the Gibbs free energies using a loss function to obtain a calibrated free energy includes:
[0018] Calibrate each of the Gibbs free energies using a loss function and the distance from the matching subsequence of each pair of primers to the 3'-end of the primer to obtain a calibrated free energy.
[0019] Optionally, the comparing the loss function value corresponding to each of the calibrated free energies with a calibrated free energy threshold includes:
[0020] Compare the loss function value corresponding to each of the calibrated free energies with a calibrated free energy threshold, mark the initial primer pairs with calibrated free energies greater than the calibrated free energy threshold as 1, and mark the initial primer pairs with calibrated free energies less than or equal to the calibrated free energy threshold as 0;
[0021] Construct a calibrated free energy matrix based on the markings.
[0022] Optionally, each pair of primers includes a P1 primer and a P2 primer, and the method further includes:
[0023] In the case where the P1 primer and the P2 primer are homologous primers, if the subsequence of the single-read length of the product generated only by the P1 primer is used to verify the specificity of the P1 primer and the verification is passed, the primer pair is marked as 0; if the subsequence does not exist or the specificity verification fails, the primer pair is marked as 1.
[0024] In the case where the P1 primer and the P2 primer are different primer combinations, if both the P1 primer and the P2 primer have single-read length subsequences, the primer pair is marked as 0; if one of the P1 primer and the P2 primer does not have a single-read length subsequence or the specificity verification fails, the primer pair is marked as 1.
[0025] Construct a specificity matrix according to the marks.
[0026] Determining the target primer set by using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix includes:
[0027] Determine the target primer set by using the intersection of the primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix.
[0028] Optionally, determining the target primer set by using the intersection of the primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix includes:
[0029] Screen the primer pairs marked as 1 according to the dimer score matrix to obtain the candidate primer set S1 that meets the score matrix;
[0030] Screen the primer pairs marked as 1 according to the Gibbs free energy matrix to obtain the candidate primer set S2 of the Gibbs free energy matrix;
[0031] Screen the primer pairs marked as 1 according to the corrected free energy matrix to obtain the candidate primer set S3 that meets the corrected free energy matrix;
[0032] Extract the intersection of the primer pairs from S1, S2, and S3 to obtain the co-existing candidate primer set of the dimer;
[0033] Screen the primer pairs marked as 0 through specificity verification according to the specificity matrix to obtain the candidate primer set S4 with high specificity;
[0034] Determine the target primer set according to the intersection of the primer pairs extracted from the co-existing candidate primer set of the dimer and the high-specificity candidate primer set S4.
[0035] In a second aspect, the present application also provides a primer set determination device, which includes:
[0036] An acquisition unit for acquiring a target template sequence;
[0037] A design unit for designing an initial primer set according to the target template sequence;
[0038] A first determination unit for determining the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and constructing a dimer score matrix by comparing each dimer score with a score threshold, and constructing a Gibbs free energy matrix by comparing each Gibbs free energy with a Gibbs free energy threshold. The larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer;
[0039] A comparison unit for correcting each Gibbs free energy using a loss function to obtain a corrected free energy, and constructing a corrected free energy matrix by comparing each corrected free energy with a corrected free energy threshold. The corrected free energy incorporates the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer;
[0040] A second determination unit for determining a target primer set using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix.
[0041] Optionally, the first determination unit is specifically configured to compare each dimer score with a score threshold, mark the initial primer pairs with a dimer score greater than the score threshold as 1, and mark the initial primer pairs with a dimer score less than or equal to the score threshold as 0;
[0042] Construct a dimer score matrix according to the markings.
[0043] Optionally, the first determination unit is specifically configured to compare each Gibbs free energy with a Gibbs free energy threshold, mark the initial primer pairs with a Gibbs free energy greater than the Gibbs free energy threshold as 1, and mark the initial primer pairs with a dimer score less than or equal to the Gibbs free energy threshold as 0;
[0044] Construct a Gibbs free energy matrix according to the markings.
[0045] Optionally, the comparison unit is specifically configured to correct each Gibbs free energy using a loss function and the distance from the matching subsequence of each pair of primers to the 3' end of the primer to obtain a corrected free energy.
[0046] Optionally, the comparison unit is specifically configured to compare each of the corrected free energies with a corrected free energy threshold, mark the initial primer pairs with a corrected free energy greater than the corrected free energy threshold as 1, and mark the initial primer pairs with a corrected free energy less than or equal to the corrected free energy threshold as 0;
[0047] Construct a corrected free energy matrix according to the marking.
[0048] Optionally, the apparatus further includes a marking unit, configured to, when the P1 primer and the P2 primer are homologous primers, if the subsequence of the single-read length of the product generated only by the P1 primer is used to verify the specificity of the P1 primer and the verification is passed, mark the primer pair as 0, and if the subsequence does not exist or the specificity verification fails, mark the primer pair as 1;
[0049] When the P1 primer and the P2 primer are different primer combinations, if both the P1 primer and the P2 primer have single-read length subsequences, mark the primer pair as 0, and if one of the P1 primer and the P2 primer does not have a single-read length subsequence or the specificity verification fails, mark the primer pair as 1;
[0050] Construct a specificity matrix according to the marking.
[0051] The second determination unit is specifically configured to determine a target primer set by using the intersection of the primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix.
[0052] Optionally, the second determination unit is specifically configured to screen the primer pairs marked as 1 according to the dimer score matrix filtering result to obtain a candidate primer set S1 that satisfies the score matrix;
[0053] Screen the primer pairs marked as 1 according to the Gibbs free energy matrix filtering result to obtain a candidate primer set S2 of the Gibbs free energy matrix;
[0054] Screen the primer pairs marked as 1 according to the corrected free energy matrix filtering result to obtain a candidate primer set S3 that satisfies the corrected free energy matrix;
[0055] Extract the intersection of the primer pairs from S1, S2, and S3 to obtain a coexistence candidate primer set of dimers;
[0056] Screen the primer pairs marked as 0 through specificity verification according to the specificity matrix to obtain a candidate primer set S4 with high specificity;
[0057] Determine the target primer set according to the intersection of the primer pairs extracted from the coexistence candidate primer set of dimers and the high-specificity candidate primer set S4.
[0058] In a third aspect, an embodiment of the present application provides a device, which includes a memory and a processor. The memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the method described in the foregoing first aspect.
[0059] In a fourth aspect, an embodiment of the present application provides a computer storage medium, in which codes are stored. When the codes are run, the device running the codes implements the method described in the foregoing first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] To more clearly illustrate the technical solutions in the embodiments or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a flowchart of a method for determining a primer set provided by an embodiment of the present application;
[0062] Figure 2 It is a structural block diagram of a method for determining a primer set provided by an embodiment of the present application;
[0063] Figure 3 It is a schematic structural diagram of a specific implementation manner of a primer set determination device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present application.
[0065] It should be noted that a method and a device for determining a primer set provided by the present application are used in the field of biomedical technology. The above is only an example and does not limit the application fields of the names of the method and the device provided by the present application.
[0066] Multiplex PCR amplification is a combination of ultra-multiplex PCR amplification and high-throughput sequencing technologies, which can detect dozens to hundreds of known pathogenic microorganisms and their virulence or drug resistance genes in a test sample. When detecting low-concentration pathogenic microorganisms, especially when detecting their virulence or drug resistance genes, this technology has the advantages of a clear pathogen spectrum, high detection accuracy, and low sequencing cost compared with pathogen metagenomic sequencing (mNGS). However, how to design high-quality and highly specific primers is the key point of this technology, which will have a very important impact on both the operation of wet experiments and the analysis of conventional next-generation sequencing sequences. Good primers can ensure the efficiency of the experiment, significantly improve the success rate of the experiment, and reduce the cost of repeated trial and error.
[0067] Currently, there are many methods for designing and simulating highly multiplex PCR primers. Among them, the wet experiment method has solved the primer design problem faced by targeted high-throughput sequencing to a certain extent, but it cannot effectively remove a large number of potential primer dimers. Primer dimers or non-specific binding may lead to false positive results, resulting in a relatively high probability of false positives and a relatively high error rate in diagnosis.
[0068] In view of this, an initial primer set determination method is proposed in an embodiment of the present application, including: obtaining a target template sequence, designing an initial primer set according to the target template sequence, determining the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, constructing a dimer score matrix by comparing each dimer score with a score threshold, and constructing a Gibbs free energy matrix by comparing each Gibbs free energy with a Gibbs free energy threshold. Among them, the larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer. Using a loss function to correct each Gibbs free energy to obtain a corrected free energy, constructing a corrected free energy matrix by comparing each corrected free energy with a corrected free energy threshold. The corrected free energy combines the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer. Using the intersection of the initial primer pairs in the dimer score matrix, Gibbs free energy matrix, and corrected free energy matrix to determine the target primer set. Through multi-dimensional evaluation and screening, the present application effectively reduces the formation of primer dimers, thereby reducing the false positive rate. First, by calculating the dimer score of each pair of primers and constructing a dimer score matrix, primer pairs with scores lower than the threshold are screened out, directly excluding those primer combinations that are prone to forming dimers. Secondly, further screening is carried out using the Gibbs free energy matrix. The larger the absolute value of the Gibbs free energy, the more stable the dimer. By comparing with the threshold, primer pairs with too large absolute values of free energy are excluded. Finally, the Gibbs free energy is corrected by a loss function to obtain a corrected free energy matrix. The corrected free energy combines the influence of non-canonical base pairing, further accurately evaluates the probability of dimer formation, and screens out better primer pairs again. Through the intersection of the dimer score matrix, Gibbs free energy matrix, and corrected free energy matrix, the finally determined target primer set has been strictly screened in multiple rounds, minimizing the probability of forming primer dimers to the greatest extent, thereby effectively reducing the occurrence of false positive results caused by primer dimers, and further reducing the probability of false positive rate.
[0069] The method provided in the embodiment of the present application can be executed by software on a terminal device. The terminal device can be, for example, a computer and other devices.
[0070] In order to enable those skilled in the art to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments. The following takes the method provided in the embodiment of the present application being executed by a computing device as an example for description. This embodiment can be called Embodiment 1. Figure 1 The flowchart provided in Embodiment 1. As Figure 1 shown, the method includes:
[0071] S101. Obtain a target template sequence.
[0072] A computing device can obtain a target template sequence, which can also be called a target nucleic acid sequence, that is, a DNA or cDNA sequence to be amplified or detected.
[0073] Specifically, the computing device can collect specific gene sequences from literature or use "species-specific" genomic regions as templates.
[0074] Exemplarily, for example, in the embodiments of the present application, 32 target template sequences can be collected as an example. At the same time, the relevant information of the target template sequences includes: the corresponding sequence name (Template_Sequences_ID), species name (Species_Name), custom species abbreviation name (Abbreviation), species taxonomic identifier (Taxid), and the gene name corresponding to the template sequence (Gene). These target template sequences and relevant information are stored as an Excel table. For example, the relevant information and format of the first 10 template sequences among 32 target template sequences can be shown in Tables 1-1 to 1-2:
[0075] Table 1-1
[0076]
[0077] Table 1-2
[0078]
[0079] S102. Design an initial primer set according to the target template sequence.
[0080] The computing device can design an initial primer set according to the target template sequence.
[0081] Specifically, the computing device can use 32 target template sequences as input and design an initial primer set for each templating. The parameters in the primer design step can be: the temperature can be 55 - 65°C, the primer size can be 18–30bp, the GC content (the proportion of guanine (G) and cytosine (C) base pairs in the nucleic acid sequence) can be 45 - 55%, and the size range of the amplification product can be 130bp to 230bp.
[0082] In addition, the thermodynamic parameters are preset. For example, the maximum self-complementarity can be set to 8.0 kcal / mol and the maximum 3'-end self-complementarity can be set to 3.0 kcal / mol.
[0083] For each target template sequence, a primer set that outputs 10 to 12 candidates can be designed. And use the information in Table 1 to define primer names for each primer and supplement relevant information, which may include TaxID (species classification identifier), Sciname (species scientific name), PrimerIndex (primer name), ForwardID (Forward primer ID), ForwardSeq (Forward primer sequence), ReverseID (Reverse primer ID), and ReverseSeq (Reverse primer sequence), etc. Finally, 280 primer pairs for 17 species can be obtained (taking 32 target template sequences as an example here). The first 20 primers output are shown in Tables 2-1 to 2-3, where each row can represent a pair of primers. (Only the results of the first two genes are selected for display):
[0084] Table 2-1
[0085]
[0086] Table 2-2
[0087]
[0088] Table 2-3
[0089]
[0090] S103. Determine the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and use each of the dimer scores to compare with a score threshold to construct a dimer score matrix, and use each of the Gibbs free energies to compare with a Gibbs free energy threshold to construct a Gibbs free energy matrix.
[0091] After designing the initial primer set in step S102, as Figure 2 shown, the computing device can determine the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and use each dimer score to compare with a score threshold to construct a dimer score matrix, and use each of the Gibbs free energies to compare with a Gibbs free energy threshold to construct a Gibbs free energy matrix.
[0092] Specifically, the higher the dimer score of the initial primer, the more likely it is to form a dimer. The obtained dimer scores are compared with a score threshold preset through experiments. Then, the initial primer pairs with scores greater than the score threshold can be marked as 1, and the initial primer pairs with scores less than or equal to the score threshold can be marked as 0. The relationship between the dimer scores of each initial primer pair in the initial primer set and the score threshold is compared in turn, so as to construct a dimer score matrix based on the marks. Then, the initial primer pairs marked as 1 in the screening and filtering results can be selected based on the dimer score matrix to obtain the candidate primer set S1 that meets the dimer score matrix.
[0093] The more negative the Gibbs free energy value (the higher the absolute value), the more stable the dimer can be characterized, and the more likely the dimer is to form. Similarly, the Gibbs free energy obtained for each initial primer pair can be compared with the Gibbs free energy threshold preset through experiments. The initial primer pairs with Gibbs free energy greater than the Gibbs free energy threshold are marked as 1, and the initial primer pairs with Gibbs free energy less than or equal to the Gibbs free energy threshold are marked as 0. The relationship between the Gibbs free energy of all initial primers in the initial primer set and the Gibbs free energy threshold is compared in turn, and then a Gibbs free energy matrix can be obtained based on the marks. Then, the initial primer pairs marked as 1 in the screening and filtering results can be selected according to the Gibbs free energy matrix to obtain the candidate primer set S2 of the Gibbs free energy matrix.
[0094] S104. Use a loss function to correct each of the Gibbs free energies to obtain corrected free energies, and construct a corrected free energy matrix by comparing each of the corrected free energies with a corrected free energy threshold.
[0095] The computing device can use a loss function to correct each Gibbs free energy to obtain a corrected free energy, and construct a corrected free energy matrix by comparing each corrected free energy with a corrected free energy threshold.
[0096] The purpose of using a loss function to correct the Gibbs free energy is to more accurately evaluate the tendency and stability of primer dimer formation. Although the Gibbs free energy is an important indicator for measuring the stability of primer dimers, relying solely on its original value may not fully reflect the actual dimer formation situation, because the actual primer dimer formation is also affected by factors such as non-canonical base pairing and local sequence structure. By introducing a loss function, the influence of these complex factors can be taken into account, the Gibbs free energy can be adjusted and corrected, so as to obtain a corrected free energy closer to the actual situation. The corrected free energy can more accurately reflect the probability and stability of primer dimer formation. Furthermore, by comparing with the corrected free energy threshold, the primer pairs that are not likely to form dimers can be more precisely screened out, thereby effectively reducing the occurrence probability of non-specific amplification and false positive results caused by primer dimers.
[0097] Specifically, taking a pair of initial primers including primer P1 and primer P2 as an example, the loss function is a function for calculating the formation of primer dimers, which is calculated based on the Gibbs free energy between pairs of primers and the distances between the subsequences matched by each primer and the 3'-end of the primer. The calculation formula is as follows:
[0098]
[0099] Among them, d1 and d2 respectively represent the distances from the matched subsequences to the 3'-ends of primer p1 and primer p2, with the unit of base number. The absolute value of the Gibbs free energy (i.e., the numerator of the formula) is used for calculation in the formula. Since different matched subsequences will result in different combinations of d1 and d2, different loss functions will be calculated. Finally, the minimum value among these loss functions is taken as the corrected free energy, that is, correction free energy = min(loss function), where "min" is a function used to find the minimum value from a set of numerical values.
[0100] The level of the corrected free energy is closely related to the possibility of primer dimer formation. The higher the corrected free energy, the greater the probability of primer dimer formation; conversely, the lower the corrected free energy, the smaller the probability of primer dimer formation. Therefore, when the loss function obtains the minimum value, it can be used to determine whether primer P1 and primer P2 form a real primer pair.
[0101] It should be noted that for each pair of primers containing matched subsequences, the larger the loss function, the more likely a dimer will form. The corrected free energy is compared with the loss function threshold preset for the experiment. The initial primer pairs with a corrected free energy greater than the corrected free threshold are marked as 1, and the initial primer pairs with a corrected free energy less than or equal to the corrected free threshold are marked as 0. The relationship between the corrected free energy of all initial primer pairs in the initial primer set and the corrected free threshold is compared in turn to obtain the corrected free energy matrix. Finally, the primer combinations with the filtering result marked as 1 in the corrected free energy matrix can be obtained, and the candidate primer set S3 that satisfies the corrected free energy matrix can be obtained.
[0102] S105. Determine the target primer set by using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix.
[0103] The computing device can determine the target primer set by using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix.
[0104] Specifically, the computing device can screen for compatible primer combinations based on the three matrices constructed in the above steps. Each matrix can be an n×n 01 symmetric matrix, where n is the total number of initial primer pairs generated in the primer design step. A primer pair with a value of 1 indicates that the dimer binding is more stable, the probability of the primers binding to each other in PCR is higher, and the risk of decreased amplification efficiency or failure is higher, that is, these two primer sets are incompatible. A primer pair with a value of 0 indicates that the dimer binding is less stable, the probability of the primers binding to each other in PCR is lower, and the possibility of higher amplification efficiency or success is higher, that is, this primer pair can coexist. Therefore, the primer combinations marked as 1 in the screening filter results based on the score matrix are used to obtain the candidate primer set S1 that satisfies the dimer score matrix; the primer combinations marked as 1 in the screening filter results based on the Gibbs free energy matrix are used to obtain the candidate primer set S2 that satisfies the Gibbs free energy matrix; the primer combinations marked as 1 in the screening filter results based on the corrected free energy matrix are used to obtain the candidate primer set S3 that satisfies the corrected free energy matrix. The intersection of the three matrices can be extracted from the data sets `S1`, `S2`, and `S3`. That is, the compatibility between primer pairs is determined by the above matrices in turn. When there are primer pairs that are compatible among all matrices, they are considered to be able to coexist. Finally, a primer set with a lower dimer is obtained, and this primer set is determined as the target primer set.
[0105] Embodiments of the present application can obtain a target template sequence, design an initial primer set according to the target template sequence, determine the dimer scores and Gibbs free energies of each pair of initial primers in the initial primer set, construct a dimer score matrix by comparing each dimer score with a score threshold, and construct a Gibbs free energy matrix by comparing each Gibbs free energy with a Gibbs free energy threshold. Among them, the larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer. The loss function is used to correct each of the Gibbs free energies to obtain the corrected free energy. A corrected free energy matrix is constructed by comparing each of the corrected free energies with a corrected free energy threshold. The corrected free energy combines the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer. The intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix is used to determine the target primer set. Through multi-dimensional evaluation and screening, the present application effectively reduces the formation of primer dimers, thereby reducing the false positive rate. First, by calculating the dimer scores of each pair of primers and constructing a dimer score matrix, primer pairs with scores lower than the threshold are screened out, directly excluding those primer combinations that are prone to forming dimers. Secondly, the Gibbs free energy matrix is used for further screening. The larger the absolute value of the Gibbs free energy, the more stable the dimer. By comparing with the threshold, primer pairs with too large absolute values of free energy are excluded. Finally, the Gibbs free energy is corrected by the loss function to obtain a corrected free energy matrix. The corrected free energy combines the influence of non-canonical base pairing and further accurately evaluates the probability of dimer formation, and again screens out better primer pairs. Through the intersection of the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix, the finally determined target primer set, after multiple rounds of strict screening, minimizes the probability of forming primer dimers, thereby effectively reducing the occurrence of false positive results caused by primer dimers, and further reducing the probability of the appearance of false positive rates.
[0106] In addition, the present application also provides an embodiment, which can be called Embodiment 2 here. In Embodiment 2, the initial primer set can also be specifically evaluated and a specificity matrix can be established to ensure that the primers can accurately and specifically bind to the target sequence in the PCR reaction, avoid non-specific amplification, and thus improve the accuracy and reliability of the experiment.
[0107] Specifically, in this embodiment, the primer pairs in the initial primer set can be compared with the complete nucleic acid database (NT) to evaluate the specificity of the primer pairs and construct a specificity matrix. The specificity evaluation is based on the following two aspects. One is the "single read length" specificity of base extension, and the other is the "species-level classification consistency" specificity of the PCR product. For each product generated by the primer combination of p1 and p2 (taking an initial primer pair including primer P1 and primer P2 as an example), subsequences with two terminal single read lengths (for example, taking 50bp as an example here) will be used for specificity evaluation. If these subsequences can be identified as a single species, it indicates that the specificity evaluation is passed. That is to say, a pair of effective specific primers can generate one or more different 50bp fragments at the end of the PCR amplification product, and at least one set of 50bp sequences can be recognized as a unique microbial species (species level). When p1 is the same as p2, the initial primer pair can be marked as 0, indicating that the subsequences with the single read length (default 50bp) of the product generated only by primer set p1 are used to verify the specificity of p1 and the verification is passed, and marked as 1 indicating that such effective subsequences are missing in P1 and P2. In the case where p1 and p2 are different, if both primer pairs p1 and p2 can maintain effective single read length subsequences, the primer pair is marked as 0, indicating that the subsequences with the single read length (default 50bp) of the products generated by p1 and p2 are used to verify the specificity of p1 and p2 and the verification is passed, and marked as 1, which can indicate that the primer pairs p1 and p2 lack effective single read length subsequences or the specificity verification fails. Thus, the construction of the specificity matrix can be achieved by using the marks.
[0108] Subsequently, high-specificity candidate primer pairs are screened, that is, based on the specificity matrix, the primer combinations marked as 0 through specificity verification are screened to obtain a high-specificity candidate primer set s4.
[0109] Finally, the intersection of the data of the two can be extracted from the primer set with lower dimers (the primer set determined according to S1, S2, and S3 in Embodiment 1) and the high-specificity candidate primer set s4, and finally a high-specificity coexisting primer data set is obtained. Furthermore, the high-specificity coexisting primer data set can be determined as the target primer set.
[0110] The above are some specific implementation manners of the primer set determination method provided by the embodiments of the present application. Based on this, the present application also provides a corresponding device. Next, the device provided by the embodiments of the present application will be introduced from the perspective of functional modularization, and the device corresponds to the primer set determination method described above and can be referred to each other.
[0111] Figure 3 It is the structural block diagram of the primer set determination device provided by the embodiments of the present invention, which is called the third specific implementation manner. Refer to Figure 3 The device may include:
[0112] An acquisition unit 300, configured to acquire a target template sequence;
[0113] A design unit 310, configured to design an initial primer set according to the target template sequence;
[0114] A first determination unit 320, configured to determine the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and use each dimer score to compare with a score threshold to construct a dimer score matrix, and use each Gibbs free energy to compare with a Gibbs free energy threshold to construct a Gibbs free energy matrix. The larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer;
[0115] A comparison unit 330, configured to correct each Gibbs free energy by using a loss function to obtain a corrected free energy, and use each corrected free energy to compare with a corrected free energy threshold to construct a corrected free energy matrix. The corrected free energy combines the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer;
[0116] A second determination unit 340, configured to determine a target primer set by using the intersection of the initial primer pairs in the dimer score matrix, the Gibbs free energy matrix, and the corrected free energy matrix.
[0117] Optionally, the first determination unit is specifically configured to compare each dimer score with a score threshold, mark the initial primer pairs with dimer scores greater than the score threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the score threshold as 0;
[0118] Construct a dimer score matrix according to the markings.
[0119] Optionally, the first determination unit is specifically configured to compare each Gibbs free energy with a Gibbs free energy threshold, mark the initial primer pairs with Gibbs free energies greater than the Gibbs free energy threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the Gibbs free energy threshold as 0;
[0120] Construct a Gibbs free energy matrix according to the markings.
[0121] Optionally, the comparison unit is specifically configured to correct each Gibbs free energy by using a loss function and the distance from the matching subsequence of each pair of primers to the 3' end of the primer to obtain a corrected free energy.
[0122] Optionally, the comparison unit is specifically configured to compare the loss function value corresponding to each of the corrected free energies with a corrected free energy threshold, mark the initial primer pairs with a corrected free energy greater than the corrected free energy threshold as 1, and mark the initial primer pairs with a corrected free energy less than or equal to the corrected free energy threshold as 0;
[0123] Construct a corrected free energy matrix based on the marking.
[0124] Optionally, the apparatus further includes a marking unit, configured to, in the case where the P1 primer and the P2 primer are homologous primers, if the subsequence of the single read length of the product generated only by the P1 primer is used to verify the specificity of the P1 primer and the verification is passed, mark the primer pair as 0, and if the subsequence does not exist or the specificity verification fails, mark the primer pair as 1;
[0125] In the case where the P1 primer and the P2 primer are different primer combinations, if both the P1 primer and the P2 primer have single read length subsequences, mark the primer pair as 0, and if one of the P1 primer and the P2 primer does not have a single read length subsequence or the specificity verification fails, mark the primer pair as 1;
[0126] Construct a specificity matrix based on the marking.
[0127] The second determination unit is specifically configured to determine a target primer set by using the intersection of the primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix.
[0128] Optionally, the second determination unit is specifically configured to screen the primer pairs marked as 1 according to the dimer score matrix filtering result to obtain a candidate primer set S1 that satisfies the score matrix;
[0129] Screen the primer pairs marked as 1 according to the Gibbs free energy matrix filtering result to obtain a candidate primer set S2 of the Gibbs free energy matrix;
[0130] Screen the primer pairs marked as 1 according to the corrected free energy matrix filtering result to obtain a candidate primer set S3 that satisfies the corrected free energy matrix;
[0131] Extract the intersection of the primer pairs from S1, S2, and S3 to obtain a coexistence candidate primer set of dimers;
[0132] Screen out the primer pairs marked as 0 through specificity verification according to the specificity matrix to obtain a candidate primer set S4 with high specificity;
[0133] Determine the target primer set according to the intersection of the primer pairs extracted from the coexistence candidate primer set of dimers and the high specificity candidate primer set S4.
[0134] The embodiments of the present application also provide corresponding devices and computer storage media for implementing the solutions provided by the embodiments of the present application.
[0135] Among them, the device includes a memory and a processor. The memory is used to store instructions or codes, and the processor is used to execute the instructions or codes so that the device executes the method described in any embodiment of the present application.
[0136] The computer storage medium stores codes. When the codes are run, the device running the codes implements the method described in any embodiment of the present application.
[0137] In the embodiments of the present application, the "first", "second" (if any) in the names such as "first" and "second" are only used as name identifiers and do not represent the first and second in order.
[0138] From the description of the above embodiments, those skilled in the art can clearly understand that all or part of the steps in the above embodiment methods can be implemented by means of software plus a general hardware platform. Based on such an understanding, the technical solution of the present application can be embodied in the form of a software product. The computer software product can be stored in a storage medium, such as read-only memory (ROM) / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network communication device such as a router) to execute the methods described in each embodiment or some parts of the embodiments of the present application.
[0139] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and the relevant parts can refer to the partial description of the method embodiments. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0140] The above is only an exemplary embodiment of the present application and is not used to limit the protection scope of the present application.
Claims
1. A method for determining a primer set, characterized in that Including: Obtain a target template sequence; Design an initial primer set according to the target template sequence; Determine the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and use each dimer score to compare with a score threshold to construct a dimer score matrix, and use each Gibbs free energy to compare with a Gibbs free energy threshold to construct a Gibbs free energy matrix. The larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer; Use a loss function to correct each Gibbs free energy to obtain a corrected free energy, and use each corrected free energy to compare with a corrected free energy threshold to construct a corrected free energy matrix. The corrected free energy incorporates the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer; Each pair of primers includes a P1 primer and a P2 primer; In the case where the P1 primer and the P2 primer are homologous primers, if the specificity of the P1 primer is verified and passed by using only the subsequence of the single-read length of the product generated by the P1 primer, the primer pair is marked as 0. If there is no such subsequence or the specificity verification fails, the primer pair is marked as 1; In the case where the P1 primer and the P2 primer are different primer combinations, if both the P1 primer and the P2 primer have single-read length subsequences, the primer pair is marked as 0. If one of the primers in the primer pair does not have a single-read length subsequence or the specificity verification fails, the primer pair is marked as 1; Construct a specificity matrix according to the markings; Use the intersection of the primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix to determine the target primer set.
2. The method according to claim 1, wherein The constructing the dimer score matrix by comparing each dimer score with a score threshold includes: Compare each dimer score with a score threshold, mark the initial primer pairs with dimer scores greater than the score threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the score threshold as 0; Construct a dimer score matrix according to the markings.
3. The method according to claim 1, characterized in that, The constructing the Gibbs free energy matrix by comparing each Gibbs free energy with a Gibbs free energy threshold includes: Compare each Gibbs free energy with a Gibbs free energy threshold, mark the initial primer pairs with Gibbs free energies greater than the Gibbs free energy threshold as 1, and mark the initial primer pairs with dimer scores less than or equal to the Gibbs free energy threshold as 0; Construct a Gibbs free energy matrix according to the markings.
4. The method according to claim 1, wherein The obtaining the corrected free energy by using a loss function to correct each Gibbs free energy includes: Use a loss function and the distance from the matching subsequence of each pair of primers to the 3' end of the primer to correct each Gibbs free energy to obtain a corrected free energy.
5. The method according to claim 4, wherein The constructing the corrected free energy matrix by comparing each corrected free energy with a corrected free energy threshold includes: Compare each of the corrected free energies with a corrected free energy threshold, mark the initial primer pairs with a corrected free energy greater than the corrected free energy threshold as 1, and mark the initial primer pairs with a corrected free energy less than or equal to the corrected free energy threshold as 0; Construct a corrected free energy matrix based on the markings.
6. The method according to claim 5, wherein Determining a target primer set by using the intersection of primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix, includes: Screen the primer pairs marked as 1 according to the dimer score matrix filtering result to obtain a candidate primer set S1 that meets the score matrix; Screen the primer pairs marked as 1 according to the Gibbs free energy matrix filtering result to obtain a candidate primer set S2 of the Gibbs free energy matrix; Screen the primer pairs marked as 1 according to the corrected free energy matrix filtering result to obtain a candidate primer set S3 that meets the corrected free energy matrix; Extract the intersection of primer pairs from S1, S2, and S3 to obtain a coexistence candidate primer set of dimers; Screen the primer pairs marked as 0 through specificity verification according to the specificity matrix to obtain a high-specificity candidate primer set S4; Determine the target primer set according to the intersection of primer pairs extracted from the coexistence candidate primer set of dimers and the high-specificity candidate primer set S4.
7. A primer set determination device, characterized in that, The device includes: An acquisition unit for acquiring a target template sequence; A design unit for designing an initial primer set according to the target template sequence; A first determination unit for determining the dimer score and Gibbs free energy of each pair of initial primers in the initial primer set, and constructing a dimer score matrix by comparing each dimer score with a score threshold, and constructing a Gibbs free energy matrix by comparing each Gibbs free energy with a Gibbs free energy threshold. The larger the dimer score, the higher the probability of forming a dimer, and the larger the absolute value of the Gibbs free energy, the more stable the dimer; A comparison unit for correcting each Gibbs free energy by using a loss function to obtain a corrected free energy, and constructing a corrected free energy matrix by comparing each corrected free energy with a corrected free energy threshold. The corrected free energy combines the influence of non-canonical base pairing, and the larger the value, the higher the probability of forming a dimer. Each pair of primers includes a P1 primer and a P2 primer; A marking unit for, in the case where the P1 primer and the P2 primer are homologous primers, if the specificity of the P1 primer is verified and passed by using the subsequence of the single-read length of the product generated only by the P1 primer, then mark the primer pair as 0, and if there is no such subsequence or the specificity verification fails, then mark the primer pair as 1; In the case where the P1 primer and the P2 primer are different primer combinations, if both the P1 primer and the P2 primer have single-read length subsequences, then mark the primer pair as 0, and if one of the P1 primer and the P2 primer does not have a single-read length subsequence or the specificity verification fails, then mark the primer pair as 1; Construct a specificity matrix according to the markings; A second determination unit, configured to determine a target primer set by using an intersection of primer pairs in the dimer score matrix, the Gibbs free energy matrix, the corrected free energy matrix, and the specificity matrix.
8. The device according to claim 7, characterized in that, The first determination unit is specifically configured to compare each of the dimer scores with a score threshold, mark an initial primer pair with a dimer score greater than the score threshold as 1, and mark an initial primer pair with a dimer score less than or equal to the score threshold as 0; Construct a dimer score matrix according to the markings.
9. The device according to claim 7, characterized in that, The first determination unit is specifically configured to compare each of the Gibbs free energies with a Gibbs free energy threshold, mark an initial primer pair with a Gibbs free energy greater than the Gibbs free energy threshold as 1, and mark an initial primer pair with a dimer score less than or equal to the Gibbs free energy threshold as 0; Construct a Gibbs free energy matrix according to the markings.
Citation Information
Patent Citations
Primer design method and system based on primer template binding capacity and storage medium
CN116665777A
Method and device for predicting primer dimer and storage medium
CN118248215A
Method for the design of oligonucleotides for molecular biology techniques
US20070059743A1