Method, system and storage medium for designing Taqman SNP genotyping probes and primer sets

Through the automated design of Taqman SNP genotyping probes and primer sets, the problems of complexity and high error rate of existing tools are solved, and the rapid and simple acquisition of probe sets targeting two alleles is achieved. It is suitable for ordinary personal computers and is suitable for researchers of different levels of experience.

CN118866099BActive Publication Date: 2025-07-29INT CENT FOR GENETIC ENG & BIOTECHNOLOGY TAIZHOU REGIONAL RES CENT
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411228170.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-03
Publication Date
2025-07-29
Estimated Expiration
2044-09-03

AI Technical Summary

Technical Problem

Existing Taqman SNP genotyping detection design tools such as Primer3 and Edesign are complex, have unfriendly user interfaces, and cannot automatically generate probe sets for two alleles, making the design process time-consuming and error-prone.

Method used

Provide a method and system to automatically design Taqman SNP genotyping probes and primer sets, and automatically generate probe sets for two alleles of target SNPs and simplify the user interface. Users only need to enter relevant information to obtain complete probes and primer sets.

Benefits of technology

It achieves rapid and easy access to Taqman SNP genotyping probes and primer sets, suitable for ordinary personal computers, no high-end hardware required, user-friendly interface, suitable for beginners and experienced researchers, and supports difficult SNP and DNA region designs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118866099B_ABST
    Figure CN118866099B_ABST
Patent Text Reader

Abstract

The present invention relates to a method, system and storage medium for designing Taqman SNP genotyping probes and primer sets, comprising the following steps: (1) A first program, according to the DNA fragment sequence containing the target SNP reference allele and the alternative allele of the target SNP input by the user, to obtain a probe set; the probe set includes a first probe targeting the reference allele of the target SNP and a second probe targeting the alternative allele of the target SNP; (2) A second program, according to the information input by the user, namely the DNA fragment sequence containing the target SNP reference allele, the alternative allele of the target SNP, the sequences and melting temperature values of the selected first probe and second probe, to obtain a primer set. The present invention is tailor-made for the Taqman SNP genotyping detection method, is easy to install and use, has flexibility, and enables users to conveniently and quickly obtain a complete set of required probes and primer sets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of bioinformatics and computer technology, and particularly relates to a method, a system and a storage medium for designing Taqman SNP genotyping probes and primer sets. Background Art

[0002] The Taqman SNP genotyping assay (also known as the 5'-nuclease allele discrimination assay) is a widely used method for genotyping single nucleotide polymorphisms (SNPs). Based on the polymerase chain reaction (PCR), in the presence of two probes, a pair of primers (i.e., a forward primer and a reverse primer) is used to amplify the flanking DNA region of the SNP. Each Taqman probe is specific for a certain allele of the SNP. It has a fluorescent reporter group dye at its 5'-end and a quencher at its 3'-end, which can absorb the fluorescence from the reporter group dye. When the Taqman probe is free in the solution, it does not emit fluorescence. During the PCR process, the Taqman probe binds to the target DNA sequence, and the Taq DNA polymerase will cleave the Taqman probe, releasing the reporter group dye from the quencher, resulting in an increase in fluorescence. Therefore, the fluorescence emitted by each probe is a measure of the specific binding of the probe to its target allele. The two probes (i.e., a probe pair) in the Taqman SNP genotyping assay are labeled with different fluorescent dyes, and two alleles of the target SNP can be detected in one PCR reaction. Therefore, the design of the probes and primer sets is the key to the success of Taqman SNP genotyping.

[0003] Currently, researchers usually use free and open-source software to design probes and primer sets for this detection method, such as Primer3 (https: / / primer3.ut.ee / ) and Edesign (https: / / pubmed.ncbi.nlm.nih.gov / 26863543 / ). However, Primer3 is not specifically designed for Taqman SNP genotyping detection and has the following main drawbacks: (1) When using Primer3 for analysis, users must correctly set many parameters. For this purpose, users must have in-depth knowledge of the internal working principle of Primer3 and the design principle of Taqman SNP probes and primer sets, but many researchers lack this knowledge. (2) The user interface of Primer3 is very complex and often overwhelms new or inexperienced users. (3) Primer3 can only design probes for one allele of the target SNP each time and cannot automatically generate the Taqman probe set required for Taqman SNP genotyping detection (generally including probes targeting two alleles of the target SNP). Therefore, users must manually pair probes for the two alleles of the target SNP, which is both time-consuming and error-prone. Edesign is a computer program derived from Primer3 and can design probes and primers for various PCR-based experimental methods. Although Edesign can be used for Taqman SNP genotyping detection, it is not specifically designed and optimized for this detection method. Thus, Edesign has the same drawbacks as Primer3, requiring users to set many design parameters and having a complex user interface. Edesign can run on a local computer, but installing Edesign requires compiling its source code (written in C language), which is difficult and impossible for most researchers in the fields of biomedicine and life sciences.

[0004] In view of this, the present invention provides a method, a system, and a storage medium for designing Taqman SNP genotyping probes and primer sets. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method, a system, and a storage medium for designing Taqman SNP genotyping probes and primer sets. The purpose is to automatically and efficiently design probe and primer sets, enabling researchers to conveniently and quickly complete Taqman SNP genotyping detection.

[0006] The technical solution of the present invention to solve the above technical problems is as follows:

[0007] In the first aspect, a method for designing Taqman SNP genotyping probes and primer sets includes the following steps:

[0008] (1) First procedure: obtaining a probe set based on a DNA fragment sequence containing a reference allele of a target SNP and an alternative allele sequence of the target SNP input by a user; the probe set includes a probe targeting the reference allele of the target SNP (a first probe) and a probe targeting the alternative allele of the target SNP (a second probe);

[0009] (2) Second procedure: A primer set is obtained based on the sequence of the DNA fragment containing the target SNP reference allele input by the user, as well as the sequence and melting temperature (Tm) value of the first and second probes input by the user (included in the result returned to the user by the first procedure).

[0010] Furthermore, the first procedure includes the following specific steps:

[0011] (1-1) obtaining a DNA fragment sequence containing a target SNP reference allele and an alternative allele of the target SNP; the DNA fragment sequence containing the target SNP reference allele includes the reference allele of the target SNP and flanking regions;

[0012] (1-2) a short fragment containing part of the bases of the flanking region and the reference allele of the target SNP is cut from the DNA fragment sequence containing the target SNP reference allele to obtain a short DNA sequence; the total number of bases C and C and the total number of bases G and G in the short DNA sequence are counted; if the total number of bases G and G is less than or equal to the total number of bases C and C, the reverse complementary sequence is used for subsequent analysis; if the total number of bases G and G is greater than the total number of bases C and C, a reverse complementary sequence of the short DNA sequence is generated and used for subsequent analysis;

[0013] (1-3) setting parameters of the first probe or the second probe; automatically generating a first input file for Primer3 analysis;

[0014] (1-4) performing Primer3 analysis using the first input file to obtain a first probe set targeting the reference allele of the target SNP and its flanking DNA region;

[0015] (1-5) replacing the reference allele of the target SNP in the short DNA sequence with the alternative allele of the target SNP, and then performing steps (1-3) to (1-4) in sequence to obtain a second probe set targeting the alternative allele of the target SNP and its flanking DNA region;

[0016] (1-6) Setting screening conditions, screening the first probe set and the second probe set to obtain a probe group.

[0017] Further, in step (1-1), the reference allele of the target SNP is represented by a capital letter, the flanking regions are represented by lowercase letters, and nucleotides in the DNA fragment sequence containing the reference allele of the target SNP that are not included in the probe and primer set are represented by n.

[0018] Further, the parameters of the first probe or the second probe in step (1-3) are set as follows:

[0019] Set the melting temperature Tm value of the first probe or the second probe to be 66°C to 70°C;

[0020] Set the sequence length of the first probe or the second probe to be 17 bp to 35 bp;

[0021] Set the percentage of GC (including G, C, g, and c) in the sequence of the first probe or the second probe to be 20% to 80%;

[0022] Set the maximum allowable length of mononucleotide repeats in the first probe or the second probe to be 5.

[0023] Further, step (1-6) includes the following specific steps:

[0024] (1-6-1) Set the screening conditions to screen the first probe set and the second probe set to obtain the screened first probe set and the screened second probe set; the set screening conditions are as follows:

[0025] Set the SNPs to be located at the middle one-third of the first probe and the second probe;

[0026] Set that the 5′ ends of the first probe and the second probe are not g;

[0027] Set that the sequences of the first probe and the second probe do not contain more than 3 consecutive g or G. For example, sequences containing consecutive gggG, gggg, gGgg are not satisfied with the requirements of the first probe and the second probe;

[0028] Set that the total number of C and c contained in the sequences of the first probe and the second probe is more than or equal to the total number of G and g;

[0029] (1-6-2) From the screened first probe set and the screened second probe set, find the first probe and the second probe with a melting temperature Tm value difference not exceeding 1°C as the probe set;

[0030] (1-6-3) Calculate the total ranking score of each probe set and sort according to the total ranking score.

[0031] Furthermore, the second procedure includes the following specific steps:

[0032] (2-1) obtaining a DNA fragment sequence containing a reference allele of the target SNP, an alternative allele of the target SNP, sequences of a first probe and a second probe selected by the user, and a melting temperature (Tm) value (contained in an output file returned to the user by the first program); the DNA fragment sequence containing the reference allele of the target SNP includes the reference allele of the target SNP and flanking regions;

[0033] (2-2) The forward primer and the reverse primer in the primer set meet the following parameters:

[0034] The forward primer sequence is set to not overlap with the sequences of the first probe and the second probe, and the reverse primer sequence is set to not overlap with the sequences of the first probe and the second probe;

[0035] The melting temperature Tm value of the forward primer is set to be 7±1.5°C lower than the melting temperature Tm value of the first probe or the melting temperature Tm value of the second probe, and the melting temperature Tm value of the reverse primer is set to be 7±1.5°C lower than the melting temperature Tm value of the first probe or the melting temperature Tm value of the second probe;

[0036] Setting the maximum difference between the melting temperature Tm values of the forward primer and the reverse primer;

[0037] Setting the size of the PCR product of the primer set to 50 bp to 150 bp;

[0038] The forward primer and the reverse primer are both set to have a length of 18 bp to 30 bp;

[0039] Setting the gc percentages in the forward primer and the reverse primer to be 30% to 80%;

[0040] (2-3) Automatically generate a second input file for Primer3 analysis; perform Primer3 analysis using the second input file to obtain a primer pair targeting the DNA fragment sequence containing the target SNP reference allele.

[0041] Furthermore, in step (2-2), the sequence of the forward primer is set to not overlap with the sequences of the first probe and the second probe, and the sequence of the reverse primer is set to not overlap with the sequences of the first probe and the second probe, which specifically includes the following steps:

[0042] Counting the length values of the DNA fragment sequence containing the target SNP reference allele, and sequentially numbering the bases of the DNA fragment sequence containing the target SNP reference allele to obtain a numbered DNA fragment sequence containing the target SNP reference allele;

[0043] Comparing the sequence of the first probe with the sequence of the numbered DNA fragment containing the target SNP reference allele, obtaining the numbers of the first base and the last base of the sequence of the first probe on the numbered DNA fragment containing the target SNP reference allele, and assigning the numbers to the first variable group respectively;

[0044] Replacing the reference allele of the target SNP in the numbered DNA fragment sequence containing the target SNP reference allele with the alternative allele of the target SNP to obtain a numbered DNA fragment sequence containing the alternative allele; comparing the sequence of the second probe with the numbered DNA fragment sequence containing the alternative allele to obtain the numbers of the first base and the last base of the sequence of the second probe in the numbered DNA fragment sequence containing the alternative allele, and assigning the numbers to the second variable group respectively;

[0045] By using the first variable group and the second variable group, the sequence of the forward primer is made not to overlap with the sequences of the first probe and the second probe, and the sequence of the reverse primer is made not to overlap with the sequences of the first probe and the second probe.

[0046] Further, the sequence of the first probe is compared with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele. If the sequence of the first probe is completely matched with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele, the numbers of the first base and the last base of the sequence of the first probe on the numbered DNA fragment sequence containing the target SNP reference allele are obtained, and the numbers are assigned to the first variable group respectively; if the sequence of the first probe is not completely matched with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele, the sequences of the first probe and the second probe are reverse complemented to obtain the numbers of the first base and the last base of the complementary sequence of the first probe on the numbered DNA fragment sequence containing the target SNP reference allele, and the numbers are assigned to the first variable group respectively.

[0047] In a second aspect, a system for designing Taqman SNP genotyping probes and primer sets comprises:

[0048] (1) The first program system obtains a probe set according to the DNA fragment sequence containing the reference allele of the target SNP and the alternative allele sequence of the target SNP input by the user; the probe set includes a probe targeting the reference allele of the target SNP (the first probe) and a probe targeting the alternative allele of the target SNP (the second probe);

[0049] (2) The second program system obtains a primer set according to the DNA fragment sequence containing the reference allele of the target SNP input by the user, as well as the sequences and melting temperature Tm values of the first probe and the second probe input by the user (contained in the result returned by the first program to the user).

[0050] The combined use of the two program systems can design a complete set of probes and primer sets for Taqman genotyping detection for a target SNP.

[0051] In a third aspect, a computer-readable storage medium stores a computer program thereon, and when the computer program is executed by a processor, all steps of the method for designing Taqman SNP genotyping probes and primer sets are implemented.

[0052] The beneficial effects of the present invention are:

[0053] (1) The method of the present invention enables relevant researchers to conveniently and quickly obtain the required Taqman SNP genotyping probes and primer sets.

[0054] (2) Easy to use: The system for designing Taqman SNP genotyping probes and primer sets of the present invention has the following characteristics, so that researchers can conveniently and quickly obtain the required probes and primer sets: (A) The user interface is simple and clear; (B) The user does not need to manually set any design parameters, and only needs to provide information related to the target SNP; (C) This program can run quickly on an ordinary personal computer without connecting to the Internet.

[0055] (3) Easy to install: The system for designing Taqman SNP genotyping probes and primer sets of the present invention only depends on Primer3 and Emacs, and these two free and open-source software (executable binary programs) can be easily installed on a local computer. The method and system of the present invention can run on all major computer operating systems (i.e., Windows, MacOS, and Linux), and do not require any high-end computer hardware, and an ordinary personal laptop can be used.

[0056] (4) Flexible: The method and system for designing Taqman SNP genotyping probes and primer sets of the present invention use reasonable default values, which are particularly useful for novice and inexperienced researchers. At the same time, the method and system described in the present invention also allow users to change the default values (optional functions), so as to suit experienced researchers and can design probes and primer sets for difficult SNPs and DNA regions. Description of the Drawings

[0057] Figure 1 It is a flow chart of the method and steps for designing Taqman SNP genotyping probe sets of the present invention.

[0058] Figure 2 It is a flow chart of the method and steps for designing Taqman SNP genotyping primer sets of the present invention.

[0059] Figure 3 It is the first flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0060] Figure 4 It is the second flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0061] Figure 5 It is the third flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0062] Figure 6 It is the fourth flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0063] Figure 7 It is the fifth flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0064] Figure 8 It is the sixth flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0065] Figure 9 It is the seventh flow chart for designing Taqman genotyping probe sets for rs2476601 (a human SNP) using the method, system and storage medium of the present invention.

[0066] Figure 10 An eighth flow chart for designing a Taqman genotyping probe set for rs2476601, a human SNP, using the methods, systems, and storage media of the present invention.

[0067] Figure 11 A ninth flow chart for designing a Taqman genotyping probe set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0068] Figure 12 A tenth flow chart for designing a Taqman genotyping probe set for rs2476601, a human SNP, using the methods, systems, and storage media of the present invention.

[0069] Figure 13 An eleventh flow chart for designing a Taqman genotyping probe set for rs2476601, a human SNP, using the methods, systems, and storage media described herein.

[0070] Figure 14 A first flow chart for designing a Taqman genotyping primer set for rs2476601, a human SNP, using the methods, systems, and storage media described herein.

[0071] Figure 15 A second flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media described herein.

[0072] Figure 16 A third flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0073] Figure 17 A fourth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0074] Figure 18 A fifth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0075] Figure 19 A sixth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0076] Figure 20 A seventh flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0077] Figure 21 An eighth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0078] Figure 22 A ninth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0079] Figure 23 A tenth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0080] Figure 24 An eleventh flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0081] Figure 25 A twelfth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0082] Figure 26 A thirteenth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0083] Figure 27 A fourteenth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0084] Figure 28 A fifteenth flow chart for designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the methods, systems, and storage media of the present invention.

[0085] Figure 29Partial screenshot of the results obtained by designing a Taqman genotyping primer set for rs2476601 (a human SNP) using the method, system, and storage medium of the present invention.

[0086] Figure 30 Genotyping results of rs2476601 obtained by performing a Taqman SNP genotyping detection experiment using the Taqman genotyping probe and primer set designed for rs2476601 (a human SNP) according to the present invention. Detailed implementation mode

[0087] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention. For those not specifying specific techniques or conditions in the examples, the techniques or conditions described in the literature in this field or according to the product specifications are followed. For reagents or instruments not indicating the manufacturer, they are all conventional products that can be purchased through regular channels.

[0088] Example

[0089] This example relates to a method for designing Taqman SNP genotyping probes and primer sets, including the following steps:

[0090] (1) First program ( Figure 1 ) : According to the DNA fragment sequence containing the reference allele of the target SNP and the alternative allele sequence of the target SNP input by the user, a probe set is obtained; the probe set includes a probe targeting the reference allele of the target SNP (the first probe) and a probe targeting the alternative allele of the target SNP (the second probe);

[0091] (2) Second program ( Figure 2 ) : According to the DNA fragment sequence containing the reference allele of the target SNP input by the user, as well as the sequences and Tm values of the first probe and the second probe input by the user (included in the results returned to the user by the first program), a primer set is obtained. Preferably in this example, the first program includes the following specific steps ( Figure 1 ) :

[0092] The first program includes the following specific steps ( Figure 1 ) :

[0093] (1-1) Please ask the user to input the following information:

[0094] ① DNA fragment sequence containing the reference allele of the target SNP: This sequence contains the reference allele sequence of the target SNP and its flanking regions (i.e., the flanking regions ≥200 nucleotides upstream of the SNP and ≥200 nucleotides downstream of the SNP). This sequence must be from the plus (+) strand of the reference genomic sequence. The reference allele sequence of the target SNP is capitalized, the flanking region sequence is lowercase, and nucleotides in the DNA fragment sequence that cannot be included in the probe sequence and primer sequence (e.g., common SNPs other than the target SNP) are represented by n.

[0095] ②Alternative allele sequences of the target SNP.

[0096] ③The name of the target SNP.

[0097] (1-2) Delete characters other than a, c, g, t, A, C, G, T, and n from the DNA fragment sequence containing the target SNP reference allele, and then extract a short fragment, namely, the 24 nucleotides upstream of the SNP (this number can be optionally changed by the user), the reference allele sequence of the SNP, and the 24 nucleotides downstream of the SNP. Count the total number of bases C and C and the total number of bases G and G in the short DNA sequence; if the total number of bases G and G is less than or equal to the total number of bases C and C, use it for subsequent analysis; if the total number of bases G and G is greater than the total number of bases C and C, generate a reverse complement sequence of the short DNA sequence, and use it for subsequent analysis.

[0098] (1-3) Automatically set the Primer3 parameters so that both the first probe and the second probe meet the following conditions:

[0099] ① The Tm value is 66°C to 70°C, where the Tm value is calculated using a commonly used nearest neighbor thermodynamic model.

[0100] ②The sequence length of the probe is 17bp to 35bp.

[0101] ③ The GC (including G, C, g and c) percentage in the probe sequence is 20% to 80%.

[0102] ④ The maximum allowed length of single nucleotide repeats in the probe sequence is 5 (for example, aaaaaa does not meet this requirement).

[0103] (1-4) Automatically generate an input file for Primer3 analysis (including the above Primer3 parameters and the above short DNA sequence containing the target SNP reference allele). Use this input file for Primer3 analysis to obtain candidate first probes (targeting the reference allele of the target SNP) for subsequent steps.

[0104] (1-5) Replace the target SNP reference allele in the above short DNA sequence containing the target SNP reference allele with the alternative allele of the target SNP: If the DNA short sequence is reverse-complemented (see step 1-2 above), this step uses the reverse-complemented sequence of the alternative allele sequence input by the user; otherwise, use the alternative allele sequence input by the user.

[0105] (1-6) Automatically generate an input file for Primer3 analysis (including the above Primer3 parameters and the short DNA sequence containing the target SNP alternative allele generated in step 1-5 above). Use this input file for Primer3 analysis to obtain candidate second probes (targeting the alternative allele of the target SNP) for subsequent steps.

[0106] (1-7) Further screen out candidate first probes and candidate second probes that meet the following requirements from the above candidate probes:

[0107] ① The target SNP is located in the middle third of the probe.

[0108] ② The 5′ end of the probe cannot be g.

[0109] ③ The probe sequence does not contain more than 3 consecutive g or G, for example, sequences containing consecutive gggG, gggg, ggGg, gGgg do not meet the requirements of the first probe and the second probe.

[0110] ④ The total number of C and c contained in the probe sequence is greater than or equal to the total number of G and g contained.

[0111] (1-8) Further find all probe pairs with a Tm value difference of no more than 1℃ from the candidate first probes and candidate second probes screened in step 1-7 (including a first probe targeting the reference allele of the target SNP and a second probe targeting the alternative allele of the target SNP). The user can choose how many (default value is 10) candidate first probes and candidate second probes to perform this pairing step.

[0112] (1-9) Calculate the total sum of the probe rankings given by Primer3 for each probe pair, sort all the obtained probe pairs according to this value (from low to high), generate an output file, and return the result to the user.

[0113] Preferably in this embodiment, the second program includes the following specific steps ( Figure 2 ):

[0114] (2-1) Ask the user to input the following information:

[0115] ① The DNA fragment sequence containing the reference allele of the target SNP (the same as the DNA sequence input when designing the Taqman probe, see ① in the above step 1-1);

[0116] ② The name of the target SNP;

[0117] ③ The sequences of the first probe (targeting the reference allele of the target SNP) and the second probe (targeting the alternative allele of the target SNP) selected by the user; these probe sequences are included in the output file returned to the user by the first program (see the above step 1-9). The reference allele and alternative allele sequences of the target SNP are in uppercase letters, and all other nucleotides are in lowercase letters.

[0118] ④ The Tm values of the first probe and the second probe selected by the user (included in the output file returned to the user by the first program, see the above step 1-9).

[0119] ⑤ The maximum allowable difference between the Tm values of the primer pair (ranging from 0 to 5.0; the default value is 2.0).

[0120] (2-2) Delete the characters other than a, c, g, t, A, C, G, T, and n from the DNA fragment sequence containing the reference allele of the target SNP input by the user above, and use this new sequence (hereinafter referred to as the "DNA template sequence") for the subsequent steps.

[0121] (2-3) Determine whether the sequence of the first probe input by the user exactly matches the DNA template sequence. If it exactly matches, use the sequences of the first probe and the second probe input by the user for the subsequent steps; if it does not exactly match, perform reverse complementation on the sequences of the first probe and the second probe (the two allele sequences are still in uppercase letters, and all other nucleotides are in lowercase letters), and then use these new probe sequences for the subsequent steps.

[0122] (2-4) Obtain the sequences of the reference allele and alternative allele of the target SNP from the above sequences of the first probe and the second probe.

[0123] (2-5) Automatically set the Primer3 parameters so that the primer pair and the forward primer and reverse primer it contains meet the following requirements:

[0124] ① The sequence of the forward primer does not overlap with the sequences of the above-mentioned first probe and second probe; the sequence of the reverse primer does not overlap with the sequences of the above-mentioned first probe and second probe.

[0125] ② The Tm value of the forward primer is 6°C to 8°C lower than the average of the Tm values of the above-mentioned first probe and second probe input by the user; the Tm value of the reverse primer is 6°C to 8°C lower than the average of the Tm values of the above-mentioned first probe and second probe input by the user.

[0126] ③ The maximum difference between the Tm values of the forward primer and the reverse primer is the value obtained in ⑤ of the above step (2-1).

[0127] ④ The size of the PCR product obtained from the primer pair is 50bp to 150bp, and the optimal value is 50bp.

[0128] ⑤ The lengths of both the forward primer and the reverse primer are 18bp to 30bp.

[0129] ⑥ The gc percentages in the sequences of both the forward primer and the reverse primer are 30% to 80%.

[0130] (2-6) Automatically generate an input file for Primer3 analysis (including the above-mentioned Primer3 parameters and the DNA template sequence), use this input file for Primer3 analysis, and return the results of Primer3 (including all primer pairs compatible with the first probe and second probe input by the user) to the user.

[0131] Furthermore, ① in the above step (2-5) specifically includes the following steps:

[0132] ① Calculate the length of the DNA template sequence and assign it to the variable "L".

[0133] ② Align the sequence of the first probe with the DNA template sequence (the first base is numbered 0), obtain the numbers of the first base and the last base of the first probe sequence on the DNA template, and assign these numbers to the variables "p1_start" and "p1_end" respectively.

[0134] ③ Replace the target SNP reference allele sequence in the DNA template sequence with the target SNP alternative allele sequence. Align this new DNA template sequence (the first base is numbered 0) with the second probe sequence, obtain the numbers of the first base and the last base of the second probe sequence on this DNA template, and assign these numbers to the variables "p2_start" and "p2_end" respectively.

[0135] ④Automatically set the Primer3 parameter "SEQUENCE_EXCLUDED_REGION" as shown below:

[0136] SEQUENCE_EXCLUDED_REGION=a,b;

[0137] Where, a=min(p1_start,p2_start);

[0138] b=max(p1_end,p2_end)-min(p1_start,p2_start)+1.

[0139] ⑤Automatically set the Primer3 parameter "SEQUENCE_PRIMER_PAIR_OK_REGION_LIST" as shown below:

[0140] SEQUENCE_PRIMER_PAIR_OK_REGION_LIST=0,x,y,z;

[0141] Where, x = min(p1_start, p2_start);

[0142] y=max(p1_end,p2_end)+1;

[0143] z=L-max(p1_end,p2_end)-1.

[0144] Furthermore, the above steps (2-5) ② specifically include the following steps:

[0145] ① Assign the Tm value of the first probe and the Tm value of the second probe to variables "Tm1" and "Tm2" respectively, and let the variable Tm_avg = (Tm1+Tm2-14) / 2.

[0146] ②Automatically set Primer3 parameters as follows:

[0147] PRIMER_MIN_TM=Tm_avg-1;

[0148] PRIMER_OPT_TM=Tm_avg;

[0149] PRIMER_MAX_TM=Tm_avg+1.

[0150] Furthermore, step ③ in the above steps (2-5) specifically includes the following steps:

[0151] ① Assign the maximum allowable difference between the Tm values of the primer pairs entered by the user (see ⑤ in the above step (2-1)) to the variable "Tm_diff".

[0152] ②Automatically set the Primer3 parameter "PRIMER_PAIR_MAX_DIFF_TM" as follows:

[0153] PRIMER_PAIR_MAX_DIFF_TM = Tm_diff.

[0154] This embodiment also relates to a system for designing Taqman SNP genotyping probes and primer sets, including:

[0155] (1) The first program system, which obtains a probe set according to the DNA fragment sequence containing the target SNP reference allele and the alternative allele sequence of the target SNP input by the user; the probe set includes a probe (the first probe) targeting the reference allele of the target SNP and a probe (the second probe) targeting the alternative allele of the target SNP;

[0156] (2) The second program system, which obtains a primer set according to the DNA fragment sequence containing the target SNP reference allele input by the user, and the sequences and Tm values of the first probe and the second probe input by the user (included in the result returned to the user by the first program).

[0157] The combined use of the two program systems can design a complete set of probes and primer sets for Taqman genotyping detection for a target SNP.

[0158] This embodiment also relates to a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, all steps of the method for designing Taqman SNP genotyping probes and primer sets are implemented.

[0159] Specifically in this embodiment, using the method, system and storage medium for designing Taqman SNP genotyping probes and primer sets of this embodiment above, a complete set of probes and primer sets for Taqman genotyping are designed for rs2476601 (the name of a human SNP), specifically as follows:

[0160] (1) Design the probe set

[0161] Run the above first program and input the following information provided by the user ( Figure 1 、 Figures 3 to 13 ):

[0162] (A) DNA fragment sequence (SEQ ID NO: 1) containing the reference allele of the target SNP: ctacacatgcatgctgctattgctctgctaaattaaaaaaaaaaaattagagatgaggtctcactatgttgngcaggctagtcttgaactcctggcctcaatgaactcctcaaactcaaggctcacacatcagcttcccaaagtgctggaattaaaggcatgagccaccatgcccatcccacactttattttatacttactgaactgtactcaccagcttcctcaaccacaataaatgattcaggtgtccAtacaggaagtggaggggggatttcatcatctatccttggagcagttgctatccaaaatgtcaaaaatattgtaacaattgttaattagaacaatccaaaggaaattcttatattctaatattaaatataaatttaccataatttatatttaaattccgttgaagcaacattatcagtaaagttgacacttgttcattcaangaaaaagcaaaataaattctcttaaggtacaaacccaggaggtttttgct。

[0163] (B) Alternative allele sequence of the target SNP: G.

[0164] (C) Name of the target SNP: rs2476601.

[0165] Part of the results returned by this program (i.e., part of the output file) is shown in Table 1:

[0166] Table 1 Taqman genotyping probe sets designed by the first program of the present invention for rs2476601 (the top 10 groups)

[0167]

[0168] From the results in Table 1, the first group of probes was selected and synthesized for subsequent primer design and experimental verification:

[0169] First probe (targeting the reference allele of rs2476601): FAM-cccctccacttcctgtaTggacacctg-BHQ1;

[0170] Second probe (targeting the alternative allele of rs2476601): HEX-ccctccacttcctgtaCggacacctg-BHQ1.

[0171] (2) Design primer sets

[0172] For the above-selected first and second probes, run the second program of the present invention and input the following information provided by the user ( Figure 2 、 Figures 14 to 29 ):

[0173] (A) DNA fragment sequence containing the reference allele of the target SNP: as shown in the above sequence SEQ ID NO: 1.

[0174] (B) Name of the target SNP: rs2476601.

[0175] (C) Sequence of the first probe selected by the user (targeting the reference allele of the target SNP): ccccctccacttcctgtaTggacacctg. Where FAM is used as the fluorescent reporter group dye and BHQ1 is used as the quencher.

[0176] (D) Sequence of the second probe selected by the user (targeting the alternative allele of the target SNP): ccctccacttcctgtaCggacacctg. Where HEX is used as the fluorescent reporter group dye and BHQ1 is used as the quencher.

[0177] (E) Tm value of the first probe selected by the user: 67.3.

[0178] (F) Tm value of the second probe selected by the user: 67.5.

[0179] Part of the results returned by this program (i.e., part of the output file) is as Figure 29 shown. Select the 190th forward and reverse primers for subsequent experimental verification:

[0180] Forward primer: caccagcttcctcaaccaca (SEQ ID NO: 22);

[0181] Reverse primer: caactgctccaaggatagatgatga (SEQ ID NO: 23).

[0182] (3) Perform Taqman SNP genotyping detection experiment verification

[0183] A. Using the above probes and primer sets, DNA samples with the following rs2476601 genotypes were identified:

[0184] ①Reference allele homozygous (A / A);

[0185] ②Heterozygous (A / G);

[0186] ③Alternative allele homozygous (G / G);

[0187] B. The PCR reaction (total volume 10 μL) used in this Taqman SNP genotyping assay contains 200 nM of the first probe, 200 nM of the second probe, 300 nM of the forward primer, 300 nM of the reverse primer, 0.8 mM dNTP, 50 mM KCl, 1.5 mM MgCl2, 0.5 units of Taq DNA polymerase (JumpStart TM Taq DNA polymerase, MilliporeSigma), and 20 ng of DNA sample (or water, as a negative control).

[0188] Table 2 PCR program used in this Taqman SNP genotyping assay on a Roche 480 real-time PCR instrument

[0189]

[0190] As Figure 30 (where the values on the X-axis and Y-axis are the FAM fluorescence intensity and HEX fluorescence intensity obtained in this experiment) shows that the genotype of each sample was correctly determined in this experiment, indicating that the Taqman probe and primer set designed for rs2476601 in the present invention is successful.

[0191] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for designing Taqman SNP genotyping probes and primer sets, characterized in that, The steps include: (1) A first program, based on a DNA fragment sequence containing a target SNP reference allele and an alternative allele of the target SNP input by a user, obtains a probe set; the probe set includes a first probe and a second probe, the first probe targeting the reference allele of the target SNP, and the second probe targeting the alternative allele of the target SNP; (2) A second program, which obtains a primer set based on the sequence of the DNA fragment containing the target SNP reference allele input by the user, and the sequence and melting temperature Tm value of the first probe and the second probe input by the user; The second procedure includes the following specific steps: (2-1) Obtaining a DNA fragment sequence containing a target SNP reference allele, an alternative allele of the target SNP, sequences of a first probe and a second probe selected by a user, and a melting temperature (Tm); the DNA fragment sequence containing the target SNP reference allele includes the target SNP reference allele and flanking regions; (2-2) The forward primer and the reverse primer in the primer set meet the following parameters: The forward primer sequence is set to not overlap with the sequences of the first probe and the second probe, and the reverse primer sequence is set to not overlap with the sequences of the first probe and the second probe; The melting temperature Tm value of the forward primer is set to be 7±1.5°C lower than the melting temperature Tm value of the first probe or the melting temperature Tm value of the second probe, and the melting temperature Tm value of the reverse primer is set to be 7±1.5°C lower than the melting temperature Tm value of the first probe or the melting temperature Tm value of the second probe; Setting the maximum difference between the melting temperature Tm values of the forward primer and the reverse primer; Setting the size of the PCR product of the primer set to 50 bp to 150 bp; The lengths of the forward primer and the reverse primer are both set to 18 bp to 30 bp; The gc percentages in the forward primer and the reverse primer are set to be 30% to 80%; (2-3) Automatically generate a second input file for Primer3 analysis; perform Primer3 analysis using the second input file to obtain a primer pair targeting the DNA fragment sequence containing the target SNP reference allele.

2. The method for designing Taqman SNP genotyping probes and primer sets according to claim 1, characterized in that, The first procedure includes the following specific steps: (1-1) Obtaining a DNA fragment sequence containing a target SNP reference allele and an alternative allele of the target SNP; the DNA fragment sequence containing the target SNP reference allele includes the reference allele of the target SNP and flanking regions; (1-2) extracting a short fragment containing part of the bases in the flanking region and the reference allele of the target SNP from the DNA fragment sequence containing the target SNP reference allele to obtain a short DNA sequence; Count the total number of bases C and c and the total number of bases G and g in the short DNA sequence; if the total number of bases G and g is less than or equal to the total number of bases C and c, it is used for subsequent analysis; if the total number of bases G and g is greater than the total number of bases C and c, generate the reverse complementary sequence of the short DNA sequence and then use it for subsequent analysis; (1-3) Set the parameters of the first probe or the second probe; automatically generate the first input file for Primer3 analysis; (1-4) Perform Primer3 analysis using the first input file to obtain the reference allele targeting the target SNP and the first probe set for the flanking region; (1-5) Replace the reference allele of the target SNP in the short DNA sequence with the alternative allele of the target SNP, and then sequentially execute steps (1-3) to (1-4) to obtain the alternative allele targeting the target SNP and the second probe set for the flanking region; (1-6) Set screening conditions to screen the first probe set and the second probe set to obtain a probe group.

3. The method for designing Taqman SNP genotyping probes and primer sets according to claim 2, wherein In step (1-1), the reference allele of the target SNP uses uppercase letters, the flanking region uses lowercase letters, and the nucleotides in the DNA fragment sequence containing the reference allele of the target SNP that cannot be included in the probe and primer set use n.

4. The method for designing Taqman SNP genotyping probes and primer sets according to claim 2, characterized in that, Step (1-3) sets the parameters of the first probe or the second probe as follows: Set the melting temperature Tm of the first probe or the second probe to be 66 °C to 70 °C; Set the sequence length of the first probe or the second probe to be 17 bp to 35 bp; Set the GC percentage in the sequence of the first probe or the second probe to be 20% to 80%; Set the maximum allowable length of single nucleotide repeats in the first probe or the second probe to be 5.

5. The method for designing Taqman SNP genotyping probes and primer sets according to claim 2, characterized in that, Step (1-6) includes the following specific steps: (1-6-1) Set screening conditions to screen the first probe set and the second probe set to obtain the screened first probe set and the screened second probe set; the set screening conditions are as follows: Set that the SNPs are all located in the middle third of the first probe and the second probe; Set that the 5′ ends of the first probe and the second probe are not g; Set that the sequences of the first probe and the second probe do not contain more than 3 consecutive g or G; Set that the total number of C and c contained in the sequences of the first probe and the second probe is more than or equal to the total number of G and g; (1-6-2) From the screened first probe set and the screened second probe set, find the first probe and the second probe with a melting temperature Tm value difference not exceeding 1 °C as the probe group; (1-6-3) Count the total ranking scores of each probe group and sort according to the total ranking scores.

6. The method for designing Taqman SNP genotyping probes and primer sets according to claim 1, wherein In step (2-2), it is set that the sequence of the forward primer does not overlap with the sequences of the first probe and the second probe, and it is set that the sequence of the reverse primer does not overlap with the sequences of the first probe and the second probe, which specifically includes the following steps: Statistically analyze the length value of the DNA fragment sequence containing the target SNP reference allele, and sequentially number the bases of the DNA fragment sequence containing the target SNP reference allele to obtain the numbered DNA fragment sequence containing the target SNP reference allele; Compare the sequence of the first probe with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele, then obtain the numbers of the first base and the last base of the sequence of the first probe on the numbered DNA fragment sequence containing the target SNP reference allele, and assign the numbers to the first variable group respectively; Replace the reference allele of the target SNP in the numbered DNA fragment sequence containing the target SNP reference allele with the alternative allele of the target SNP to obtain a numbered DNA fragment sequence containing the alternative allele; compare the sequence of the second probe with the numbered DNA fragment sequence containing the alternative allele, obtain the numbers of the first base and the last base of the sequence of the second probe on the numbered DNA fragment sequence containing the alternative allele, and assign the numbers to the second variable group respectively; Using the first variable group and the second variable group, make the sequence of the forward primer not overlap with the sequences of the first probe and the second probe, and make the sequence of the reverse primer not overlap with the sequences of the first probe and the second probe.

7. The method for designing a Taqman SNP genotyping probe and primer set according to claim 6, wherein Compare the sequence of the first probe with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele. If the sequence of the first probe is completely matched with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele, then obtain the numbers of the first base and the last base of the sequence of the first probe on the numbered DNA fragment sequence containing the target SNP reference allele, and assign the numbers to the first variable group respectively; if the sequence of the first probe is not completely matched with the sequence of the numbered DNA fragment sequence containing the target SNP reference allele, then perform reverse complementation on the sequences of the first probe and the second probe, obtain the numbers of the first base and the last base of the complementary sequence of the first probe on the numbered DNA fragment sequence containing the target SNP reference allele, and assign the numbers to the first variable group respectively.

8. A system for designing Taqman SNP genotyping probes and primer sets, characterized in that, The system adopted by the method for designing a Taqman SNP genotyping probe and primer set according to claim 1 includes: A first program system for obtaining a probe set according to a DNA fragment sequence containing a reference allele of a target SNP and an alternative allele of the target SNP input by a user; the probe set includes a probe targeting the reference allele of the target SNP and a probe targeting the alternative allele of the target SNP. A second program system for obtaining a primer set according to a DNA fragment sequence containing a reference allele of a target SNP input by a user, and the sequences and melting temperature Tm values of a first probe and a second probe input by the user.

9. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, all steps of the method for designing Taqman SNP genotyping probes and primer sets according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Allele variant detection method, kit and composition

    CN103215361A

  • Probe melting curve analysis -based multiplex nucleic acid detection method and kit

    WO2022179419A1