Real-time fluorescent quantitative PCR (polymerase chain reaction) primer and probe design method based on panogenomics
By using pan-genomics analysis and primer design software to screen target genes, the problems of low efficiency and insufficient specificity in primer design for pathogenic microorganisms have been solved, and primer design with high sensitivity and specificity has been achieved, which is applicable to microorganisms such as bacteria, fungi and nematodes.
Patent Information
- Application Number
- CN202511773618.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-28
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies suffer from low efficiency and insufficient specificity in primer design for pathogenic microorganisms, especially in the screening of multi-copy genes, resulting in long primer design cycles and insufficient sensitivity.
Using a pangenomics-based approach, we analyzed the genomic data of target microorganisms with OrthoFinder software to screen for target genes that are both conserved and specific. Primers and probes were then designed using qPrimer software to ensure efficient screening and validation.
It improves primer sensitivity and specificity, shortens the design cycle, is applicable to a variety of microorganisms, simplifies the operation process, and reduces the professional knowledge requirements of bioinformatics.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics and molecular biology, specifically to a method for designing primers and probes for real-time quantitative PCR targeting pathogenic microorganisms based on pan-genomics analysis. Background Technology
[0002] Real-time quantitative PCR technology has been widely used in fields such as pathogen detection, population dynamics monitoring, and disease risk early warning due to its advantages of high sensitivity, good specificity, and high throughput. One of the core aspects of this technology is the design of primers for pathogenic microorganisms with high sensitivity and high specificity.
[0003] A key technical bottleneck in primer design lies in the difficulty of rapidly screening for high-quality target gene sequences with high interspecificity and intraspecific conservation. Traditional methods rely on multiple sequence alignment of known conserved genes to identify specific regions within these sequences. However, due to differences in the degree of variation among conserved genes in different species, multiple conserved genes need to be selected for primer design, resulting in long design cycles and low efficiency. Furthermore, for some closely related species, insufficient interspecific variation in conserved genes leads to stagnation in primer design.
[0004] With the rapid development of sequencing technology, third-generation sequencing technology and genomics are widely used in microbial genome assembly. Among these, pan-genome sequencing provides strong data support for the development of pathogen-specific primers and probes. However, this technology still faces the following technical bottlenecks in practical applications: Existing publicly available primer design and target sequence mining methods mainly employ two technical routes. The first route, as mentioned in the design method and system for targeted pathogen sequencing primers disclosed in Chinese patent application CN202411253978.9, involves identifying common conserved genes of species through pan-genome analysis, then designing primers on these conserved genes, and screening them by evaluating their coverage and specificity. This method, after identifying the conserved genome of the target microorganism, suffers from a lack of efficient specific primer screening mechanisms, resulting in a large workload for primer design and a large number of non-specific primers during subsequent qPCR validation, leading to low primer design efficiency. The second technical route, as disclosed in Chinese patent application CN202411805012, involves a primer design method for multiplex targeted sequencing technology. After determining the core genome, an optimal reference genome is selected, and all primers are designed and validated by traversing the optimal reference genome with a fixed step size. This method has the problems of large data computation, high requirements for users' bioinformatics expertise and computer hardware configuration, and also faces a lot of primer design and repeated screening of primer specificity and coverage in the later stage.
[0005] Existing technologies mostly focus on primer specificity and coverage, without considering the important role of multi-copy genes in improving primer sensitivity, and without conducting related work on screening multi-copy target genes. Summary of the Invention
[0006] To address the problems of low efficiency, excessive repetition, and insufficient specificity in existing technologies for specific primer screening, this invention provides the following technical solution: This application discloses a method for designing primer-probe compositions based on pangenomics, the method comprising the following steps: S1) Collect genomic data of target microorganisms to construct a target microorganism genome database, and select the best reference genome from the target microorganism genome database; S2) Analyze the target microbial genome database to screen and obtain the set of conserved genes of the target microorganisms; S3) Collect genome data of target microorganisms similar to those species and construct a genome database of target microorganisms similar to those species; S4) Combine the best reference genome from S1) with the target microbial similar species genome database from S3) for joint analysis to screen and obtain a set of genes specific to the target microorganisms; S5) Take the intersection of the conserved gene set of the target microorganism in S2) and the unique gene set of the target microorganism in S4) as the target gene. S6) Use primer design software to design primers and probes for the target gene sequences obtained in S5), and determine the final primer-probe combination based on the penalty score of the primer design software, the coverage of the primers or probes, the specificity of the primers or probes, the theoretically feasible primer-probe combination, and experimental verification.
[0007] As used in this article, "pan-genome" is the collective term for all the genes of a microorganism. A pan-genome includes a core genome and non-essential genomes. The core genome consists of genes that are universally present in a population of microorganisms; the non-essential genome consists of genes present in some populations. In practical research, pan-genomes can also be divided into core genomes (genes present in all populations), non-essential genomes (genes present in two or more populations), and strains-specific genes (genes present only in a specific population).
[0008] In this application, the target microorganism may be a prokaryotic microorganism or a eukaryotic microorganism.
[0009] In this application, the target microorganism may be bacteria, nematodes, or fungi.
[0010] In the method described in S1), the method for collecting the genomic data of the target microorganism is as follows: S11) Prioritize downloading genomic data from the NCBI Refseq database; if the data in the Refseq database is insufficient, select to download data from the GenBank database. S12) Referring to the NCBI website's evaluation of genome integrity, discard genome data with integrity less than 90%. S13) Prioritize genomic data with a complete assembly degree. If the number of complete data is insufficient, select chromosome, scaffold, and conting data in sequence.
[0011] In the method described above, the method for determining the optimal reference genome data in S1) is as follows: S101) Prioritize data from the RefSeq database in the NCBI database that comes from the type material, and determine the best reference genome in the order of priority: reference genomes > complete > chromosome > scaffold > conting; If no data meets the conditions of S101) in S102), select the best reference genome from non-type strains (databases not marked "From typematerial") and determine the best reference genome in the following priority order: reference genomes > complete > chromosome > scaffold > conting.
[0012] In the method described in S2), the set of conserved genes of the target microorganism is obtained by comparative genomic analysis of the target microorganism genome database.
[0013] In this application, the conserved gene set of the target microorganism is obtained by comparative genomic analysis of the target microorganism genome database using Orthofinder 2.5 software.
[0014] In some embodiments of this application, the steps for obtaining the conserved gene set using Orthofinder 2.5 software are as follows: (1) Enter orthoFinder -f<target microbial genome database file name>-d -og in the command line and run the program; (2) Select the file Orthogroups.tsv in the results folder "Orthogroups" to filter and obtain the orthologous genome file names commonly contained in all species; (3) Merge the genome files obtained in the previous step in the results folder "Orthogroup Sequences" to obtain the conserved gene set of the target microorganism.
[0015] As used in this article, “conserved genes” are defined and output by OrthoFinder based on the distribution patterns of orthologs in all species. They are a class of genes whose DNA sequences are highly similar or remain unchanged in target microorganisms.
[0016] In the method described above, the approximate species are microorganisms that belong to the same genus as the target microorganism but are different species in taxonomy, and the database genome screening principles are the same as those described in S11 to S13 above.
[0017] In the method described above, the method for obtaining the target microorganism-specific genes in S4) is to add the best reference genome data of the target microorganism in S1) into the approximate species genome database established in S3) for comparative genomic analysis to obtain the combination of target microorganism-specific genes.
[0018] As used in this article, “unique genes” are DNA sequences unique to the best reference genome of the target microorganism, defined and output by OrthoFinder based on the distribution patterns of orthologs across all species.
[0019] The comparative genomics analysis described in S4) was performed using Orthofinder 2.5 software.
[0020] In some embodiments of this application, the steps of obtaining the target microorganism-specific gene combination using Orthofinder2.5 software include: (1) running the Orthofinder software; (2) obtaining the target microorganism-specific multi-copy gene set (hereinafter referred to as specific gene set 1 for convenience), selecting the file Orthogroups.tsv in the results folder "Orthogroups" to filter only the genome number specific to the target microorganism; merging the filtered genome files in the results folder "OrthogroupSequences" to obtain specific gene set 1; (3) obtaining the target microorganism-specific single-copy gene set (hereinafter referred to as specific gene set 2): selecting the file Orthogroups_UnassignedGenes.tsv in the results folder "Orthogroups" to filter the file name of the target microorganism-specific genome, merging the filtered genome files in the results folder "Orthogroup Sequences" to obtain specific gene set 2.
[0021] Unique genes consist of two parts: one part is the homologous genes of the target microorganism itself, that is, the multi-copy genes unique to the target microorganism; the other part is the genes in the target microorganism that are not classified into the ortholog group (UnassignedGenes).
[0022] In this application, the method includes the step of comparing the obtained specific gene set 1 and specific gene set 2 with genes in the conserved gene set of the target microorganism, respectively.
[0023] In this application, the method steps for comparing specific gene set 1 and specific gene set 2 with conserved gene set are as follows: (1) Use the seq -n command in the seqkit software toolkit to extract the names of all gene files in specific gene set 1 and specific gene set 2 respectively; (2) Use the grep command in the seqkit software toolkit to screen the conserved gene set files in step "2. Screening to obtain conserved gene set of target microorganism" to screen gene files containing the file names obtained in the previous step. The result is the target gene sequence with both conservation and specificity; (3) Use the seq -m command in the seqkit software to screen and remove genes with a length of less than 300bp obtained in the previous step.
[0024] In the method described in S5), the target gene satisfies the following condition: it is longer than 300 bp and exists simultaneously in the conserved gene set of the target microorganism described in S2) and the unique gene set described in S4).
[0025] In this application, the target gene may be a multi-copy gene or a single-copy gene.
[0026] The multicopy gene is a gene that contains two or more paralogous genes (genes generated by gene replication within the same species) in the target microorganism.
[0027] In this application, the number of target genes is n, where n ≥ 1 and n is a natural number.
[0028] In this application, the primer design conditions described in S6) include an amplification product length of 80-120 bp, a primer annealing temperature of 52-60℃, a primer length of 18-23 nt, and a GC content of 30%-70%.
[0029] In this application, the primer design conditions described in S6) include a primer annealing temperature of 56°C, a primer length of 20 nt, and a GC content of 50%.
[0030] In this application, the probe design conditions described in S6) include a probe annealing temperature of 62-70°C, a length of 13-30 nt, and a GC content of 20%-80%.
[0031] In this application, the probe design conditions described in S6) include a probe annealing temperature of 66°C, a length of 20 nt, and a GC content of 50%.
[0032] In this application, the method further includes arranging the selected primer-probe compositions in ascending order according to their penalty scores, with the top m compositions selected as candidate primer-probe compositions, where m is a natural number and 10≤m≤50, and the penalty score of the primer-probe composition = forward primer penalty score + reverse primer penalty score + probe penalty score.
[0033] In this application, the method further includes the step of evaluating the coverage and specificity of the candidate primer-probe composition to obtain the target primer-probe composition.
[0034] In this application, the penalty values for each primer can be calculated using the default parameters of the qPrimer software.
[0035] In some embodiments of this application, the qPrimer software (https: / / github.com / swu1019lab / qPrimer), developed based on Primer3, was used to design primers and probes for the target gene sequences. Primers were annealed at 52-60°C, with a length of 18-23 nt and a GC content of 30%-70%. Probes were annealed at 62-70°C, with a length of 13-30 nt and a GC content of 20%-80%. The annealing temperature was predicted using the `calculate_tm()` function of Primer3.
[0036] Specifically, qPrimer is used to design primers and probes for the target gene sequence. After running the program, or by writing the input command `qPrimer design –seq_file<target gene sequence>--ini_file<primer design parameter file>--csv –out_name<output primer design result file name>`, locate the columns for forward primer penalty value, reverse primer penalty value, and probe penalty value in the output file, i.e., the columns named "PENALTY_F", "PENALTY_R", and "PENALTY". Calculate the penalty value for each primer-probe combination according to the formula: Penalty value = Forward primer penalty value + Reverse primer penalty value + Probe penalty value. Evaluate and select the top 10 primer pairs with the lowest penalty values based on coverage and specificity. The primer sequences used for primer coverage and specificity evaluation are shown in Table 1 below.
[0037] The penalty values for each primer were calculated using the software's default parameters.
[0038] In this application, the coverage evaluation method is as follows: using BLAST software, the amplicon sequences or probe sequences obtained by primer amplification are compared with the target microbial genome database, and the proportion of the number of qualified genomes that the amplicon or probe can match is calculated to the total number of genomes in the target microbial genome database.
[0039] Primer coverage = Number of genomes with valid amplicon matches / Total number of genomes in the target microorganism genome database × 100%; Probe coverage = (Number of genomes with valid probe sequence matches / Total number of genomes in the target microorganism genome database) × 100%; The amplicon referred to here is the nucleic acid region defined and amplified by the forward and reverse primers during PCR. The start end is defined by the 5' end of the forward primer, and the end end is defined by the 5' end of the reverse primer. A successful match means that the following conditions must be met simultaneously: the amplicon alignment matching length is greater than 95% of the total amplicon length, the alignment accuracy is greater than 95%, both the forward and reverse primers can achieve a match, and the 3' ends are not mismatched.
[0040] The standard for qualified probe sequence matching is that the probe sequence is completely matched, that is, the entire probe length is completely matched and the accuracy rate is 100%.
[0041] The specificity evaluation method for the primers is to use the primer-blast function provided by NCBI to compare the candidate primer pairs in the nt database. The standard for passing the specificity evaluation is that the primer pairs whose blast-matched amplification products are all gene sequences of the target microorganism and whose amplicon length matches the design value.
[0042] The specificity evaluation method for the probes involves performing BLAST alignment of the probe sequences in the NCBI nt database. The criterion for successful probe specificity evaluation is that all alignment results correspond to the gene sequences of the target microorganism.
[0043] The amplicon described in this application is a nucleic acid region defined and amplified by forward and reverse primers during PCR, with the start end properly matched to the 5' end of the forward primer.
[0044] The standard for qualified probe sequence matching described in this application is that the probe sequence is completely matched, that is, the probe is completely matched in its entire length and the accuracy rate is 100%.
[0045] This application also provides primer-probe compositions or primer compositions obtained by the above methods, wherein the primer composition includes a forward primer and a reverse primer; and the primer-probe composition includes a forward primer, a reverse primer, and a probe.
[0046] In the primer-probe composition or primer composition, the target microorganism may be dispersible ubiquitous. The forward primer can be a single-stranded DNA nucleotide sequence as shown in SEQ ID NO:1; The reverse primer can be a single-stranded DNA nucleotide sequence as shown in SEQ ID NO:2; The nucleotide sequence of the probe is shown in SEQ ID NO:3; or The target microorganism may be *Russula styracifolium*, and the primer-probe composition or primer composition is specific to *Russula styracifolium*. The forward primer can be a single-stranded DNA nucleotide sequence as shown in SEQ ID NO:4; The reverse primer can be a single-stranded DNA nucleotide sequence as shown in SEQ ID NO:5; The nucleotide sequence of the probe is shown in SEQ ID NO:6.
[0047] In the primer composition, the ratio of the amount of the forward primer to the amount of the reverse primer is 1:1.
[0048] In the primer-probe composition, the molar ratio of the forward primer, the reverse primer, and the probe is 1:1:1.
[0049] In this invention, the probe may be a DNA probe. The 5' end of the DNA probe may be labeled with a reporter group, and the 3' end may be labeled with a quencher group.
[0050] In this invention, the fluorescent group is labeled on the 5' end of the probe.
[0051] In this invention, the quenching group is labeled at the 3' end of the probe.
[0052] In this invention, the fluorescent group is selected from at least one of FAM, VIC, HEX, TRT, CY3, CY5, ROX, JOE, FITC, TET, NED, TAMRA, LC RED640, LC RED705, Quasar705 or Texas Red.
[0053] In this invention, the quenching group may be selected from at least one of TAMRA, BHQ1, BHQ2, BHQ3, MGB and Dabcy1.
[0054] In some embodiments of the present invention, the fluorescent group is 6-FAM (6-carboxyfluorescein), and the quenching group is TAMRA-N (5(6)-carboxytetramethylrhodamine N-hydroxysuccinimide ester).
[0055] In some embodiments of the present invention, the first T at the 5' end of the probe is modified with 6-FAM; the first T at the 3' end is modified with TAMRA-N.
[0056] This application also provides any of the following applications: D1) The above-described primer-probe composition specific to the dispersible pantotheca, or the use of the primer composition in at least one of the following: Application of D1-1 in the detection of dispersible pantothenia; Application of D1-2) in the preparation of products for detecting dispersible pantothenia; Application of D1-3) in the identification or auxiliary identification of dispersible pantothenia; D1-4) Application in the preparation of products for the identification or auxiliary identification of dispersible pantothenia; D1-5) Application in identifying or assisting in the identification of whether a sample to be tested is or contains dispersible pantothenia; D1-6) Application in the preparation of products for identifying or assisting in the identification of whether a test sample is or contains dispersible pantothenia; D2) The above-described primer-probe composition specific to the *Stripetrata rust*, or the use of the primer composition in at least one of the following: Application of D2-1 in the detection of striped stalk rust fungus; Application of D2-2) in the preparation of products for detecting striped stalk rust fungus; Application of D2-3 in the identification or auxiliary identification of striped stalk rust fungi; Application of D2-4) in the preparation of products for the identification or auxiliary identification of *Stripetra rust*; D2-5) Application in identifying or assisting in the identification of whether a sample to be tested is or contains *Strombyx mori*; D2-6) Application in the preparation of products for identifying or assisting in the identification of whether a sample is or contains *Strombus styracifolius*.
[0057] This application also provides an apparatus for designing primer and probe compositions based on pangenomics, the apparatus comprising the following modules: M1) Target microbial genome data receiving and analysis module: used to collect the genome data of target microorganisms, construct a target microbial genome database, and screen from the target microbial genome database to obtain the best reference genome; M2) Target Microorganism Conserved Gene Analysis Module: Used to screen and obtain the set of conserved genes of target microorganisms based on the target microorganism genome database; M3) Target microbial similar species genome data receiving module, used to collect target microbial similar species genome data to construct a target microbial similar species genome database; The M4) target microbial specific gene analysis module is used to jointly analyze the best reference genome of M1) and the target microbial similar species genome database of M3) to screen and obtain the target microbial specific gene set; The M5) target gene analysis module is used to take the intersection of the conserved gene set of the target microorganism in M2) and the unique gene set of the target microorganism in M4) as the target gene. M6) Primer-probe composition or primer composition design module, used to design primers and / or probes for the target gene sequences obtained by screening in M5) using primer design software, and to determine the final primer-probe composition or primer composition based on the theoretically feasible primer-probe combination selected by the primer 3 penalty score, primer and / or probe coverage, primer and / or probe specificity, and experimental verification.
[0058] Compared with existing technologies, the advantages and beneficial effects of this invention include: 1) This invention flexibly utilizes the Orthofinder homologous genome analysis function to ensure that the screened target genes simultaneously meet the requirements of conservation and specificity, guaranteeing high-quality target sequences, reducing the risk of unqualified primer and probe verification in the later stages, and greatly shortening the primer development time. 2) Considering the important role of multi-copy genes in improving primer sensitivity, this invention specifically designs a program to screen multi-copy target genes, enabling primers designed using this invention to have higher sensitivity. 3) This invention has a wide range of applications and can be applied to the development of specific primers for various microorganisms such as bacteria, fungi, and nematodes. 4) This invention mainly relies on Orthofinder software for analysis and uses qprimer software developed based on Primer3 for batch primer design. The operation is simple and does not require the use of a large number of complex bioinformatics analysis software. It can be easily used even by users with relatively weak bioinformatics knowledge. Furthermore, a streamlined software tool for applying this invention has been developed later, further simplifying the operation of designing specific primers and probes using the method of this invention. Attached Figure Description
[0059] Figure 1 Flowchart for designing fluorescent quantitative primers for targeted pathogens.
[0060] Figure 2 This is the result of conventional PCR detection for the specificity of dispersive pan-based primers. Figure 2 The Marker DL500 was used, and the DNA samples for each lane were sourced from 1: dispersed pantothenia ( Pantoea dispersa ), 2: Pantothecin in pineapple ( P. ananatis ), 3: P. endophytic 4: Bacillus belye ( Bacillus velezensis ), 5: P. endophytic 6: Clumping of pan-bacteria ( P. agglomerans ), 7: Bacillus belye ( B. velezensis ), 8: Bacillus amyloliquefaciens ( B. amyloliquefaciens ), 9: Colloidal anthrax bacteria ( Colletotrichum gloeosporioides ), 10: Fusarium oxysporum ( Fusarium oxysporum ), 11: DNA from strawberry leaf tissue, 12: H2O.
[0061] Figure 3 To disperse the sensitivity of pantothenic primers, the results of conventional PCR detection were obtained.
[0062] Figure 4 This is the result of conventional PCR detection using primers specific to *Strombus styracifolius*. M in the figure represents Marker DL1000; the DNA samples from each lane originated as follows: 1. *Strombus styracifolius* (wheat leaf rust). Puccinia triticina), 2: Multiple piles of rust fungi ( P. polysora ), 3: Powdery mildew of wheat ( Blumeria graminis f. sp. tritici ), 4: Rhizoctonia solani AG-1 ( Rhizoctonia solani ), 5: Rhizoctonia solani AG-4 ( R. solani ), 6: Fusarium pseudograss ( Fusarium pseudograminearum FP1359, 7: Fusarium graminearum ( F. graminearum ), 8: Verticillium dahliae ( Verticillium dahliae WX-1, 9: Fusarium oxysporum ( F. oxysporum f. sp. lycopersici FQ143, 10: Wheat leaf tissue DNA, 11: H2O, 12: Stranded rust fungus ( P.striiformis f. sp. tritici ).
[0063] Figure 5 The results are from routine PCR detection of the primer sensitivity of *Strombus styracifolius*. Detailed Implementation
[0064] I. Terminology in this application: Examples of resources describing many of the molecular biology-related terms used in this article can be found in the following literature: Alberts et al., Molecular Biology of The Cell, 5th ed., Garland Science Publishing, Inc.: New York, 2007; Rieger et al., Glossary of Genetics: Classical and Molecular, 5th ed., Springer-Verlag: New York, 1991; King et al., A Dictionary of Genetics, 6th ed., Oxford University Press: New York, 2002; and Lewin, GenesIX, Oxford University Press: New York, 2007.
[0065] Any references cited in this article, including, for example, all patents, published patent applications and non-patent publications, are incorporated in their entirety by reference.
[0066] For ease of understanding this application, several terms and abbreviations used herein are defined as follows: When used in a list of two or more items, the term "and / or" means that any of the listed items can be used alone or in combination with any one or more of the listed items. For example, the expression "A and / or B" is intended to mean either or both of A and B, i.e., A alone, B alone, or a combination of A and B. The expression "A, B and / or C" means A alone, B alone, C alone, a combination of A and B, a combination of A and C, a combination of B and C, or a combination of A, B and C.
[0067] The term "comprising" is not intended to be restrictive, but rather inclusive and implies the presence of other elements besides those listed, and can be interpreted as "including but not limited to". The term "comprising" also encompasses the terms "consisting of" and "substantially consisting of". In this document, the terms "including" and "comprise" are used interchangeably.
[0068] A species, also called a genus, is the basic unit of biological classification. It refers to a group of organisms that can recognize each other as potential mating partners and share a common mating recognition system. According to biological taxonomy, a species is located at the lowest level of the taxonomic hierarchy, below the genus. It is usually composed of a group of morphologically similar organisms that are capable of mating and producing fertile offspring.
[0069] II. Implementation Examples The present application will now be described in further detail with reference to specific embodiments. The embodiments given are merely illustrative of the present application and are not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the present application in any way.
[0070] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.
[0071] Unless otherwise specified, the data processing tools involved in the following examples are all based on default parameters.
[0072] Pantoea dispersa is owned by the applicant and is described in the literature "First report of Pantoea dispersacausing strawberry root rot in China", J. Wang, YF Wang, JY Wang, HYWu, XF Zhang, and JH Guo. Plant Disease. 2023, 107(12):4031. DOI: 10.1094 / PDIS-04-23-0638-PDN, PubMed ID: 37401551. Bacillus berleis and Fusarium oxysporum are owned by the applicant and are described in the literature "Effects of ComX on the biocontrol traits of Bacillus berleis B31 and its inhibitory activity against Fusarium wilt of tomato", Bulletin of Microbiology. 2023, 50(8):3392-3405. DOI: 10.13344 / j.microbiol.china.220932. Bacillus amyloliquefaciens is owned by the applicant and is described in the literature "Antibacterial substances of Bacillus amyloliquefaciens HMB33604 and their control effect on potato black scurf", Shen Mengjie, Wang Mengliang, Tan Yinghui, et al. Chinese Journal of Biological Control. 2023, 39(1): 132-141.10.16409 / j.cnki.2095-039x.2022.03.015. Fusarium moniliforme and Fusarium nucleatum are owned by the applicant and are described in the literature "Early detection of southern rust in maize based on PCR and nested PCR technology", Yang Xiaonan, Gao Peng, Ji Lijing, et al. Maize Science. 2023, 31(4): 172-178.10.13597 / j.cnki.maize.science.20230423. The wheat powdery mildew fungus is owned by the applicant and is described in the literature "The effect of reducing application and increasing efficacy of the mixture of amino oligosaccharide and pyraclostrobin in the control of wheat powdery mildew", Li Dan, Wang Yan, Lu Junjiao, et al. Pesticides. 2023, 62(6): 441-444.10.16820 / j.nyzz.2023.6002. Rhizoctonia solani AG-1 and AG-4 are owned by the applicant and are described in the literature "Study on the mycelial fusion group of Rhizoctonia solani in cotton in Hebei Province and its pathogenicity". Sun Mengwei, Wang Xiaonan, Zhang Yalin, et al. Cotton Science. 2022, 34(5):425-435.10.11963 / cs20220045. Fusarium graminearum is owned by the applicant and is described in the literature “Effects of fludioxonil and tebuconazole combined with the growth of Fusarium graminearum mycelium and the diseases caused by it”, Liu Di, Li Congcong, Yuan Hongxia, et al. Acta Phytopathologica Sinica. 2022, 52(6): 1003-1011.10.13926 / j.cnki.apps.000705.Wheat leaf rust and stripe stalk rust were provided by the Plant Disease Epidemiology Laboratory of China Agricultural University and are described in the literature "Study on the Influence of Different Concentrations of Mixed Strains on the Development of Wheat Rust". Verticillium dahliae is owned by the applicant and is described in the literature "Preliminary Analysis of Antimicrobial Proteins of Bacillus subtilis NCD_2 Strain", Guo Qinggang, Lu Xiuyun, Li Shezeng, et al. North China Journal of Agricultural Sciences. 2008, 23(5): 191-194.10.7668 / hbnxb.2008.05.042. The public can apply to obtain the above biological materials from the applicant. The obtained materials can only be used for the verification of the technical solution of this application and cannot be used for other purposes.
[0073] All experimental materials used by the applicant were legally sourced, all operations were legally performed, and all tools used were legally sourced.
[0074] Example 1: A method for designing primers and probes for real-time quantitative PCR targeting pathogenic microorganisms. A method for designing primers and probes for targeted microbial real-time quantitative PCR includes the following steps: Step 1: Collect target microbial genome data, screen and establish a high-quality target microbial genome database, and determine the best reference genome; Step 2: Based on the target microbial genome database established in Step 1, screen to obtain the set of conserved genes of the target microorganisms; Step 3: Collect genomic data of target microorganisms similar to other species, and screen and construct a high-quality genomic database of target microorganisms similar to other species; Step 4: Combine the optimal reference genome of the target microorganism determined in Step 1 with the target microorganism similar species genome database established in Step 3 for joint analysis, and screen to obtain the unique gene set of the target microorganism. Step 5: Combine the conserved gene set of the target microorganism obtained in Step 2 with the target microorganism-specific gene set obtained in Step 4 for joint analysis, and screen to obtain target microorganism gene sequences that have both conservation and specificity.
[0075] Step 6: Use primer design software to design primers and probes for the target gene sequences obtained in Step 5. Based on the penalty score of Primer3 and the coverage and specificity of primers and probes, theoretically feasible primer and probe combinations are selected. Finally, the optimal combination is determined by experimental verification.
[0076] Specifically, the methods for collecting and screening target microbial genome data that can be used to establish a high-quality target microbial genome database in step one are as follows: 1) Prioritize downloading genomic data from the NCBI RefSeq database; if the data in the RefSeq database is insufficient, download data from the GenBank database. 2) Referring to the NCBI website's evaluation of genome integrity, discard genome data with an integrity score of less than 90%; 3) Prioritize genomic data with a complete assembly degree. If the number of complete data is insufficient, select chromosome, scaffold, and conting data in that order.
[0077] The method for determining the best reference genome data in step one is as follows: firstly, select whole genome sequencing data from the type strain (From type material) in the RefSeq database of NCBI database; if there is no genome data that meets the criteria, determine the best reference genome in the RefSeq database according to the priority order: reference genomes > complete > chromosome > scaffold > conting. If there is no genome data that meets the criteria in the RefSeq database, determine the best reference genome in the GenBank database according to the above priority.
[0078] Further, in step two, the conserved gene set of the target microorganism is obtained by screening the target microorganism genome database established in step one. The specific method used is to use Othofinder 2.5 software for comparative genomic analysis to determine the orthologous genome of the target microorganism.
[0079] Furthermore, in step three, the target microbial similar species are mainly microorganisms that belong to the same genus as the target microorganism in taxonomy but are different species. The database genome screening principle is the same as in step one.
[0080] Furthermore, step four involves obtaining the target microorganism-specific genes by adding the optimal reference genome data of the target microorganism from step one to the approximate species genome database established in step three, and then using Orthofinder 2.5 software for comparative genomics analysis to obtain the target microorganism-specific genome. The specific genes consist of two parts: one part is the target microorganism's own homologous genes, i.e., multi-copy genes unique to the target microorganism; the other part is genes in the target microorganism that are not classified into orthologous groups (UnassignedGenes).
[0081] Furthermore, in step five, gene sequences with a length greater than 300 bp that exist simultaneously in the target microorganism orthologous genome obtained in step two and the target microorganism-specific genome obtained in step four are the specific gene sequences that can be used for primer design.
[0082] Furthermore, the primer design conditions used in step six are as follows: amplified product length 80-120 bp; set parameters: primer annealing temperature 52-60℃, length 18-23 nt, GC content 30%-70%; probe annealing temperature 62-70℃, length 13-30 nt, GC content 20%-80%. Preferred primer annealing temperatures are 56℃, primer length 20 nt, and GC content 50%; probe annealing temperatures are 66℃, length 20 nt, and GC content 50%.
[0083] Step Six: Further screen the designed primers according to the following criteria: 1. Primers were designed in batches for the identified target genes using qprimer software developed with primer3 as the kernel. The designed primers were sorted in order of increasing penalty score calculated by primer3. The top 10 primer and probe pairs with the lowest penalty scores were selected for coverage and specificity testing.
[0084] Penalty score = Forward primer penalty score + Reverse primer penalty score + Probe penalty score 2. The method for evaluating primer and probe coverage in step six is as follows: Use BLAST software to compare the amplicon sequences or probe sequences obtained by primer amplification with the high-quality target microbial genome database established in step one, and calculate the proportion of the number of qualified genomes that the amplicon or probe can match to the total number of genomes in the genome database.
[0085] Primer coverage = (Number of genomes with valid amplicon matches / Total number of genomes in the high-quality target microbial genome database) × 100% Probe coverage = (Number of genomes with valid probe sequence matches / Total number of genomes in the high-quality target microbial genome database) × 100% The amplicon referred to here is the nucleic acid region defined and amplified by the forward and reverse primers during PCR. The start end is defined by the 5' end of the forward primer, and the end end is defined by the 5' end of the reverse primer. A successful match means that the following conditions must be met simultaneously: the amplicon alignment matching length is not less than 95% of the total amplicon length, the alignment accuracy is not less than 95%, both the forward and reverse primers can achieve a match, and the 3' ends are not mismatched.
[0086] The standard for qualified probe sequence matching mentioned here is that the probe sequence is completely matched, that is, the entire probe length is completely matched and the accuracy rate is 100%.
[0087] 3. The specificity evaluation method described in step six is to use the primer-blast function provided by NCBI to compare the candidate primer pairs in the nt database. The standard for qualified specificity evaluation is that the primer pairs whose blast-matched amplification products are all gene sequences of the target microorganism and whose amplicon length matches the design value.
[0088] Step six describes a method for evaluating probe specificity by performing a BLAST alignment of the probe sequence in the NCBI nt database. The criterion for passing the probe specificity evaluation is that all alignment results match the gene sequences of the target microorganism.
[0089] Example 2: Using dispersed pantothecin ( Pantoea dispersa Using the target microorganism as an example, explain the design and validation process of specific primers. 2.1 Design and Validation Process of Specific Primers This embodiment uses Panamax, a bacterium that causes root rot in Chinese cabbage and strawberry. Pantoea dispersa For the target microorganism, explain in detail the design process of specific primers.
[0090] 1. Collect genomic data of target species (microorganisms), screen and establish a high-quality target microbial genome database, and determine the best reference genome.
[0091] (1) Collect target microbial genome data A total of 71 whole genome sequences of *Umbrella dispersae* were retrieved from the RefSeq database and GeneBank database on the NCBI website (https: / / www.ncbi.nlm.nih.gov / ) (retrieval date: May 31, 2025). The specific method was to select "Genome" as the search resource type on the NCBI website homepage and enter the Latin name of the target microorganism in the search box. Pantoea dispersa ".
[0092] Prioritize downloading the CDS data (58 genome data files) of 58 genomes from the RefSeq database. Specifically, the database source and its number can be displayed by checking "RefSeq" and "GenBank" on the "Select columns" page.
[0093] The RefSeq numbers of the 58 sequences are: GCF_019890955.1, GCF_009362975.1, GCF_022220865.1, GCF_019890975.1, GCF_037039485.1, GCF_025853815.1, GCF_028993655.1, GCF_040207715.1, GCF_025919605.1, GCF_025490495.1, GCF_047325125.1, GCF_018798945.1, GCF_018798925.1, GCF_014155765.1, GCF_008692915.1, GCF_031453715.1, GCF_004009945.1, GCF_035783155.1, GCF_040065815.1, GCF_037145435.1, GCF_025882115.1, GCF_025882125.1, GCF_025882675.1, GCF_037145515.1, GCF_037144915.1, GCF_037144975.1, GCF_037145415.1, GCF_037145455.1, GCF_037145495.1, GCF_041386385.1, GCF_046638605.1, GCF_006546345.1, GCF_030017095.1, GCF_025245605.1, GCF_035783935.1, GCF_001476295.1, GCF_025460745.1, GCF_001476715.1, GCF_001477165.1, GCF_001476745.1, GCF_039710145.1, GCF_047395865.1, GCF_001476735.1, GCF_022627855.1, GCF_037150095.1, GCF_001477155.1, GCF_001476765.1, GCF_001476325.1, GCF_018257675.1, GCF_030056125.1, GCF_001477195.1, GCF_032334275.1, GCF_037150035.1, GCF_037150215.1, GCF_018257625.1, GCF_003936175.1, GCF_009866445.1, GCF_000465555.2。
[0094] (2) Based on the NCBI website's evaluation of genome integrity, sequences with integrity higher than 95% were selected. All 58 sequences meet the genome integrity requirement (integrity greater than 95%), and together they constitute a high-quality dispersed pantothenic genome database. Specifically, you can check "CheckMcompleteness" in the "Select columns" section of the page to display the sequence integrity.
[0095] (3) Determining the optimal reference genome: Prioritize the whole genome sequence of the type species as the reference sequence, and then select sequences with higher assembly levels based on their assembly degree. Prioritize genome data with a complete assembly level. If the number of complete data is insufficient, select chromosome, scaffold, and conting data in that order. Specifically, check "From type material" in the "Filters" section to obtain the genome sequence of the type strain. In this example, there are two whole genome sequences from the type species in the *Ureaplasma dispersans* genome. Check "Level" in the "Select columns" section to display the assembly degree of the sequences. Select the genome sequence with a higher assembly level (Scaffold) as the optimal reference genome, with the genome number GCF_014155765.1.
[0096] 2. Screening to obtain the conserved gene set of target microorganisms The high-quality genome database was analyzed using Orthofinder 2.5 software. A total of 3399 homologous genomes coexisting in 58 strains of dispersed pantothenia were obtained and merged to form a set of conserved genes of dispersed pantothenia. Specifically, the steps for obtaining homologous genomes using Orthofinder 2.5 software were as follows: (1) Enter orthoFinder -f <target microorganism genome database file name> -d -og in the command line to run the program; (2) Select the file Orthogroups.tsv in the results folder "Orthogroups" to filter and obtain the file names of orthologous genomes commonly contained in all species; (3) Merge the genome files obtained in the previous step in the results folder "Orthogroup Sequences" to obtain the set of conserved genes of the target microorganisms.
[0097] Conserved genes are a class of genes defined and output by Orthofinder based on the distribution patterns of orthologs across all species, and whose DNA sequences are highly similar or remain unchanged in target microorganisms.
[0098] 3. Construction of a database of similar species genomes Search the NCBI database (specifically, select "Genome" in the search resource type on the NCBI website homepage, and enter the Latin genus name of the target microorganism in the search box). Pantoea” A total of 1130 genome records for the genus *Pantheraea* were found in the Refseq database. The data was then filtered to remove entries with ambiguous classifications (e.g., those with "organism_name"). Pantoea The genome sequences of sp. were obtained, and then, referring to the NCBI website's evaluation of genome integrity, genome data with integrity less than 95% were discarded, and genome data with assembly degree of completeness were selected first. Finally, a total of 78 fully assembled whole genome sequences from 17 species of the genus Pantotheca were selected to form the Dispersed Pantotheca Approximate Species Genome Database.
[0099] 4. Extracting unique gene sets The best reference genome of *Ureaplasma dispersans* was added to the near-species genome database for analysis using Orthofinder 2.5 software, and the unique gene set of the best reference genome was obtained through screening.
[0100] The steps for obtaining the unique gene set of the best reference genome using Orthofinder 2.5 software are as follows: (1) Run the Orthofinder software; (2) Obtain the set of multi-copy genes unique to the target microorganism (hereinafter referred to as specific gene set 1 for convenience), select the file Orthogroups.tsv in the results folder "Orthogroups" to filter only the genome numbers unique to the target microorganism; merge the filtered genome files in the results folder "Orthogroup Sequences" to obtain specific gene set 1. In this example, Pantothenia gravis has 7 unique multi-copy genes, totaling 17 genes; (3) Obtain the set of single-copy genes unique to the target microorganism (hereinafter referred to as specific gene set 2): select the file Orthogroups_UnassignedGenes.tsv in the results folder "Orthogroups" to filter the file names of the genomes unique to the target microorganism, merge the filtered genome files in the results folder "Orthogroup Sequences" to obtain specific gene set 2. In this example, Pantothenia gravis has 801 genes.
[0101] The specific gene is a DNA sequence specific to the optimal reference genome of the target microorganism, defined and output by Orthofinder based on the distribution pattern of orthologs across all species. In this embodiment, the specific gene is a DNA sequence specific to the strain with genome number GCF_014155765.1.
[0102] Since using multi-copy genes as targets to design primers can improve primer sensitivity, such genes should be given priority in primer design while ensuring conservation. Considering the advantages of multi-copy genes, this invention will perform conservation comparisons and primer design for multi-copy genes separately with other genes in subsequent steps, so that primers and probes designed using multi-copy genes as templates are preferentially selected.
[0103] 5. Extract target gene sequences that possess both conservation and specificity. The specific gene set 1 and specific gene set 2 obtained in step "4. Extracting the specific gene set" were compared with the conserved gene set in step "2. Screening to obtain the conserved gene set of the target microorganism". In this example, there was no matching result between the specific gene set 1 of Pantotheca dispersa and its conserved gene set; the specific gene set 2 of Pantotheca dispersa and the conserved gene set had 266 matching genes. After screening and removing genes with a length of less than 300 bp from the matching genes, 194 target genes were finally obtained and can be used for primer design in the next step.
[0104] Specifically, the steps for comparing specific gene set 1 and specific gene set 2 with the conserved gene set are as follows: (1) Use the seq -n command in the seqkit software toolkit to extract the names of all gene files in specific gene set 1 and specific gene set 2 respectively; (2) Use the grep command in the seqkit software toolkit to filter gene files containing the file names obtained in the previous step in the conserved gene set file in step "2. Screening to obtain the conserved gene set of the target microorganism". The result is the target gene sequence that has both conservation and specificity; (3) Use the seq -m command in the seqkit software to filter and remove genes with a length of less than 300bp obtained in the previous step.
[0105] 6. Primer design and screening Primers and probes for the target gene sequences obtained in step 5 were designed using the qPrimer software (https: / / github.com / swu1019lab / qPrimer) based on Primer3. Primers were annealed at 52-60℃, with a length of 18-23 nt and a GC content of 30%-70%. Probes were annealed at 62-70℃, with a length of 13-30 nt and a GC content of 20%-80%. The annealing temperature was predicted using the `calculate_tm()` function in Primer3.
[0106] Specifically, qPrimer is used to design primers and probes for the target gene sequence obtained in step 5. After running the program, or by writing the input command `qPrimer design –seq_file<target gene sequence obtained in step 5>--ini_file<primer design parameter file>--csv –out_name<output primer design result file name>`, the columns for forward primer penalty value, reverse primer penalty value, and probe penalty value are found in the output result file. These are the columns named "PENALTY_F", "PENALTY_R", and "PENALTY". The penalty value for each primer-probe combination is calculated using the formula: Penalty Value = Forward Primer Penalty Value + Reverse Primer Penalty Value + Probe Penalty Value. The top 10 primer pairs with the lowest penalty values are then evaluated for coverage and specificity and selected. The primer sequences used for primer coverage and specificity evaluation are shown in Table 1 below.
[0107] The penalty values for each primer were calculated using the software's default parameters.
[0108] Table 1. Primer and probe sequence information for *Plasmodium dispersans* used for specificity and coverage evaluation.
[0109] The coverage evaluation method is as follows: The coverage of the amplicon or probe sequences of the candidate primer pairs in the high-quality target microbial genome database established in step 1 is calculated. The amplicon or probe sequences are then compared with those in the high-quality target microbial genome database established in step 1 using BLAST software, and the proportion of genomes that can be matched by the amplicon or probe to the total number of genomes in the database is calculated.
[0110] The conditions for qualified coverage matching are: both upstream and downstream primers can match the target genome; amplicon matching length ≥ 95% of total amplicon length; matching similarity ≥ 95%; and Evalue < 1e-5, i.e., 0.00001.
[0111] Primer coverage = (Number of genomes with valid amplicon matches / Total number of genomes in the high-quality target microbial genome database) × 100% Probe coverage = (Number of genomes with valid probe sequence matches / Total number of genomes in the high-quality target microbial genome database) × 100% Specifically, taking the coverage of the amplicon or probe sequence of the primer pair corresponding to Padis7 in the high-quality target microbial genome database established in step 1 as an example, the evaluation steps of coverage are as follows: (1) Use ncbi-blast2.17 software to establish a local database for the high-quality target microbial genome database established in step 1; (2) Perform BLAST alignment on the candidate primer pairs and probes in the database established in the previous step, and the alignment results are shown in Table 2.
[0112] The specificity evaluation method is as follows: For primer specificity, the candidate primer pairs are compared in the nt database using the Primer-BLAST function provided by NCBI. Primer pairs whose amplification products are all target microorganisms and whose amplicon length matches the design value are selected. For probe specificity, the probe sequences are compared in the nt database using the BLAST program. Alignment results with both 100% query coverage and percentage identity must all be target microorganisms.
[0113] Specifically, taking the primer specificity analysis method of Padis7 as an example, the primer specificity analysis steps are as follows: Using the Primer-BLAST function on the BLAST page of the NCBI website (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi), enter the candidate primer pair sequences in the "Primer Parameters" tab, and modify the parameter settings in the "Primer Pair Specificity Checking Parameters" tab: select "nt" for the "Database" option, leave "Organism" blank, and finally click "Get Primers". Check the results page to ensure that all aligned species are target microorganisms and that the amplicon length matches the design value. The alignment results are shown in Table 2.
[0114] Taking the probe specificity analysis method of Padis7 as an example, the steps of probe specificity analysis are as follows: Access the BLAST page on the NCBI website (https: / / blast.ncbi.nlm.nih.gov / Blast.cgi), click the "Nucleotide BLAST" tab, enter the candidate probe sequences in the "Enter Query Sequence" tab, select "Standard Database" for "Database" in "Choose Search Set," and then click "BLAST." On the results page, all alignments with 100% query coverage and percentage identity must be for the target microorganism. The specific results in this example are shown in Table 2 below. Table 2. Evaluation results of coverage and specificity of dispersive pantothenic acid primers and probes.
[0115] 2.2 Validation of the sensitivity and specificity of specific primers In this embodiment, the sensitivity and specificity of the primers designed in Example 1 were verified by conventional PCR and TaqMan real-time quantitative PCR.
[0116] 2.2.1 Verification of specificity and sensitivity using conventional PCR experiments (1) Verification of specificity using conventional PCR test The specificity of the selected primers was verified using conventional PCR. The reaction volume was 50.0 μL, including 1.0 μL of DNA template (genomic DNA from each sample in Table 3), and forward and reverse primers (10 μmol·L⁻¹). -1 1.0 μL each of the following: 2×Ex Taq Mix (Dalian Takara Bio Co., Ltd.), 25.0 μL, and 22.0 μL of ddH2O. Reaction program: 95℃ pre-denaturation for 5 min; 95℃ denaturation for 20 s, 58℃ annealing for 10 s, 72℃ extension for 10 s, for a total of 35 cycles; final extension at 72℃ for 5 min. PCR products were detected by 2% agarose gel electrophoresis (120 V, 30 min), and the electrophoretic bands were observed and analyzed using a gel imaging system.
[0117] Dispersed pantothenic acid (PPAN) results showed that primer pairs numbered padis3 and padis7 did not exhibit nonspecific amplification. Amplification results are shown in Table 3 and the appendix to the instruction manual. Figure 2 Further synthesis of TaqMan probes can be used for the detection of fluorescence quantitative reactions.
[0118] Table 3 Results of conventional PCR screening of dispersible pantothenia
[0119] (2) Verification of sensitivity using conventional PCR To further screen for more sensitive specific primer pairs and synthesize corresponding probes, the sensitivity of each primer pair for Padis3 and Padis7 was verified using conventional PCR. Target genomic DNA was extracted, and after determining the DNA concentration using a Nanodrop ND-2000, the DNA concentration was adjusted to 50 ng / μL. -1 10-fold serial dilution to a concentration of 0.05 pg·μL -1 -50 ng·μL -1 Using the DNA at the aforementioned gradient concentrations as templates, conventional PCR was performed to test primer sensitivity.
[0120] Both the Padis3 and Padis7 have been verified to have a sensitivity of 5 pg·μL. -1 However, the primer pair Padis7 was effective at 5 pg·μL. -1 The amplified bands were clearer, and the corresponding probe was finally synthesized using Padis7. The detection results are attached to the instruction manual. Figure 3 .
[0121] 2.2.2 Verification of sensitivity using quantitative real-time PCR assay The quantitative real-time PCR experiment was validated. The quantitative real-time reaction system was 25.0 μL, containing 5.0 μL of 0.1% BSA and 10× Buffer (Mg). 2+ free)2.5 μL, MgCl23.0 μL, dNTP mix(2.5 mmol·L -1 2.0 μL, rTaq(5U·μL) -1 0.5 μL, primers padis7-F / padis7-R (10 µmol·L⁻¹) -1 1.0 μL each, probe padis7-P (10 µmol·L⁻¹) -1 1.0 μL of DNA template and 5.0 μL of ddH2O were added to make up the remaining volume. All reagents used in the system were from Dalian Takara Bio Co., Ltd. The reaction program was: 95℃ pre-denaturation for 5 min; 95℃ denaturation for 30 s, 58℃ annealing for 45 s, for 45 cycles. The standard for passing TaqMan quantitative PCR was a Ct value of less than or equal to 35 for positive templates, and a Ct value of greater than or equal to 35 or no Ct value for control and blank templates.
[0122] The results of TaqMan real-time PCR for dispersible pantothenia are shown in Table 4. The results show that only in the reaction system containing dispersible pantothenia was the Ct value 18.52; in other reaction systems, the Ct value was greater than 35 or there was no Ct value. This indicates that the padis7 primer and probe combination selected in this example can specifically detect the target microorganism, meeting practical needs.
[0123] Table 4. Results of Padis7 TaqMan real-time PCR using primer and probe combinations
[0124] The method provided by this invention screened and obtained a specific primer and probe composition for targeting dispersed pantothenia, designated Padis7. The target gene of Padis7 is designated NZ_SCKT010000010.1_cds_WP_021506876.1_373, and contains one copy in the target microorganism.
[0125] The sequences of the specific primer-probe composition numbered padis7 are as follows: padis7-F: GACACAACGGTGGTTATGGTT (SEQ ID NO: 1); padis7-R: AAACTCCTGTGGTGACAATG (SEQ ID NO: 2); padis7-P: TGGCGGTTTTATGCAGCAAATGGAGCA (SEQ ID NO: 3).
[0126] The first T at the 5' end of the probe is modified with 6-FAM; the first A at the 3' end is modified with TAMRA-N.
[0127] Example 3: Illustrate the design and validation process of specific primers using *Puccinia striiformis* f. sp. *tritici* as an example. This embodiment uses the rust fungus *Strombus styracifolius*, which causes wheat stripe rust, as an example. Puccinia striiformis f.sp. tritici For the target microorganism, explain in detail the design process of specific primers.
[0128] 3.1. Stranded rust fungus ( Puccinia striiformis f. sp. tritici For the target microorganism, explain in detail the design process of specific primers.
[0129] 1. Construction of high-quality genome databases.
[0130] (1) The genome sequences of *Strombus styracifolius* were retrieved from the Refseq database on the NCBI website (https: / / www.ncbi.nlm.nih.gov / ) and the GeneBank database. Only one *Strombus styracifolius* sequence was found in the Refseq database, while eight sequences were found in the GeneBank database. Finally, the CDS data of nine genomes were downloaded to construct a high-quality *Strombus styracifolius* genome database. The sequence reference numbers for the nine genomes are GCA_001191645.1, GCA_002920065.1, GCA_002920205.1, GCA_021901655.1, GCA_025169535.1, GCA_025169555.1, GCA_025869475.1, GCA_025869495.1, and GCF_021901695.1.
[0131] The operation method is the same as in Example 2, except that: (i) Change the Latin name entered in the search box to " Puccinia striiformis f. sp. tritici ”; (ii) When the Refseq database is insufficient, supplement it with genomic data from the GeneBank database.
[0132] (2) Determine the best reference genome: After screening using the “From type material” option, the genome of *Stripetra rust* in the Refseq database was selected as the best reference genome, numbered GCF_021901695.1.
[0133] 2. Extracting the set of conserved genes Following the method described in Implementation 2, the constructed high-quality genome database was analyzed using Orthofinder 2.5 software. A total of 5173 homologous genomes coexisting in 9 strains of *Striga styracifolium* were obtained and merged to form a conserved gene set for *Striga styracifolium*.
[0134] 3. Construction of a database of similar species genomes Following the method described in Example 2, genomic data for the genus *Stemonae* were queried from the NCBI database and the Refseq database. The data was then filtered, removing entries with ambiguous classifications ("organism_name"). Puccinia The genome sequence of sp. was obtained, and then, referring to the NCBI website's evaluation of genome integrity, genome data with integrity less than 95% were discarded, and genome data with assembly degree of completeness were selected first. Finally, 35 whole genome sequences were obtained from the genome data of the genus Stylostella to form a genome database of similar species.
[0135] 4. Extracting unique gene sets Following the method described in Example 2, the optimal reference genome of *Striga styracifolium* was added to a database of similar species genomes for analysis using Orthofinder 2.5 software. This process screened for genes unique to the optimal reference genome. This gene set consists of two parts: one part is the homologous genome unique to the target bacterium (hereinafter referred to as specific gene set 1 for convenience), in which *Striga styracifolium* has 2415 unique multicopy genes, totaling 7063 genes; the other part is the unassigned genes unique to the target bacterium that were not included in the homologous genome (hereinafter referred to as specific gene set 2), in which *Striga styracifolium* has 5311 genes in this example.
[0136] 5. Extract target gene sequences that possess both conservation and specificity. Referring to the method in Example 2, the specific gene set 1 and specific gene set 2 obtained in step 4 were compared with the conserved gene set, respectively. In this example, the specific gene set 1 of *Rhizoctonia solani* and its conserved gene set have 2611 shared genes. After screening and removing genes with a length of less than 300 bp, 2488 target genes were finally obtained, which can be used for primer design in the next step. The specific gene set 2 of *Rhizoctonia solani* and its conserved gene set have 1907 shared genes. After screening and removing genes with a length of less than 300 bp, 1825 target genes were finally obtained, which can be used for primer design in the next step.
[0137] Since using multi-copy genes as targets to design primers can improve primer sensitivity, such genes should be given priority in primer design while ensuring conservation. Considering the advantages of multi-copy genes, this invention will perform conservation comparisons and primer design for multi-copy genes separately with other genes in subsequent steps, so that primers and probes designed using multi-copy genes as templates are preferentially selected.
[0138] 6. Primer design and screening Following the method described in Example 2, the qprimer software (https: / / github.com / swu1019lab / qPrimer), developed based on Primer3, was used to design primers and probes for the 2488 target gene sequences obtained from the *Rhizoctonia solani*-specific gene set 1 and its conserved gene set obtained in step 5. The primer design results were then output in ascending order of penalty score.
[0139] Following the method described in Example 2, the top 10 primer pairs with the lowest penalty scores were evaluated for coverage and specificity, and then screened. The primer sequences used for coverage and specificity evaluation are shown in Table 5 below. The primer coverage and specificity results in this example are shown in Table 6 below.
[0140] Table 5 Primer and probe sequence information for *Rhizoctonia solani* used for specificity and coverage evaluation.
[0141] Table 6. Evaluation results of primer and probe coverage and specificity for *Rhizoctonia solani*.
[0142] 3.2 Verification of the specificity and sensitivity of primers and probes The sensitivity and specificity of the primers designed in 3.2 were verified by conventional PCR and TaqMan real-time quantitative PCR, following the method described in 2.2 of Example 2.
[0143] Results for *Stripetrata rust* showed that primer pair Pst6 did not provide nonspecific amplification. Amplification results are shown in Table 7 and the instruction manual appendix. Figure 4 It can be further synthesized into TaqMan probes for detection in real-time quantitative PCR reactions.
[0144] The sensitivity of the Pst6 primer is 500 pg / μL. -1 The test results are attached to the instruction manual. Figure 5 .
[0145] The TaqMan real-time PCR results for *Strombus styracifolius* are shown in Table 8. The results show that the Ct value was 28.21 only when the sample was *Strombus styracifolius*. In other reaction systems, the Ct value was greater than 35 or there was no Ct value. This indicates that the Pst6 primer and probe combination selected in this example can specifically detect the target microorganism, meeting practical needs.
[0146] Table 7 Results of routine PCR screening of *Rhizoctonia solani*
[0147] Table 8. Primer and probe combination pst6 TaqMan real-time PCR results
[0148] The target gene of the *Strombus styracifolius*-specific primer-probe composition, numbered Pst6, obtained by screening using the method in Example 1, is NC_063031.1_cds_XP_047807403.1_6305, which has 3 copies in *Strombus styracifolius*. Its nucleotide sequences are as follows: pst6-F:GCTCAGCAATTGAAGCAATC (SEQ ID NO:4); pst6-R: AGCCTCAATTTGTTGGGTTT (SEQ ID NO: 5); pst6-P: CAAACAGGCCAATTCCGCTCC (SEQ ID NO: 6).
[0149] The first C at the 5' end of the probe is modified with 6-FAM; the first C at the 3' end is modified with TAMRA-N.
[0150] The present application has been described in detail above. Those skilled in the art will recognize that the present application can be implemented in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. Although specific embodiments are given in this application, it should be understood that further modifications can be made to the present application. In summary, in accordance with the principles of this application, this application is intended to include any changes, uses, or improvements to the present application, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.
Claims
1. A method for designing a primer probe combination or a primer combination based on pan-genomics, characterized by: The method comprises the following steps: S1) Collecting genomic data of target microorganisms to construct a target microorganism genomic database, and screening an optimal reference genome from the target microorganism genomic database; S2) Analyzing the target microorganism genomic database to screen a target microorganism conserved gene set; S3) Collecting genomic data of target microorganism proximate species to construct a target microorganism proximate species genomic database; S4) Jointly analyzing the optimal reference genome of S1) and the target microorganism proximate species genomic database of S3) to screen a target microorganism specific gene set; S5) Taking the intersection of the target microorganism conserved gene set of S2) and the target microorganism specific gene set of S4) as a target gene; S6) Using primer design software to design primers and probes for the target gene sequence screened in S5), and screening a theoretically feasible primer / probe combination according to the penalty value of the primer design software, the coverage of the primer or probe, and the specificity of the primer or probe, and determining the final primer / probe combination by combining experimental verification.
2. The method of claim 1, wherein: The collection method of the genomic data of the target microorganism in S1) is as follows: S11) Preferentially selecting genomic data downloaded from the NCBI Refseq database, and selecting data downloaded from the GenBank database when the data in the Refseq database is insufficient; S12) Discarding genomic data with an integrity less than 90 according to the evaluation of the integrity of the genome on the NCBI website; S13) Preferentially selecting genomic data with an assembly degree of complete, and sequentially selecting data with an assembly degree of chromosome, scaffold, and contig if the number of complete data cannot meet the requirement; Or / and The determination method of the optimal reference genome data in S1) is as follows: S101) Preferentially selecting data from type material in the RefSeq database in the NCBI database, and determining the optimal reference genome according to the priority order reference genomes > complete > chromosome > scaffold > contig; S102) Selecting data from non-type material (not marked "From type material" in the database) and determining the optimal reference genome according to the priority order reference genomes > complete > chromosome > scaffold > contig if there is no data meeting the condition of S101).
3. The method of claim 1, wherein: The target microorganism conserved gene set in S2) is obtained by comparative genomics analysis of the target microorganism genomic database.
4. The method of claim 1, wherein: The proximate species is a microorganism that is taxonomically the same genus but different species as the target microorganism, and the database genomic screening principle is the same as S1) in claim 1 or 2.
5. The method of claim 1, wherein: S4) the method for obtaining the target microorganism-specific genes in S4) is to add the target microorganism optimal reference genome data in S1) into the approximate species genome database established in S3) to perform comparative genomics analysis, and obtain the combination of the target microorganism-specific genes.
6. The method of claim 1, wherein: The target genes in S5) satisfy the following conditions: the length is greater than 300 bp, and the target genes are simultaneously present in the target microorganism-conserved gene set in S2) and the specific gene set in S4).
7. The primer probe combination or primer combination obtained by the method according to any one of claims 1 to 6, wherein the primer combination comprises a forward primer and a reverse primer; and the primer probe combination comprises a forward primer, a reverse primer and a probe.
8. The primer probe composition or primer composition according to claim 7, characterized in that: The target microorganism is Pantoea dispersa, The forward primer is a single-stranded DNA with a nucleotide sequence as shown in SEQ ID NO: 1; The reverse primer is a single-stranded DNA with a nucleotide sequence as shown in SEQ ID NO: 2; The nucleotide sequence of the probe is as shown in SEQ ID NO: 3; Or The target microorganism is Puccinia horiana, and the primer probe combination or primer combination is specific to the Puccinia horiana; The forward primer is a single-stranded DNA with a nucleotide sequence as shown in SEQ ID NO: 4; The reverse primer is a single-stranded DNA with a nucleotide sequence as shown in SEQ ID NO: 5; The nucleotide sequence of the probe is as shown in SEQ ID NO:
6.
9. The use of any one of the following: D1), the use of the primer probe combination or primer combination specific to the Pantoea dispersa according to claim 8 in at least one of the following: D1-1) for detecting the Pantoea dispersa; D1-2) for preparing a product for detecting the Pantoea dispersa; D1-3) for identifying or assisting in identifying the Pantoea dispersa; D1-4) for preparing a product for identifying or assisting in identifying the Pantoea dispersa; D1-5) for identifying or assisting in identifying whether a test sample is or contains the Pantoea dispersa; D1-6) for preparing a product for identifying or assisting in identifying whether a test sample is or contains the Pantoea dispersa; D2), the use of the primer probe combination or primer combination specific to the Puccinia horiana according to claim 8 in at least one of the following: D2-1) for detecting the Puccinia horiana; D2-2) for preparing a product for detecting the Puccinia horiana; D2-3) for identifying or assisting in identifying the Puccinia horiana; D2-4) for preparing a product for identifying or assisting in identifying the Puccinia horiana; D2-5) for identifying or assisting in identifying whether a test sample is or contains the Puccinia horiana; D2-6) for preparing a product for identifying or assisting in identifying whether a test sample is or contains the Puccinia horiana.
10. An apparatus for designing a primer probe composition or a primer composition based on pan-genomics, characterized by: The device comprises the following modules: M1) target microorganism genome data receiving and analyzing module: for collecting the genome data of a target microorganism to construct a target microorganism genome database, and screening an optimal reference genome from the target microorganism genome database; M2) Target microorganism conserved gene analysis module: used to obtain a target microorganism conserved gene set from the target microorganism genome database; M3) Target microorganism approximate species genome data receiving module, used to collect target microorganism approximate species genome data to construct a target microorganism approximate species genome database; M4) Target microorganism specific gene analysis module, used to jointly analyze the best reference genome of M1) and the target microorganism approximate species genome database of M3) to obtain a target microorganism specific gene set; M5) Target gene analysis module, used to take the intersection of the target microorganism conserved gene set of M2) and the target microorganism specific gene set of M4) as a target gene; M6) Primer probe composition or primer composition design module, used to design primers or / and probes for the target gene sequence obtained by M5) using primer design software, and to screen theoretically feasible primer probe combinations according to the penalty value of primer3, the coverage of primers or / and probes, the specificity of primers or / and probes, and to determine the final primer probe composition or primer composition by combining experimental verification.
Citation Information
Patent Citations
A method and system for designing sequencing primers targeting pathogenic microorganisms
CN118762752B
A primer design method for multiplex PCR targeted sequencing technology
CN119296644B