Methods, systems, devices, and media for designing PCR primers for targeted sequencing

CN122551891APending Publication Date: 2026-08-11SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

这种策略仅适用于基因组小、菌株数量少的病毒/细菌,难以应对真菌、寄生虫等大基因组物种

Benefits of technology

本发明从基因组筛选到引物设计、引物质量评估、引物特异度评估和引物综合排序,提供了一套严谨的设计方法和评估体系,能快速设计出扩增性能好、覆盖度高和特异度高的引物。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122551891A_ABST
    Figure CN122551891A_ABST
Patent Text Reader

Abstract

This application discloses a method, system, device, and medium for designing PCR primers for targeted sequencing, belonging to the field of molecular biology technology. The design method of this application covers the entire process from genome screening to primer design, primer quality assessment, primer specificity assessment, and primer sequencing, providing a rigorous design methodology and evaluation system that enables the rapid design of primers with good amplification performance, high coverage, and high specificity. The design method and system of this application further establish a penalty mechanism related to amplification efficiency, thereby improving the quality and success rate of primer design. It also establishes a new evaluation index for primer specificity, and through simulation analysis of sequence differences between sequencing products and the establishment of a corresponding rating mechanism, significantly improves the sensitivity of primer design for monophyletic species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of molecular biology technology, and in particular to a method, system, device and medium for designing PCR primers for targeted sequencing. Background Technology

[0002] In fields such as clinical diagnosis, food safety monitoring, and prevention and control of emerging infectious diseases, rapid and accurate parallel detection of multiple pathogenic microorganisms is a core requirement. Traditional multiplex PCR technology, due to problems such as primer interference and uneven amplification efficiency, usually limits its multiplex detection capability to less than ten, making it difficult to handle the simultaneous screening of dozens or even hundreds of targets in complex samples.

[0003] In recent years, ultra-multiplex PCR targeted next-generation sequencing (tNGS) has emerged, which can increase the multiplet number to thousands or even tens of thousands, thus enabling the detection of hundreds or thousands of species at once. However, this is accompanied by strong interference between primers, and challenges include severe dimers, nonspecific amplification, and hairpin structures. This also means that a sufficient set of qualified primers needs to be designed for each species to provide replacements in case of primer failure or conflicts.

[0004] Current primer design integration software, including MultiPrime and DeGenPrime, employs a strategy of first performing multiple sequence alignment, then identifying conserved regions, and finally designing primers. This strategy is only suitable for viruses / bacteria with small genomes and a limited number of strains, and is ill-suited for large-genome species such as fungi and parasites. Furthermore, the explosive growth in genome sequences brought about by the rapid development of high-throughput sequencing in recent years means that the multiple sequence alignment step alone can take at least several days, resulting in low efficiency and insufficient throughput to meet the aforementioned challenges.

[0005] Furthermore, clinical applications typically use single-end 50bp sequencing, which means that after removing the primer region, only a very small segment of the product region remains to distinguish between specific and non-specific amplified sequences, which undoubtedly exacerbates the challenges of primer design.

[0006] Existing technologies only include the non-specific amplification rate of primer pairs in their evaluation indicators for primer specificity. They do not further analyze the distinguishability (number of different bases) between specific and non-specific amplification sequences when a large number of non-specific amplifications occur. Therefore, they do not have higher primer design sensitivity and cannot design specific primers for species within a monophyletic group (all species within a genus are highly similar, such as Streptococcus). Summary of the Invention

[0007] To solve at least one of the above-mentioned technical problems, the technical solution adopted in this application is as follows.

[0008] The first aspect of this application provides a method for designing PCR primers for targeted sequencing of a target species, comprising the following steps: Genome library construction: Obtain all known genomic information of the target species, generate a genome library, and obtain a representative genome; Primer design: Based on the representative genome, multiple primer pairs were designed to obtain candidate primer combinations; Coverage filtering: Based on the genome library, species coverage is calculated for each primer pair in the candidate primer combinations. Species coverage = (Number of genomes that the primer pair can cover / Number of genomes in the genome library) × 100%. Primer pairs with species coverage ≤30% were filtered out.

[0009] In some embodiments of this application, the target species is a microorganism, which may be, but is not limited to, pathogenic microorganisms, including bacteria, fungi, viruses, mycoplasma, etc.

[0010] In some embodiments of this application, the genome of the target species is obtained from the GenBank library, or the genome of the target species obtained by local sequencing may be added.

[0011] In some embodiments of this application, filtering of anomalous genomes is also included, wherein the anomalous genomes are those marked as Failed or Inconclusive in the taxonomy-check-status column of ANI_report_prokaryotes.txt provided by NCBI.

[0012] In some embodiments of this application, a step of filtering genomes based on assembly quality level is also included, specifically, retaining only genomes whose assembly_level column is labeled Chromosome, Complete Genome, or Scaffold.

[0013] In some embodiments of this application, the genome is further clustered and redundancy is removed, with the redundancy removal thresholds set as identity=0.95 and coverage=0.95, thereby obtaining a non-redundant genome library.

[0014] In this application, those skilled in the art can use any primer design software, such as Primer3, Primer Premier, etc., or not limited to one, to obtain a sufficient number of candidate primers.

[0015] In some embodiments of this application, the primer design parameters are set as follows: Primer length: 18bp~25bp; Primer temperature value Tm: 58~62℃; Primer GC content: 40%~60%; Amplification product length: 200bp~400bp.

[0016] Those skilled in the art will know that for primer design, the optimal primer length is 22 bp, the optimal primer Tm value is 60℃, and the optimal primer GC content is 50%.

[0017] In some embodiments of this application, the primer pair being able to cover means: i. The genome contains an upstream primer-binding region and a downstream primer-binding region that are sequentially located from the 3' end to the 5' end. The upstream primer-binding region can be aligned with the upstream primer of the primer pair and the number of mismatched bases is ≤6. The reverse complementary sequence of the downstream primer-binding region can be aligned with the downstream primer of the primer pair and the number of mismatched bases is ≤6. ii. The length of the genome from the 3' end of the upstream primer-binding region to the 5' end of the downstream primer-binding region is 50bp to 1200bp.

[0018] In some embodiments of this application, the number of primer pairs with species coverage ≥95%, ≥90%, ≥80%, and ≥70% is further counted and denoted as a, b, c, and d. If a < 2, b < 5, c < 15, and d < 40, then primer redesign is initiated. For each genome in the genome library, the number of primer pairs that can cover that genome is counted, i.e., the number of targets in that genome; Select several genomes with the lowest target count, design multiple primer pairs for each, and merge them with candidate primers designed based on the representative genomes for coverage evaluation again.

[0019] In some embodiments of this application, the design method further includes a primer pair redundancy removal step: if the amplification products of multiple primer pairs overlap, the primer pairs with low species coverage are filtered out one by one until the amplification products of all primer pairs do not overlap.

[0020] In some embodiments of this application, the design method further includes at least one of the following filtering steps: Multiple copy filtering: If one or more alignment results exist for a primer in a specified range other than the target position, the corresponding primer pair will be filtered out. Primer dimer filtering: If a primer pair poses a risk of dimerization, it is filtered out. Background genome nonspecific amplification filtering: If the primer pair can be aligned to the background genome, it is filtered out.

[0021] In some specific embodiments of this application, the specified range refers to a range of 800 bp upstream and downstream of the target location.

[0022] In some embodiments of this application, dimer risk analysis software is used to determine whether there is a dimer risk in the primer pair. For example, the dimer module of MFEprimer-3.2.6 is used to analyze the dimer risk of each primer pair and obtain a score value. If the score is ≥10, then there is a dimer risk.

[0023] In some specific embodiments of this application, the background genome refers to the genomes of other species that can be obtained simultaneously or frequently and with high probability when obtaining the genome of the target species. For example, if the target species is a microorganism, then the background genome is the host genome. In this case, if the primer pair can match the host genome, there may be a risk of host non-specific amplification, and the primer pair needs to be filtered out.

[0024] In some specific embodiments of this application, "able to align to the background genome" means: i. Number of base matches ≥ 13bp; ii. The number of mismatched bases in the 3' terminal 5bp is ≤2bp; iii. The length of the amplified product is 50bp~1200bp.

[0025] In some embodiments of this application, the above-mentioned multicopy filtering, primer dimer filtering, and background genome nonspecific filtering steps can be performed before or after the primer pair redundancy removal step, and that is, they can be performed before or after the coverage filtering step.

[0026] In some embodiments of this application, the aforementioned multicopy filtering, primer dimer filtering, and background genome nonspecific filtering steps can also be separated into other steps. However, since the primer pair redundancy removal step needs to be based on species coverage, the primer pair redundancy removal step must be placed after the species coverage calculation step, and the order of the remaining steps is not particularly restricted.

[0027] In some specific embodiments of this application, the steps are performed in the following order: coverage filtering, primer pair redundancy removal, multiple copy filtering, primer dimer filtering, and background genome non-specific filtering. In other specific embodiments, the steps are performed in the following order: multiple copy filtering, primer dimer filtering and background genome non-specific filtering, coverage filtering, and primer pair redundancy removal. In still other specific embodiments, the steps are performed in the following order: coverage filtering, multiple copy filtering, primer dimer filtering, primer pair redundancy removal, and background genome non-specific filtering. In this application, the primer pair redundancy removal step may also be omitted, and the other steps can be performed in the order described above.

[0028] In some embodiments of this application, for any primer pair, the following steps are also included: The primer pairs were aligned to the NCBI NT database to obtain the total aligned sequence of the primer pairs. Based on the species information of the target sequence, the total aligned sequence was divided into specific amplification sequences and non-specific amplification sequences. The primer alignment regions on the specific amplification sequence and the non-specific amplification sequence were removed to obtain the effective analytical sequences of the specific amplification sequence and the non-specific amplification sequence, respectively. These sequences were then compared, and the sequence differences were analyzed and the number of differing bases was counted. If the number of non-specific amplified sequences with ≤1 differential bases is greater than 2, then the primer pair should be filtered out.

[0029] In this application, the primer alignment region refers to the region that binds to the primer.

[0030] In some embodiments of this application, each sequence in the total alignment sequence is taken as the read length of next-generation sequencing and then subjected to subsequent excision alignment analysis. For example, if the read length is 50 bp, then both the specific amplification sequence and the non-specific amplification sequence are 50 bp in length, extending from the primer comparison region towards the other primer until the length reaches 50 bp. If the primer length is 22 bp, then the extension length is 28 bp, and in this case, the extended 28 bp is the valid analysis sequence.

[0031] In some embodiments of this application, the PCR primers are used for targeted sequencing, such as next-generation sequencing. In this case, only the sequencing lengths at both ends of the total alignment sequence are taken as the specific amplification sequence or the non-specific amplification sequence.

[0032] In some embodiments of this application, the primer pairs are required to be the same as those in the NCBI NT library: i. Number of base matches ≥ 13bp; ii. The number of mismatched bases in the 3' terminal 5bp is ≤2bp; iii. The length of the amplified product is 50bp-1200bp.

[0033] In some embodiments of this application, the design method further includes at least one of the following analytical steps: Primer sequence conservation analysis: The diversity sequence of each primer pair is obtained. If there are ≥3 diverse sequences with more than 2 mismatched bases, a penalty of m points is applied to the corresponding primer pair. If there are ≥3 diverse sequences with more than 1 mismatched base at the 3' end, a penalty of n points is applied to the corresponding primer pair. The diversity sequence count and primer sequence conservation penalty are accumulated to obtain the primer pair's conservatism score, where m... <n; Primer sequence feature analysis: ① If a primer contains a dinucleotide repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a dinucleotide repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ② If a primer contains a continuous base repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a continuous base repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ③ If the number of G / C bases in the 5 bp region at the 3' end is ≥3, a penalty of p points is applied to the corresponding primer pair. ④ If the last base at the 3' end is A / T, a penalty of p points is applied to the corresponding primer pair. The primer sequence feature penalty points for each primer pair are accumulated, where p... <q。

[0034] In some embodiments of this application, the diverse sequences of the primer pairs are obtained based on the following steps: For any primer pair, for each genome it can cover, the sequence of the upstream primer binding region is recorded as the diversity sequence of the upstream primer of the primer pair; the reverse complementary sequence of the downstream primer binding region is recorded as the diversity sequence of the downstream primer of the primer pair. The diversity sequences of the upstream primer and the diversity sequences of the downstream primer together constitute the diversity sequence of the primer pair.

[0035] In some embodiments of this application, before the primer sequence conservation analysis and primer sequence feature analysis are performed, the primer sequence conservation penalty and primer sequence feature penalty of the primer pair are also used for primer pair redundancy removal.

[0036] For primer pairs with overlapping amplification products, they are sorted sequentially by species coverage (from high to low), primer sequence feature penalty (from low to high), and primer sequence conservation penalty (from low to high). Primer pairs ranked lower are removed first until all primer pairs have no overlapping amplification products. In some preferred embodiments of this application, the redundancy removal step ensures that the distance between the amplification products of all primer pairs is ≥200 bp.

[0037] In this application, the order of species coverage from high to low, primer sequence feature penalty from low to high, and primer sequence conservation penalty from low to high means that, firstly, primer pairs are sorted by species coverage from high to low; for pairs with the same coverage, they are sorted by primer sequence feature penalty from low to high; furthermore, for pairs with the same primer sequence feature penalty, they are sorted by primer sequence conservation penalty from low to high.

[0038] In some embodiments of this application, the diversity sequence number, primer sequence conservation penalty, and primer sequence feature penalty of the primer pair are also used for the final selection of primer pairs, specifically: Primer pairs were ranked in descending order of species coverage, primer sequence feature penalty score, primer sequence conservation penalty score, number of diverse sequences, and FDR (free radix) score. The primer pairs with the highest rankings were selected as PCR primers for targeted sequencing of the target species. Wherein, FDR = number of non-specific amplified sequences / total number of aligned sequences.

[0039] In this application, sorting by PDR from smallest to largest has the same meaning as sorting by PPV from largest to smallest, where PPV = number of specifically amplified sequences / total number of aligned sequences.

[0040] A second aspect of this application provides a PCR primer design system for targeted sequencing of a target species, comprising the following modules: The genome library input storage module is used to obtain all known genome information of the target species and store it as a genome library, while marking representative genomes; The primer design module, connected to the genome library input storage module, is used to design multiple primer pairs based on the representative genome to obtain candidate primer combinations; Evaluation module: Connected to both the genome library input storage module and the primer design module, and used for: Based on the aforementioned genome library, species coverage was calculated for each primer pair in the candidate primer combinations: Species coverage = Number of genomes that the primer pair can cover / Number of genomes in the genome library × 100% Primer pairs with species coverage ≤30% were filtered out.

[0041] In some embodiments of this application, the evaluation module is further used to count the number of primer pairs with species coverage ≥95%, ≥90%, ≥80%, and ≥70%, denoted as a, b, c, and d. If a < 2, b < 5, c < 15, and d < 40, then: For each genome in the genome library, the number of primer pairs that can cover the genome is counted, i.e. the number of targets of the genome. The genomes with the lowest number of targets are selected and fed back to the primer design module. The primer design module is also used to design multiple primer pairs based on the several genomes respectively, and after merging them with candidate primers designed based on the representative genome, they are output to the evaluation module again for coverage evaluation.

[0042] In some embodiments of this application, the evaluation module is also used for primer redundancy removal: if the amplification products of multiple primer pairs overlap, primer pairs with low species coverage are filtered out one by one until the amplification products of all primer pairs do not overlap.

[0043] In some embodiments of this application, the evaluation module is also used to implement at least one of the following filters: Multiple copy filtering: If one or more alignment results exist for a primer in a specified range other than the target position, the corresponding primer pair will be filtered out. Primer dimer filtering: If primer pairs are present, they are filtered out; Background genome nonspecific amplification filtering: If the primer pair can be aligned to the background genome, it is filtered out.

[0044] In some specific embodiments of this application, the specified range refers to a range of 800 bp upstream and downstream of the target location.

[0045] In some embodiments of this application, the evaluation module is further configured to perform the following steps for any primer pair: The primer pairs were aligned to the NCBI NT database to obtain the total aligned sequence of the primer pairs. Based on the species information of the target sequence, the total aligned sequence was divided into specific amplification sequences and non-specific amplification sequences. The primer alignment regions on the specific amplification sequence and the non-specific amplification sequence were removed to obtain the effective analytical sequences of the specific amplification sequence and the non-specific amplification sequence, respectively. These sequences were then compared, and the sequence differences were analyzed and the number of differing bases was counted. If the number of non-specific amplified sequences with ≤1 differential bases is greater than 2, then the primer pair should be filtered out.

[0046] In some embodiments of this application, the evaluation module is also used to perform at least one of the following analysis steps: Primer sequence conservation analysis: The diversity sequence of each primer pair is obtained. If there are ≥3 diverse sequences with more than 2 mismatched bases, a penalty of m points is applied to the corresponding primer pair. If there are ≥3 diverse sequences with more than 1 mismatched base at the 3' end, a penalty of n points is applied to the corresponding primer pair. The diversity sequence count and primer sequence conservation penalty are accumulated to obtain the primer pair's conservatism score, where m... <n; Primer sequence feature analysis: ① If a primer contains a dinucleotide repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a dinucleotide repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ② If a primer contains a continuous base repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a continuous base repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ③ If the number of G / C bases in the 5 bp region at the 3' end is ≥3, a penalty of p points is applied to the corresponding primer pair. ④ If the last base at the 3' end is A / T, a penalty of p points is applied to the corresponding primer pair. The primer sequence feature penalty points for each primer pair are accumulated, where p... <q。

[0047] In some embodiments of this application, the dinucleotide repeat refers to a dinucleotide repeat that is repeated at least 3 times, and the continuous base repeat refers to a continuous base repeat that is repeated 4 times or more.

[0048] In some embodiments of this application, the evaluation module is further configured to sort primer pairs sequentially by coverage from largest to smallest, primer sequence feature penalty from smallest to largest, primer sequence conservation penalty from smallest to largest, number of diverse sequences from smallest to largest, and FDR from smallest to largest, and select the top-ranked primer pairs as PCR primers for targeted sequencing of the target species and output them. Wherein, FDR = number of non-specific amplified sequences / total number of aligned sequences.

[0049] A third aspect of this application provides a computer device, comprising: a memory for storing a computer program; and a processor for executing the steps of any of the design methods described in the first aspect of this application when executing the computer program.

[0050] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of any of the design methods described in the first aspect of this application.

[0051] Compared with the prior art, this application has the following advantages: This invention provides a rigorous design method and evaluation system from genome screening to primer design, primer quality assessment, primer specificity assessment, and primer comprehensive sequencing, enabling the rapid design of primers with good amplification performance, high coverage, and high specificity.

[0052] The design method and system of this application further establish a primer sequence conservation penalty mechanism related to amplification efficiency, and select primer pairs with less penalty (i.e. more conserved sequences), thereby improving the quality of primer design.

[0053] The design method and system of this application also establish a primer sequence feature penalty mechanism related to amplification efficiency, and select primer pairs with less penalty, thereby improving the quality and success rate of primer design.

[0054] The design method and system of this application also establish a new evaluation index for primer specificity. By simulating and analyzing the sequence differences between sequencing products and establishing a corresponding rating mechanism, the sensitivity of primer design for monophyletic species is significantly improved.

[0055] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0056] The above and other objects, features, and advantages of exemplary embodiments of this application will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. Several embodiments of this application are illustrated in the drawings by way of example and not limitation, in which: Figure 1 A schematic diagram of the primer design process in Embodiment 1 of this application is shown. Detailed Implementation

[0057] Unless otherwise stated, implied from the context, or as is customary in the art, all parts and percentages in this application are based on weight, and all testing and characterization methods used are concurrent with the filing date of this application. Where applicable, any patent, patent application, or disclosure relating to this application is incorporated herein by reference in its entirety, and its equivalent patent families are also incorporated herein by reference, particularly the definitions of relevant terms in the art disclosed in such documents. If any definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition provided in this application shall prevail.

[0058] To make the technical problems, technical solutions and beneficial effects solved by this application clearer, the following detailed description is provided in conjunction with embodiments.

[0059] The following examples are used to illustrate preferred embodiments of this application. Those skilled in the art will understand that the techniques disclosed in the examples represent technologies discovered by the inventors that can be used to implement this application, and therefore can be considered preferred embodiments of this application. However, those skilled in the art should understand from this specification that many modifications can be made to the specific embodiments disclosed herein, still yielding the same or similar results, without departing from the spirit or scope of this application.

[0060] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains, and all materials cited herein and referenced by them are incorporated herein by reference.

[0061] Those skilled in the art will recognize, or can learn through routine experimentation, many equivalents of the specific embodiments of the invention described herein. These equivalents will be included in the claims.

[0062] Unless otherwise specified, the experimental methods used in the following examples are conventional methods. Unless otherwise specified, the instruments and equipment used in the following examples are all conventional laboratory instruments and equipment; unless otherwise specified, the experimental materials used in the following examples were all purchased from conventional biochemical reagent stores.

[0063] Example 1: Design and validation of multiplex PCR primers for Staphylococcus aureus refer to Figure 1 This embodiment provides a multiplex PCR primer design process.

[0064] I. Construction of Genome Library (1) Determine the Latin name and taxid of the target species (obtained from the taxonomy database). For Staphylococcus aureus, its Latin name is 金黄色葡萄球菌 The taxid is 1280; (2) Genome download: The initial genome of Staphylococcus aureus was obtained from the GenBank repository (ftp: / / ftp.ncbi.nlm.nih.gov / genomes), totaling 120,510 entries; (3) Genome filtering: ① Abnormal Genome Removal: The ANI_report_prokaryotes.txt file provided by NCBI (https: / / ftp.ncbi.nlm.nih.gov / genomes / ASSEMBLY_REPORTS / ANI_report_prokaryotes.txt) records detailed quality control information for microbial strain genome data. Genomes marked as Failed or Inconclusive are removed based on the taxonomy-check-status column. ② Assembly quality filtering: Only genomes labeled "Chromosome," "Complete Genome," and "Scaffold" in the `assembly_level` column of the GenBank library are included, resulting in a genome library containing 120,510 genomes. Furthermore, genomes labeled "reference genome" in the `refseq_category` column are used as representative genomes (items) for the target species. In this embodiment, the reference genome is 2.81 MB in length.

[0065] (4) De-redundancy of genomes: Use psi-cd-hit to perform genome clustering to remove redundancy. The thresholds are: identity = 0.95 and coverage = 0.95, resulting in a non-redundant genome library (unit) containing 1439 genomes.

[0066] II. Primer Design and Filtering (1) Primer design Primer design was performed using Primer3 based on the item, with the following parameters set: i. Primer length: 22bp; ii. Primer temperature value Tm: 58-62℃; iii. Primer GC content: 40%~60%; iv. Amplification product length: 200bp~400bp.

[0067] Output 1000 initial primer pairs.

[0068] (2) Primer pair filtration ①Multiple copy filtering Each primer pair was aligned to the item using blast v2.16. The alignment threshold was set at ≤2 mismatched bases. If the upstream and / or downstream primers in the primer pair had one or more alignment results within 800bp upstream and downstream of the target position, the primer pair was considered to have an extremely high risk of off-target amplification and was filtered out.

[0069] ② Primer dimer filtration The dimer module of MFEprimer-3.2.6 was used to analyze the dimer risk of each primer pair, obtain the score value, and filter primer pairs with a score ≥ 10.

[0070] ③ Low coverage filtration Input "unit" and use usearch v11 to perform genome coverage analysis on primer pairs. For a genome, the requirement for primer pairs to cover it is: i. The genome contains an upstream primer-binding region and a downstream primer-binding region that are sequentially located from the 3' end to the 5' end. The upstream primer-binding region can be aligned with the upstream primer of the primer pair and the number of mismatched bases is ≤6. The reverse complementary sequence of the downstream primer-binding region can be aligned with the downstream primer of the primer pair and the number of mismatched bases is ≤6. ii. The length of the amplified product from the 3' end of the upstream primer binding region to the 5' end of the downstream primer binding region on the genome is 50bp~1200bp.

[0071] Calculate the species coverage of each primer pair (the percentage of genomes that the primer pair can cover out of the total number of genomes in the unit), filter out primer pairs with species coverage ≤30%, leaving 653 primer pairs.

[0072] Count the number of primer pairs with species coverage ≥95%, ≥90%, ≥80%, and ≥70%, respectively, and denote them as a, b, c, and d. If a < 2, b < 5, c < 15, and d < 40, primer redesign needs to be initiated.

[0073] In this embodiment, there are a total of 85 primer pairs with species coverage ≥95%, i.e., a=85, b, c, and d are at least 85, so no primer redesign is required.

[0074] III. Primer Pair Quality Assessment (1) Primer pair conservation analysis For any primer pair, for each genome it covers, the sequence of the upstream primer binding region is recorded as the diversity sequence of the upstream primer of that primer pair, and the number of mismatched bases is counted; the reverse complementary sequence of the downstream primer binding region is recorded as the diversity sequence of the downstream primer of that primer pair, and the number of mismatched bases is counted. For that primer pair, the diversity sequence number of the upstream primer is added to the diversity sequence number of the downstream primer to obtain the diversity sequence number of the primer pair.

[0075] For any primer pair, ① if there are ≥3 diverse sequences with more than 2 mismatched bases, then the primer pair is penalized 1 point; ② if there are ≥3 diverse sequences with more than 1 mismatched base at the 3' end, then the primer pair is penalized 2 points. Thus, the primer sequence conservation penalty score of the primer pair is obtained.

[0076] The number of mismatched bases mentioned above refers to the number of bases in the diverse sequence that differ from the corresponding bases in the original primer.

[0077] (2) Primer pair sequence characteristic analysis For any primer pair: ① Identify whether there are more than 3 dinucleotide repeats in each primer (such as ATATAT, CGCGCGCG). If so, further determine the distance (number of bases) from the dinucleotide repeat to the 3' end. If it exists but the distance from the 3' end is >5bp, then penalize the primer pair with 1 point. If it exists and the distance from the 3' end is ≤5bp, then penalize the primer pair with 2 points. ② Identify whether there are more than 4 consecutive base repeats in each primer (such as GGGG, AAAAAA). If so, further determine the distance (number of bases) from the 3' end of the consecutive base repeat. If it exists but the distance from the 3' end is >5bp, then penalize the primer pair with 1 point. If it exists and the distance from the 3' end is ≤5bp, then penalize the primer pair with 2 points. ③ Determine whether the number of G / C bases in the 5bp at the 3' end is ≥3. If so, deduct 1 point from the primer pair. ④ Determine whether the last base at the 3' end is A / T. If it is, deduct 1 point from the primer pair.

[0078] Finally, all penalties for the primer pair are accumulated, thus obtaining the primer sequence feature penalty for the primer pair.

[0079] IV. Primer pair redundancy removal To analyze the overlap between amplification products, primer pairs with overlapping amplification products were sorted in order of species coverage from high to low, primer sequence feature penalty from low to high, and primer sequence conservation penalty from low to high. Primer pairs with lower rankings were removed first until the spacing between the amplification products of all primer pairs was ≥200bp.

[0080] In this embodiment, after removing the last 332 primer pairs, the amplification products of the remaining 321 primer pairs all have a spacing of ≥200bp, which meets the requirements.

[0081] V. Primer Pair Specificity Assessment (1) Assessment of host nonspecific amplification The primers were aligned to the host reference genome (in this example, the human reference genome hg38+ human transcript library refMrna) using the software blast v2.16 (MFEprimer v3.2.6 can also be used). The alignment threshold was set as follows: i. Number of base matches ≥ 13bp; ii. The number of mismatched bases in the 3' terminal 5bp is ≤2bp; iii. The length of the amplified product is 50bp~1200bp.

[0082] If the primer pair can match the host genome, there is a risk of non-specific host amplification, and it should be filtered out.

[0083] (2) Assessment of nonspecific amplification of microorganisms ① Further align the remaining primers to the NCBI NT library (ftp: / / ftp.ncbi.nlm.nih.gov / blast / db / FASTA / nt.gz), and set the matching threshold as follows: i. Number of base matches ≥ 13bp; ii. The number of mismatched bases in the 3' terminal 5bp is ≤2bp; iii. The length of the amplified product is 50bp~1200bp.

[0084] For each primer pair, analyze its simulated amplification results to obtain the total contigs of that primer pair.

[0085] ② Based on the species information of the target sequences being compared, the total compared sequences are divided into: specific amplification sequences (TP_match) and non-specific amplification sequences (FP_match). ③ Statistical results were obtained for specificity (PPV) and nonspecificity (FDR), where PPV = TP_match / total contigs and FDR = FP_match / total contigs.

[0086] (3) Evaluation of the discriminative power of nonspecific amplification products For each primer pair: ① The primer alignment region (primer binding region) (22bp) was removed from both the specific amplification sequence set and the non-specific amplification sequence set to obtain the effective analysis sequence set of the specific amplification sequence and the effective analysis sequence set of the non-specific amplification sequence, respectively. ② By setting the parameters SE50 (default), SE75 or SE100, only the first 28bp, 53bp or 78bp of the effective analysis sequence are retained, and the sequence differences between the effective analysis sequence set of the above specific amplification sequence and the effective analysis sequence set of the non-specific amplification sequence are analyzed. ③ Count the number of valid analysis sequences of non-specific sequences with a difference of ≤1 between the valid analysis sequences and the specific amplified sequences. If the number of sequences is >2, mark the primer pair as discard; otherwise, mark it as pass.

[0087] After filtering out the 210 primer pairs marked as discard, 111 primer pairs remained.

[0088] VI. Primer Output Primer pairs were sorted in the following order: species coverage (from largest to smallest), primer sequence feature penalty (from smallest to largest), primer sequence conservation penalty (from smallest to largest), number of diverse sequences (from smallest to largest), and FDR (from smallest to largest). The earlier the ranking, the better the performance.

[0089] After sorting, the top 5 primer pairs were selected (Table 1) to form a primer combination, which was then used for multiplex PCR.

[0090] Table 1: Information on the designed multiplex PCR primers for Staphylococcus aureus ; VII. Primer Performance Verification DNA extraction, PCR amplification, library preparation, and sequencing on a next-generation sequencer (BGIseq) were performed using three known Staphylococcus aureus positive samples and three closely related Staphylococcus epidermidis samples.

[0091] Analysis of the sequencing data showed that all primers could detect Staphylococcus aureus positive samples, with a sensitivity of 100%. Of the other three Staphylococcus epidermidis samples, only Primer03 primers detected two of them, and the reads accurately identified them as Staphylococcus epidermidis, with no false positives and a specificity of 100%.

[0092] Example 2: Design and Validation of Multiplex PCR Primers for Dengue Virus The design and evaluation methods described in Example 1 are followed, but excluding the "primer redundancy removal" step, and the amplification product length is set to 150bp~200bp.

[0093] I. Construction of Genome Library The Latin name for dengue virus is... 登革病毒 The taxid is 12637. A total of 4519 genomes were initially retrieved from the GenBank database. After genome filtering (including abnormal genome removal and assembly quality filtering) and genome redundancy removal, 113 genomes were left, forming a non-redundant genome unit. The reference genome (item) is 10.7KB in length.

[0094] II. Primer Design and Evaluation Referring to the method in Example 1, after primer design and filtering, 714 primer pairs were obtained. Species coverage was statistically analyzed, where a=0, b=0, c=1, and d=10. These pairs satisfy the conditions a<2, b<5, c<15, and d<40, requiring primer redesign. Specifically: Count the number of primer pairs that hit the target genome in each unit (a primer pair that can cover the genome is considered to be hit), and sort them from low to high frequency of hits. The 10 genomes with the lowest number of targets were used as items. Primers were designed and filtered using the same steps and parameters as those used for the representative genomes. Finally, all the obtained primer pairs were merged and primer quality was evaluated and redundancy was removed using the method in Example 1, resulting in 241 primer pairs.

[0095] Finally, primer pair specificity was evaluated according to Example 1. After filtering out primer pairs marked as discard, 152 primer pairs remained.

[0096] III. Primer Output and Performance Verification Referring to Example 1, the above 152 primer pairs were sorted, and the top 5 primer pairs (Table 2) were selected to form a primer combination, which was then used for multiplex PCR.

[0097] Table 2: Information on the designed multiplex PCR primers for dengue virus ; In Table 2, since the original primers have multiple diverse sequences, degenerate bases are assigned to specific positions according to the degenerate base rules to obtain degenerate primers; only the degenerate bases within 10 bp of the 3' end are recorded and retained.

[0098] Using five known dengue virus positive samples, RNA extraction, reverse transcription, PCR amplification, library preparation, and sequencing were performed.

[0099] The sequencing data were analyzed, and the results are shown in Table 3.

[0100] Table 3: Statistical Table of Sample Detection ; RPM (Reads per million mapped reads) refers to data homogenization based on a total sequencing data volume of 1Mb reads. The specific calculation formula is: RPM = Number of species annotation reads × 10 6 / Total number of reads for this sample; Criteria for positive detection: RPM ≥ 3.

[0101] As shown in Table 3, primer pair 7 was detected in all 5 samples (sensitivity 100%), and the number of reads detected was significantly higher than the other 4 primer pairs. Primer pairs 9 and 10 had low coverage, resulting in some missed detections of viral strains, with 2 cases detected (sensitivity 40%) and 1 positive case detected (sensitivity 20%), respectively. Although primer pair 6 had the highest coverage, its diversity sequence count was as high as 10, which was an important reason for the decline in amplification performance.

[0102] Based on the above results, the evaluation indicators affecting primer performance in the primer design methods described above are relatively objective and accurate.

[0103] Example 3: Design and validation of multiplex PCR primers for oral streptococci Similarly, primer design and evaluation were performed with reference to Example 1.

[0104] I. Construction of Genome Library The Latin name for oral streptococci is 口腔链球菌 The taxid is 1303. The initial genomes retrieved from the GenBank database totaled 328. After genome filtering (including abnormal genome removal and assembly quality filtering) and genome redundancy removal, 105 genomes remained, forming a non-redundant genome unit. The reference genome (item) is 1.93 MB in length.

[0105] II. Primer Design and Evaluation Analysis Following the method in Example 1, 376 primer pairs were obtained through primer design and filtering, of which 37 pairs had a coverage of ≥95%, eliminating the need for primer redesign.

[0106] Primer pair quality assessment and redundancy removal were performed according to Example 1, leaving 84 primer pairs.

[0107] Finally, primer pair specificity was evaluated according to Example 1, and primer pair sets E_pass and E_discard, labeled as pass and discard, were obtained and sorted respectively.

[0108] III. Primer Output and Performance Verification The top 3 primer pairs from E_pass and E_discard were selected to form a primer combination (Table 4), which was then used for multiplex PCR.

[0109] Table 4: Information on the designed multiplex PCR primers for oral streptococci ; Three known clinically positive samples of oral streptococci, including one clinically positive sample each of closely related species such as Streptococcus pneumoniae and Streptococcus pyogenes, were used. DNA extraction, PCR amplification, library preparation, and sequencing were performed.

[0110] The sequencing data were analyzed, and the results are shown in Table 5.

[0111] Table 5: Statistical Table of Sample Detection ; In this diagram, "+" indicates a positive result (RPM≥3), and "-" indicates a negative result; gray shading indicates a false positive result.

[0112] As shown in Table 5, primer pairs categorized as "pass" did not result in false positives (detection of non-target species) in samples 4 and 5. Primer pairs 11 and 13 failed to amplify Streptococcus pneumoniae reads in sample 4, consistent with the nonspecificity (FDR%) assessment results. However, primer pairs categorized as "discard" resulted in false positives in all samples, and the frequency of these false positives was highly consistent with the nonspecificity (FDR%) assessment results.

[0113] The results above demonstrate that the method is accurate in assessing primer specificity and non-specific amplification product differentiation, and significantly improves the ability to design specific primers for species within monophyletic groups.

[0114] Furthermore, it should be understood that after reading the foregoing contents of this application, those skilled in the art can make various alterations or modifications to this application, and these equivalent forms also fall within the scope defined by the appended claims.

Claims

1. A method for designing PCR primers for targeted sequencing of a species of interest, characterized in that, Includes the following steps: Genome library construction: Obtain all known genomic information of the target species, generate a genome library, and obtain a representative genome; Primer design: Based on the representative genome, multiple primer pairs were designed to obtain candidate primer combinations; Coverage filtering: Based on the genome library, species coverage is calculated for each primer pair in the candidate primer combinations. Species coverage = (Number of genomes that the primer pair can cover / Number of genomes in the genome library) × 100%, Primer pairs with species coverage ≤30% were filtered out.

2. The design method of claim 1, wherein, Further count the number of primer pairs with species coverage ≥95%, ≥90%, ≥80%, and ≥70%, denoted as a, b, c, and d. If a < 2, b < 5, c < 15, and d < 40, then initiate primer redesign. For each genome in the genome library, the number of primer pairs that can cover that genome is counted, i.e., the number of targets in that genome; Select several genomes with the lowest target count, design multiple primer pairs for each, and merge them with candidate primers designed based on the representative genomes for coverage evaluation again.

3. The design method according to claim 1 or 2, characterized in that, Also includes: Primer pair redundancy removal: If the amplification products of multiple primer pairs overlap, the primer pairs with low species coverage are filtered out one by one until the amplification products of all primer pairs do not overlap.

4. The design method according to claim 1 or 2, characterized in that, It also includes at least one of the following filtering steps: Multiple copy filtering: If one or more alignment results exist for a primer in a specified range other than the target position, the corresponding primer pair will be filtered out. Primer dimer filtering: If a primer pair poses a risk of dimerization, it is filtered out. Background genome nonspecific amplification filtering: If the primer pair can be aligned to the background genome, it is filtered out.

5. The design method according to claim 1 or 2, characterized in that, For any primer pair, the following steps are also included: The primer pairs were aligned to the NCBI NT database to obtain the total aligned sequence of the primer pairs. Based on the species information of the target sequence, the total aligned sequence was divided into specific amplification sequences and non-specific amplification sequences. The primer alignment regions on the specific amplification sequence and the non-specific amplification sequence were removed to obtain the effective analytical sequences of the specific amplification sequence and the non-specific amplification sequence, respectively. These sequences were then compared, and the sequence differences were analyzed and the number of differing bases was counted. If the number of non-specific amplified sequences with ≤1 differential bases is greater than 2, then the primer pair should be filtered out.

6. The design method according to claim 5, characterized in that, It also includes at least one of the following analytical steps: Primer sequence conservation analysis: The diversity sequence of each primer pair is obtained. If there are ≥3 diverse sequences with more than 2 mismatched bases, a penalty of m points is applied to the corresponding primer pair. If there are ≥3 diverse sequences with more than 1 mismatched base at the 3' end, a penalty of n points is applied to the corresponding primer pair. The diversity sequence count and primer sequence conservation penalty are accumulated to obtain the primer pair's conservatism score, where m... <n; Primer sequence feature analysis: ① If a primer contains a dinucleotide repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a dinucleotide repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ② If a primer contains a continuous base repeat but the number of bases from the 3' end is >5, a penalty of p points is applied to the corresponding primer pair; if a primer contains a continuous base repeat and the number of bases from the 3' end is ≤5 bp, a penalty of q points is applied to the corresponding primer pair. ③ If the number of G / C bases in the 5 bp region at the 3' end is ≥3, a penalty of p points is applied to the corresponding primer pair. ④ If the last base at the 3' end is A / T, a penalty of p points is applied to the corresponding primer pair. The primer sequence feature penalty points for each primer pair are accumulated, where p... <q。 7. A PCR primer design system for targeted sequencing of a target species, characterized in that, Includes the following modules: The genome library input storage module is used to obtain all known genome information of the target species and store it as a genome library, while marking representative genomes; A primer design module, connected to the genome library input storage module, is used to design multiple primer pairs based on the representative genome to obtain candidate primer combinations; Evaluation module: Connected to both the genome library input storage module and the primer design module, and used for: Based on the aforementioned genome library, species coverage was calculated for each primer pair in the candidate primer combinations: Species coverage = (Number of genomes that the primer pair can cover / Number of genomes in the genome library) × 100%, Primer pairs with species coverage ≤30% were filtered out.

8. The design system according to claim 7, characterized in that, The evaluation module is also used to count the number of primer pairs with species coverage ≥95%, ≥90%, ≥80%, and ≥70%, denoted as a, b, c, and d. If a < 2, b < 5, c < 15, and d < 40, then: For each genome in the genome library, the number of primer pairs that can cover the genome is counted, i.e. the number of targets of the genome. The genomes with the lowest number of targets are selected and fed back to the primer design module. The primer design module is also used to design multiple primer pairs based on the several genomes respectively, and after merging them with candidate primers designed based on the representative genome, they are output to the evaluation module again for coverage evaluation.

9. A computer device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the steps of the design method according to any one of claims 1 to 6 when executing the computer program.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the design method according to any one of claims 1 to 6.